API asset processing method, system and equipment based on sample entropy and load entropy, and medium

By combining sample entropy and load entropy-based methods with URL prefix matching and traffic packet analysis, the problem of inaccurate interface identification in API asset management was solved, achieving efficient merging and security detection, and improving the accuracy and security of API asset management.

CN120915593AActive Publication Date: 2025-11-07南京中孚信息技术有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511384349.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-11-07
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing technologies for API asset management suffer from inaccurate interface identification, leading to redundancy and misjudgment of the asset database, and making it difficult to distinguish between normal dynamic interfaces and abnormal traffic.

Method used

By employing a method based on sample entropy and load entropy, the load entropy and sample entropy values ​​of HTTP interface information are calculated, and combined with URL prefix matching and traffic packet analysis, interface merging and anomaly detection are achieved.

Benefits of technology

It improves the accuracy and efficiency of API asset management, reduces resource consumption, enhances security, and can accurately identify similar dynamic interfaces and isolate abnormal traffic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120915593A_ABST
    Figure CN120915593A_ABST
Patent Text Reader

Abstract

The invention provides an API asset processing method, system and device based on sample entropy and load entropy and a medium, and belongs to the technical field of API asset processing.The method comprises the steps that firstly, HTTP interface information is collected, and load entropy quantification content randomness is calculated; if the load entropy is smaller than the threshold value, judging that the interface is a static interface and directly finding or collecting; otherwise, comparing the load entropy mean value difference with the same URL prefix interface, and if the difference is small, merging. And for residual dynamic interfaces, calculating sample entropy quantification calling sequence time regularity, and performing classification processing according to entropy values: forced merging when the entropy value is smaller than 1.2, parameterization when the entropy value is smaller than 1.2, and isolation alarm when the entropy value is larger than 2.5. Static interfaces are quickly filtered through load entropy, and the subsequent calculation amount is reduced; the sample entropy is combined with time law analysis, so that the limitation of path matching is avoided, the dynamic interface merging accuracy is improved, and the repeated record of assets is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of API asset processing, and particularly relates to an API asset processing method, system, device and medium based on sample entropy and load entropy. BACKGROUND

[0002] In the field of network security, API has become an important entry for the expansion of attack surface. The API asset list established by general enterprises supports security capabilities such as vulnerability scanning, access control, and abnormal behavior detection. For large enterprises, usually hundreds to thousands of microservices are operated, and each service exposes multiple API endpoints. Unified registration, version management, call monitoring and permission control of these interfaces are the premise of realizing efficient service governance.

[0003] At present, Shannon entropy is introduced to realize network security and traffic analysis, which is used to measure the randomness and uncertainty of data. However, it depends on URL complete matching. If the static interface URL contains dynamic interfaces similar to time stamp, it will be misjudged as multiple independent interfaces, resulting in asset library redundancy. For example, the fixed picture interface of CDN service will be recorded as 3 different interfaces due to the addition of random parameters in the front end.

[0004] There are also interface upgrades, but the URL prefix changes, and the core function does not change. URL matching cannot identify the same logical interface, resulting in scattered asset library. From the version iteration of the user information interface of an e-commerce system, the URL prefixes of version 1 and version 2 are different, and the traditional method will regard them as two independent interfaces. Moreover, malicious scanning requests and normal dynamic interfaces are difficult to distinguish, which may misjudge abnormal traffic as normal interface, or mistakenly delete normal interface due to random content, causing interface misjudgment and affecting normal use. SUMMARY

[0005] The application provides an API asset processing method based on sample entropy and load entropy. The method optimizes resource use efficiency while ensuring identification accuracy, and improves the standardization and security of API asset management.

[0006] The method comprises the following steps: S101, acquiring HTTP interface information; S102, analyzing the load entropy value of the HTTP interface information, wherein the load entropy value is used to quantify the randomness of the interface call content; S103, making a filtering decision according to the load entropy value, and if the load entropy value is greater than or equal to the dynamic critical threshold, performing S104; S104, comparing the load entropy value of the to-be-processed interface with the average load entropy of the interfaces with the same URL prefix in the existing assets, and if the merging condition is not met, performing S105; S105. Calculate the sample entropy value of the traffic packets of the interface to be processed. The sample entropy value is used to quantify the temporal regularity of the interface call sequence. S106. Filtering decision rules are made based on sample entropy values ​​to achieve API interface discovery and merging; The decision rules are as follows: if the sample entropy value is less than the lower limit of the first interval, it is determined to be a fixed interface and forced to be merged; if the sample entropy value is between the upper and lower limits of the first interval, it is determined to be a dynamic parameter interface and path parameterization is performed; if the sample entropy value is greater than the upper limit of the second interval, it is determined to be abnormal traffic and alarm isolation is performed.

[0007] It should be further noted that the formula for calculating the HTTP interface information load entropy value in step S102 is as follows:

[0008] Where, p k represents the frequency of occurrence of byte values, and k represents all byte values ​​that appear in the HTTP interface payload data.

[0009] It should be further noted that in step S105, the step of calculating the entropy value of the traffic packet sample is as follows: S1051. Define the initial parameters for calculating sample entropy, set the embedding dimension m, and set the similarity tolerance r; S1052. Extract the length values ​​of the request body or response body from the traffic packets of the interface to be processed, and arrange them in chronological order to form a time series T={t1,t2,...,t...} n}, where n is the length of the time series; S1053. Construct m-dimensional message length subsequences. Based on the time series T, generate all consecutive m-dimensional length subsequences. , 1≤i≤n−m, where each subsequence Xm i Corresponding time window The length of the interface call within; S1054. Calculate the number of similarity vectors for each m-dimensional subsequence Xm. i Iterate through all other m-dimensional subsequences Xm j , j≠i, calculate Xm i With Xm j Calculate the Chebyshev distance; count the number of subsequences whose Chebyshev distance is less than or equal to the similarity tolerance r; S1055, For each subsequence Xm i Calculate the proportion of their similarity vectors = × Number of similar vectors, where the denominator is the number of similar vectors. The total number of comparisons excluding itself; S1056, calculating an m-dimensional overall similarity ratio, which is a ratio of similarity vectors of all m-dimensional sub-sequences averaging to obtain , represents an average self-similarity of all m-dimensional sub-sequences, reflecting the overall regularity of the length sequence under m dimensions; S1057, extending to m+1-dimensional sub-sequences and repeating steps S1054 to S1056 to calculate an m+1-dimensional overall similarity ratio B m +1 (r); S1058, calculating a sample entropy value based on the m-dimensional and m+1-dimensional overall similarity ratios by the formula SampEn(m, r) = -ln(B m+1 (r) / B m (r)), wherein the sample entropy value is used to quantify the temporal regularity of the interface call length sequence.

[0010] It should be further explained that step S104 specifically includes: retrieving a historical interface set matching the current to-be-processed interface URL prefix from the API asset library, wherein the URL prefix matching refers to that the first M characters of the current interface URL are the same as the first M characters of the historical interface URL, and M is a preset matching length; calculating the average load entropy of all interfaces in the historical interface set; calculating the absolute difference between the load entropy value of the current to-be-processed interface and the average value to obtain an entropy difference degree; comparing the entropy difference degree with a preset difference degree threshold value, if the difference degree is less than the threshold value, determining that the current interface and the historical interface are the same logical interface, and merging the current interface into the API asset to which the historical interface belongs; if the difference degree is greater than or equal to the threshold value, ending step S104 and entering step S105.

[0011] It should be further explained that step S103 specifically includes: obtaining a dynamic critical threshold value preset by the system, which is generated by statistical load entropy values of historical static interfaces; verifying the validity of the load data of the current to-be-processed interface, excluding abnormal short messages or non-HTTP protocol messages caused by network packet loss or truncation; numerically comparing the valid load entropy value after verification with the dynamic critical threshold value to determine whether the load entropy value is less than the dynamic critical threshold value; If the load entropy value is less than the dynamic critical threshold, a secondary verification is performed in combination with the structural characteristics of the URL, and after confirming that the static interface characteristics are met, new API asset discovery or matching of the same URL and content of the interface in the asset library is performed. If the less than condition is not met or the secondary verification fails, step S103 is ended and step S104 is entered.

[0012] Further, step S106 specifically includes: obtaining sample entropy interval parameters preset by the system, including a first interval lower limit value, a first interval upper limit value, and a second interval upper limit value, wherein the first interval upper limit value and the second interval upper limit value coincide, and are used to distinguish between dynamic parameter interfaces and abnormal traffic; extracting the sample entropy value of the current interface to be processed from the output result of step S105; performing multi-level comparison of the sample entropy value with the preset interval parameters: first, determining whether the sample entropy value is less than the first interval lower limit value, and if yes, marking the interface as a fixed interface; if not, continuing to determine whether the sample entropy value is less than or equal to the first interval upper limit value, and if yes, marking the interface as a dynamic parameter interface; if not, marking the interface as abnormal traffic; performing classification processing operations according to the marking results: for a fixed interface, verifying the consistency of the request body and the response body content with the same URL interface in the existing asset, and after passing, forcibly merging the interface into the corresponding API asset; for a dynamic parameter interface, extracting the position of the length sequence change in the traffic message, parameterizing the corresponding position in the URL path, and updating the path template in the API asset; for abnormal traffic, recording the URL, timestamp, and sample entropy value, triggering an alarm, and storing the related information in an abnormal log.

[0013] Further, in S103, if the load entropy value is less than the dynamic critical threshold, it is determined that the static interface characteristics are met, and the process is directly discovered as a new API asset or collected into an existing API asset, and the process is ended. In S104, if the entropy difference degree is less than the difference degree threshold, it is determined that the same logical interface is met, and the interface is merged and collected into the API asset, and the process is ended.

[0014] The application also provides an API asset processing system based on sample entropy and load entropy, which includes: an interface information acquisition module for acquiring HTTP interface information; a load entropy calculation module for analyzing the load entropy value of the HTTP interface information, the load entropy value being used to quantify the randomness of the interface call content; an interface filtering module for performing filtering decisions according to the load entropy value, and if the load entropy value is greater than or equal to the dynamic critical threshold, an entropy difference comparison and merging module is executed. The entropy difference comparison and merging module is configured to compare the load entropy value of the interface to be processed with the average load entropy value of the interfaces with the same URL prefix in the existing assets, and if the merging condition is not met, the sample entropy calculation module is executed; The sample entropy calculation module is configured to calculate the sample entropy value of the interface to be processed based on the traffic message of the interface, and the sample entropy value is used to quantify the time regularity of the interface call sequence. The sample entropy decision module is configured to make a filtering decision rule according to the sample entropy value to realize the discovery and merging of the API interface.

[0015] According to another embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to realize the steps of the API asset processing method based on sample entropy and load entropy.

[0016] According to another embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to realize the steps of the API asset processing method based on sample entropy and load entropy.

[0017] From the above technical solutions, it can be seen that the present application has the following advantages: The API asset processing method based on sample entropy and load entropy provided by the present application determines static interfaces by calculating load entropy, directly collects them into existing assets, and avoids repeated records caused by URL parameter changes. By extracting historical interfaces with the same URL prefix, the average load entropy is calculated, and if the difference between the entropy value of the interface to be processed and the average value is less than a threshold, it is determined that they are the same logical interface and are merged into the same asset. The sample entropy is calculated to determine abnormal traffic and isolate alarms; and the parameterized path of the normal dynamic interface avoids misjudgment.

[0018] The present application quickly filters static interfaces by load entropy, improves processing efficiency, accurately merges similar dynamic interfaces by comparing the average load entropy of interfaces with the same URL prefix, reduces asset redundancy, analyzes from the time dimension by combining sample entropy, enhances the accuracy of dynamic parameter interface identification, and effectively distinguishes abnormal traffic. While reducing the consumption of computing resources, the present application realizes high-precision discovery, efficient merging and security protection of API assets, and improves the efficiency and reliability of API asset management. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the present application, the drawings needed in the description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0020] Figure 1 Flow chart of API asset processing method based on sample entropy and payload entropy; Figure 2 Flow chart of API asset processing method based on sample entropy and payload entropy; Figure 3 Flow chart of another embodiment of API asset processing method based on sample entropy and payload entropy; Figure 4 Schematic diagram of API asset processing system based on sample entropy and payload entropy; Figure 5 Schematic diagram of electronic device. DETAILED DESCRIPTION

[0021] The API asset processing method based on sample entropy and payload entropy provided in the present application is established on the mathematical framework of Shannon information entropy, and quantifies the uncertainty in HTTP communication, thereby realizing API asset discovery and merging.

[0022] Among them, sample entropy (Sample Entropy, SampEn) is a nonlinear dynamics index for quantifying the complexity and randomness of time series, which evaluates the disorder of the sequence by measuring the probability of new patterns in the signal. The higher the value, the higher the complexity of the sequence and the lower the self-similarity; otherwise, it indicates that the regularity is stronger. In terms of API interface assets, fixed URL interface has strong regularity and presents low entropy value, while dynamic parameter interface has strong randomness and presents high entropy value.

[0023] Payload entropy (Payload Entropy) is a specific application of Shannon entropy theory in the field of network security, which is used to quantify the randomness of the payload part of network data packets. Its essence is an index for measuring data disorder, and the higher the entropy value, the closer the payload content to random distribution, and the lower the entropy value, the more regularity or predictable pattern. In API asset discovery, the payload entropy value of normal API request presents stable distribution, and the dynamic parameter part significantly increases the local entropy value, and the entropy value curve of the same interface has similar waveform characteristics.

[0024] Both sample entropy and payload entropy can improve the asset identification accuracy based on the traditional URL path matching scheme, and the two entropy values are different. Sample entropy can quantify the time regularity of interface call sequence, which is suitable for identifying call frequency / interval pattern, and payload entropy can measure the randomness of request content, which can effectively distinguish parameter dynamics. From the perspective of resources, payload entropy is a single interface traffic content analysis and calculation, and single-step calculation, with small performance consumption; sample entropy is based on time series, which needs to extract multiple interface traffic within a certain time range, construct a multi-dimensional matrix, and the calculation process is complex, with larger performance consumption than payload entropy.

[0025] The application proposes two kinds of entropy value fusion processing ideas. Through load entropy fast filtering decision and load entropy difference comparison, about 80% of interfaces can be identified and merged. Then, a small amount of undecided interfaces are identified in depth through sample entropy, and the API interface identification accuracy can be significantly improved by combining content dimension and time dimension cross verification.

[0026] The API asset processing method based on sample entropy and load entropy related to the application will be described in detail below. In order to illustrate but not to limit, specific details such as specific system structure, technology, etc. are proposed so as to thoroughly understand the embodiments of the application. However, it should be clear to those skilled in the art that the application can also be implemented in other embodiments without these specific details.

[0027] The phrase "one embodiment" or "some embodiments" appearing in the specification means that a specific feature, structure or characteristic described in the embodiment is included in one or more embodiments of the application. Therefore, the phrases "in one embodiment", "in some embodiments", "in other some embodiments", "in further some embodiments" appearing in the specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized.

[0028] In the embodiments of the present application, computer program code for carrying out operations of the present disclosure can be written in one or more programming languages or combinations of them, including but not limited to object oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Program code can be executed entirely on a user computer, partially on a user computer, as an independent software package, partially on a user computer and partially on a remote computer, or entirely on a remote computer. In the case of remote computer, the remote computer can be connected to the user computer through any kind of network, including local area network (LAN) or wide area network (WAN), or can be connected to external computer (for example, connected to the Internet through an Internet service provider).

[0029] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0030] Please refer to Figure 1Fig. 1 shows a flowchart of an API asset processing method based on sample entropy and payload entropy in an embodiment, the method comprising: S101: Collect HTTP interface information from the network or query from a data source, the HTTP interface information including a URL.

[0031] In some embodiments, the basic information of the HTTP interface is obtained by active collection or passive query. The main sources include capturing HTTP request / response packets through switch port mirroring, API gateway logs, Nginx, Apache access logs, etc. The obtained information mainly includes URL, and also includes request method GET / POST, User-Agent in request header, response status code 200 / 404, etc.

[0032] S102: Analyze the payload entropy value of the HTTP interface information, which is used to quantify the randomness of the interface call content.

[0033] In some embodiments, the payload entropy is used to quantify the randomness of the interface call content, and the calculation is based on the byte stream of the HTTP request body or response body. First, the payload content is extracted, and the content is converted into a byte sequence, such as each character corresponding to a byte value of 0-255, and the frequency p of each byte value is counted k .

[0034] The Shannon entropy formula is used for specific application:

[0035] where p k is the frequency of byte value occurrence, and k is all byte values appearing in the HTTP interface payload data.

[0036] Specifically, each byte consists of 8 binary bits, and its value range is 0 to 255. Here, k is sequentially valued from 0 to 255, corresponding to the 256 possible byte values. For example, k=0 corresponds to the case where the byte value is 0, k=1 corresponds to the case where the byte value is 1, and so on, until k=255 corresponds to the case where the byte value is 255. By traversing all values of k from 0 to 255, all possible bytes in the payload can be covered, and the entropy value of the entire payload data can be calculated to quantify the randomness of the interface call content.

[0037] For example, if the payload content is repeated abcabcabc, the frequencies of byte values 'a', 'b', and 'c' are 3 / 9, 3 / 9, and 3 / 9, respectively, and the rest are 0. The calculation gives Hpayload≈1.58; if the payload content is random binary data, the frequency of each byte is close to 1 / 256, and Hpayload≈8.

[0038] It can be seen that for the API interface, the return of the fixed picture or the load content of the configuration file has high repetition and concentrated byte distribution, and thus has low entropy. The JSON return of the user real-time data has large change in load content, and the byte contains random fields such as user ID and timestamp, and thus has high entropy. By calculating the load entropy, the interface can be divided into two categories of low entropy with highly fixed content and high entropy with dynamically changed content. The time regularity can be further analyzed to avoid misjudgment caused by content randomness.

[0039] S103: making a rapid filtering decision according to the calculated load entropy, if the load entropy is less than a dynamic critical threshold, it is determined to meet the static interface characteristics, and a new API asset is directly discovered or collected into an existing API asset according to the determination, and the process ends; if the load entropy is greater than or equal to the dynamic critical threshold, S104 is performed.

[0040] In some embodiments, the preset dynamic critical threshold can be optionally set to 3.0, 4.0, etc., and the load entropy calculated in S102 is compared with the threshold. If Hpayload<3.0, it is determined to be a static interface, and asset discovery or collection is directly performed; if Hpayload≥3.0, subsequent steps are entered.

[0041] The threshold setting is based on historical data statistics. For example, if the load entropy of an interface is 2.5, which is less than 3.0, it is considered that the content is highly fixed, the URL is matched with an existing interface in the asset library, and if the matching is successful, the interface is collected, and if the matching fails, the interface is added as an asset. The statistical setting of the threshold ensures the universality of the classification and avoids the deviation caused by subjective setting.

[0042] For the present application, step S103 also involves a specific implementation mode as shown in Figure 2 S103 specifically further includes the following steps: S1031: obtaining a dynamic critical threshold preset by the system, and the dynamic critical threshold is generated by statistics of load entropies of historical static interfaces.

[0043] Optionally, the calculation of the dynamic critical threshold uses the 95th percentile of the load entropies of the historical static interfaces, and the calculation method is to arrange the load entropies of the historical static interfaces in ascending order, and take the value corresponding to the 95th percentile. The calculation is based on statistical distribution, and ensures that the threshold can cover 95% of the static interfaces, and only 5% of the static interfaces can exceed the value, balancing the strictness and coverage of the threshold.

[0044] S1032: verifying the validity of the load data of the current interface to be processed, and excluding abnormal short messages or non-HTTP protocol messages caused by network packet loss or truncation.

[0045] It should be noted that setting the message length threshold and the protocol identifier is to filter the payload data through logical judgment (length ≥ 100 bytes and protocol header matching). The logical judgment can be set as length ≥ 100 bytes and protocol header matching. Based on the basic characteristics of the HTTP protocol, invalid messages caused by network failure or attack are excluded.

[0046] S1033, the checked payload entropy value is compared with the dynamic critical threshold value, and it is judged whether the payload entropy value is strictly less than the dynamic critical threshold value; S1034, if the condition that the payload entropy value is less than the dynamic critical threshold value is met, secondary verification is performed in combination with the structural characteristics of the URL, and after it is confirmed that the static interface characteristics are met, new API asset discovery or matching of the interfaces with the same URL and content in the asset library is performed; if the condition of being less than is not met or the secondary verification fails, the step S103 is ended, and the step S104 is entered.

[0047] As can be seen, the step S103 realizes rapid and accurate identification of the static interface through threshold setting driven by statistics, data validity verification and multi-dimensional feature verification. Abnormal messages are excluded through data validity verification, so as to avoid invalid data interference with the judgment of the static interface; after the entropy value comparison passes, the structural characteristics of the URL are further combined, so as to ensure rapid filtering of the dynamic interface and reduce the probability of misjudgment of the static interface.

[0048] S104: the load entropy value of the interface to be processed is compared with the average load entropy value of the interfaces with the same URL prefix in the existing asset, if the entropy difference is less than the difference threshold value, it is determined that they are the same logical interface, and the interfaces are merged and collected into the API asset, and the process is ended; if the merging condition is not met, S105 is executed.

[0049] In some embodiments, for the interface to be processed determined as a dynamic interface in S103, all historical interfaces with the same URL prefix are first retrieved from the asset library.

[0050] In this embodiment, the average load entropy value of the historical interfaces is first calculated, and then the absolute difference between the entropy value of the interface to be processed and the average value is calculated, such as the entropy value 3.4 of the interface to be processed, and the difference is |3.4-3.35|=0.05.

[0051] If the difference is less than the difference threshold value, it is determined that they are the same logical interface, and are merged into the asset; otherwise, S105 is executed.

[0052] It can be seen that the URL prefix matching is based on the assumption that "different versions or variants of the same logical interface usually share the same prefix". The average of the load entropy reflects the average level of randomness of the content of the historical interface, and the difference degree measures the deviation of the current interface from the historical group. If the difference degree is small, it means that the randomness of the content of the current interface is consistent with that of the historical interface, which belongs to the variant of the same logical interface, and therefore can be combined, thereby improving the accuracy of the combination and avoiding the repeated recording of assets caused by the slight difference of the URL.

[0053] In S105, a sample entropy value of the to-be-processed interface is calculated based on the flow message of the to-be-processed interface, and the sample entropy value is used to quantify the time regularity of the interface call sequence.

[0054] In some embodiments, the sample entropy is used to quantify the time regularity of the interface call sequence, and the calculation is based on the time sequence of the interface call.

[0055] Specifically, the sample entropy quantifies the complexity of the sequence by comparing the self-similarity of different dimension sub-sequences. The m-dimensional sub-sequence reflects the pattern in a short time window, and the m+1-dimensional sub-sequence reflects the pattern in a longer window. If the similarity ratio does not decrease significantly after increasing the dimension, it means that the sequence still maintains regularity in a longer window, that is, the entropy value is small; if the similarity ratio decreases significantly, it means that the sequence pattern quickly collapses with the increase of the dimension, that is, the entropy value is large. In this way, the calculation of the sample entropy only depends on the request / response length sequence, without the need to parse the specific content, thereby reducing the calculation complexity.

[0056] In S106, a filtering decision rule is made according to the sample entropy value, so as to realize the discovery and combination of the API interface. The decision rule is that: if the sample entropy value is less than the lower limit value of the first interval, it is determined as a fixed interface and is forced to be combined; if the sample entropy value is between the upper and lower limit values of the first interval, it is determined as a dynamic parameter interface and is subjected to path parameterization processing; and if the sample entropy value is greater than the upper limit value of the second interval, it is determined as abnormal flow and is subjected to alarm isolation.

[0057] In some embodiments, a preset sample entropy interval is provided, and optionally, the lower limit of the first interval is 1.2 and the upper limit is 2.5; the upper limit of the second interval is 2.5, and the sample entropy value calculated in S105 is classified and processed. If SampEn<1.2: it is determined as a fixed interface and is combined into the same URL interface in the asset library. If 1.2≤SampEn≤2.5: it is determined as a dynamic parameter interface, the position of the length change is extracted, and the corresponding position in the URL path is parameterized. If SampEn>2.5: it is determined as abnormal flow, the URL, timestamp and length sequence are recorded, an alarm is triggered and the storage is isolated.

[0058] The embodiment is based on sample entropy for filtering decision, that is, discovery and merging of API interfaces. The merging decision rule is shown in Table 1.

[0059] Table 1: Merging decision rule table based on sample entropy

[0060] It can be seen that the interval division of sample entropy is based on the distribution statistics of historical normal interfaces. The entropy value of the fixed interface is extremely low, indicating that the call sequence is highly regular, and it can be directly merged to avoid duplication; the entropy value of the dynamic parameter interface is moderate, indicating that the call sequence has predictable changes, and the path template can be unified by parameterization extraction; the entropy value of the abnormal traffic is extremely high, indicating that the call sequence is irregular and beyond the normal business range, and needs to be isolated for detection to prevent attacks. In this way, the introduction of sample entropy avoids repeated recording of interfaces due to parameter changes, and at the same time improves the security of the system by isolating abnormal traffic.

[0061] It needs to be further explained that step S106 also involves the following specific steps, specifically including: S1061, acquiring sample entropy interval parameters preset by the system, including a first interval lower limit value, a first interval upper limit value and a second interval upper limit value, wherein the first interval upper limit value coincides with the second interval upper limit value, and is used to distinguish dynamic parameter interfaces and abnormal traffic; S1062, extracting the sample entropy value of the current interface to be processed from the output result of step S105; S1063, comparing the sample entropy value with the preset interval parameters in multiple levels: first, judging whether it is less than the first interval lower limit value, if yes, marking it as a fixed interface; if not, continue to judge whether it is less than or equal to the first interval upper limit value, if yes, mark it as a dynamic parameter interface; if both are not satisfied, mark it as abnormal traffic.

[0062] Optionally, the sample entropy value comparison method can adopt a multi-level threshold judgment method, which first compares with the lower limit value to determine the fixed interface, then compares with the upper limit value to determine the dynamic parameter interface, and the remaining cases are determined as abnormal traffic, avoiding misjudgment of a single threshold.

[0063] S1064, performing classification processing operation according to the marking result: for the fixed interface, verifying the consistency of the request body and response body content with the same URL interface in the existing asset, and forcibly merging it into the corresponding API asset after passing; for the dynamic parameter interface, extracting the position of the length sequence change in the traffic message, parameterizing the corresponding position in the URL path, and updating the path template in the API asset; for abnormal traffic, recording its URL, timestamp and sample entropy value, triggering alarm and isolating and storing related information to abnormal log.

[0064] Optionally, the fixed interface processing method can be: when verifying the consistency of the request body and response body content, extract key fields for hash matching. If the hash values ​​are the same, it is determined to be the same interface, triggering a forced merge.

[0065] The dynamic parameter interface can be processed as follows: when extracting the position of length sequence change, compare the difference between adjacent length sequences through a sliding window, count the positions of high frequency change, and replace the fixed value of the corresponding position in the URL path with parameter placeholders.

[0066] For abnormal traffic handling, the URL, timestamp, sample entropy value, and length sequence of the last 10 calls of the interface are recorded during isolated storage. This information is used for subsequent manual analysis or machine learning model training to optimize threshold parameters.

[0067] It is evident that by using interval classification of sample entropy and multi-dimensional processing logic, the API interface achieves refined discovery, merging, and anomaly detection, improving asset identification accuracy while also considering processing efficiency and business interpretability.

[0068] In one embodiment of the present invention, based on step S105, the following is a possible embodiment and its specific implementation is described in a non-limiting manner. In step S105, the step of calculating the entropy value of the traffic packet sample is as follows: S1051. Define the initial parameters for calculating sample entropy, set the embedding dimension m, and set the similarity tolerance r.

[0069] Optionally, the embedding dimension m is set to 2 to capture the second-order time dependency of the API call length sequence. The similarity tolerance r is set to 0.1 to 0.25 times the standard deviation of the API time series dataset to be processed, in order to balance the strictness and robustness of the similarity judgment.

[0070] S1052. Extract the length values ​​of the request body or response body from the traffic packets of the interface to be processed, and arrange them in chronological order to form a time series T={t1,t2,...,t...} n}, where n is the length of the time series.

[0071] S1053. Construct m-dimensional message length subsequences. Based on the time series T, generate all consecutive m-dimensional length subsequences. , where each subsequence Xm i Corresponding time window The length of the interface call within.

[0072] Optionally, the m-dimensional subsequence X m ᵢ represents a vector consisting of m consecutive length values ​​extracted from the time series T, used to analyze length patterns within a short time window.

[0073] S1054, calculate the number of similar vectors of the m-dimensional subsequence, for each m-dimensional subsequence Xm i , traverse all other m-dimensional subsequences Xm j , j≠i, calculate the Chebyshev distance of Xm i and Xm j ; count the number of subsequences satisfying the Chebyshev distance less than or equal to the similarity tolerance r, denoted as Ci.

[0074] It should be noted that the fluctuation of the length of the API interface call may contain local random noise, such as slight changes caused by network delay, but does not affect the overall pattern. The Chebyshev distance focuses on the most significant difference dimension by taking the maximum value, and is more suitable for capturing the core pattern of the interface call, which is a targeted improvement on the similarity measure.

[0075] S1055, for each subsequence Xm i , calculate its similar vector proportion = × similar vector number, where the denominator is the total number of comparisons after excluding itself.

[0076] S1056, calculate the m-dimensional overall similarity proportion, average the similar vector proportions of all m-dimensional subsequences to obtain , , which represents the average self-similarity of all m-dimensional subsequences, reflecting the overall regularity of the length sequence in m dimensions.

[0077] S1057, extend to m+1-dimensional subsequences and repeat the calculation.

[0078] This embodiment increases the embedding dimension to m+1, constructs m+1-dimensional message length subsequences , repeats steps S1054 to S1056, and calculates the m+1-dimensional overall similarity proportion B m+1 (r).

[0079] S1058, based on the m-dimensional and m+1-dimensional overall similarity proportions, calculate the sample entropy value by the formula SampEn(m, r) = -ln(B m+1 (r) / B m (r)), where the sample entropy value is used to quantify the time regularity of the interface call length sequence. The smaller the entropy value, the stronger the sequence regularity; the larger the entropy value, the stronger the sequence randomness.

[0080] It should be noted that the embedding dimension m represents the number of consecutive time points used when constructing sub-sequences, which is used to capture the second-order time dependence of the interface call length sequence. The similarity tolerance r represents the threshold for judging the similarity of two sub-sequences, which is used to balance the identification of noise interference and real regularity. The time sequence T represents the sequence of request / response length of the interface to be processed in a time sequence, reflecting the dynamic behavior trajectory of the interface call.

[0081] Step S105 deeply integrates sample entropy calculation with the time regularity of API interface call sequence through parameter definition, multi-dimensional sub-sequence construction, similarity measurement and progressive dimension analysis, solving the problem that path matching cannot quantify time patterns.

[0082] In an embodiment of the present application, based on step S104, a possible embodiment will be given below to illustrate the specific implementation thereof. As shown in Figure 3 Step S104 specifically includes: S1041, retrieving a set of historical interfaces matching the URL prefix of the current interface to be processed from the API asset library, the URL prefix matching indicating that the first M characters of the current interface URL are the same as the first M characters of the historical interface URL, M being a preset matching length; S1042, calculating the mean of the load entropy of all interfaces in the set of historical interfaces, the mean being the arithmetic mean of the load entropy values of the historical interfaces, i.e. the sum of all load entropy values of the historical interfaces divided by the number of historical interfaces; S1043, calculating the absolute difference between the load entropy value of the current interface to be processed and the mean, obtaining the entropy difference degree, the absolute difference being the absolute value of the current load entropy value minus the mean; S1044, comparing the entropy difference degree with a preset difference degree threshold, if the difference degree is less than the threshold, determining that the current interface and the historical interface are the same logical interface, and merging the current interface into the API asset to which the historical interface belongs; if the difference degree is greater than or equal to the threshold, ending step S104 and entering step S105.

[0083] It should be noted that the entropy difference degree is calculated by the absolute difference method, specifically Sc=|H c -μ|, where μ is the mean of the load entropy. H c is the load entropy value of the current interface to be processed, and the absolute difference avoids the inapplicability of the relative difference when the mean is zero, directly reflecting the deviation degree of the current interface from the group mean.

[0084] Optionally, the difference degree threshold can be set to 0.2, and the difference degree threshold is used to distinguish the normal fluctuation of the same logical interface and the significant difference of different logical interfaces, and generally the difference degree is greater than or equal to 0.2.

[0085] The step S104 of the embodiment filters the historical interface through the semantically associated URL prefix, and realizes accurate merging of the dynamic interface variants by combining the statistical mean and quantitative difference analysis of the load entropy. The deviation degree of the current interface from the historical group is quantified by calculating the absolute difference between the current interface and the mean, and the difference threshold is used to determine whether the current interface belongs to the normal dynamic change of the same logical interface, thereby reducing invalid comparison and improving the merging accuracy.

[0086] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the application.

[0087] The following is an embodiment of an API asset processing system based on sample entropy and load entropy provided by the embodiment of the disclosure. The system and the API asset processing method based on sample entropy and load entropy described above belong to the same inventive concept. Details not described in the embodiment of the API asset processing system based on sample entropy and load entropy can be referred to the embodiment of the API asset processing method based on sample entropy and load entropy described above.

[0088] As shown in Figure 4 , the system comprises: An interface information acquisition module 201 configured to acquire HTTP interface information. A load entropy calculation module 202 configured to analyze the load entropy value of the HTTP interface information, the load entropy value being used to quantify the randomness of the interface call content. An interface filtering module 203 configured to make a filtering decision according to the load entropy value, and if the load entropy value is greater than or equal to a dynamic critical threshold, execute an entropy difference comparison and merging module. The entropy difference comparison and merging module 204 is configured to compare the load entropy value of the interface to be processed with the mean load entropy value of the interface with the same URL prefix in the existing asset, and if the merging condition is not met, execute a sample entropy calculation module. The sample entropy calculation module 205 is configured to calculate the sample entropy value of the to-be-processed interface based on the traffic message of the to-be-processed interface, and the sample entropy value is used to quantify the time regularity of the interface call sequence. The sample entropy decision module 206 is configured to make a filtering decision according to the sample entropy value to realize the discovery and merging of the API interface.

[0089] As shown in Figure 5As shown, the present application also provides an electronic device, comprising a display module 103, a memory 102, a processor 101 and a computer program stored in the memory and capable of running on the processor 101, wherein the processor 101 implements the steps of the API asset processing method based on sample entropy and load entropy when executing the program.

[0090] In embodiments of the present application, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workstation, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections, and relationships, and their functions, are shown as examples only and are not meant to limit implementations of the embodiments described and / or claimed herein.

[0091] In embodiments of the present application, the processor 101 can be implemented by using at least one of an application-specific integrated circuit, a programmable logic device, a field programmable gate array, a processor, a controller, a microcontroller, a microprocessor, an electronic unit designed to perform the functions described herein, and in some cases, such implementation can be implemented in a controller. For software implementation, the implementation of such as processes or functions can be implemented with separate software modules allowing at least one function or operation to be performed, and the software code can be implemented by a software application (or program) written in any appropriate programming language and stored in a memory and executed by a controller.

[0092] The display module 103 is used to display information input by a user or information provided to a user. The display module 103 can include a display panel, which can be configured in the form of a liquid crystal display, an organic light emitting diode, etc.

[0093] The memory 102 can be used to store software programs and various data. The memory 102 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device.

[0094] The present application also provides a storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the API asset processing method based on sample entropy and load entropy.

[0095] The storage medium can be any available medium that can be accessed by a general purpose or special purpose computer. By way of example, and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code means in the form of instructions or data structures and that can be accessed by a general-purpose or special-purpose computer, or a general-purpose or special-purpose processor. Also, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or other

[0096] In this document, the terms "computer-readable medium" or "computer- readable media" is used to generally refer to media such as removable storage, volatile memory, non-volatile memory, or any other storage medium readable by a computer or a general purpose or special purpose computer. Exemplary computer-readable storage media includes storage mediums such as a magnetic, optical, or semiconductor storage medium, or any other medium which is used to provide storage "instructions" to a computer or a general purpose or special purpose computer system. Exemplary embodiments of the present application can be implemented as a method, apparatus, or article of manufacture using standard programming and / or engineering techniques. The term "article of manufacture" as used herein is intended to encompass a computer program accessible from a computer-readable medium or storage media providing program instructions to a computer or a general purpose or special purpose computer system. In addition, the article of manufacture can take many forms of media, for example, but not limited to, a

[0097] The foregoing description of the disclosed embodiments enables a person skilled in the art to implement or use the application. Numerous modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An API asset processing method based on sample entropy and load entropy, characterized in that, The method comprises: S101, acquiring HTTP interface information; S102, parsing a load entropy value of the HTTP interface information, the load entropy value being used to quantify randomness of interface call content; S103, making a filtering decision according to the load entropy value, if the load entropy value is greater than or equal to a dynamic critical threshold, executing S104; S104, comparing the load entropy value of the to-be-processed interface with a load entropy average value of an interface with the same URL prefix in an existing asset, if the comparison does not meet a merging condition, executing S105; S105, calculating a sample entropy value of the to-be-processed interface based on a traffic message of the to-be-processed interface, the sample entropy value being used to quantify time regularity of an interface call sequence; S106, making a filtering decision rule according to the sample entropy value, to realize discovery and merging of API interfaces; The decision rule is: judging that the sample entropy value less than a lower limit value of a first interval is a fixed interface and performing forced merging, judging that the sample entropy value between upper and lower limit values of the first interval is a dynamic parameter interface and performing path parameterization processing, and judging that the sample entropy value greater than an upper limit value of a second interval is abnormal traffic and performing alarm isolation.

2. The API asset processing method based on sample entropy and load entropy according to claim 1, characterized in that, In step S102, a calculation formula of the load entropy value of the HTTP interface information is: where p k is the frequency of byte values, k is all byte values that appear in the HTTP interface load data.

3. The API asset processing method based on sample entropy and load entropy according to claim 1, characterized in that, In step S105, a step of calculating the sample entropy value of the traffic message is: S1051, defining initial parameters for sample entropy calculation, setting an embedding dimension m, and setting a similar tolerance r; S1052, extract the length value of the request body or response body from the traffic message of the interface to be processed, arrange in time sequence to form a time sequence T={t1, t2,...,t n n}, wherein n is the length of the time sequence; S1053, constructing m-dimensional message length subsequences, generating all continuous m-dimensional length subsequences based on time series T where each subsequence Xm i corresponding to the interface call length within the time window within the time window S1054、Calculate the number of similar vectors of m-dimensional subsequence, for each m-dimensional subsequence Xm i , traverse all other m-dimensional subsequence Xm j , j≠i, calculate the Chebyshev distance between Xm i and Xm j ; count the number of subsequence that satisfies the Chebyshev distance less than or equal to the similarity tolerance r; S1055、for each sub-sequence Xm i , calculate its similar vector proportion = × similar vector number, where denominator is the total comparison number after excluding itself; S1056, calculate m-dimensional overall similarity proportion, the similarity vector proportion of all m-dimensional subsequences average, get , represent the average self-similarity of all m-dimensional subsequences, reflecting the overall regularity of the length sequence under m dimensions; S1057, extend to m+1 dimensional sub-sequence and repeat steps S1054 to S1056, calculate m+1 dimensional overall similarity ratio B m+1 (r); S1058, based on the m-dimensional and m+1-dimensional overall similarity ratios, a sample entropy value is calculated by the formula SampEn(m, r) = -ln(B m+1 (r) / B m (r)), wherein the sample entropy value is used to quantify the temporal regularity of the interface call length sequence.

4. The API asset processing method based on sample entropy and load entropy according to claim 1, characterized in that, Step S104 specifically comprises: Retrieving a historical interface set matching a current to-be-processed interface URL prefix from an API asset library, the URL prefix matching indicating that the first M characters of the current interface URL are the same as the first M characters of a historical interface URL, M being a preset matching length; Calculating a load entropy average value of all interfaces in the historical interface set; Calculating an absolute difference value of the load entropy value of the current to-be-processed interface and the average value, to obtain an entropy difference degree; Comparing the entropy difference degree with a preset difference degree threshold value, if the difference degree is less than the threshold value, judging that the current interface and the historical interface are the same logical interface, and merging the current interface into an API asset to which the historical interface belongs; if the difference degree is greater than or equal to the threshold value, ending step S104 and entering step S105.

5. The API asset processing method based on sample entropy and load entropy according to claim 1, characterized in that, Step S103 specifically comprises: Acquiring a system preset dynamic critical threshold value, the dynamic critical threshold value being generated by load entropy value statistics of historical static interfaces; Verifying validity of load data of the current to-be-processed interface, to exclude abnormal short messages or non-HTTP protocol messages caused by network packet loss or truncation; Comparing the verified valid load entropy value with the dynamic critical threshold value, to judge whether the load entropy value is less than the dynamic critical threshold value; If the condition that the load entropy value is less than the dynamic critical threshold value is met, performing secondary verification combined with a structure feature of the URL, confirming that the static interface feature is met, and executing new API asset discovery or matching an interface with the same URL and content in an asset library; if the condition is not met or the secondary verification is not passed, ending step S103 and entering step S104.

6. The API asset processing method based on sample entropy and load entropy according to claim 1, characterized in that, Step S106 specifically comprises: Obtain the sample entropy interval parameters preset by the system, including a first interval lower limit value, a first interval upper limit value, and a second interval upper limit value, wherein the first interval upper limit value coincides with the second interval upper limit value, and is used to distinguish between dynamic parameter interfaces and abnormal traffic; Extract the sample entropy value of the current interface to be processed from the output result of step S105; Compare the sample entropy value with the preset interval parameters in multiple stages: first, determine whether it is less than the first interval lower limit value, if yes, mark it as a fixed interface; if not, continue to determine whether it is less than or equal to the first interval upper limit value, if yes, mark it as a dynamic parameter interface; if neither condition is met, mark it as abnormal traffic; According to the marking result, perform classification processing operations: for a fixed interface, verify the consistency of the request body and response body content with the same URL interface in the existing asset, and after passing, forcibly merge it into the corresponding API asset; for a dynamic parameter interface, extract the position of the length sequence change in the traffic message, parameterize the corresponding position in the URL path, and update the path template in the API asset; for abnormal traffic, record its URL, timestamp, and sample entropy value, trigger an alarm, and store the related information in the abnormal log.

7. The API asset processing method based on sample entropy and load entropy according to claim 1, characterized in that, In S103, if the load entropy value is less than the dynamic threshold, it is determined to meet the characteristics of a static interface, and a new API asset is directly discovered or collected into an existing API asset based on this, and the process ends; In S104, if the entropy difference is less than the difference threshold, it is determined to be the same logical interface, and the interfaces are merged and collected into the API asset, and the process ends.

8. An API asset processing system based on sample entropy and load entropy, characterized in that, The system is used to implement the API asset processing method based on sample entropy and load entropy according to any one of claims 1 to 7. The system comprises: An interface information acquisition module for obtaining HTTP interface information; A load entropy calculation module for analyzing the load entropy value of the HTTP interface information, which is used to quantify the randomness of the interface call content; An interface filtering module for filtering decision based on the load entropy value, and if the load entropy value is greater than or equal to the dynamic threshold, an entropy difference comparison and merging module is executed; An entropy difference comparison and merging module for comparing the load entropy value of the interface to be processed with the average load entropy of the interfaces with the same URL prefix in the existing asset, and if the merging condition is not met, a sample entropy calculation module is executed; A sample entropy calculation module for calculating the sample entropy value of the traffic message of the interface to be processed, which is used to quantify the time regularity of the interface call sequence; A sample entropy decision module for filtering decision based on the sample entropy value to realize the discovery and merging of API interfaces.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the steps of the API asset processing method based on sample entropy and load entropy according to any one of claims 1 to 7.

10. A storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the steps of the API asset processing method based on sample entropy and load entropy according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • A network traffic anomaly detection and prevention method

    CN109274673A

  • Cloud information system index acquisition method, system, equipment and medium

    CN120475028A

  • Application behavior detection using network traffic

    US11985152B1

  • System and method for determining data entropy to identify malware

    US20080184367A1