A method for processing classification and grading of response data and an electronic device

CN122845281APending Publication Date: 2026-09-29BEIJING ANHUA JINHE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611219241.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-12
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0005]本申请实施例提供了一种响应数据的分类分级处理方法及电子设备,以至少解决现有技术中依赖于数据接口文档来获取该数据接口的输出字段所导致的分类分级偏差大的问题

Benefits of technology

在本申请实施例中,采用了从HTTP流量日志中获取HTTP响应;获取所述HTTP响应的应用标识、接口标识和响应体;查找所述HTTP响应对应模板哈希,其中,所述模板哈希是对响应体的结构特征进行抽取,将抽取到的结构特征经过排序和/或去重之后,再进行哈希运算得到的;所述模板哈希用于对响应体的不同的结构特征进行标识;将所述应用标识、所述接口标识和所述模板哈希组合成桶键;以所述桶键在采样桶的已分类分级的模板集合中查询快照,并将查找到的快照作为对所述HTTP响应进行分类分级处理的依据。通过本申请解决了现有技术中依赖于数据接口文档来获取该数据接口的输出字段所导致的偏差大的问题,从而提高了数据接口输出字段获取的效率和准确度,进一步提高了数据分类分级的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845281A_ABST
    Figure CN122845281A_ABST
Patent Text Reader

Abstract

The application discloses a classification and grading processing method of response data and an electronic device. The method comprises the following steps: obtaining an HTTP response from an HTTP traffic log; obtaining an application identifier, an interface identifier and a response body of the HTTP response; searching for a template hash corresponding to the HTTP response, wherein the template hash is obtained by performing hashing operation on the structural features of the response body after extracting the structural features and performing sorting and / or deduplication; combining the application identifier, the interface identifier and the template hash into a bucket key; querying a snapshot in a classified and graded template set of a sampling bucket by using the bucket key, and taking the found snapshot as a basis for classifying and grading processing of the HTTP response. The application solves the problem of large deviation caused by the dependence on the data interface document to obtain the output field of the data interface in the prior art, thereby improving the efficiency and accuracy of the data interface output field acquisition, and further improving the accuracy of data classification and grading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network data, and more specifically, to a method and electronic device for classifying and grading response data. Background Technology

[0002] In a Data Security Governance (DSG) scenario, enterprises need to classify and categorize the data elements output by their internal data interfaces (APIs) (such as personal information, sensitive personal information, important data, and general data), and implement differentiated access control, anonymization, and auditing accordingly. The prerequisite for API classification and categorization is a true understanding of which fields each API actually outputs and the sample values ​​for those fields.

[0003] In existing technologies, manual interface documentation (such as OpenAPI, Swagger, WSDL) is typically relied upon to obtain data interface output field and value samples from these documentation. However, these documentation needs to be updated in real time and cannot be lost. In actual production environments, documentation is often missing, outdated, or inconsistent with actual traffic, resulting in large deviations in classification and grading.

[0004] There is currently no suitable solution to the aforementioned problems in the existing technology. Summary of the Invention

[0005] This application provides a method and electronic device for classifying and grading response data, which at least solves the problem of large classification and grading deviations caused by relying on data interface documents to obtain the output fields of the data interface in the prior art.

[0006] According to one aspect of this application, a method for classifying and grading response data is provided, comprising: obtaining an HTTP response from an HTTP traffic log; obtaining an application identifier, an interface identifier, and a response body of the HTTP response; searching for a template hash corresponding to the HTTP response, wherein the template hash is obtained by extracting structural features of the response body, sorting and / or deduplicating the extracted structural features, and then performing a hash operation; the template hash is used to identify different structural features of the response body; combining the application identifier, the interface identifier, and the template hash into a bucket key; querying a snapshot in the classified and graded template set of the sampling bucket using the bucket key, and using the found snapshot as the basis for classifying and grading the HTTP response.

[0007] Further, finding the template hash corresponding to the HTTP response includes: scanning the response body to obtain a fingerprint hash based on the type of the HTTP response, wherein the fingerprint hash is a hash value used to identify the response body; using the application identifier, the interface identifier, the response body hash value carried in the HTTP response, and the fingerprint hash combined into a four-tuple hash as a cache key; and using the cache key to find the template hash corresponding to the HTTP response in the cache of template hashes.

[0008] Furthermore, if the template hash corresponding to the HTTP response is not found, the method further includes: flattening the response body to obtain a flattened key-value pair set according to the content type of the HTTP response, wherein the flattening extraction is to extract the structural features of the response body; sorting and / or deduplicating the extracted keys, and then performing a hash operation to obtain the template hash corresponding to the HTTP response; and writing the mapping between the quadruple hash and the obtained template hash into a template hash cache.

[0009] Furthermore, if no corresponding snapshot is found in the classified and graded template set using the bucket key, the method further includes: determining whether the response body is a valid sample; if so, using the bucket key as the sampling bucket key, placing the hash value of the response body into the sampling bucket; performing quality filtering on the valid sample, and generating a snapshot based on the quality-filtered valid sample and saving it in the classified and graded template set.

[0010] Further, determining whether the response body is a valid sample includes: when the response body hash value does not exist in the sampling bucket, or when the number of response body hash values ​​collected in the sampling bucket has not reached the sample counting threshold and the current response body hash value is not in the sampling bucket, it is added and determined to be a valid sample; otherwise, it is determined to be an invalid sample and the current processing ends.

[0011] Furthermore, during the system startup phase, the sampling bucket queries the field sample records collected within the current time period from the persistent data storage, groups and aggregates them according to the application identifier, the interface identifier, and the template hash, and fills the corresponding response body hash value set of each group back into the corresponding sampling bucket in memory; the sampling bucket is then cleared according to the time period by a scheduled task.

[0012] Furthermore, the quality filtering of the effective samples includes at least one of the following: response status code filtering, minimum field count filtering, and dual-condition cross-filtering of error keywords and sensitive data.

[0013] Furthermore, the dual-condition cross-filtering of error keywords and sensitive data includes: maintaining a configurable set of error keywords; determining whether there is an intersection between the list of values ​​in the flattened key-value pair set and the set of error keywords to obtain an error hit identifier; determining whether the response location contains sensitive data based on the sensitive data identification results carried in the HTTP traffic log to obtain a sensitive hit identifier; when the error hit identifier is true and the sensitive hit identifier is false, it is determined to be an error response and discarded; when the error hit identifier is true and the sensitive hit identifier is true, it is determined to be a normal response containing sensitive fields and retained; when the error hit identifier is false, it is retained normally.

[0014] According to another aspect of this application, an electronic device is also provided, including a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the above-described method steps.

[0015] According to another aspect of this application, a readable storage medium is also provided, on which computer instructions are stored, wherein the computer instructions, when executed by a processor, implement the above-described method steps. In this embodiment, the method involves obtaining HTTP responses from HTTP traffic logs; acquiring the application identifier, interface identifier, and response body of the HTTP response; searching for the template hash corresponding to the HTTP response, wherein the template hash is obtained by extracting structural features of the response body, sorting and / or deduplicating the extracted structural features, and then performing a hash operation; the template hash is used to identify different structural features of the response body; combining the application identifier, the interface identifier, and the template hash into a bucket key; using the bucket key to query a snapshot in the categorized and graded template set of the sampling bucket, and using the found snapshot as the basis for classifying and grading the HTTP response. This application solves the problem of large deviations caused by relying on data interface documents to obtain the output fields of the data interface in the prior art, thereby improving the efficiency and accuracy of obtaining data interface output fields, and further improving the accuracy of data classification and grading. Attached Figure Description

[0016] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a structural block diagram of a processing system according to an embodiment of this application; Figure 2 This is a flowchart of the template-based extraction and classification / grading preprocessing method according to an embodiment of this application; Figure 3This is a timing diagram for looking up a two-layer fingerprint hash cache according to an embodiment of this application; Figure 4 This is a schematic diagram of feedback closed-loop atom snapshot replacement according to an embodiment of this application; Figure 5 is a schematic diagram of sampling bucket status management according to an embodiment of this application; Figure 6 is a logic diagram of multiple quality filtering determination according to an embodiment of this application. Detailed Implementation

[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0018] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0019] The following embodiments provide a method for classifying and grading response data, including: obtaining an HTTP response from an HTTP traffic log; obtaining the application identifier, interface identifier, and response body of the HTTP response; finding the template hash corresponding to the HTTP response, wherein the template hash is obtained by extracting structural features of the response body, sorting and / or deduplicating the extracted structural features, and then performing a hash operation; the template hash is used to identify different structural features of the response body; combining the application identifier, the interface identifier, and the template hash into a bucket key; querying a snapshot in the classified and graded template set of the sampling bucket using the bucket key, and using the found snapshot as the basis for classifying and grading the HTTP response.

[0020] The above steps solve the problem of large deviations caused by relying on data interface documents to obtain the output fields of the data interface in the existing technology, thereby improving the efficiency and accuracy of obtaining the output fields of the data interface, and further improving the accuracy of data classification and grading.

[0021] There are several methods for finding template hashes. For example, finding the template hash corresponding to the HTTP response includes: scanning the response body to obtain a fingerprint hash based on the type of the HTTP response, wherein the fingerprint hash is a hash value used to identify the response body; using the application identifier, the interface identifier, the response body hash value carried in the HTTP response, and the fingerprint hash as a four-tuple hash as a cache key; and using the cache key to find the template hash corresponding to the HTTP response in the template hash cache.

[0022] If the template hash corresponding to the HTTP response is not found, the method further includes: flattening the response body to obtain a flattened key-value pair set according to the content type of the HTTP response, wherein the flattening extraction is to extract the structural features of the response body; sorting and / or deduplicating the extracted keys, and then performing a hash operation to obtain the template hash corresponding to the HTTP response; and writing the mapping between the quadruple hash and the obtained template hash into a template hash cache.

[0023] If no corresponding snapshot is found in the classified and graded template set using the bucket key, the method further includes: determining whether the response body is a valid sample; if so, using the bucket key as the sampling bucket key, putting the hash value of the response body into the sampling bucket; performing quality filtering on the valid sample, and generating a snapshot based on the quality-filtered valid sample and saving it in the classified and graded template set.

[0024] There are several ways to determine whether the response body is a valid sample. For example, if the response body hash value does not exist in the sampling bucket, or if the number of response body hash values ​​collected in the sampling bucket has not reached the sample counting threshold and the current response body hash value is not in the sampling bucket, it is added and determined to be a valid sample; otherwise, it is determined to be an invalid sample and the current processing ends.

[0025] During system startup, the sampling bucket queries the field sample records collected within the current time period from the persistent data storage, groups and aggregates them according to the application identifier, the interface identifier, and the template hash, and fills the corresponding response body hash value set of each group back into the corresponding sampling bucket in memory; the sampling bucket is cleared according to the time period by a scheduled task.

[0026] Quality filtering can flexibly select quality filtering strategies as needed. For example, quality filtering of the valid samples includes at least one of the following: response status code filtering, minimum field count filtering, and dual-condition cross-filtering of error keywords and sensitive data. Specifically, dual-condition cross-filtering of error keywords and sensitive data includes: maintaining a configurable set of error keywords; determining whether there is an intersection between the list of values ​​in the flattened key-value pair set and the set of error keywords to obtain an error hit identifier; determining whether the response location contains sensitive data based on the sensitive data identification results carried in the HTTP traffic log to obtain a sensitive hit identifier; when the error hit identifier is true and the sensitive hit identifier is false, it is determined to be an error response and discarded; when both the error hit identifier and the sensitive hit identifier are true, it is determined to be a normal response containing sensitive fields and retained; when the error hit identifier is false, it is retained normally.

[0027] The above technical solution will now be described with reference to the accompanying drawings and optional embodiments. The following figures contain the following reference numerals: The following embodiments provide a pre-processing method for templated extraction and classification of interface response fields based on two-layer fingerprint hashing and feedback loop. Figure 2 This is a flowchart of the template-based extraction and classification / grading preprocessing method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps: Step S1, extract and clean the response data; Step S2, calculate fingerHash and query the L1 cache using the quadruple cache key; if a match is found, return; otherwise, proceed to Step S3; Step S3, perform flattening extraction; Step S4, use TreeSet to calculate templateHash and write it back to the L1 cache, i.e., when calculating the template hash, first perform deduplication and sorting, and then calculate the hash value; Step S5, calculate cacheKeyHash (i.e., the hash value of the quadruple cache key) and query the classified and graded template set; if a match is found, stop sampling; Step S6, perform putIfAbsent adaptive sampling on the sampling bucket; Step S7, perform multiple quality filtering; Step S8, output in dual channels. Figure 2 The three early return branches in the illustrated process (cache hit, classified, and sampling saturation) together determine that this embodiment can keep the processing cost at a level that is approximately linear with the number of interfaces.

[0028] The above steps involve fingerprinting (also known as response body fingerprint), which is a unique identifier obtained by hashing the response body (the fingerprint of other data is the unique identifier obtained by hashing the response body). TemplateHash refers to the hash value corresponding to a message template. In this template hash, a fixed-format response body (fixed JSON structure, fixed HTML template) is used as the standard template; the hash value obtained by calculating the hash of the template bytes is the templateHash. Template hashing allows for quick matching of whether the interface's returned data conforms to the preset template format, identifying similar interface responses.

[0029] The above steps also involve L1 caching, where L stands for Level. L1 caching is the first-level cache in a multi-level caching architecture. It's the cache closest to the computation logic, hence its fastest speed. There are also typically L2 (second-level cache) and L3 (third-level cache) caches. The higher the level number, the slower the access speed, but the larger the storage capacity is usually.

[0030] The above steps involve TreeSet, a Java collection utility class based on a red-black tree (an ordered binary balanced tree). TreeSet has the following characteristics: automatic element sorting: the stored data will be arranged in ascending order according to the rules; automatic deduplication: duplicate elements are not allowed, and only one copy of the duplicate content will be kept.

[0031] Comparison: A regular HashSet is unordered and only removes duplicates; a TreeSet removes duplicates and sorts the data.

[0032] In one scenario, the JSON structure of responses from the same type of API is fixed, but the order of the fields may be scrambled each time they are returned. For example, two responses: json {"name":"test","id":1}{"id":1,"name":"test"}.

[0033] The content structure is completely identical, only the field order is different. They belong to the same template in terms of business logic, so they should calculate the same templateHash.

[0034] If you directly calculate the hash on the raw bytes, the hash will be different if the field order changes or the byte stream is different, making it impossible to identify that they are from the same template.

[0035] The workflow of TreeSet is as follows: Parse the JSON response body to extract all field names (keys); store all keys into a TreeSet; automatically remove duplicate fields; automatically sort by alphabetical order; retrieve the sorted, unique field sequence; calculate the hash of the ordered field string to obtain the templateHash. Even if the original JSON fields are out of order, after being sorted by TreeSet, the generated strings will be identical, and the templateHash values ​​will be consistent, achieving template matching. Of course, other methods can also be used for sorting and deduplication, which will not be elaborated upon here.

[0036] The above steps also involve sampling buckets and putIfAbsent, which will be explained below. A sampling bucket is a grouped storage unit used for traffic and / or packet sampling. For example, cacheKeyHash and templateHash can be used as distinguishing identifiers for buckets; traffic with the same identifier is grouped into the same sampling bucket. A bucket stores statistical counts, sampling tags, sampling thresholds, and other data for the same type of packet.

[0037] The sampling bucket can aggregate similar packets into the same bucket and control the sampling frequency separately; the sampling bucket can also record the cumulative number of hits for that type of packet; in addition, each bucket can be configured with an independent sampling strategy, such as sampling fewer high-frequency packets and sampling all low-frequency packets.

[0038] `putIfAbsent` is a commonly used concurrent Map method in Java. The logic of `putIfAbsent(key, value)` is as follows: if the key does not exist in the container: store the value and return null; if the key already exists in the container: do not modify it and directly return the original value. `putIfAbsent` implements the principle of only adding a value if it does not already exist.

[0039] The following is about Figure 2 The steps are explained in detail.

[0040] Step S1: Extract the application identifier, interface identifier, response body, response body hash value, and response content type from the Hypertext Transfer Protocol (HTTP) traffic log.

[0041] In this step, the English word for application is Application, or APP for short. In a broad sense, APP refers to any application that can run on a computer or mobile device, as long as it is not the operating system itself.

[0042] This step involves the response body, which will be explained below. After a client (e.g., a browser and / or an app) sends an HTTP request to the server, the server returns a complete HTTP response through the interface. The entire response consists of two parts: Response Header: containing metadata such as status code, server information, encoding, and cookies. The response header only contains descriptive information and does not store business data; Response Body: containing the actual data content that the server returns to the client, and is the core carrier of the response.

[0043] Step S2: Based on the response content type, perform a single-pass byte scan on the response body to obtain the first-level fingerprint hash (fingerHash); use the application identifier (appID), the interface identifier (apiID), the response body hash value (respBodyMD5), and the first-level fingerprint hash (fingerHash) to form a four-tuple hash as the cache key, and query the template hash cache; if a match is found, return the second-level template hash obtained from the match as the template hash of this response, for example, it can be written to the HTTP log (HttpLog) and returned; if no match is found, proceed to step S3.

[0044] There are many ways to obtain the first-layer fingerprint hash. In this step, a single-pass byte scan is used, that is, the entire binary bytes of the HTTP response body are read in a streaming manner. The data is traversed only once, and hash calculation is performed as it is read to calculate the hash value representing the entire response content. The calculated hash value is the first-layer fingerprint hash.

[0045] The quadruple hash in step S2 can be generated by the following formula: Quadruple hash = H(appId + separator + apiId + separator + respBodyMd5 + separator + fingerHash), where H is any hash function among CRC32 variant, MurmurHash3, and xxHash; the template hash cache is implemented using a local cache based on access time expiration, and the maximum number of entries and expiration time are dynamically loaded through the configuration center.

[0046] Step S3: Based on the response content type, the response body is flattened and extracted (for example, for key-value pair combinations in JSON or XML) to obtain a flattened key-value pair set consisting of a list of path-based key names and corresponding values, and all keys are collected into an ordered key set container.

[0047] Optionally, the response content type in step S3 includes JSON and XML, and their flattening extraction respectively adopts the following methods: (a) when the response content type is JSON, a streaming key-value pair extractor is used to traverse the response body in a streaming manner, and nested objects are flattened into flat key names in the form of "abc" using path-based key names, and arrays are merged and flattened; (b) when the response content type is XML, key business nodes are automatically identified to strip the envelope and namespace wrapping layer, and after removing escape sequences in CDATA segments, they are flattened into a homogeneous flattened key-value pair set; and each value is truncated according to a configurable single-value maximum length threshold.

[0048] Optionally, the ordered key set container in step S3 can be an ordered set based on a red-black tree, so that the concatenation of the template key string in step S4 is independent of the order in which the fields appear in the original response body, so that structurally equivalent responses can obtain the same second-level template hash despite differences in field order.

[0049] Step S4: Concatenate the keys in the ordered key set container into a template key string (templateKey) in lexicographical order, hash the template key string (e.g., CRC32) to obtain the second-level template hash (templateHash), and write the mapping of the four-tuple hash to the second-level template hash into the template hash cache.

[0050] Step S5: Using the bucket key (cacheKeyHash=CRC32(appID+apiID+templateHash)) of the combination of the application identifier, the interface identifier, and the second-layer template hash, query the immutable snapshot of the classified and hierarchical template set maintained based on atomic references; if the bucket key matches the snapshot, end the current process; otherwise, proceed to step S6.

[0051] Optionally, the classified and graded template set can adopt the following concurrent structure: (a) the current snapshot is held in a primitive type collection container with a primitive long type bucket key; (b) the snapshot is held in a reference with an atomic reference; (c) when the downstream classification and grading module completes template governance, a callback function is triggered, the callback function constructs a brand new primitive type collection container, and finally replaces the old snapshot once through the atomic reference; (d) after the query in step S5 obtains the current snapshot through the atomic reference, it performs an inclusion determination on the immutable set, and the entire query path is lock-free.

[0052] Step S6: Using the bucket key as the sampling bucket key, put the response body hash value into the sampling bucket; when the sampling bucket does not exist or the number of response body hash values ​​collected in the sampling bucket has not reached the sample count threshold and the current response body hash value is not in the sampling bucket, add it and determine it as a valid sample, i.e., putIfAbsent(cacheKeyHash,respBodyMD5,N_MAX); otherwise, determine it as an invalid sample and end this process.

[0053] Optionally, the sampling bucket may have a cross-restart state retention mechanism, including: (a) during the system startup phase, querying the field sample records collected on the current day from the persistent OLAP data storage, grouping and aggregating them according to the application identifier, the interface identifier and the second-layer template hash, and backfilling the set of response body hash values ​​corresponding to each group into the corresponding sampling bucket in memory; (b) clearing the sampling bucket daily through a scheduled task, so that the sampling bucket follows the state lifecycle of "daily accumulation, cross-day rotation, and cross-restart continuation".

[0054] Figure 5 is a schematic diagram of the sampling bucket state management according to an embodiment of this application. Figure 5 shows the state change curve of the sampling bucket (122) during the system operation cycle. Figure 5 In the diagram, the horizontal axis represents time, the vertical axis represents bucket capacity, and the dashed horizontal line represents the N_MAX threshold.

[0055] T0 is the system startup time, where "cold start backfilling" occurs: the aggregated results of samples collected on the same day are retrieved from the OLAP data storage and backfilled into the memory sampling bucket, so that the bucket state continues to the level before the restart (solid line for high-frequency interfaces, approximately 80% capacity; dashed line for low-frequency interfaces, approximately 20% capacity).

[0056] The period from T0 to T_clear (end of Day 1) is the "daily accumulation": high-frequency interfaces quickly reach N_MAX and then become saturated and no longer consume resources; low-frequency interfaces accumulate slowly.

[0057] When T_clear (cron triggered) triggers a "cross-day clear": the bucket is completely cleared, and accumulation starts again on the new day.

[0058] Step S7: Perform quality filtering on valid samples. The quality filtering includes at least: response status code filtering, minimum field count filtering, and dual-condition cross-filtering of error keywords and sensitive data. Samples that pass the quality filtering proceed to step S8.

[0059] The dual-condition cross-filtering of error keywords and sensitive data in step S7 may include the following steps: (a) maintaining a configurable set of error keywords; (b) determining whether there is an intersection between the list of values ​​in the flattened key-value pair set and the set of error keywords to obtain an error hit identifier; (c) determining whether the response location contains sensitive data based on the sensitive data identification result carried in the HTTP traffic log to obtain a sensitive hit identifier; (d) when the error hit identifier is true and the sensitive hit identifier is false, determining it as an error response and discarding it; when the error hit identifier is true and the sensitive hit identifier is true, determining it as a normal response containing sensitive fields and retaining it; when the error hit identifier is false, retaining it normally.

[0060] Figure 6 is a logic diagram of multiple quality filtering determination according to an embodiment of this application, such as... Figure 6 As shown, the logic proceeds sequentially as follows: S7-1 Response status code = 200; S7-2 Number of fields ≥ minKeyCount; S7-3 Detection of incorrect keyword hit; S7-4 Detection of sensitive data in the response body. Figure 6 The system employs a dual-condition cross-determination mechanism based on S7-3 and S7-4: a response is discarded only if both conditions of "hitting an incorrect word" and "not containing sensitive data" are met simultaneously; if an incorrect word is hit but the response body contains sensitive data, it is retained to avoid missing real sensitive fields.

[0061] This multi-layered filtering achieves a balance between filtering out erroneous responses and protecting sensitive fields.

[0062] Step S8: For each key-value pair in the flattened key-value pair set, write a field sample record to the OLAP data storage; when the outbound switch is enabled, further deliver the key-value pair as a message to the message queue.

[0063] Optionally, the field sample record in step S8 may include at least one of the following: application identifier, interface identifier, uniform resource identifier, response body hash value, hostname, application path, second-level template hash, response content type code, key, value list, insertion time and update version number; the update version number is generated by a globally unique monotonically increasing identifier generator and is used for incremental synchronization and change data capture of downstream modules.

[0064] Through the above steps, a first-layer fast fingerprint hash is calculated on the HTTP response body. The hash of the quadruple formed by the application ID, interface ID, response body MD5, and the first-layer fingerprint hash is used as the cache key. The template hash cache is queried; if a match is found, it is returned directly. If a match is missed, the response body is flattened to obtain a set of key-value pairs. An ordered key set container is used to collect all keys and calculate the second-layer template hash. A bucket key is formed using the application ID, interface ID, and the second-layer template hash. The immutable snapshot set maintained based on atomic references is queried to determine if classification and grading have been completed. If completed, it is skipped; otherwise, adaptive sampling and dual-condition cross-filtering of erroneous keywords and sensitive data are performed. Finally, the field samples are written to the OLAP wide table and message queue. The above steps outperform existing technologies in at least one of the following aspects: parsing performance, template merging accuracy, cross-restart state maintenance, and feedback loop concurrency.

[0065] Figure 1 This is a structural block diagram of the processing system according to an embodiment of this application, such as... Figure 1 As shown, in Figure 1 The solid arrows in the text indicate the direction of data or control flow. Figure 1The diagram illustrates a feedback loop: governance completion 200 - notification 210 - callback 133 - atomic replacement 131 / 132 - next zero-overhead stop sampling. The following section... Figure 1 The system will be explained in detail.

[0066] exist Figure 1 In this module, 170 is the traffic entry point; 110 is the traffic parsing module (scheduling the entire process according to steps S1 to S8); 140 is the fingerprint algorithm module, including the JSON fingerprint algorithm (FSM) labeled 141 and the XML fingerprint algorithm labeled 142; 120 is the two-level caching module, including the template hash cache (L1) labeled 121 and the sampling bucket (L2) labeled 122; 130 is the feedback loop module, including the immutable snapshot set labeled 131, the atomic reference labeled 132, and the cache update callback interface labeled 133; 150 is the configuration loading module; 160 is the output module, including the OLAP writer labeled 161 and the message queue delivery device labeled 162; 180 is the OLAP data storage; 190 is the message queue; 200 is the downstream classification and grading module; and 210 is the configuration caching framework. The feedback loop workflow is as follows: After template governance is completed (200), the configuration cache framework (210) detects the change and triggers a callback (133); the callback function constructs a new set (131) and atomically replaces it (132); (110) reads the snapshot held by (132) without locks during the next processing, achieving zero-overhead stop sampling for the governed template.

[0067] Figure 1 The system is also known as a pre-processing system for templated extraction and classification of interface response fields based on two-layer fingerprint hashing and feedback loop. The modules included in the system are described below.

[0068] The traffic parsing module is used to execute steps S1, S3, S4, S7, and S8 in the above method; the two-level caching module is used to maintain the template hash cache and the sampling bucket, and supports steps S2 and S6 in the above method; the feedback loop module is used to maintain the classified and graded template set in an immutable snapshot atomic replacement method based on atomic references, and supports step S5 in the above method; the fingerprint algorithm module is used to generate the first-level fingerprint hash on the response body in a single-pass byte scan manner; the configuration loading module is used to dynamically load the error keyword set, the sample count threshold, the single-value maximum length threshold, and the outgoing switch; the output module includes an OLAP data storage writer and a message queue delivery unit, and is used to execute step S8 in the above method.

[0069] Optionally, the two-level caching module may include: a first caching submodule, which stores the mapping from the quadruple hash to the second-level template hash, using a local cache based on access time expiration; and a second caching submodule, which stores the mapping from the bucket key to the set of response body hash values, serving as the sampling bucket; which performs a cold start backfill operation during the system startup phase and performs a daily clearing operation through a scheduled task.

[0070] Optionally, the feedback closed-loop module includes an implementation class of the cache update callback interface, which is configured such that when the classified and hierarchical template configuration in the upstream configuration cache changes, the configuration cache framework triggers its callback function; the callback function reads the latest classified and hierarchical template configuration list from the configuration cache, constructs a brand new primitive type collection container, and replaces the old snapshot once through atomic references.

[0071] Optionally, the traffic parsing module is registered with a scope of independent instances per thread, and there is no shared state between the instances of the traffic parsing module when multiple threads execute concurrently.

[0072] Figure 3 This is a timing diagram for a two-layer fingerprint hash cache lookup according to an embodiment of this application, such as... Figure 3 As shown, the traffic parsing module 110 requests the fingerprint algorithm module 140 to calculate the fingerprintHash. After the fingerprint algorithm module 140 returns the fingerprintHash, it requests a get operation from the L1 template hash cache 121. If the get operation fails, it returns null. The traffic parsing module 110 requests extractKV(respBody) from the flattening extractor. The flattening extractor returns a (key, valueList) set plus an ordered key set. The module then calculates the templateHash, requests put(quadKey, templateHash) from the L1 template hash cache, and requests to write to the OLAP wide table from the OLAP writer 161. This is the complete parsing path before the first response is received.

[0073] The fast path when the second identical response arrives is as follows: request to calculate fingerHash, return fingerHash, request get(quadKey), if hit, return templateHash, skip step S3 to step S8, that is, skip all parsing steps such as flattening extraction and OLAP writing in the second time (i.e., flattening, template hash calculation, bucketing, filtering, output and other steps are all omitted).

[0074] That is, in Figure 3In the first half of the parsing process, the first response follows the complete parsing path, sequentially performing fingerHash calculation, L1 cache lookup (miss), flattening extraction, template hash calculation, L1 cache write, and OLAP wide table write. In the second half, the same response arrives via a fast path, only performing fingerHash calculation and L1 cache lookup (hit), and then immediately returns; the flattening extraction, template hash calculation, sampling bucket, filtering, and output steps are all skipped. Because the quadruple cache key (appId, apiId, respBodyMd5, fingerHash) contains the response body MD5, the same response body can directly hit the cache on the second arrival, thus maximizing reuse.

[0075] Figure 4 Figure 4 is a schematic diagram of feedback closed-loop atomic snapshot replacement according to an embodiment of this application. It shows the concurrent model of atomic snapshot replacement in feedback closed loop. Figure 4 In the first half (write side): Callbacks are triggered at steps 200-210-133; the callback function constructs a new `131' LongOpenHashSet` (containing the managed `cacheKeyHash`) and atomically replaces it once using `132AtomicReference.set()`. This process does not hold any locks. In the second half (read side): step 110 obtains the current snapshot reference via `132.get()`, performs a `contains` check on the immutable set `131`, and operates entirely without locks. During the atomic replacement on the write side, the old snapshot `131*` held by the read side can still be safely read until the read operation ends and is subsequently garbage collected; write and read are completely decoupled.

[0076] The following examples will illustrate this point.

[0077] Example 1 illustrates this with specific data. Figure 2 The steps in the process.

[0078] Step S1: Extract the application identifier appId, interface identifier apiId, uniform resource identifier uri (where uri is the interface identifier that has been merged by URL template), hostname, application path appPath, response body respBody, response body hash value respBodyMd5, and response content type Content-Type from the HTTP traffic log (HttpLog) object; at the same time, perform purification processing on the response body, removing invisible special characters such as the BOM character (\uFEFF) in the response body header.

[0079] Step S2: Determine the response content type; when it belongs to application / json or text / json, use the JSON fingerprint algorithm based on finite state machine (see the relevant algorithm patent application of this invention) to perform a single scan of the response body byte stream to obtain the first-level fingerprint hash fingerHash; when it belongs to application / xml, text / xml or application / soap+xml, use the XML fingerprint algorithm to obtain the fingerHash; for other types, end this processing.

[0080] Next, construct the four-tuple cache key as follows: cacheKey = CRC32( appId + "_" + apiId + "_" + respBodyMd5 + "_" +fingerHash ); Query the template hash cache fingerHashCache using cacheKey. If a match is found, write the obtained respTemplateHash back to HttpLog and return immediately, skipping all subsequent steps S3 to S8; if no match is found, proceed to step S3.

[0081] Step S3: Flatten the response body according to the response content type.

[0082] For JSON responses, the streaming key-value extractor DataExtractor.jsonInstance().extractKV() is used to stream the response body. In the iteration callback, path-based key names (e.g., data.user.name) are added to the TreeSet type ordered key set container respTemplateKeySet, and the values ​​are truncated to the maximum single value length threshold V_MAX and added to the corresponding key's value list flatterValueMap.

[0083] For XML responses, first use XmlParseUtils.findKeyNode() to automatically identify key business nodes (automatically stripping SOAP envelopes and other wrapping layers), then use XmlParseUtils.clearXmlCData() to clear CDATA segments, and finally use XmlParseUtils.getFlatteredXml() to flatten the XML tree into a (key, valueList) structure of the same form, and add all keys to respTemplateKeySet in the same way.

[0084] Step S4: Concatenate all keys in respTemplateKeySet into respTemplateKey in lexicographical order: respTemplateKey = sortedKeySet.join(",") respTemplateHash = CRC32(respTemplateKey).

[0085] Because TreeSet is inherently ordered, and respTemplateKey is lexicographically unique, structurally equivalent responses (regardless of field order) will always yield the same respTemplateHash. After calculation, the mapping from cacheKey to respTemplateHash is written to fingerHashCache for quick cache matching of subsequent identical responses.

[0086] Step S5: Construct the categorized and hierarchical bucket key cacheKeyHash: cacheKeyHash = CRC32( appId + "_" + apiId + "_" + respTemplateHash ); Query the current snapshot of the categorized and hierarchical template collection maintained by the feedback closed-loop component: if ( cacheKeyHashSetRef.get().contains(cacheKeyHash) ) return; Where cacheKeyHashSetRef is AtomicReference <longopenhashset>The type, LongOpenHashSet, is a primitive long type open-addressable hash set from the fastutil library. The entire reading process is lock-free: the atomic reference's get operation returns a snapshot reference at a certain point in time, and the read end performs a contains check on this immutable snapshot.

[0087] Step S6: Using cacheKeyHash as the sampling bucket key, call the putIfAbsent(key, respMd5, N_MAX) operation of the two-level cache module. This operation is modified with synchronized, and the logic is as follows: if the bucket does not exist, create it and add respMd5, then return true; if the bucket exists but the number of elements it contains has not reached the threshold N_MAX and the current respMd5 is not in the bucket, then add it and return true; otherwise, return false. Only when true is returned is this response considered a valid sample, and proceed to step S7.

[0088] Step S7: Perform the following filtering on the valid samples in sequence: (a) Response status code filtering: Only allow responses with an HTTP status code of 200; (b) Minimum field count filter: Only passes if the size of finalFlatterValueMap is not less than the configured parameter minKeyCount; (c) Dual-condition cross-filtering of error keywords and sensitive data: Maintain a configurable set of error keywords, errorKeySet (e.g., "system error", "operation failure", "service unavailable"); obtain the error hit identifier hasErrorContent by intersecting the list of values ​​in flatterValueMap with errorKeySet; simultaneously, read sensitive data information identified by the upstream in the HttpLog through isNoSensitiveDataInResponseBody(httpLog) to obtain the sensitive hit identifier isNotSensitive; when hasErrorContent=true and isNotSensitive=true, it is judged as an error response and discarded; other cases are retained. The key value of this dual-condition design is to avoid misjudging responses containing sensitive fields (where a field value literally contains an error word) as error responses and thus missing them.

[0089] Step S8: For each (key, valueList) pair in flatterValueMap, write a row to the ClickHouse field sample wide table. The column structure includes: (appId, apiId, uri, respBodyMd5, hostName, appPath, respTemplateHash, respContentType, key, values="[v1,v2,...]", insertTime, updateVersion ).

[0090] The updateVersion is generated by YitIdHelper (a variant of Snowflake ID) and serves as a monotonically increasing primary key for downstream CDC consumption. When the outbound switch is enabled, a UrlSampleTemplateLog message is further constructed for each value in valueList and delivered to Kafka.

[0091] Example 2 illustrates end-to-end JSON processing.

[0092] Assume the input response body is: {"code":0,"data":{"userName":"Alice","userAge":30}}, appId=10, apiId=20.

[0093] respBodyMd5="abc123…"; fingerHash=0x1A2B3C; cacheKey=CRC32("10_20_abc123_0x1A2B3C"), query fingerHashCache failed.

[0094] DataExtractor.extractKV outputs: code → ["0"], data.userName → ["Alice"], data.userAge → ["30"]; TreeSet collects {code, data.userAge, data.userName}; respTemplateKey="code,data.userAge,data.userName"; respTemplateHash=0x7788AABB.

[0095] Write the mapping from cacheKey to respTemplateHash to fingerHashCache. cacheKeyHash=CRC32("10_20_2005978299")=0xCCDDEEFF.

[0096] The query for `cacheKeyHashSetRef` failed. `putIfAbsent` returned true. Status code 200, number of fields 3 ≥ `minKeyCount`, and no erroneous keywords were found; all conditions were met. Finally, three lines were written to ClickHouse (one line per key) and three messages were delivered to Kafka.

[0097] Example 3 shows the same response body arriving twice.

[0098] When the same response body (with the same respBodyMd5) arrives again, the cacheKey, containing the respBodyMd5, can directly hit the fingerHashCache. In this process, only the byte scan of fingerHash needs to be performed (milliseconds), and all subsequent steps S3 to S8 are skipped, including TreeSet construction, CRC32 calculation, sampling, filtering, ClickHouse writing, and Kafka push.

[0099] Example 4 shows different responders but with the same structure.

[0100] Suppose another response from interface X, {"code":0,"data":{"userName":"Bob","userAge":25}}, arrives with the same structure as in Example 2, but different data. The respBodyMd5 value is "xyz789…", and the fingerHash value remains 0x1A2B3C due to the identical structure (and field order). However, because the respBodyMd5 value in the cacheKey is different, the fingerHashCache query fails, requiring further complete parsing. The respTemplateHash obtained after complete parsing is the same as in Example 2 (both are 0x7788AABB), proving that the template merging implemented by the ordered key set in this invention is stable and effective for structurally equivalent responses. When this response body arrives again, it can hit the fast path of the fingerHashCache.

[0101] Example 5 illustrates a feedback loop.

[0102] After system startup, cacheKeyHashSetRef points to an empty LongOpenHashSet. When a certain interface template appears for the first time, cacheKeyHash=H1, the snapshot set is not hit, and the sampling process proceeds normally.

[0103] At a certain point, the downstream classification and grading module completes the governance of the template. The configuration caching framework detects a change in CACHE_DSAS_API_CATEGORY_NO_CONFIG and calls DsasApiCategoryNoConfigUtils.callBack(). This callback function reads the latest list of governed templates from the configuration cache, constructs a new LongOpenHashSet (containing H1), and atomically replaces the old snapshot in one go using cacheKeyHashSetRef.set(). This process does not hold any locks.

[0104] After that, when all traffic from this interface enters step S5, cacheKeyHashSetRef.get().contains(H1) returns true, immediately ending the processing and no longer consuming resources in subsequent steps.

[0105] Example 6 illustrates service restart cold start backfilling.

[0106] Assume that before the system restart, the sampling bucket H1 of interface X had accumulated multiple (e.g., 80) different respBodyMd5 values ​​(with a capacity of 80 / 100). The memory bucket is lost when the system shuts down. Upon restart, the @PostConstruct method initCache() of UrlSampleCacheComponent executes the following SQL query: SELECT app_id, url_temp_id, resp_template_hash, groupUniqArray(resp_body_md5) AS resp_body_md5 FROM url_sample_template WHERE toDate(insert_time) = :Today's date GROUP BY app_id, url_temp_id, resp_template_hash.

[0107] The `respBodyMd5` values ​​of multiple responses (80 in this example) from interface X in the query results are backfilled into bucket H1. After restarting, when new traffic arrives, the bucket state continues from before the restart, and the same response bodies that have already been collected will not be duplicated; the remaining 20 slots in the bucket can continue to be used for accumulating new response bodies. Early the next morning, the `dsmpSampleCacheClear` cron task is triggered, the bucket is cleared, and a new day's sampling rotation begins.

[0108] Example 7 illustrates high-concurrency execution using multiple threads.

[0109] The traffic parsing module is registered with the Spring `@Scope("thread")` scope. Each thread holds an independent `UrlSampleTemplateParser` instance, and intermediate state variables (such as `respTemplateKeySet`, `flatterValueMap`, etc.) are not shared among multiple threads. Shared state exists only in the thread-safe `UrlSampleCacheComponent` (whose `putIfAbsent` is protected by `synchronized`) and `DsasApiCategoryNoConfigUtils` (whose atomic snapshot mode is inherently thread-safe). Therefore, this invention can scale linearly in multi-core, high-concurrency scenarios without lock contention bottlenecks.

[0110] Some other alternative implementations are also given below.

[0111] (1) The hash function can be replaced by any 64-bit hash such as MurmurHash3, xxHash, CityHash; (2) The ordered key set container can be replaced by any data structure that guarantees lexicographical uniqueness, such as an ordered set based on a skip list or a sorted array; (3) The adaptive sampling strategy can be replaced by reservoir sampling, probability decay sampling, sliding time window sampling, etc.; (4) The feedback closed-loop carrier can be replaced by any shared read and write data structure such as CopyOnWriteArraySet, Redis Set or Bloom filter; (5) The template hash cache carrier can be replaced by Guava Cache, Ehcache, Redis, etc.; (6) The cold start data source can be replaced by Doris, StarRocks, Redis or local files, etc.; (7) The output channel can be extended to support any OLAP engine and any message middleware; (8) The response content type can be extended to support protobuf, form-urlencoded, yaml, etc.

[0112] For the output fields of the data retrieval interface, other processing methods can be used, such as directly scanning the original response body for sensitive words. This approach leads to repeated scanning of responses with the same structure, wasting computational resources and failing to establish a stable field-to-data element mapping. Directly parsing the JSON and / or XML corresponding to the data interface in its entirety incurs significant overhead; tests have shown that CPU consumption becomes unbearable in scenarios with hundreds of thousands of transactions per second. If an interface is high-frequency, the number of samples generated is enormous, resulting in high global database storage costs; if an interface is low-frequency, there is a possibility of oversampling. In other cases, there is also the problem of erroneous responses contaminating samples. For example, an HTTP 200 response with a business error could contaminate the field template if stored in the database, leading to misclassification and misjudgment in downstream systems. If templates are used, differences in field order can cause template fragmentation. For instance, the field order of JSON objects is not fixed in the standard; two semantically equivalent responses might be identified as different templates if their response bodies are hashed directly.

[0113] All of the above methods lack feedback linkage with downstream classification and grading modules, resulting in continued data collection of already-managed interfaces and long-term resource waste. Besides these issues, other problems may exist, such as the lack of cross-restart state persistence. After a service restart, the original sampling state is lost, causing already sampled interfaces to be resampled, leading to duplicate sample entry and wasted computation. Another example is the introduction of read / write contention due to feedback state changes: if the feedback set allows concurrent read / write, high-concurrency reads on hot paths will be slowed down by locking, making it difficult to balance real-time performance and consistency. All of these processing methods have various problems. The processing solution provided in the above embodiments of this application differs from these solutions and can solve at least one of these problems in at least one aspect. Specifically, the above embodiments provide a technical solution for HTTP interface response data, which automatically splits fields, merges and deduplicates templates, adaptively samples, and forms a feedback loop with downstream classification and grading modules through traffic samples, thereby providing pre-processing data governance for interface classification and grading. This technical solution automatically splits interface response fields without relying on interface documentation; automatically merges and deduplicates structurally equivalent response templates; eliminates the need to parse duplicate responses through two-level caching; makes resource consumption approximately linear with the number of interfaces through adaptive sampling and cross-restart state retention; filters error samples and protects sensitive fields; and forms a feedback loop with downstream classification and grading modules in a lock-free manner.

[0114] The above embodiments can achieve at least one of the following technical effects: (1) Significantly improved parsing performance: The two-level caching of the quadruple cache key enables zero parsing return for repeated responses, and it is estimated that CPU consumption can be reduced by more than 70% in typical production repetition rate (>80%) scenarios; (2) High template merging accuracy: The template hashing algorithm based on the ordered key set container eliminates template fragmentation caused by field order; (3) Zero hot path overhead for feedback closed loop: The immutable snapshot mode of atomic references and primitive type sets makes the read end completely lock-free; (4) Cross-restart state continuation: Cold start backfilling from OLAP data storage eliminates restart jitter; (5) High sample quality: Multiple filtering based on status code, number of fields, error keywords, and sensitive data cross-judgment significantly reduces the proportion of dirty samples; (6) Unified multi-format: JSON and XML follow the same downstream storage and classification logic; (7) Hot-configurable: Thresholds, switches, and keywords are all dynamically loaded through the configuration center without restarting; (8) Naturally supports high concurrency: The parsing module is registered with an independent instance range for each thread, and there is no lock contention between multiple threads.

[0115] In this embodiment, an electronic device is provided, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to perform the methods described in the above embodiments.

[0116] The aforementioned program can run on a processor or be stored in memory (or a computer-readable medium). Computer-readable media includes both permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0117] These computer programs may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes can be implemented using different modules, and different steps can be implemented using different modules.

[0118] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.< / longopenhashset>

Claims

1. A method for classifying and grading response data, characterized in that, include: Obtain the HTTP response from the HTTP traffic log; Obtain the application identifier, interface identifier, and response body of the HTTP response; The template hash corresponding to the HTTP response is found. The template hash is obtained by extracting the structural features of the response body, sorting and / or deduplicating the extracted structural features, and then performing a hash operation. The template hash is used to identify different structural features of the response body. Combine the application identifier, the interface identifier, and the template hash into a bucket key; The snapshot is queried in the categorized template set of the sampling bucket using the bucket key, and the found snapshot is used as the basis for classifying and classifying the HTTP response.

2. The method according to claim 1, characterized in that, Finding the template hash corresponding to the HTTP response includes: Based on the type of the HTTP response, the response body is scanned to obtain a fingerprint hash, wherein the fingerprint hash is a hash value used to identify the response body; The application identifier, the interface identifier, the response body hash value carried in the HTTP response, and the fingerprint hash combined into a four-tuple hash are used as the cache key; Use the cache key to look up the template hash corresponding to the HTTP response in the template hash cache.

3. The method according to claim 2, characterized in that, If the template hash corresponding to the HTTP response is not found, the method further includes: Based on the content type of the HTTP response, the response body is flattened to obtain a flattened key-value pair set, wherein the flattening extraction is to extract the structural features of the response body. The extracted keys are sorted and / or deduplicated, and then hashed to obtain the template hash corresponding to the HTTP response; The mapping between the quadruple hash and the template hash obtained by the operation is written into the template hash cache.

4. The method according to claim 3, characterized in that, If no corresponding snapshot is found in the categorized and hierarchical template set using the bucket key, the method further includes: Determine whether the response body is a valid sample. If so, use the bucket key as the sampling bucket key and put the hash value of the response body into the sampling bucket. The valid samples are quality filtered, and snapshots are generated based on the quality-filtered valid samples and saved in the classified and graded template set.

5. The method according to claim 4, characterized in that, Determining whether the response body is a valid sample includes: If the response body hash value is not present in the sampling bucket, or if the number of response body hash values ​​collected in the sampling bucket has not reached the sample counting threshold and the current response body hash value is not in the sampling bucket, it is added and determined to be a valid sample; otherwise, it is determined to be an invalid sample and the current processing ends.

6. The method according to claim 5, characterized in that, During the system startup phase, the sampling bucket queries the field sample records collected within the current time period from the persistent data storage, groups and aggregates them according to the application identifier, the interface identifier and the template hash, and fills the set of response body hash values ​​corresponding to each group back into the corresponding sampling bucket in memory. The sampling bucket is emptied according to the specified time period via a scheduled task.

7. The method according to claim 4, characterized in that, Quality filtering of the valid samples includes at least one of the following: response status code filtering, minimum field count filtering, and dual-condition cross-filtering of error keywords and sensitive data.

8. The method according to claim 7, characterized in that, Two-condition cross-filtering of error keywords and sensitive data includes: Maintain a configurable set of error keywords; Determine whether there is an intersection between the list of values ​​in the flattened key-value pair set and the set of error keywords to obtain the error hit identifier; Based on the sensitive data identification results carried in the HTTP traffic log, it is determined whether the response location contains sensitive data to obtain a sensitive hit indicator; When the error hit flag is true and the sensitive hit flag is false, it is determined to be an error response and discarded; when the error hit flag is true and the sensitive hit flag is true, it is determined to be a normal response containing sensitive fields and retained; when the error hit flag is false, it is retained normally.

9. An electronic device comprising a memory and a processor; wherein, The memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the steps of the method according to any one of claims 1 to 8.

10. A readable storage medium having computer instructions stored thereon, wherein, When executed by a processor, the computer instructions implement the steps of the method described in any one of claims 1 to 8.