A method for real-time detection and response of endpoint security threats

CN122824501APending Publication Date: 2026-09-25HANGZHOU DISHEN SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611242737.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-17
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本发明的目的在于提供一种端点安全威胁的实时检测与响应方法,以解决上述背景技术中提出的检测模型算力适配性差、上报策略灵活性不足、防护体系自迭代能力弱的问题

Benefits of technology

[0046](1)、通过多参数筛选融合得到统一的端点综合算力指标,模型裁剪时给核心特征层设置裁剪比例上限,避免检测能力出现大幅衰减;同时按通道威胁敏感度从高到低筛选保留,在算力限制内尽可能保留更多有效检测通道。端侧推理可随实时算力调整启用的通道数量,算力充足时多启用通道提升检测精细度,业务负载升高时也能维持基础检测不中断,可适配不同硬件配置的终端与动态负载场景。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824501A_ABST
    Figure CN122824501A_ABST
Patent Text Reader

Abstract

The application discloses a kind of real-time detection and response methods of endpoint security threat, belong to network security technical field.Collect endpoint hardware operating parameter, screen core computing power influence item, weightedly calculate comprehensive computing power index upload cloud end;Train security detection mother model, divide feature anchor point layer and tailorable layer, obtain full channel threat sensitivity matrix, convert the corresponding relationship of computing power and model calculation amount.According to sensitivity, complete model tailoring quantization, obtain lightweight submodel and issue;End side matching reasoning gear completes detection grading, combines confidence and power adjustment anchor point feature report.Establish risk-power two-dimensional response decision matrix, as needed select local or edge gateway disposal;Recycle feedback update sample library, freeze anchor point layer fine-tuning mother model, iteratively issue optimized submodel.This method adapts heterogeneous terminal and dynamic load scene, balances end side resource consumption and cloud end analysis demand, and gives consideration to protection effect and business operation stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of network security technology, specifically relating to a real-time detection and response method for endpoint security threats. Background Technology

[0002] In recent years, with the continuous advancement of enterprise digital transformation and the increasing number of IoT terminals deployed, endpoint security has become a crucial aspect of cybersecurity protection. Currently, threat detection and mitigation in the industry are primarily implemented through an end-to-cloud collaborative approach.

[0003] After such solutions were actually deployed, several problems were exposed. The hardware computing power of different terminals varies significantly, and the load during business operations is constantly changing. When existing detection models are trimmed and compressed, corresponding protection restrictions are not set for the core feature layer. After compression, the detection accuracy drops significantly, and the inference process cannot flexibly adjust to changes in computing power. This problem is more pronounced on low-configuration terminals. The amount and interval of data reported are basically fixed settings. Full reporting consumes too much terminal computing resources and bandwidth, while overly concise reporting content prevents in-depth analysis in the cloud. Furthermore, the entire protection system lacks a stable and continuous optimization method. Response strategies are only matched according to risk levels, which easily consumes resources required for business operation. Model iteration may also affect the originally stable core detection capabilities, making it difficult to balance protection effectiveness and business operation. To address these problems, this invention proposes a real-time detection and response method for endpoint security threats. Summary of the Invention

[0004] The purpose of this invention is to provide a real-time detection and response method for endpoint security threats, in order to solve the problems of poor computing power adaptability of detection models, insufficient flexibility of reporting strategies, and weak self-iteration capability of protection systems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a real-time detection and response method for endpoint security threats, comprising:

[0006] Step 1: Collect endpoint hardware operating status parameters, filter core computing power impact parameters, obtain endpoint comprehensive computing power index through weighted fusion and upload it to the cloud; construct a security detection mother model, divide it into feature anchor layer and ordinary customizable layer, generate a full-channel threat sensitivity matrix, and establish a linear mapping relationship between computing power index and model computation volume;

[0007] Step 2: Based on the security detection master model, the full-channel threat sensitivity matrix and the linear mapping relationship of computing power, generate a lightweight detection sub-model adapted to the endpoint computing power and distribute it; match the corresponding inference level on the edge and perform threat detection and risk classification; combine confidence and computing power matching reporting strategy to report the intermediate feature vector of the anchor point to the cloud;

[0008] Step 3: Based on the initial assessment of effective threats and risk levels, construct a two-dimensional response decision matrix of risk and computing power, and combine it with the comprehensive computing power index of the endpoint to execute local processing or offloading at the edge gateway; recover the processing feedback to update the sample library, fine-tune the security detection master model and correct the full-channel threat sensitivity matrix, and iteratively generate a lightweight detection sub-model for distribution and update.

[0009] Preferably, the specific process of collecting endpoint hardware operating status parameters, filtering core computing power influencing parameters, obtaining the endpoint comprehensive computing power index through weighted fusion, and uploading it to the cloud is as follows:

[0010] For each target endpoint, a preset parameter acquisition period is set to collect various hardware operating status parameters that affect the real-time computing capability of the endpoint;

[0011] Invalid parameters with too small fluctuation ranges were eliminated by variance screening, and core computing power impact parameters that meet the standards of correlation with the endpoint benchmark computing performance were selected by Pearson correlation coefficient method.

[0012] The core computing power impact parameters are linearly normalized according to their positive and negative attributes, and then the endpoint comprehensive computing power index is obtained by linear weighted fusion calculation.

[0013] The metric is bound to a unique endpoint identifier and transmitted in encrypted form to the cloud-based detection node.

[0014] Preferably, the specific process of constructing a security detection master model, dividing it into a feature anchor layer and a normal customizable layer, and generating a full-channel threat sensitivity matrix is ​​as follows:

[0015] A global sample library consisting of a global endpoint security threat sample library and a normal behavior sample set is pre-constructed. After data augmentation, the standardized training dataset and the standard threat verification dataset are divided. A network architecture combining temporal feature extraction and fully connected classification is built, and a security detection master model is trained.

[0016] A fixed-ratio random channel pruning is performed sequentially on each hidden layer. The decrease in detection accuracy is used to characterize the degree of feature representation destruction of the corresponding layer. Based on this, the feature anchor layer and the ordinary pruning layer are divided, and the pruning ratio protection constraint of the feature anchor layer is set.

[0017] A historical IoT threat sample set independent of the training set is selected as input, and the mean loss gradient of each network channel is calculated as the threat sensitivity. The entire channel threat sensitivity matrix is ​​generated by traversing all pruning layer channels.

[0018] Preferably, the specific process for establishing a linear mapping relationship between computing power indicators and model computational volume is as follows:

[0019] Establish a positive proportional mapping relationship between the comprehensive computing power index corresponding to the preset benchmark endpoint and the floating-point operation volume of the benchmark detection model;

[0020] By combining the preset hardware architecture calibration coefficients for correction, the comprehensive computing power index of the current endpoint is substituted into the mapping relationship to calculate the maximum computing power threshold that the current endpoint can bear, which serves as the computing power constraint boundary for model pruning.

[0021] Preferably, the specific process of generating and distributing a lightweight detection sub-model adapted to the endpoint computing power based on the security detection master model, the full-channel threat sensitivity matrix, and the linear mapping relationship between computing power is as follows:

[0022] With the goal of maximizing the total threat sensitivity of the retained network channels, and with the constraints that the total computational cost of the model does not exceed the maximum computational cost threshold and the pruning ratio of the feature anchor layer channels does not exceed the anchor layer protection ratio threshold, a greedy selection algorithm is used to solve the optimal pruning scheme. The network channels of each layer are sorted from high to low threat sensitivity, and the high-sensitivity channels are retained in sequence while the total computational cost is accumulated. When the accumulated computational cost approaches the threshold, the pruning ratio of the feature anchor layer is checked. If it exceeds the protection constraint, the low-sensitivity channels of the ordinary pruning layer are adjusted and pruned until all constraints are met.

[0023] Based on the final set of retained channels, a pruning mask containing channel retention and pruning identifiers is generated. Based on the pruning mask, channel pruning is performed on the security detection parent model to obtain a basic sub-model. After performing low-bit-width fixed-point quantization on the basic sub-model, a lightweight detection sub-model adapted to the current endpoint's computing power is generated. The lightweight detection sub-model and the corresponding pruning mask file are synchronously distributed to the corresponding endpoint.

[0024] Preferably, the specific process of performing threat detection and risk classification by matching the corresponding inference level on the endpoint is as follows:

[0025] For the target endpoint where the model deployment is completed, load the lightweight detection sub-model delivered from the cloud, collect the real-time behavior data stream of the endpoint, and generate standardized behavior sequence data by segmenting and aggregating it through a sliding time window.

[0026] Based on the full-channel threat sensitivity matrix corresponding to the security detection mother model, the channel ranking relationship is obtained. The retained channels in the ordinary pruning layer of the sub-model are extracted and the sensitivity is arranged from high to low using the same ranking. After testing the average inference computing power consumption of a single channel, multiple ordinary pruning layer channel groups with equal computing power are divided according to the preset single-group benchmark inference computing power consumption. The remaining channels that are less than one group are grouped independently. With all feature anchor layers fully enabled as the base state, high-sensitivity channel groups are superimposed in sequence and multiple sets of computing power accuracy correlation data are obtained through testing. Based on this, multiple levels of inference tiers are divided and the benchmark computing power step size of a single tier is determined.

[0027] Read the current endpoint's comprehensive computing power index to match the corresponding inference level, enable the corresponding number of high-sensitivity channel groups and keep the feature anchor layer fully enabled, input the standardized behavior sequence to perform inference, and output the threat classification result, confidence level and risk level; after filtering by the effective judgment threshold, the effective threat preliminary judgment result is obtained.

[0028] Preferably, the specific process of reporting the intermediate feature vector of the anchor point to the cloud is as follows:

[0029] During the forward propagation calculation of the lightweight detection sub-model, the feature tensors at the output of each feature anchor layer are extracted as the intermediate feature vectors of the anchor points; confidence intervals are divided based on historical false alarm statistics, and computing power intervals are divided based on the edge computing power reservation standard. The two types of intervals intersect to form multiple sets of confidence-computing power combinations, and a reporting strategy configuration library is preset for each set of combinations corresponding to the reporting granularity and reporting cycle parameters.

[0030] Read the threat confidence level and endpoint comprehensive computing power index of the current window, match them to obtain the corresponding reporting granularity and reporting cycle; process the anchor intermediate feature vector according to the corresponding granularity, package it together with the initial threat judgment result and the endpoint unique identifier, upload it to the cloud detection node according to the corresponding cycle and store it in the global endpoint security threat sample library.

[0031] Preferably, based on the initial assessment of effective threats and risk levels, a two-dimensional risk-computing power response decision matrix is ​​constructed. The specific process for executing local handling or edge gateway offloading, combined with the endpoint's comprehensive computing power index, is as follows:

[0032] Pre-test the peak computing power consumption of each security response strategy on the endpoint device, and establish a corresponding mapping relationship between security response strategies and peak computing power consumption;

[0033] Historical threat handling samples were collected and divided into risk level and computing power range combinations. Matching strategies for each combination were selected with the upper limit of the corresponding computing power range as a constraint and the highest handling efficiency as the objective. A two-dimensional response decision matrix of risk level and computing power range was constructed.

[0034] Based on the initial assessment of effective threats, determine the corresponding risk level, read the current endpoint's comprehensive computing power index, substitute it into the two-dimensional response decision matrix to obtain the target security response strategy, and query the corresponding peak computing power consumption.

[0035] If the endpoint's overall computing power index meets the requirements, the corresponding security response policy is executed locally; otherwise, a response unload request carrying a threat identifier, risk level, and handling instructions is sent to the pre-bound edge gateway.

[0036] The edge gateway matches the corresponding handling execution authority based on the risk level carried in the request, executes the corresponding security response action, and then returns a handling result receipt to the endpoint.

[0037] Preferably, the recycling and disposal feedback updates the sample library, and the specific process is as follows:

[0038] The endpoint records the handling results and generates local security logs, verifies the handling effect to obtain false alarm and missed alarm confirmation information, and reports the handling results to the cloud detection node and updates the handling records of the corresponding threat events;

[0039] The cloud performs cleaning on the reported anchor point intermediate feature vectors, threat judgment results and handling feedback data, removes invalid data and labels three types of samples based on confidence intervals and verification results;

[0040] Add valid threat samples and unknown threat features to the global endpoint security threat sample library, and update the feature dimensions and coverage of the sample library.

[0041] Preferably, the specific process of fine-tuning the security detection master model and correcting the full-channel threat sensitivity matrix, and iteratively generating a lightweight detection sub-model for updates is as follows:

[0042] Based on the expanded global endpoint security threat sample library and handling feedback data, incremental fine-tuning training was performed on the security detection master model with the goal of reducing the false alarm and false negative rates. During training, all feature anchor layer parameters were frozen, and only the weight parameters of the ordinary pruning layer were updated.

[0043] The channel contribution is calculated based on the channel activation response of false alarm and false alarm samples, and the sensitivity value of the corresponding channel in the full channel threat sensitivity matrix is ​​corrected in reverse to optimize the priority ranking of channel pruning.

[0044] Based on the preset iteration cycle and the latest comprehensive computing power indicators reported by each endpoint, the optimized lightweight detection sub-model is generated by reusing the channel pruning and low bit-width fixed-point quantization processing flow, and then sent to the corresponding endpoint to complete the local model update.

[0045] Compared with the prior art, the beneficial effects of the present invention are:

[0046] (1) A unified endpoint comprehensive computing power index is obtained through multi-parameter screening and fusion. When pruning the model, an upper limit is set for the pruning ratio of the core feature layer to avoid a significant decrease in detection capability. At the same time, channels are selected and retained from high to low according to their threat sensitivity, so as to retain as many effective detection channels as possible within the computing power limit. The number of channels enabled by the edge inference can be adjusted according to the real-time computing power. When the computing power is sufficient, more channels are enabled to improve the detection precision. When the business load increases, the basic detection can be maintained without interruption. It can adapt to terminals with different hardware configurations and dynamic load scenarios.

[0047] (2) The reporting process is flexibly adjusted based on the threat confidence level and the computing power of the endpoints, and the feature data output by the anchor layer is directly extracted as the reporting content. When the confidence level is high and the computing power is sufficient, the complete features are uploaded to meet the needs of cloud-based source tracing analysis; when the computing power is tight and the confidence level is low, the reporting content is reduced to reduce the computing and bandwidth consumption on the endpoint, and a reasonable balance is achieved between the needs of both ends.

[0048] (3) This method establishes a complete process from threat detection and graded handling to model optimization. The response and handling simultaneously considers the risk level and endpoint computing power. When the endpoint computing power is insufficient, the operation is transferred to the edge gateway to avoid crowding out terminal business operation resources. After the handling feedback is collected, the samples are classified and labeled. When fine-tuning the parent model, the anchor layer parameters are fixed, and only the weights of the ordinary pruning layer are adjusted. At the same time, the channel sensitivity ranking is updated. Under the premise of ensuring the stability of the core detection capabilities, the false alarm and false negative rates are gradually reduced, and the operational stability of the entire protection system is improved. Attached Figure Description

[0049] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] Example 1;

[0052] Please see Figure 1 This invention provides a real-time detection and response method for endpoint security threats, comprising:

[0053] Step 1: Collect endpoint hardware operating status parameters, filter core computing power impact parameters, obtain the endpoint comprehensive computing power index through weighted fusion, and upload it to the cloud; construct a security detection master model, divide it into feature anchor layer and ordinary customizable layer, generate a full-channel threat sensitivity matrix, and establish a linear mapping relationship between computing power index and model computational load. The specific process is as follows:

[0054] For each target endpoint, a corresponding parameter acquisition period is preset, and various hardware operating status parameters that affect the real-time computing capability of the endpoint are collected according to the parameter acquisition period.

[0055] Operating status parameters include, but are not limited to: processor load rate, available memory percentage, remaining battery percentage, disk I / O utilization, CPU clock speed scheduling level, and number of system processes; and all collected parameters are temporarily stored in the local cache on the device side.

[0056] For each collected operating status parameter, obtain the parameter time series value within the preset statistical window, and calculate the variance of the numerical fluctuation of each parameter; preset a minimum effective fluctuation threshold. If the variance of the numerical fluctuation of a parameter is less than the preset minimum effective fluctuation threshold, the corresponding parameter is removed; otherwise, it is marked as a valid parameter to be verified.

[0057] For all valid parameters to be verified, the correlation coefficient between each parameter and the endpoint benchmark is calculated using the Pearson correlation coefficient method.

[0058] A preset correlation determination threshold is set, and the absolute value of the correlation coefficient corresponding to each valid parameter to be verified is compared with the correlation determination threshold. Parameters with an absolute value of correlation coefficient greater than or equal to the preset correlation determination threshold are retained and marked as core computing power impact parameters.

[0059] All core computing power impact parameters are classified according to their index attributes and linear normalization is performed, mapping all parameters to a unified numerical range of 0 to 1. The normalized parameter values ​​are then organized into a set of normalized core computing power impact parameters.

[0060] Among them, parameters with larger values ​​represent stronger available computing power and are marked as positive parameters, while parameters with larger values ​​represent weaker available computing power and are marked as negative parameters.

[0061] When performing linear normalization, positive linear normalization is used for positive parameter classes, and negative linear normalization is used for negative parameter classes.

[0062] Obtain the normalized values ​​of each core computing power influence parameter in the normalized core computing power influence parameter set for the current collection period, and use the linear weighted fusion formula:

[0063] The endpoint comprehensive computing power index S is obtained;

[0064] Where i is the index of the core computing power impact parameter, and m is the total number of core computing power impact parameters. This represents the normalized value corresponding to the parameter affecting the i-th core computing power. The preset weighting coefficient corresponding to the i-th core computing power influence parameter;

[0065] The comprehensive computing power index corresponding to the current endpoint is bound to the unique identifier of the endpoint and transmitted to the cloud detection node using conventional encryption methods in this field.

[0066] The security detection master model is constructed to classify and detect security threats in various endpoint behavior data. The specific process of building the security detection master model is as follows:

[0067] A global endpoint security threat sample library and a normal behavior sample set are pre-built, which together constitute the global sample library;

[0068] Extract endpoint behavior samples from a global sample library across multiple scenarios, covering normal business behavior samples and various types of endpoint security threat samples;

[0069] Endpoint security threat samples include, but are not limited to: malicious process execution, abnormal network connection, configuration file tampering, fileless attack, and privilege escalation behavior samples.

[0070] All endpoint behavior samples are labeled, cleaned, and subjected to data augmentation processes such as temporal perturbation and fragment masking. They are then divided into two parts according to a preset ratio to construct a standardized training dataset and a standard threat verification dataset, respectively.

[0071] A deep neural network architecture combining a multi-layer temporal feature extraction network and a fully connected classification layer was constructed.

[0072] Deep neural networks consist of: an input layer, multiple hidden layers, and an output layer;

[0073] The network front-end extracts local spatial features and long temporal dependency features of the behavioral sequence through multi-layer feature extraction units;

[0074] Feature extraction units include, but are not limited to: convolutional units and recurrent units;

[0075] The network backend outputs threat classification results and corresponding confidence levels through a fully connected layer and a normalized output layer.

[0076] The method for calculating the confidence level is a conventional technique in this field and is not further limited.

[0077] Using the classification loss function as the optimization objective, an iterative optimization algorithm is employed to train the network parameters of the deep neural network;

[0078] The network parameters are the trainable weights and biases of the neural network.

[0079] After each round of training, the threat detection accuracy and recall are tested on the standard threat verification dataset.

[0080] A preset detection accuracy threshold is set. If the core detection index reaches the preset detection accuracy threshold, training stops and the trained security detection master model is output.

[0081] For all hidden layers of the security detection parent model, select one hidden layer at a time as the target compression layer;

[0082] In this context, the hidden layer is an intermediate network layer located between the input layer and the output layer in a deep neural network.

[0083] Perform a preset fixed proportion of random channel pruning on the target compression layer, while keeping the structure and network parameters of the remaining hidden layers unchanged;

[0084] After channel pruning, the security detection master model is subjected to inference testing on a standard threat verification dataset. The decrease in the model's threat detection accuracy is calculated, and this decrease is used as the degree of feature destruction of the corresponding network layer.

[0085] A preset anchor point judgment threshold is set, and all hidden layers of the security detection parent model are traversed. Network layers whose feature representation is more damaged than the preset anchor point judgment threshold are marked as feature anchor layers, and the remaining layers are marked as ordinary pruning layers.

[0086] The anchor point determination threshold is preset based on the maximum allowable drop in detection accuracy.

[0087] Set a preset anchor point layer protection ratio threshold and set structural protection constraints for all feature anchor point layers;

[0088] The structural protection constraint is specifically: the channel pruning ratio of the feature anchor layer shall not exceed the preset anchor layer protection ratio threshold;

[0089] The ordinary cuttable layer is not constrained by the anchor point layer protection ratio threshold.

[0090] After completing the labeling and constraint settings for all network layers, output the security detection parent model network topology and feature anchor layer set with structural protection constraint labels.

[0091] A historical IoT threat sample set, independent of the standardized training dataset, is selected as the computational input for the channel sensitivity metric.

[0092] Input the historical IoT threat sample set into the security detection master model, perform forward propagation calculation, and obtain the model prediction results and classification loss function value.

[0093] Perform the backpropagation process to obtain the loss gradient of all network parameters for each network channel in each network layer;

[0094] The average of the absolute values ​​of the gradients of all network parameters within each network channel is taken, and this average value is used as the threat sensitivity value of the corresponding network channel.

[0095] The formula for calculating single-channel threat sensitivity is as follows:

[0096] ;

[0097] in, Let L be the threat sensitivity value of the k-th network channel in layer l, and L be the classification loss function. This is the set of all network parameters for the k-th network channel in layer l. For partial derivative operators, This is to calculate the average of the absolute values ​​of the gradients of all network parameters within the channel;

[0098] Traverse all network channels of all customizable layers in the security detection master model, store the threat sensitivity values ​​of each network channel in the order of network layer, and generate a full-channel threat sensitivity matrix.

[0099] The benchmark comprehensive computing power index value corresponding to the benchmark endpoint is preset, as well as the floating-point operation volume of the benchmark detection model that the benchmark endpoint can bear, and a linear mapping relationship between the comprehensive computing power index and the maximum allowable computing volume of the model is established.

[0100] The linear mapping relationship is the direct proportional relationship between the endpoint computing power level and the maximum computing power of the model that can be supported.

[0101] Substitute the current comprehensive computing power index reported by the endpoint into the linear mapping relationship to calculate the maximum computing power threshold that the current endpoint can bear, which serves as the computing power constraint boundary for model pruning.

[0102] The mapping calculation formula is as follows:

[0103] ;

[0104] in, The maximum computational cost allowed for the model at the current endpoint. S represents the floating-point computational cost of the benchmark detection model, and S represents the overall computing power index of the current endpoint. The baseline comprehensive computing power index value is set for the preset baseline endpoint. This is the preset hardware architecture calibration coefficient.

[0105] It should be noted that this step serves as the underlying foundation of the entire solution—encompassing five core components: computational power optimization, master model construction, hierarchical division, sensitivity calibration, and computational power mapping. The design of each component is deeply tied to the actual gain.

[0106] First, invalid parameters without fluctuation are eliminated through initial screening of fluctuation variance. Then, the core computing power influencing parameters that are strongly correlated with the endpoint computing performance are identified through secondary screening using Pearson correlation coefficient. Combined with positive and negative differential normalization and weighted fusion, the endpoint comprehensive computing power index is output (which not only eliminates redundant parameters to reduce computing overhead, but also provides a unified computing power evaluation benchmark for subsequent model pruning, tier matching, and response decision-making).

[0107] Through actual measurement of the accuracy reduction of single-layer random channel pruning, the feature anchor layer and ordinary pruning layer are objectively divided, and the anchor layer pruning ratio protection constraint is set simultaneously—delineating the lightweight security boundary from the network structure level, and fundamentally avoiding the mispruning of core threat features when the end-side model is compressed.

[0108] Using historical IoT threat samples independent of the training set, the threat sensitivity of a single channel is calculated through inverse gradient and a full-channel threat sensitivity matrix is ​​generated, providing an objective priority basis for subsequent channel pruning and ensuring that threat detection capabilities are maximized under the same computing power constraints;

[0109] By introducing hardware architecture calibration coefficients to correct architectural differences, a linear mapping relationship between comprehensive computing power indicators and the maximum computing power of the model is established, transforming the abstract computing power status into pruning constraints that can be directly implemented. This not only improves the solution's adaptability and coverage to heterogeneous endpoints, but also provides a reusable computing power conversion benchmark for subsequent model iterations.

[0110] Step 2: Based on the security detection master model, the full-channel threat sensitivity matrix, and the linear mapping relationship with computing power, generate a lightweight detection sub-model adapted to the endpoint computing power and distribute it; the endpoint matches the corresponding inference level and performs threat detection and risk classification; combining confidence and computing power matching reporting strategies, the intermediate feature vector of the anchor point is reported to the cloud. The specific process is as follows:

[0111] With the optimization objective of maximizing the total threat sensitivity of preserved network channels, and with constraints that the total computational cost of the model does not exceed the maximum computational cost threshold and the channel pruning ratio of the feature anchor layer does not exceed the anchor layer protection ratio threshold, a greedy selection algorithm is used to solve for the optimal pruning scheme. The execution process is as follows:

[0112] For all network channels within each network layer, sort them from highest to lowest threat sensitivity value;

[0113] Based on the sorting results, high-sensitivity network channels are retained sequentially from each layer, while the total computational cost of the model corresponding to the currently retained network channel is accumulated.

[0114] When the cumulative computational load approaches the maximum computational load threshold, verify the channel pruning ratio of all feature anchor layer layers.

[0115] If the pruning ratio of a feature anchor layer channel exceeds the anchor layer protection ratio threshold, the pruning target is adjusted to prune low-sensitivity network channels in ordinary pruning layers until the total computational cost meets the threshold requirement and all feature anchor layers meet the protection constraints.

[0116] Based on the final set of retained channels, a pruning mask is generated for each layer. The mask contains the retention and pruning identifiers for each layer's channels, which are used to mark the locations of channels that need to be retained and removed in each layer of the network.

[0117] Channel pruning is performed on the security detection parent model based on the pruning mask to obtain the pruned basic sub-model.

[0118] Low-bit-width fixed-point quantization (including but not limited to 8-bit integer quantization) is performed on the pruned basic sub-model to compress the storage space of model parameters and reduce the computational resource consumption of the inference process, and finally a lightweight detection sub-model adapted to the current endpoint computing power is generated.

[0119] The lightweight detection sub-model and the corresponding pruning mask file are simultaneously sent to the corresponding endpoints.

[0120] Furthermore, the specific implementation methods of confidence calculation, forward propagation operation, back propagation operation, parameter quantization, etc. involved in this step are all conventional technical means in this field, and are not further elaborated or limited.

[0121] For each target endpoint where the model deployment is completed, load the corresponding lightweight detection sub-model delivered from the cloud and collect various behavioral data generated during the real-time operation of the endpoint;

[0122] Real-time behavioral data streams include, but are not limited to: process calls, network connections, configuration file changes, and system call sequences;

[0123] The sliding time window length and sliding step size are preset, and the real-time behavior data stream is segmented and aggregated according to the sliding time window to generate standardized behavior sequence data of fixed length.

[0124] Based on the full-channel threat sensitivity matrix corresponding to the security detection master model, the threat sensitivity ranking relationship of each network channel is obtained;

[0125] For the lightweight detection sub-model corresponding to the target endpoint, extract all network channels actually retained in its ordinary pruning layer;

[0126] Using the ranking relationship in the full-channel threat sensitivity matrix, the ordinary customizable layer network channels retained in the sub-model are ranked from high to low sensitivity.

[0127] Pre-testing of the lightweight detection sub-model reveals the average inference computational cost of a single ordinary scalable layer network channel processing a single window behavior sequence.

[0128] The baseline inference computing power consumption of a single channel is preset as the target computing power increment when a single channel is enabled.

[0129] Divide the baseline inference computing power consumption of a single channel by the average inference computing power consumption of a single network channel per window, and round down the result to obtain the number of network channels contained in a single ordinary scalable layer channel group.

[0130] Based on the calculated number of channels, the sorted ordinary customizable layer network channels within the sub-model are sequentially divided into multiple ordinary customizable layer channel groups. Channels that are less than one group are grouped independently, so that the average inference computing power consumption per single window for each group is basically the same.

[0131] With all feature anchor layers in the lightweight detection sub-model fully enabled as the base state, starting from the first group of ordinary pruning layer channels with the highest sensitivity, each group of ordinary pruning layer channels is enabled sequentially.

[0132] The detection accuracy and average inference computational power consumption per single window of the model under each superimposed state were tested on the standard threat verification dataset to obtain multiple sets of computational power-accuracy correlation data.

[0133] Based on multiple sets of computing power-precision correlation data obtained from the test, multiple levels of inference tiers are defined.

[0134] Each level corresponds to a fixed number of ordinary customizable layer channels, a corresponding computing power range, and a corresponding accuracy range. The computing power range and accuracy range are determined by the actual measured data under the corresponding level.

[0135] Construct mapping rules between multi-level inference levels and computing power ranges, and establish a one-to-one correspondence between inference levels and the number of ordinary customizable layer channel groups.

[0136] The measured computing power consumption of each gear is converted into a normalized proportion of the total computing power of the endpoint. The average normalized computing power difference between adjacent gears is calculated, and this average value is set as the benchmark computing power step size for a single gear, serving as the unified calculation unit for online gear matching.

[0137] Read the endpoint comprehensive computing power index of the current collection period and match it with the corresponding inference level. The level matching calculation formula is as follows:

[0138] ;

[0139] Where G is the inference tier number and S is the current endpoint's overall computing power index. This refers to the single-level baseline computing power step size; both are normalized relative values. This is for floor function.

[0140] Based on the inference level obtained from the matching, activate the corresponding number of high-sensitivity ordinary pruning layer channel groups, and keep all feature anchor layers fully enabled.

[0141] Input the standardized behavior sequence of the current window of the target endpoint into the corresponding lightweight detection sub-model, perform forward propagation calculation, and output the threat classification result, corresponding confidence level and risk level;

[0142] Forward propagation computation is the forward computation process from input data to output result of the model.

[0143] The risk level is calculated using a quantitative formula, as follows:

[0144] ;

[0145] Where R is the quantified value of the threat risk level, is the preset weight coefficient corresponding to the t-th type of threat, c is the confidence level of the threat determination result, and t is the threat type number.

[0146] The system is set to three risk levels: high, medium, and low. Each risk level corresponds to a range of quantitative values ​​for threat risk.

[0147] The threat risk level quantification value calculated from the target endpoint is matched with the corresponding interval of each level to output the corresponding risk level.

[0148] A preset threshold for valid threat determination is set. The output confidence level is compared with the threshold. If the confidence level is greater than or equal to the corresponding preset threshold, the corresponding result is determined to be a valid initial threat determination result. If the confidence level is lower than the threshold, it is determined to be a false alarm and discarded.

[0149] During the forward propagation calculation of the lightweight detection sub-model, when the behavioral sequence data stream flows through the output of each feature anchor layer, the feature tensor output by the corresponding layer is extracted as the intermediate feature vector of the anchor point.

[0150] Confidence intervals are defined based on historical false alarm statistics, and computing power intervals are defined based on edge computing power reservation standards. The intersection of these two types of intervals forms multiple sets of confidence-computing power combinations.

[0151] A pre-defined reporting strategy configuration library is provided, in which each confidence level-computing power combination corresponds to a set of reporting granularity and reporting cycle parameters.

[0152] Among them, the reporting granularity includes two types: complete anchor intermediate feature vector and feature hash digest. The two types of granularity correspond to different information completeness and single data volume.

[0153] The reporting cycle can be set to multiple durations, with each duration corresponding to a different reporting frequency.

[0154] The parameter correspondence in the reporting strategy configuration library can be constructed in the following way: enumerate candidate reporting schemes containing different reporting granularities and reporting periods;

[0155] For each confidence-computing power combination, the data tracing benefit is determined by the proportion of real threats in the corresponding confidence interval and the information completeness of the reporting granularity, and the resource cost is determined by the total resource consumption of edge computing and transmission per unit time of the candidate solution.

[0156] Using the upper limit of the tolerable cost for the corresponding computing power range as a constraint, candidate reporting schemes that meet the cost constraints are selected, and the scheme with the highest data traceability benefit is selected as the reporting parameter corresponding to the combination.

[0157] The reporting parameter correspondence of all combinations is summarized to form a reporting strategy configuration library.

[0158] Read the threat confidence level and endpoint comprehensive computing power index of the current window, determine the interval to which they belong, and match the reporting strategy configuration library to obtain the corresponding reporting granularity and reporting period;

[0159] The intermediate feature vector of the anchor point is processed according to the reporting granularity obtained by matching, and the feature data is reported according to the reporting period obtained by matching.

[0160] The processed feature data, initial threat assessment results, and endpoint unique identifiers are packaged and uploaded to the cloud detection node according to the corresponding reporting cycle.

[0161] After receiving all the data reported by the endpoints, the cloud-based detection node stores it in the global endpoint security threat sample library for subsequent in-depth threat tracing and iterative optimization of the detection model.

[0162] It should be noted that this step follows the previous parent model (including the results of hierarchical division and sensitivity ranking) and computing power mapping results, and sequentially completes three processes: model lightweight pruning and distribution, edge inference level matching, and anchor point feature backhaul.

[0163] When pruning channels, they are retained sequentially from highest to lowest threat sensitivity, while ensuring that the total computational load of the model (including floating-point operations and memory usage) does not exceed the endpoint's capacity limit. The pruning ratio of the feature anchor layer also cannot exceed a preset protection value (i.e., the pruning ratio protection threshold). If the anchor layer pruning ratio reaches its limit, the pruning process switches to pruning channels with lower sensitivity in the ordinary pruning layers until all constraints (including computational constraints and anchor layer protection constraints) are met. After pruning, a corresponding pruning mask is generated, followed by low-bit-width fixed-point quantization compression (primarily 8-bit integer quantization), ultimately resulting in a lightweight detection sub-model adapted to the target endpoint's computing power. This pruning mask can be reused in subsequent model iterations without needing to be regenerated.

[0164] When the endpoint runs the model, it refers to the original channel sensitivity ranking of the parent model and divides the retained channels into several groups, with each group having a similar computing power consumption (the computing power deviation of a single group is controlled within a preset range). The endpoint tests by stacking groups one by one, recording the computing power consumption and detection accuracy (including the two core indicators of detection accuracy and false alarm rate) corresponding to different channel combinations. Based on this, multiple inference levels are defined. Throughout the test, the feature anchor layer remains fully enabled and does not participate in level adjustments. In actual operation, the endpoint automatically matches the corresponding level based on the current computing power load. If there is surplus computing power, more high-sensitivity channels are enabled to improve detection precision; if the load increases, the level is lowered (adjusted step-by-step according to preset increments) to reduce the number of enabled channel groups, avoiding a significant drop in detection accuracy or direct interruption of the detection service due to computing power fluctuations.

[0165] The reporting method for intermediate features at anchor points is not fixed and will be adjusted based on the current threat confidence level (including the confidence score of the detection output) and the endpoint computing power. Referring to historically accumulated false alarm data and the reserved computing power margin on the endpoint, multiple combinations of confidence levels and computing power are divided (by interval), each corresponding to different levels of reported content detail and reporting intervals. During inference, the output tensor of the feature anchor layer is directly extracted and uploaded (including both feature hash digest and full feature modes), eliminating the need to transmit the full original data. This reduces the computational and transmission overhead on the endpoint, and the features obtained by the cloud are sufficient for threat attribution and sample database expansion.

[0166] Step 3: Based on the initial assessment of effective threats and risk levels, construct a two-dimensional risk-computing power response decision matrix. Combine this with the endpoint comprehensive computing power index to execute local handling or edge gateway offloading. Recover handling feedback to update the sample library, fine-tune the security detection master model and correct the full-channel threat sensitivity matrix, iteratively generate a lightweight detection sub-model and distribute it for updates. The specific process is as follows:

[0167] Pre-test the peak computing power consumption of each security response strategy on the endpoint device, and establish a corresponding mapping relationship between security response strategies and peak computing power consumption.

[0168] Historical threat handling samples were collected, covering threat risk level, endpoint computing power conditions during the handling period, security response strategies adopted, handling results, and actual computing power usage data. Threat risk level and endpoint computing power ranges were divided, and the intersection of the two types of ranges formed multiple risk-computing power combinations.

[0169] For each risk-computing power combination, extract all historical handling samples within the corresponding range and group them according to the type of security response strategy adopted;

[0170] For each set of security response strategies, the handling results and computing power usage data are statistically analyzed, and the average handling efficiency and average computing power usage are calculated.

[0171] Using the upper limit of computing power in the corresponding computing power range as the acceptable computing power threshold constraint, the strategy that meets the computing power constraint and has the highest average handling efficiency is selected from all security response strategy groups and used as the matching strategy corresponding to the risk-computing power combination of that group.

[0172] By summarizing the matching strategies of all risk-computing power combinations, a two-dimensional response decision matrix of risk level and computing power range is constructed.

[0173] Security response strategies include, but are not limited to: behavior auditing, process blocking, IP blocking, access control, and environment isolation.

[0174] For the initial threat assessment results that are deemed valid, the corresponding threat risk level is determined, the endpoint comprehensive computing power index of the current collection period is read, and the target security response strategy is obtained by substituting it into the two-dimensional response decision matrix; the peak computing power consumption corresponding to the target security response strategy is queried from the corresponding mapping relationship.

[0175] When the endpoint's overall computing power index is greater than or equal to the peak computing power consumption of the target security response strategy, the endpoint executes the corresponding security response strategy locally.

[0176] When the endpoint's overall computing power index is less than the peak computing power consumption of the target security response strategy, the endpoint suspends the local high computing power consumption response execution task and sends a response offload request to the pre-bound edge gateway with an established secure communication channel; the response offload request carries the threat identifier, risk level, and matching disposal instructions.

[0177] The edge gateway is pre-configured with a set of handling execution permissions and actions corresponding to each risk level. After receiving the response unload request, it matches the corresponding handling permissions according to the risk level carried in the request, executes the corresponding security response action, and returns the handling result receipt to the corresponding endpoint after completion.

[0178] The endpoint synchronously records the execution process and final result of the handling, and generates a local security log; combined with the verification of the handling effect, the false alarm and missed alarm confirmation information is obtained and reported to the cloud detection node along with the handling result, and the handling record of the corresponding threat event in the cloud is updated.

[0179] For the anchor point intermediate feature vectors, threat determination results and handling feedback data reported by all endpoints, data cleaning is first performed to remove invalid data with duplicate reports and missing fields; then, the data is labeled and classified based on the threat confidence interval and the handling effect verification results: samples whose threats are confirmed to exist after handling verification are marked as valid threat samples.

[0180] Samples whose judgment results are found to be inconsistent with the actual situation after verification are marked as false alarm or missed alarm samples;

[0181] Samples with confidence levels within a preset fuzzy range and for which final verification has not yet been completed are marked as samples with unknown features.

[0182] Add valid threat samples and unknown threat features to the global endpoint security threat sample library, and update the threat feature dimensions and sample coverage of the sample library.

[0183] Based on the expanded sample library and processing feedback data, with the optimization goal of reducing false positive and false negative rates, incremental fine-tuning training was performed on the security detection master model; during the training process, the network parameters of all feature anchor layers were frozen, and only the weight parameters of ordinary pruning layers were updated.

[0184] After training, the channel contribution is calculated based on the channel activation response of false positive and false negative samples in each network layer. The sensitivity values ​​of the corresponding channels in the full-channel threat sensitivity matrix are then corrected in reverse, and the priority ranking of subsequent channel pruning is optimized.

[0185] According to the preset iteration cycle, combined with the latest comprehensive computing power indicators reported by each endpoint, the channel pruning and low bit width fixed-point quantization processing flow is reused to generate an iteratively optimized lightweight detection sub-model, which is then sent to the corresponding endpoint to complete the local model update.

[0186] It should be noted that this step is mainly responsible for implementing threat response and model iteration optimization. The front end receives the threat results obtained from detection and handles them in a graded manner, while the back end uses the real data generated from the handling to optimize the detection model in reverse.

[0187] When compiling risk and computing power-related handling plans, the peak computing power consumption (including peak CPU and memory usage) of each security response strategy is first measured and recorded. Then, statistical analysis is performed using historical handling samples, prioritizing the strategy with the highest handling efficiency based on the endpoint's computing power limit. This identifies the optimal handling method for different risk levels and computing power ranges, and a handling decision table is compiled. In actual operation, if the endpoint's current computing power meets the strategy's consumption requirements, the handling is executed locally. If the computing power is insufficient, the handling task is transferred to a pre-bound edge gateway, which completes the handling according to the corresponding risk level's operation permissions, minimizing the impact of the handling action on the terminal's business operations.

[0188] After the handling is completed, the endpoint sends the handling results back to the cloud. The cloud first cleans and deduplicates the reported data, removing invalid, duplicate, and abnormal data. Then, based on the threat confidence interval and the handling verification results, the samples are divided into three categories (valid threat samples, false positives and false negatives, and samples with unknown characteristics). The valid threat samples and samples with unknown characteristics are added to the global sample library to expand the feature types and actual application scenarios covered by the samples while ensuring data quality.

[0189] The model iteration and adjustment follow the established approach: the core feature layers remain unchanged, and only the pruning parts are adjusted. During incremental fine-tuning training, the parameters of all feature anchor layers are frozen, and only the weights of ordinary pruning layers are updated to avoid compromising the stability of core threat feature identification. Then, based on the channel activation status of false positives and false negatives, the contribution of each channel is calculated, and the ranking priority of the full-channel threat sensitivity matrix is ​​adjusted. Finally, the existing channel pruning and low-bit-width quantization process is reused (using the original pruning mask framework) to generate an optimized lightweight detection sub-model for distribution. This eliminates the need to rebuild the pruning logic, resulting in higher efficiency in iterative adjustments.

[0190] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for real-time detection and response to endpoint security threats, characterized in that, include: Step 1: Collect endpoint hardware operating status parameters, filter core computing power influencing parameters, obtain the endpoint comprehensive computing power index through weighted fusion, and upload it to the cloud; Construct a security detection master model, divide it into a feature anchor point layer and a normal customizable layer, generate a full-channel threat sensitivity matrix, and establish a linear mapping relationship between computing power indicators and model computational volume. Step 2: Based on the security detection master model, the full-channel threat sensitivity matrix and the linear mapping relationship of computing power, generate a lightweight detection sub-model adapted to the endpoint computing power and distribute it; match the corresponding inference level on the edge and perform threat detection and risk classification; combine confidence and computing power matching reporting strategy to report the intermediate feature vector of the anchor point to the cloud; Step 3: Based on the initial assessment of effective threats and risk levels, construct a two-dimensional response decision matrix of risk and computing power, and combine it with the comprehensive computing power indicators of the endpoints to execute local processing or offloading at the edge gateway; The system updates the sample library based on the feedback from recycling and disposal, fine-tunes the security detection master model and corrects the full-channel threat sensitivity matrix, and iteratively generates a lightweight detection sub-model for distribution and updates.

2. The real-time detection and response method for endpoint security threats according to claim 1, characterized in that: The specific process of collecting endpoint hardware operating status parameters, filtering core computing power impact parameters, obtaining the endpoint comprehensive computing power index through weighted fusion, and uploading it to the cloud is as follows: For each target endpoint, a preset parameter acquisition period is set to collect various hardware operating status parameters that affect the real-time computing capability of the endpoint; Invalid parameters with too small fluctuation ranges were eliminated by variance screening, and core computing power impact parameters that meet the standards of correlation with the endpoint benchmark computing performance were selected by Pearson correlation coefficient method. The core computing power impact parameters are linearly normalized according to their positive and negative attributes, and then the endpoint comprehensive computing power index is obtained by linear weighted fusion calculation. The metric is bound to a unique endpoint identifier and transmitted in encrypted form to the cloud-based detection node.

3. The real-time detection and response method for endpoint security threats according to claim 2, characterized in that: The specific process of constructing a security detection master model, dividing it into a feature anchor layer and a general customizable layer, and generating a full-channel threat sensitivity matrix is ​​as follows: A global sample library consisting of a global endpoint security threat sample library and a normal behavior sample set is pre-constructed. After data augmentation, the standardized training dataset and the standard threat verification dataset are divided. A network architecture combining temporal feature extraction and fully connected classification is built, and a security detection master model is trained. A fixed-ratio random channel pruning is performed sequentially on each hidden layer. The decrease in detection accuracy is used to characterize the degree of feature representation destruction of the corresponding layer. Based on this, the feature anchor layer and the ordinary pruning layer are divided, and the pruning ratio protection constraint of the feature anchor layer is set. A historical IoT threat sample set independent of the training set is selected as input, and the mean loss gradient of each network channel is calculated as the threat sensitivity. The entire channel threat sensitivity matrix is ​​generated by traversing all pruning layer channels.

4. The real-time detection and response method for endpoint security threats according to claim 3, characterized in that: The specific process of establishing a linear mapping relationship between computing power indicators and model computational volume is as follows: Establish a positive proportional mapping relationship between the comprehensive computing power index corresponding to the preset benchmark endpoint and the floating-point operation volume of the benchmark detection model; By combining the preset hardware architecture calibration coefficients for correction, the comprehensive computing power index of the current endpoint is substituted into the mapping relationship to calculate the maximum computing power threshold that the current endpoint can bear, which serves as the computing power constraint boundary for model pruning.

5. The real-time detection and response method for endpoint security threats according to claim 4, characterized in that: The specific process of generating and distributing a lightweight detection sub-model adapted to endpoint computing power, based on the security detection master model, the full-channel threat sensitivity matrix, and the linear mapping relationship between computing power, is as follows: With the goal of maximizing the total threat sensitivity of the retained network channels, and with the constraints that the total computational cost of the model does not exceed the maximum computational cost threshold and the pruning ratio of the feature anchor layer channels does not exceed the anchor layer protection ratio threshold, a greedy selection algorithm is used to solve the optimal pruning scheme. The network channels of each layer are sorted from high to low threat sensitivity, and the high-sensitivity channels are retained in sequence while the total computational cost is accumulated. When the accumulated computational cost approaches the threshold, the pruning ratio of the feature anchor layer is checked. If it exceeds the protection constraint, the low-sensitivity channels of the ordinary pruning layer are adjusted and pruned until all constraints are met. Based on the final set of retained channels, a pruning mask containing channel retention and pruning identifiers is generated. Based on the pruning mask, channel pruning is performed on the security detection parent model to obtain a basic sub-model. After performing low-bit-width fixed-point quantization on the basic sub-model, a lightweight detection sub-model adapted to the current endpoint's computing power is generated. The lightweight detection sub-model and the corresponding pruning mask file are synchronously distributed to the corresponding endpoint.

6. The real-time detection and response method for endpoint security threats according to claim 5, characterized in that: The specific process of matching the corresponding inference level on the endpoint and performing threat detection and risk classification is as follows: For the target endpoint where the model deployment is completed, load the lightweight detection sub-model delivered from the cloud, collect the real-time behavior data stream of the endpoint, and generate standardized behavior sequence data by segmenting and aggregating it through a sliding time window. Based on the full-channel threat sensitivity matrix corresponding to the security detection mother model, the channel ranking relationship is obtained. The retained channels in the ordinary pruning layer of the sub-model are extracted and the sensitivity is arranged from high to low using the same ranking. After testing the average inference computing power consumption of a single channel, multiple ordinary pruning layer channel groups with equal computing power are divided according to the preset single-group benchmark inference computing power consumption. The remaining channels that are less than one group are grouped independently. With all feature anchor layers fully enabled as the base state, high-sensitivity channel groups are superimposed in sequence and multiple sets of computing power accuracy correlation data are obtained through testing. Based on this, multiple levels of inference tiers are divided and the benchmark computing power step size of a single tier is determined. Read the current endpoint's comprehensive computing power index to match the corresponding inference level, enable the corresponding number of high-sensitivity channel groups and keep the feature anchor layer fully enabled, input the standardized behavior sequence to perform inference, and output the threat classification result, confidence level and risk level; The initial assessment of effective threats was obtained by filtering based on the effective threshold.

7. The real-time detection and response method for endpoint security threats according to claim 6, characterized in that: The specific process of reporting the intermediate feature vector of the anchor point to the cloud is as follows: During the forward propagation calculation of the lightweight detection sub-model, the feature tensors at the output of each feature anchor layer are extracted as the intermediate feature vectors of the anchor points; confidence intervals are divided based on historical false alarm statistics, and computing power intervals are divided based on the edge computing power reservation standard. The two types of intervals intersect to form multiple sets of confidence-computing power combinations, and a reporting strategy configuration library is preset for each set of combinations corresponding to the reporting granularity and reporting cycle parameters. Read the threat confidence level and endpoint comprehensive computing power index of the current window, match them to obtain the corresponding reporting granularity and reporting cycle; process the anchor intermediate feature vector according to the corresponding granularity, package it together with the initial threat judgment result and the endpoint unique identifier, upload it to the cloud detection node according to the corresponding cycle and store it in the global endpoint security threat sample library.

8. The real-time detection and response method for endpoint security threats according to claim 7, characterized in that: Based on the initial assessment of effective threats and risk levels, a two-dimensional risk-computing power response decision matrix is ​​constructed. The specific process of executing local processing or edge gateway offloading, combined with the endpoint comprehensive computing power index, is as follows: Pre-test the peak computing power consumption of each security response strategy on the endpoint device, and establish a corresponding mapping relationship between security response strategies and peak computing power consumption; Historical threat handling samples were collected and divided into risk level and computing power range combinations. Matching strategies for each combination were selected with the upper limit of the corresponding computing power range as a constraint and the highest handling efficiency as the objective. A two-dimensional response decision matrix of risk level and computing power range was constructed. Based on the initial assessment of effective threats, determine the corresponding risk level, read the current endpoint's comprehensive computing power index, substitute it into the two-dimensional response decision matrix to obtain the target security response strategy, and query the corresponding peak computing power consumption. If the endpoint's overall computing power index meets the requirements, the corresponding security response policy is executed locally; otherwise, a response unload request carrying a threat identifier, risk level, and handling instructions is sent to the pre-bound edge gateway. The edge gateway matches the corresponding handling execution authority based on the risk level carried in the request, executes the corresponding security response action, and then returns a handling result receipt to the endpoint.

9. The real-time detection and response method for endpoint security threats according to claim 8, characterized in that: The process of updating the sample library based on recycling and disposal feedback is as follows: The endpoint records the handling results and generates local security logs, verifies the handling effect to obtain false alarm and missed alarm confirmation information, and reports the handling results to the cloud detection node and updates the handling records of the corresponding threat events; The cloud performs cleaning on the reported anchor point intermediate feature vectors, threat judgment results and handling feedback data, removes invalid data and labels three types of samples based on confidence intervals and verification results; Add valid threat samples and unknown threat features to the global endpoint security threat sample library, and update the feature dimensions and coverage of the sample library.

10. A real-time detection and response method for endpoint security threats according to claim 9, characterized in that: The specific process of fine-tuning the security detection master model, correcting the full-channel threat sensitivity matrix, and iteratively generating a lightweight detection sub-model for updates is as follows: Based on the expanded global endpoint security threat sample library and handling feedback data, incremental fine-tuning training was performed on the security detection master model with the goal of reducing the false alarm and false negative rates. During training, all feature anchor layer parameters were frozen, and only the weight parameters of the ordinary pruning layer were updated. The channel contribution is calculated based on the channel activation response of false alarm and false alarm samples, and the sensitivity value of the corresponding channel in the full channel threat sensitivity matrix is ​​corrected in reverse to optimize the priority ranking of channel pruning. Based on the preset iteration cycle and the latest comprehensive computing power indicators reported by each endpoint, the optimized lightweight detection sub-model is generated by reusing the channel pruning and low bit-width fixed-point quantization processing flow, and then sent to the corresponding endpoint to complete the local model update.