Model evaluation method and apparatus
By dividing signal data into hierarchical units and standardizing the processing, combined with security and trustworthiness analysis models and threat intelligence data assessment, the problem of data overload in network security monitoring has been solved, the accuracy and efficiency of model assessment have been improved, and the ability to identify and warn of potential threats has been enhanced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-04
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies in network security and monitoring suffer from network data overload due to the large and complex sources of information, resulting in low efficiency in model training and analysis, and increased resource consumption when processing high-volume data models.
By acquiring signal data and performing hierarchical unit division and standardization, the credibility of the signal data is evaluated using a security credibility analysis model. The model is evaluated in conjunction with threat intelligence data. The hierarchical unit approach alleviates the problem of large data volume and difficulty in processing, and improves the accuracy and efficiency of credibility analysis results.
It effectively alleviates the problems of large data volume and difficulty in processing, improves the accuracy and efficiency of model evaluation, and enhances the ability to identify and warn of potential threats.
Smart Images

Figure CN118631576B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a model evaluation method and apparatus. Background Technology
[0002] In recent years, the development of Internet applications has been accelerating, leading to a growing demand for cybersecurity, and cybersecurity and monitoring technologies have been evolving accordingly.
[0003] Current methods for information collection and analysis in cybersecurity and monitoring include machine learning and model analysis. However, the sheer volume and complexity of data sources in reality can easily overload network data, leading to inefficient model training and analysis. On the other hand, using models that handle large amounts of data can increase resource consumption. Summary of the Invention
[0004] In view of the above problems, embodiments of this application provide a model evaluation method and apparatus that overcomes or at least partially solves the above problems.
[0005] In a first aspect, embodiments of this application provide a model evaluation method, the method comprising:
[0006] Acquire the first signal data and the second signal data in each level unit of the hierarchical space;
[0007] Based on the first signal data, the second signal data in each level unit is standardized to obtain the signal standardization result corresponding to each level unit.
[0008] The standardized signal results corresponding to each level unit are input into the security and reliability analysis model to obtain the reliability analysis results corresponding to each level unit.
[0009] The credibility analysis results for each level of unit are evaluated based on threat intelligence data to obtain the model evaluation results.
[0010] Optionally, acquiring the first signal data and the second signal data in each hierarchical unit of the hierarchical space includes:
[0011] Acquire the third signal data, and remove noise and filter from the third signal data to obtain the second signal data;
[0012] The second signal data is stored in column storage in each hierarchical unit of the hierarchical space;
[0013] The first signal data received after acquiring the second signal data from all hierarchical units in the hierarchical space.
[0014] Optionally, the step of standardizing the second signal data in each level unit based on the first signal data to obtain the signal standardization result corresponding to each level unit includes:
[0015] Calculate the similarity metric between the first signal data and the second signal data in each level unit;
[0016] If the similarity metric between the first signal data and at least one second signal data is greater than a preset metric, the second signal data with the largest similarity metric to the first signal data is replaced with the first signal data.
[0017] If the similarity metric values of the first signal data and all the second signal data are less than or equal to the preset metric value, the second signal data in each level unit remains unchanged.
[0018] The second signal data in each level unit is standardized to obtain the signal standardization result corresponding to each level unit.
[0019] Optionally, the step of standardizing the second signal data in each level unit to obtain the signal standardization result corresponding to each level unit includes:
[0020] The characteristic value of the second signal data in each level unit is calculated using the following formula:
[0021]
[0022] The standard deviation of the second signal data in each level unit is calculated using the following formula:
[0023]
[0024] in, This represents the mean value of each level of unit;
[0025] x1, x2…x m This represents m second signal data in each level unit, where i represents any value from 1 to m;
[0026] m represents the number of second signal data in each level unit, and m is a positive integer;
[0027] ME represents the eigenvalue of each level unit;
[0028] SD represents the standard deviation of each level of cells;
[0029] Q represents the coefficient of the current level unit.
[0030] Optionally, the signal standardization result corresponding to each level unit is input into the security reliability analysis model to obtain the reliability analysis result corresponding to each level unit, which is specifically calculated using the following formula:
[0031]
[0032] Where n represents the number of hierarchical units, and n is a positive integer;
[0033] k j-1 This represents the weight corresponding to the j-th level unit;
[0034] Q j This represents the signal normalization result of the j-th level unit;
[0035] i represents the i-th model strategy;
[0036] v represents the dimension of the security and trustworthiness analysis model, where v is an integer from 1 to 3;
[0037] P represents user behavior characteristics;
[0038] P i This represents the user behavior features obtained using the i-th model strategy;
[0039] score() represents a score for user behavior;
[0040] t represents the duration of time;
[0041] S (i,v) This represents the credibility analysis result calculated using the i-th model strategy in the v-th dimension.
[0042] Optionally, the step of evaluating the credibility analysis results corresponding to each level unit based on threat intelligence data to obtain model evaluation results includes:
[0043] Obtain real-time threat intelligence data;
[0044] The credibility analysis results corresponding to each level unit are evaluated based on the threat intelligence data to obtain the model evaluation results.
[0045] Optionally, in the process of evaluating the credibility analysis results corresponding to each level unit based on the threat intelligence data, the method further includes:
[0046] In cases where the model evaluation results cannot be assessed, the second signal data is identified as threat data, and a storage evaluation score is applied to the second signal data to obtain the storage evaluation results.
[0047] Secondly, embodiments of this application also provide a model evaluation apparatus, the apparatus comprising:
[0048] The acquisition module is used to acquire the first signal data and the second signal data in each level unit of the hierarchical space;
[0049] The first processing module is used to standardize the second signal data in each level unit according to the first signal data to obtain the signal standardization result corresponding to each level unit.
[0050] The second processing module is used to input the signal standardization results corresponding to each level unit into the security and reliability analysis model to obtain the reliability analysis results corresponding to each level unit.
[0051] The evaluation module is used to evaluate the credibility analysis results corresponding to each level unit based on threat intelligence data, and obtain the model evaluation results.
[0052] Thirdly, embodiments of this application also provide an electronic device, including a memory, a transceiver, and a processor:
[0053] A memory for storing computer programs; a transceiver for sending and receiving data under the control of a processor; and a processor for reading the computer programs from the memory and executing the method described in the first aspect above.
[0054] Fourthly, embodiments of this application also provide a processor-readable storage medium storing a computer program for causing the processor to perform the method described in the first aspect above.
[0055] In the embodiments described above, by acquiring first signal data and second signal data from each hierarchical unit in the hierarchical space, and standardizing the second signal data in each hierarchical unit based on the first signal data, a signal standardization result corresponding to each hierarchical unit is obtained. This standardization result is then input into a security reliability analysis model to obtain a reliability analysis result for each hierarchical unit. This hierarchical unit-based standardization effectively alleviates the problem of large data volume and processing difficulty in practical applications. Furthermore, obtaining the reliability analysis result for each hierarchical unit through the security reliability analysis model improves the accuracy and efficiency of the reliability analysis results. Moreover, the reliability analysis result for each hierarchical unit is evaluated based on threat intelligence data to obtain a model evaluation result. This comparison with external threat intelligence data verifies the model's practicality, thereby improving the model's ability to identify and warn of potential threats. Attached Figure Description
[0056] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 A flowchart of the model evaluation method provided in the embodiments of this application;
[0058] Figure 2 This is a deployment architecture diagram of the model evaluation system provided in the embodiments of this application;
[0059] Figure 3 Standardized processing flowchart provided for embodiments of this application;
[0060] Figure 4 A comparison chart of the analysis efficiency between the security and reliability analysis model provided in this application embodiment and the traditional model;
[0061] Figure 5 Structural block diagram of the model evaluation device provided in the embodiments of this application
[0062] Figure 6 A structural block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0063] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0064] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0065] The model evaluation method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0066] Specifically, this application provides a model evaluation method applied to a server, such as... Figure 1 As shown, the specific steps may include the following:
[0067] Step 101: Obtain the first signal data and the second signal data in each level unit of the hierarchical space.
[0068] Specifically, hierarchical space refers to the logical or physical space used to divide data into levels. Signal data (first signal data, second signal data) includes, but is not limited to, network traffic data and network security logs. Network traffic data and network security logs are combined into a dataset through a security information management system 21, such as... Figure 2 As shown.
[0069] It should be noted that the Security Information Management System 21 integrates a User and Entity Behavior Analysis (UEBA) system on top of the Extended Detection and Response (XDR) network security platform. The UEBA system can analyze user behavior to meet the needs of the security trustworthiness analysis model. XDR aims to help organizations detect, respond to, and remediate threats when security incidents occur. The XDR platform typically integrates data from different security products and utilizes advanced analytics to automatically detect potential security threats.
[0070] Network traffic data comprises key characteristics such as network connection records, network protocol types, packet size and throughput, packet identifiers, and network session models, and can be used to reflect network activity over a period of time. Network security logs consist of user access information, firewall logs, security event trigger logs, and system logs.
[0071] Step 102: Based on the first signal data, the second signal data in each level unit is standardized to obtain the signal standardization result corresponding to each level unit.
[0072] Specifically, the second signal data in each level unit is updated based on the first signal data, and then the updated second signal data in each level unit is standardized to obtain the signal standardization result corresponding to each level unit.
[0073] Step 103: Input the signal standardization results corresponding to each level unit into the security and reliability analysis model to obtain the reliability analysis results corresponding to each level unit.
[0074] Specifically, such as Figure 2As shown, a security trustworthiness analysis model 26 is established. The standardized signal results corresponding to each level unit are input into the security trustworthiness analysis model 26. After processing by the security trustworthiness analysis model 26, the trustworthiness analysis results corresponding to each level unit are obtained. These trustworthiness analysis results are not immediately used for iterative updates of the security trustworthiness analysis model 26, but are stored (external storage or cloud storage). Then, these stored trustworthiness analysis results are periodically retrieved and used for further training of the security trustworthiness analysis model. This approach is to accumulate data over a period of time so that the performance of the security trustworthiness analysis model can be more comprehensively evaluated, and more effective iterative updates can be performed based on this. In short, the security trustworthiness analysis model does not iterate immediately after obtaining the trustworthiness analysis results, but waits for a period of time to accumulate more data before iterating. This improves the stability and effectiveness of the security trustworthiness analysis model training, maintaining the training effect of the model.
[0075] Step 104: Evaluate the credibility analysis results corresponding to each level unit based on the threat intelligence data to obtain the model evaluation results.
[0076] like Figure 2 As shown, threat intelligence data 27 is obtained, and the security trust analysis model is evaluated by comparing the threat intelligence data 27 with the trust analysis results corresponding to each layer of several units, thereby obtaining the model evaluation results.
[0077] In the embodiments described above, by acquiring first signal data and second signal data from each hierarchical unit in the hierarchical space, and standardizing the second signal data in each hierarchical unit based on the first signal data, a signal standardization result corresponding to each hierarchical unit is obtained. This standardization result is then input into a security reliability analysis model to obtain a reliability analysis result for each hierarchical unit. This hierarchical unit-based standardization effectively alleviates the problem of large data volume and processing difficulty in practical applications. Furthermore, obtaining the reliability analysis result for each hierarchical unit through the security reliability analysis model improves the accuracy and efficiency of the reliability analysis results. Moreover, the reliability analysis result for each hierarchical unit is evaluated based on threat intelligence data to obtain a model evaluation result. This comparison with external threat intelligence data verifies the model's practicality, thereby improving the model's ability to identify and warn of potential threats.
[0078] As an optional specific embodiment, step 101, acquiring the first signal data and the second signal data in each hierarchical unit in the hierarchical space, includes:
[0079] Acquire the third signal data, and remove noise and filter from the third signal data to obtain the second signal data;
[0080] The second signal data is stored in column storage in each hierarchical unit of the hierarchical space;
[0081] The first signal data received after acquiring the second signal data from all hierarchical units in the hierarchical space.
[0082] Specifically, third-signal data (including but not limited to network traffic data and network security logs) is acquired and compiled into a dataset using a security information management system. Noise and filtering are then applied to the third-signal data within this dataset to obtain second-signal data. The second-signal data is then subjected to hierarchical partitioning, that is, its spatial hierarchy is determined based on its characteristics and processing order. This hierarchical partitioning is then stored in cloud storage 22 or database 23 using column-oriented storage. Figure 2 As shown, the second signal data, which is processed first, is placed in the highest level of the hierarchical space. The second signal data is stored cyclically from the highest level to the lowest level until each level is full. After each level is full, the received signal data is the first signal data.
[0083] It should be noted that the order in which second signal data is processed depends on the arrival time, importance, or other business rules of the second signal data, and no specific restrictions are made here.
[0084] As an optional specific embodiment, step 102, based on the first signal data, standardizes the second signal data in each level unit to obtain the signal standardization result corresponding to each level unit, including:
[0085] Calculate the similarity metric between the first signal data and the second signal data in each level unit;
[0086] If the similarity metric between the first signal data and at least one second signal data is greater than a preset metric, the second signal data with the largest similarity metric to the first signal data is replaced with the first signal data.
[0087] If the similarity metric values of the first signal data and all the second signal data are less than or equal to the preset metric value, the second signal data in each level unit remains unchanged.
[0088] The second signal data in each level unit is standardized to obtain the signal standardization result corresponding to each level unit.
[0089] Specifically, after each hierarchical unit is full, the received first signal data is compared with the second signal data stored in each hierarchical unit to calculate a similarity metric (or difference metric). If the similarity metric between the first signal data and at least one second signal data is greater than a preset metric, it indicates that the boundary of the first signal data contains an outlier. The second signal data with the highest similarity metric is then identified and updated as the first signal data. Conversely, if the similarity metric between the first signal data and all second signal data is less than or equal to the preset metric, the second signal data in each hierarchical unit remains unchanged, and the first signal data is discarded. This method can be used in scenarios such as anomaly detection, data compression, or pattern recognition.
[0090] After obtaining the latest second signal data for each level unit, the second signal data in each level unit is standardized. The signal standardization result for each level unit is then displayed on the remote signal monitoring terminal 24. Figure 2 As shown.
[0091] It should be noted that the remote signal monitoring terminal 24 is located at the signal standardization and model training stage, and is used for close-range data processing and rapid display of results.
[0092] like Figure 3 As shown below, the standardization process described above will be explained through a specific workflow:
[0093] Step 301: Acquire the first signal data and the second signal data in each level unit of the hierarchical space.
[0094] Step 302: Obtain the third signal data, remove noise and filter out the third signal data to obtain the second signal data.
[0095] Step 303: Perform hierarchical division processing on the second signal data.
[0096] Step 304: Store the data in columnar storage to obtain a distribution diagram.
[0097] Step 305: If the similarity metric between the first signal data and at least one second signal data is greater than a preset metric, the second signal data with the largest similarity metric is updated to the first signal data.
[0098] Step 306: If the similarity metric values of the first signal data and all the second signal data are less than or equal to the preset metric value, then the first signal data is discarded.
[0099] Step 307: Standardize the second signal data in each level unit to obtain the signal standardization result.
[0100] Step 308: Display the signal standardization results on the remote signal monitoring terminal.
[0101] As an optional specific embodiment, step 102 standardizes the second signal data in each level unit to obtain the signal standardization result corresponding to each level unit, including:
[0102] The characteristic value of the second signal data in each level unit is calculated using the following formula:
[0103]
[0104] The standard deviation of the second signal data in each level unit is calculated using the following formula:
[0105]
[0106] in, This represents the mean value of each level of unit;
[0107] x1, x2…x m This represents m second signal data in each level unit, where i represents any value from 1 to m;
[0108] m represents the number of second signal data in each level unit, and m is a positive integer;
[0109] ME represents the eigenvalue of each level unit;
[0110] SD represents the standard deviation of each level of cells;
[0111] Q represents the coefficient of the current level unit.
[0112] As an optional specific embodiment, step 103 inputs the signal standardization result corresponding to each level unit into the security reliability analysis model to obtain the reliability analysis result corresponding to each level unit, which is specifically calculated using the following formula:
[0113]
[0114] Where n represents the number of hierarchical units, and n is a positive integer;
[0115] k j-1 This represents the weight corresponding to the j-th level unit;
[0116] Q j This represents the signal normalization result of the j-th level unit;
[0117] i represents the i-th model strategy;
[0118] v represents the dimension of the security and trustworthiness analysis model, where v is an integer from 1 to 3;
[0119] P represents user behavior characteristics;
[0120] P i This represents the user behavior features obtained using the i-th model strategy;
[0121] score() represents a score for user behavior;
[0122] t represents the length of time, expressed in units such as seconds, minutes, and hours; t can also represent a point in time, i.e., a specific timestamp.
[0123] S (i,v) This represents the credibility analysis result calculated using the i-th model strategy in the v-th dimension.
[0124] It should be noted that 'i' represents the index or number of the model strategy. The selection of a model strategy is typically based on a specific scenario, requirement, or objective. In practical applications, multiple model strategies are available, each targeting different security monitoring objectives or network environment characteristics. The choice of which model strategy to use depends on the following factors:
[0125] 1. Network Environment: Different network environments may require different security monitoring strategies. For example, the security requirements of an enterprise intranet are different from those of a public network.
[0126] 2. Security objectives: Depending on the types of data to be protected, the criticality of the system, and the types of potential threats, a more refined or broader model strategy needs to be selected.
[0127] 3. Historical and Real-Time Data: Past cybersecurity incidents and current network traffic patterns influence the choice of model strategies. For example, if an increase in a certain type of attack is detected recently, strategies need to be adjusted to better address this type of threat.
[0128] 4. Performance considerations: Some model strategies are more efficient in terms of computing resources and are suitable for resource-constrained environments.
[0129] 5. Compliance and policy requirements: Specific industries or regions have specific safety standards and policies. These can also influence the choice of model strategy.
[0130] Therefore, the selection of a model strategy is a decision-making process that takes into account multiple factors, involving input from security experts, network administrators, and relevant policymakers.
[0131] Additionally, it should be noted that user behavior scores are typically derived through a series of evaluation methods designed to quantify the safety and risk of user behavior. The evaluation methods are explained below:
[0132] 1. Rule-based scoring: User behavior is scored according to predefined security rules. For example, certain network activities or access patterns are considered high-risk and therefore assigned lower scores.
[0133] 2. Machine learning models: Machine learning models are trained using historical data to predict the security of user behavior. These models can score users based on multiple characteristics of their behavior, such as access frequency, time, and location.
[0134] 3. Anomaly detection algorithm: Use statistical or machine learning algorithms to detect user behaviors that are significantly different from normal behavior patterns and score them according to the degree of anomaly.
[0135] 4. Expert system or knowledge base: The security of user behavior is assessed based on the knowledge and experience of security experts, which can be achieved through rule engines or expert systems.
[0136] 5. User Behavior Analysis: In-depth analysis of user behavior patterns, including browsing habits and transaction history, to identify potential risky behaviors and assign corresponding scores.
[0137] In practical applications, multiple assessment methods can be used in combination to obtain more comprehensive and accurate scores of user behavior. These scores can then be used to trigger alerts, conduct further investigations, or take other security measures.
[0138] It should be noted that the security trustworthiness analysis model analyzes user behavior characteristics using the processed signal standardization results. This analysis process is an important intermediate step before the security trustworthiness analysis model obtains the final trustworthiness analysis result. The security trustworthiness analysis model combines these intermediate results to comprehensively evaluate the trustworthiness of user behavior and gives a score to the user behavior accordingly. The user score is used to evaluate whether the current user behavior poses a threat to the system; the user score behavior defined in the security trustworthiness analysis model is referenced in Table 1:
[0139] Table 1
[0140]
[0141] Specifically, when a user's behavior receives a low score or is deemed threatening, the system will not proceed further but will instead take appropriate security measures, such as issuing alerts or restricting user permissions, to prevent potential security risks. In other words, when a user's score falls below a preset threshold, the system will pause or restrict further user actions.
[0142] As an optional specific embodiment, step 104 evaluates the credibility analysis results corresponding to each level unit based on threat intelligence data to obtain model evaluation results, including:
[0143] Obtain real-time threat intelligence data;
[0144] The credibility analysis results corresponding to each level unit are evaluated based on the threat intelligence data to obtain the model evaluation results.
[0145] Specifically, establish partnerships with threat intelligence providers, intelligence-sharing organizations, or internal intelligence teams to obtain real-time threat intelligence data. Simultaneously, correlate the threat intelligence data with the credibility analysis results trained by the security credibility analysis model. Detect potential threat activities by comparing and identifying abnormal or suspicious data and user behavior patterns. If the two information match, it indicates that the model's identification is accurate; if they do not match, further investigation and analysis of the reasons are required.
[0146] As an optional specific embodiment, in the process of evaluating the credibility analysis results corresponding to each level unit based on the threat intelligence data in step 104, the method further includes:
[0147] In cases where the model evaluation results cannot be assessed, the second signal data is identified as threat data, and a storage evaluation score is applied to the second signal data to obtain the storage evaluation results.
[0148] Specifically, during the analysis of threat intelligence data, if unpredictable situations arise or the credibility analysis results cannot be assessed, the second signal data will be considered a potential or uncertain threat. This second signal data needs to be stored, its risk score assessed, and displayed in the remote analysis center 25, such as... Figure 2 As shown in Table 2, this information will be displayed in the remote analysis center 25 for further analysis and processing by professionals. Security analysts or relevant team members can also trace the source of this data to better understand the background and motivations of potential threats. The assessment scores for potential or uncertain threats are shown in Table 2.
[0149] Table 2
[0150]
[0151]
[0152] As shown in Table 2, the score is determined based on the nature and impact of the threat. Specifically, for a potential threat whose harm is deemed "unpredictable / unassessable," the threat level is high, with a score ranging from 0 to 0.5. This covers different levels of potential harm, and therefore the corresponding score is a range value, which needs to be further refined based on specific circumstances or needs. For identified threats, the score is determined based on the target distribution range. If the target distribution range is 0%, the score is 0; if the target distribution range is 1-15%, the score is 0.25; if the target distribution range is 16-49%, the score is 0.75; and if the target distribution range reaches 50-100%, the score is 1.0. This reflects the direct relationship between the breadth of the threat's impact and the score.
[0153] It's important to note that in threat assessment, target distribution refers to the scope of systems, networks, data, or assets affected by the threat. In cybersecurity, it can represent the breadth of targets that have been attacked or may be attacked. For example, if malware or attacks target only a specific system or service, the target distribution is relatively small; however, if the threat can affect the entire network or multiple systems, the target distribution is relatively large.
[0154] In Table 2, "Target Distribution Range" is used as a factor in determining threat level and score. Here, "Target Distribution Range" refers to the percentage or proportion affected by the threat, as shown in Table 2, and is divided into several intervals: 0%, 1–15%, 16–49%, and 50–100%. These intervals reflect the potential scope of the threat's impact, thus affecting the threat rating and score. "Target Distribution Range" is a factor considered when assessing threat severity, helping to determine the breadth and potential scope of the threat's impact.
[0155] Furthermore, after step 103, the signal standardization result is used as the input to the security and reliability analysis model, i.e., the initial state of model training. A semi-supervised learning process, obtaining the signal standardization result through the security and reliability analysis model channel, iterates the security and reliability analysis model using labeled or unlabeled data. The iterated security and reliability analysis model is then applied to the real-time analysis of the signal data stream to generate corresponding security and reliability analysis results. Labeled data refers to data that already has a label or result, while unlabeled data is data without a clear label or result. The core idea of semi-supervised learning is to use a small amount of labeled data to guide the model's learning and a large amount of unlabeled data to improve the model's generalization ability. Therefore, "labeled and unlabeled" refers to two data types used in semi-supervised learning. Labeled data is:
[0156]
[0157] Where, if f θ Belongs to D x Then calculate f according to the above formula. θ The tag value, and the f θ According to the label value, it belongs to D. x The labeled data; if f θ It does not belong to D x Then calculate f according to the above formula. θ The tag value, and the f θ Marked as not belonging to D according to the label value. x tagged data;
[0158] This indicates the signal disturbance value; it is an indicator of signal instability or noise.
[0159] f θ This represents the metric value of the second signal data; it is a function used to quantify or evaluate a certain characteristic of the data, where θ represents the parameters of the function.
[0160] C indicates the category of the second signal data, either labeled with a data number or unlabeled with a data number;
[0161] x k This represents the signal standardization result of each level unit after semi-supervised learning processing;
[0162] D x This represents the set of signal standardization results after pre-defined semi-supervised learning processing;
[0163] x represents the data value of the hierarchical unit.
[0164] The above formula describes a complex ratio or relationship related to signal perturbation, data measurement, and standardized values after hierarchical unit processing. This ratio or relationship is used to evaluate the characteristics of the data, the accuracy of classification, or the quality of the signal. The structural analysis of the entire formula is as follows:
[0165] The numerator represents a certain ratio or relationship between the calculated signal perturbation value and the data metric value, and this ratio or relationship is calculated for a specific hierarchical unit.
[0166] The denominator represents the summation of a series of values, each being the product of the logarithm of the data measure and the data value of a certain hierarchical unit, divided by the data category C. This summation is performed across all hierarchical units or all data points.
[0167] It should be noted that f θ It is derived through the analysis and calculation of data using specific algorithms or models. θIt can represent a certain characteristic, importance, degree of anomaly, or other quantitative indicators related to the second signal data.
[0168] f θ The calculation method depends on the specific application scenario and the nature of the required metric. Below are some possible algorithms or models through which f can be calculated. θ :
[0169] 1. Statistical Model: Analyze the second signal data using statistical methods, such as calculating the mean, variance, covariance, and other statistical quantities of the second signal data, or estimating the characteristics of the second signal data through probability distribution models.
[0170] 2. Machine Learning Model: This involves training a machine learning model (such as linear regression, logistic regression, support vector machine, neural network, etc.) to extract features from the second signal data, and then calculating f based on these features. θ The model's output or internal representation can be used as a data metric.
[0171] 3. Deep Learning Models: These models use deep neural networks to learn complex representations of data, such as Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). Deep neural networks can automatically extract high-level features and compute f based on these features. θ .
[0172] 4. Unsupervised learning algorithms: such as clustering algorithms or dimensionality reduction algorithms, which can reveal the inherent structure and relationships of data without labels, thereby deriving data metrics.
[0173] 5. Anomaly detection algorithms: such as Isolation Forest and a class of Support Vector Machines, these algorithms can detect outliers or isolated points in the data. θ It can indicate the degree of anomaly of a data point.
[0174] 6. Graph Algorithms: If the data can be represented as a graph structure, then graph algorithms can be used to calculate the importance or centrality of nodes, and these values can also be used as f. θ .
[0175] 7. Custom Algorithms: For specific problems, custom algorithms can be designed to calculate f. θ For example, logical judgments based on business rules, and the application of knowledge in specific domains.
[0176] Choosing a suitable algorithm or model to calculate f θWhen doing so, it is necessary to consider the nature of the data, the background of the problem, and the required measurement objectives. Different algorithms and models may be suitable for different scenarios, so it is necessary to select and adjust them according to the actual situation.
[0177] The following simulation experiment compares the security and reliability analysis model of this application with the traditional model:
[0178] In this experiment, two models were created: the security and trustworthiness analysis model of this application and the traditional method model. The main comparison was between the processing speed and computational power of the models under different complexities.
[0179] The dataset size was set to 1800 samples; the initial learning rate was 0.004, and the number of iterations was 64; the optimizer was Momentum; to further evaluate the model's performance, the complexity value was used as an independent variable, where the complexity value ranged from 1 to 6. Figure 4 As shown, when the complexity is 1, the training model efficiency of the method in this application is 0.9, while the efficiency of the traditional method model is 0.88. When the complexity decreases, the method in this application is more efficient than the traditional method model in the range of 1 to 6, indicating that the security and trustworthiness analysis model of this application is more applicable than the traditional method model.
[0180] The initial learning rate is a parameter set at the beginning of machine learning or deep learning training to control the step size of parameter updates during training. The learning rate determines the extent to which the model parameters are adjusted based on the gradient of the loss function in each iteration. A large learning rate may lead to unstable model training, while a small learning rate may result in slow training speed. Here, 0.004 is the initial learning rate value set, which can be adjusted as needed.
[0181] Complexity range values refer to a series of complexity levels used in experiments to test model performance. Complexity refers to the complexity of the dataset, the difficulty of the problem, the number of features the model needs to process, or the complexity of the model itself. A complexity range of 1 to 6 means that during the experiment, the complexity gradually increases from 1 to 6 to observe and compare the performance and efficiency of the two models at different complexities. Within this complexity range, the security and trustworthiness analysis model of this application outperforms the traditional method model.
[0182] In summary, the embodiments described above in this application firstly effectively utilize existing network resources by collecting network traffic data and network security log data to form a dataset for signal extraction. Secondly, establishing a security trustworthiness analysis model and training it using adaptive analysis methods improves the model's accuracy and efficiency. Finally, comparing the model with external threat intelligence data further verifies its practicality and enhances its ability to identify and warn of potential threats. The hierarchical division method for data comparison effectively alleviates the problem of large data volume and difficulty in processing data in practice. Establishing a security trustworthiness analysis model based on the compared data reduces the amount of model training, making the model more efficient. Furthermore, incorporating a semi-supervised learning process into the model training expands the dataset used for training, improving the model's generalization ability and adaptability.
[0183] The above describes the model evaluation method provided by the embodiments of this application. The model evaluation device provided by the embodiments of this application will be described below with reference to the accompanying drawings.
[0184] like Figure 5 As shown in the illustration, this application also provides a model evaluation device 500, the device comprising:
[0185] The acquisition module 501 is used to acquire the first signal data and the second signal data in each level unit in the hierarchical space;
[0186] The first processing module 502 is used to standardize the second signal data in each level unit according to the first signal data to obtain the signal standardization result corresponding to each level unit.
[0187] The second processing module 503 is used to input the signal standardization result corresponding to each level unit into the security and reliability analysis model to obtain the reliability analysis result corresponding to each level unit.
[0188] Evaluation module 504 is used to evaluate the credibility analysis results corresponding to each level unit based on threat intelligence data, and obtain the model evaluation results.
[0189] In the embodiments described above, by acquiring first signal data and second signal data from each hierarchical unit in the hierarchical space, and standardizing the second signal data in each hierarchical unit based on the first signal data, a signal standardization result corresponding to each hierarchical unit is obtained. This standardization result is then input into a security reliability analysis model to obtain a reliability analysis result for each hierarchical unit. This hierarchical unit-based standardization effectively alleviates the problem of large data volume and processing difficulty in practical applications. Furthermore, obtaining the reliability analysis result for each hierarchical unit through the security reliability analysis model improves the accuracy and efficiency of the reliability analysis results. Moreover, the reliability analysis result for each hierarchical unit is evaluated based on threat intelligence data to obtain a model evaluation result. This comparison with external threat intelligence data verifies the model's practicality, thereby improving the model's ability to identify and warn of potential threats.
[0190] Optionally, the acquisition module 501 is specifically used for:
[0191] Acquire the third signal data, and remove noise and filter from the third signal data to obtain the second signal data;
[0192] The second signal data is stored in column storage in each hierarchical unit of the hierarchical space;
[0193] The first signal data received after acquiring the second signal data from all hierarchical units in the hierarchical space.
[0194] Optionally, the first processing module 502 is specifically used for:
[0195] Calculate the similarity metric between the first signal data and the second signal data in each level unit;
[0196] If the similarity metric between the first signal data and at least one second signal data is greater than a preset metric, the second signal data with the largest similarity metric to the first signal data is replaced with the first signal data.
[0197] If the similarity metric values of the first signal data and all the second signal data are less than or equal to the preset metric value, the second signal data in each level unit remains unchanged.
[0198] The second signal data in each level unit is standardized to obtain the signal standardization result corresponding to each level unit.
[0199] Optionally, the first processing module 502 is specifically used for:
[0200] The characteristic value of the second signal data in each level unit is calculated using the following formula:
[0201]
[0202] The standard deviation of the second signal data in each level unit is calculated using the following formula:
[0203]
[0204] in, This represents the mean value of each level of unit;
[0205] x1, x2…f m This represents m second signal data in each level unit, where i represents any value from 1 to m;
[0206] m represents the number of second signal data in each level unit, and m is a positive integer;
[0207] ME represents the eigenvalue of each level unit;
[0208] SD represents the standard deviation of each level of cells;
[0209] Q represents the coefficient of the current level unit.
[0210] Optionally, the second processing module 503 performs the calculation using the following formula:
[0211]
[0212] Where n represents the number of hierarchical units, and n is a positive integer;
[0213] k j-1 This represents the weight corresponding to the j-th level unit;
[0214] Q j This represents the signal normalization result of the j-th level unit;
[0215] i represents the i-th model strategy;
[0216] v represents the dimension of the security and trustworthiness analysis model, where v is an integer from 1 to 3;
[0217] P represents user behavior characteristics;
[0218] P i This represents the user behavior features obtained using the i-th model strategy;
[0219] score() represents a score for user behavior;
[0220] t represents the duration of time;
[0221] S (i,v) This represents the credibility analysis result calculated using the i-th model strategy in the v-th dimension.
[0222] Optionally, the evaluation module 504 is specifically used for:
[0223] Obtain real-time threat intelligence data;
[0224] The credibility analysis results corresponding to each level unit are evaluated based on the threat intelligence data to obtain the model evaluation results.
[0225] Optionally, in the process of evaluating the credibility analysis results corresponding to each level unit based on the threat intelligence data, the method further includes:
[0226] In cases where the model evaluation results cannot be assessed, the second signal data is identified as threat data, and a storage evaluation score is applied to the second signal data to obtain the storage evaluation results.
[0227] In summary, the embodiments described above in this application firstly effectively utilize existing network resources by collecting network traffic data and network security log data to form a dataset for signal extraction. Secondly, establishing a security trustworthiness analysis model and training it using adaptive analysis methods improves the model's accuracy and efficiency. Finally, comparing the model with external threat intelligence data further verifies its practicality and enhances its ability to identify and warn of potential threats. The hierarchical division method for data comparison effectively alleviates the problem of large data volume and difficulty in processing data in practice. Establishing a security trustworthiness analysis model based on the compared data reduces the amount of model training, making the model more efficient. Furthermore, incorporating a semi-supervised learning process into the model training expands the dataset used for training, improving the model's generalization ability and adaptability.
[0228] It should be noted that the model evaluation apparatus provided in this application embodiment can implement all the method steps implemented in the above model evaluation method embodiment and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0229] It should be noted that the division of units in the embodiments of this application is illustrative and only represents one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units.
[0230] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0231] like Figure 6 As shown, embodiments of this application also provide an electronic device, including a memory 620, a transceiver 610, and a processor 600:
[0232] Memory 620 is used to store computer programs;
[0233] Transceiver 610 is used to send and receive data under the control of the processor;
[0234] The processor 600 is configured to read a computer program from memory and execute the steps of the model evaluation method as described in any of the above embodiments.
[0235] Among them, Figure 6 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 600) and memory (memory 620). The bus architecture can also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 610 can be multiple elements, including transmitters and receivers, providing a unit for communicating with various other devices over transmission media, including wireless channels, wired channels, optical fibers, etc. The processor 600 is responsible for managing the bus architecture and general processing, and the memory 620 can store data used by the processor 600 during operation.
[0236] The processor 600 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD). The processor can also adopt a multi-core architecture.
[0237] The processor executes the model evaluation method provided in the embodiments of this application according to the obtained executable instructions by calling a computer program stored in memory. The processor and memory can also be physically separated.
[0238] It should be noted that the electronic device provided in this application embodiment can implement all the method steps implemented in the above model evaluation method embodiment and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0239] Embodiments of this application also provide a processor-readable storage medium storing a computer program for causing the processor to execute the above-described model evaluation method.
[0240] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).
[0241] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0242] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-executable instructions. These computer-executable instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0243] These processor-executable instructions may also be stored in a processor-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the processor-readable memory produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0244] These processors can execute instructions that can also be loaded onto a computer or other programmable data processing device, causing a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0245] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A model evaluation method, characterized in that, The method includes: Acquiring the first signal data and the second signal data in each level unit of the hierarchical space specifically includes: storing the second signal data in each level unit of the hierarchical space in a columnar storage manner; This continues until each level unit is full of the second signal data; once each level unit is full, the received signal data is the first signal data. The hierarchical space is a logical or physical space that divides data levels; Based on the first signal data, the second signal data in each level unit is standardized to obtain the signal standardization result corresponding to each level unit. The standardized signal results corresponding to each level unit are input into the security and reliability analysis model to obtain the reliability analysis results corresponding to each level unit. The credibility analysis results for each level of unit are evaluated based on threat intelligence data to obtain the model evaluation results. The step of standardizing the second signal data in each level unit based on the first signal data to obtain the signal standardization result corresponding to each level unit includes: Calculate the similarity metric between the first signal data and the second signal data in each level unit; If the similarity metric between the first signal data and at least one second signal data is greater than a preset metric, the second signal data with the largest similarity metric to the first signal data is replaced with the first signal data. If the similarity metric values of the first signal data and all the second signal data are less than or equal to the preset metric value, the second signal data in each level unit remains unchanged. The second signal data in each level unit is standardized to obtain the signal standardization result corresponding to each level unit.
2. The method according to claim 1, characterized in that, The acquisition of the first signal data and the second signal data in each hierarchical unit in the hierarchical space includes: Acquire the third signal data, and remove noise and filter from the third signal data to obtain the second signal data.
3. The method according to claim 1, characterized in that, The step of standardizing the second signal data in each level unit to obtain the signal standardization result corresponding to each level unit includes: The characteristic value of the second signal data in each level unit is calculated using the following formula: The standard deviation of the second signal data in each level unit is calculated using the following formula: in, This represents the mean value of each level of unit; Represents each level of unit A second signal data, Indicates 1 to m Any value in the range; This indicates the number of second signal data in each level unit. It is a positive integer; Represents the eigenvalues of each level unit; This represents the standard deviation of each level of unit; This represents the coefficient of the current level unit.
4. The method according to claim 1, characterized in that, The standardized signal result corresponding to each level unit is input into the security and reliability analysis model to obtain the reliability analysis result corresponding to each level unit, which is specifically calculated using the following formula: in, Indicates the number of hierarchical units. It is a positive integer; Indicates the first The weights corresponding to each hierarchical unit; Indicates the first Signal standardization results of hierarchical units; Indicates the first Individual model strategies; This represents the dimensions of the security and trustworthiness analysis model. Integers from 1 to 3; Indicates user behavior characteristics; Indicates the use of the first User behavior characteristics obtained from a model strategy; This represents a score indicating user behavior; Indicates the length of time; Indicates the use of the first The model strategy in the first v The credibility analysis results obtained from dimensional calculation.
5. The method according to claim 1, characterized in that, The evaluation of the credibility analysis results corresponding to each level unit based on threat intelligence data yields model evaluation results, including: Obtain real-time threat intelligence data; The credibility analysis results corresponding to each level unit are evaluated based on the threat intelligence data to obtain the model evaluation results.
6. The method according to claim 5, characterized in that, In the process of evaluating the credibility analysis results corresponding to each level unit based on the threat intelligence data, the method further includes: In cases where the model evaluation results cannot be assessed, the second signal data is identified as threat data, and a storage evaluation score is applied to the second signal data to obtain the storage evaluation results.
7. A model evaluation device, characterized in that, The device includes: The acquisition module is used to acquire first signal data and second signal data in each level unit of the hierarchical space; store the second signal data in each level unit of the hierarchical space in column storage; until each level unit is full of second signal data; after each level unit is full, the signal data received is the first signal data; wherein, the hierarchical space is a logical or physical space that divides data levels. The first processing module is used to standardize the second signal data in each level unit according to the first signal data to obtain the signal standardization result corresponding to each level unit. The second processing module is used to input the signal standardization results corresponding to each level unit into the security and reliability analysis model to obtain the reliability analysis results corresponding to each level unit. The evaluation module is used to evaluate the credibility analysis results corresponding to each level unit based on threat intelligence data, and obtain the model evaluation results. The first processing module is specifically used for: Calculate the similarity metric between the first signal data and the second signal data in each level unit; If the similarity metric between the first signal data and at least one second signal data is greater than a preset metric, the second signal data with the largest similarity metric to the first signal data is replaced with the first signal data. If the similarity metric values of the first signal data and all the second signal data are less than or equal to the preset metric value, the second signal data in each level unit remains unchanged. The second signal data in each level unit is standardized to obtain the signal standardization result corresponding to each level unit.
8. An electronic device, characterized in that, Includes memory, transceiver, and processor: Memory, used to store computer programs; Transceiver, used to send and receive data under the control of the processor; A processor for reading a computer program from the memory and executing the model evaluation method as described in any one of claims 1 to 6.
9. A processor-readable storage medium, characterized in that, The processor-readable storage medium stores a computer program for causing the processor to perform the model evaluation method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Industrial control network security risk assessment method based on multilayer fuzzy system
CN112327767A
Threat complexity analysis method combining multiple levels and entropy weight method
CN114124526A