A network data asset security discovery system based on distributed machine learning
By combining distributed machine learning and support vector description models with Chebyshev polynomial acquisition technology, the problems of high concurrency and anomaly detection in network data asset discovery are solved, low-latency, efficient anomaly identification and load balancing are achieved, and the security discovery capabilities of network data assets are improved.
Patent Information
- Application Number
- CN202411841877.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-12-13
AI Technical Summary
The existing network data asset discovery process does not take into account high concurrent data flows, resulting in long task queues and high latency. In addition, anomaly detection is based on fixed thresholds, resulting in high false alarm rates and insufficient practicality.
A network data asset security discovery system based on distributed machine learning is adopted. Historical data is obtained through data center nodes to establish a support vector description model. Chebyshev polynomials are combined for real-time data collection and load balancing distribution, and the support vector description model is used for anomaly identification.
It effectively reduces task queue delays, improves data processing efficiency and anomaly detection accuracy, reduces false alarm rates, adapts to large-scale high-concurrency scenarios, and ensures the security discovery efficiency and practicality of network data assets.
Smart Images

Figure CN119728210B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing systems, and in particular relates to a network data asset security discovery system based on distributed machine learning. Background Art
[0002] Network data assets refer to virtual or physical devices and their associated resources in cyberspace, including but not limited to servers, routers, switches, virtual machines, containers, IP addresses, MAC addresses, domain names, network interfaces, protocol ports, and so on. They are the core components of network communication and data transmission, representing key nodes in data flow. Distributed machine learning is a technology that decomposes and distributes machine learning tasks across multiple computing nodes for processing, making it suitable for large-scale data computation and analysis. By distributing data, models, or tasks across multiple nodes, distributed machine learning can improve computing efficiency, processing power, and scalability, while also supporting collaborative learning in a distributed environment.
[0003] Security discovery of network data assets can help promptly identify abnormal traffic, intrusions, or configuration vulnerabilities, preventing data leaks, device loss of control, or system outages. As network scale and complexity increase, security risks also rise. Using advanced technologies to conduct security discovery of network data assets can ensure the normal operation of critical assets, enhance network security situational awareness, and mitigate the impact of security threats on business.
[0004] However, existing network data asset discovery processes fail to account for the high concurrent data flows of individual assets. This results in excessively long task queues and extremely high latency at data processing nodes, which is extremely detrimental to the security of network data assets. Furthermore, existing network data asset discovery processes often rely on fixed thresholds for anomaly detection, resulting in inaccurate anomaly detection, frequent false positives, and limited practicality. Summary of the Invention
[0005] To address the existing technical issues of network data asset discovery, which fail to consider the high concurrent data flows of individual assets, leading to excessively long task queues and high latency at data processing nodes, which are detrimental to the security of network data assets. Furthermore, existing network data asset discovery processes often rely on fixed thresholds for anomaly detection, resulting in insufficient accuracy, frequent false positives, and limited practicality, the present invention provides a network data asset security discovery system based on distributed machine learning.
[0006] The present invention provides a network data asset security discovery system based on distributed machine learning, the network data asset security discovery system comprising a data center node and a plurality of data processing nodes connected to the data center node;
[0007] The network data asset security discovery system based on distributed machine learning also includes:
[0008] An acquisition module, configured to acquire historical network data through a data center node at a first preset time period, wherein the historical network data includes asset interaction times of each network data asset and asset interaction frequencies at the asset interaction times;
[0009] Establish a module for combining the distribution center of historical network data and establishing a support vector description model based on historical network data to determine the minimum hypersphere used to describe the normal operation status of network data assets;
[0010] A first sending module is used to send the support vector description model to each data processing node;
[0011] The first acquisition module is used to perform approximate acquisition of real-time network data of each network data asset in combination with Chebyshev polynomials;
[0012] The second acquisition module is used to collect load data of each data processing node;
[0013] The second sending module is used to send real-time network data to each data processing node based on load data and with the goal of minimizing network data processing delay;
[0014] The identification module is used to safely identify the received real-time network data through the support vector description model in each data processing node, and output abnormal real-time network data, that is, output real-time network data outside the minimum hypersphere.
[0015] Compared with the prior art, the present invention has at least the following beneficial technical effects:
[0016] In the present invention, the data center node is used as the data allocation center, and data processing tasks are issued according to the load status of each data processing node, effectively avoiding the problems of overlong task queues and excessive processing delays, ensuring load balancing, improving data processing efficiency and real-time performance, making the security discovery of network data assets more accurate and efficient, and suitable for large-scale high-concurrency scenarios. In the process of secure identification of network data assets, a support vector description model based on historical network data is established in combination with the distribution center of historical network data to determine the minimum hypersphere used to describe the normal operating status of network data assets. This can not only adapt to dynamically changing network environments, but also more accurately distinguish normal from abnormal data, greatly improving the accuracy of anomaly detection. In the process of real-time data collection, Chebyshev polynomials are combined to estimate the approximate collection of real-time network data of each network data asset. This can effectively address the problems of data acquisition difficulties and excessive delays caused by data missing and high-concurrency data streams. It can quickly acquire network data in high-concurrency and data missing environments, significantly reducing delays and computational complexity, ensuring the real-time and reliability of data collection, and improving the efficiency and practicality of network data asset security discovery. Finally, at each data processing node, the established support vector description model is used to securely and quickly identify the collected data, achieving low latency and high accuracy, capable of handling large-scale data streams that meet real-world requirements. Throughout the entire process, from data collection to data identification, the system effectively reduces asset identification latency and accurately and promptly detects abnormal data states based on historical data patterns, resulting in a low false alarm rate and enhanced practicality. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The preferred embodiments will be described below in a clear and understandable manner with reference to the accompanying drawings to further illustrate the above-mentioned characteristics, technical features, advantages and implementation methods of the present invention.
[0018] Figure 1 This is a schematic diagram of the structure of a network data asset security discovery system based on distributed machine learning provided by the present invention;
[0019] Figure 2 This is a structural diagram of the relationship between a data center node and a data processing node provided by the present invention. DETAILED DESCRIPTION
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the specific embodiments of the present invention will be described below with reference to the accompanying drawings. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings and other embodiments can be obtained based on these drawings without inventive work.
[0021] To simplify the drawings, only portions relevant to the invention are schematically depicted in each figure; they do not represent the actual structure of the product. Furthermore, to simplify the drawings and facilitate understanding, in some figures, only one component with the same structure or function is schematically depicted or labeled. In this document, "one" not only means "only one" but also "more than one."
[0022] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0023] It should be noted that, unless otherwise specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they can refer to fixed connections, removable connections, or integral connections. They can also refer to mechanical connections or electrical connections. They can also refer to direct connections or indirect connections through an intermediary, or to internal communication between two components. Those skilled in the art will understand the specific meanings of these terms in the present invention based on the specific circumstances.
[0024] In addition, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0025] In one embodiment, referring to Figure 1 , shows a schematic diagram of the structure of a network data asset security discovery system based on distributed machine learning provided by the present invention. Figure 2 , showing a structural diagram of resource allocation of a medical community provided by the present invention.
[0026] The present invention provides a network data asset security discovery system based on distributed machine learning. The network data asset security discovery system includes a data center node and multiple data processing nodes connected to the data center node.
[0027] Among them, the network data asset security discovery system collects network data through the data center nodes and realizes intelligent distribution through the collaborative work of data center nodes and multiple data processing nodes (1 to n). By processing network data in a distributed manner, it can discover anomalies in real time and realize efficient and secure data asset management.
[0028] The network data asset security discovery system based on distributed machine learning also includes:
[0029] The acquisition module 1 is configured to acquire historical network data through a data center node with a first preset time period as a period.
[0030] The historical network data includes the asset interaction time of each network data asset and the asset interaction frequency at the asset interaction time.
[0031] Network data assets are various virtual or physical devices within the monitored network, including IP addresses, routers, and so on. The interaction moment is the moment when data transmission occurs for this network data asset, which can be uploading or receiving data. The asset interaction frequency is the total number of data transmissions at that interaction moment, with each data transmission being recorded as an asset interaction. All historical data collected in step S1 is normal operating data.
[0032] It should be noted that those skilled in the art can set the first preset duration according to actual needs, and the present invention does not limit this. Specifically, the preset duration can be one minute, one hour, or one day, etc.
[0033] Optionally, the first preset duration is determined by first performing a Fourier transform on the asset interaction frequency time series in the historical network data to obtain a spectrum. Then, the square of the modulus of the spectrum is used as the contribution intensity of different occurrence times to the original signal, i.e., the spectrum, to obtain a frequency spectrum density that reflects the signal intensity of each frequency, including each asset interaction frequency, wherein the peak frequency f peak As the most significant periodic component, and based on this periodic component, calculate the period The time series is then divided using this period as the base window length W. A simple k-means clustering algorithm is then used to cluster the asset interaction frequency time series, obtaining multiple candidate windows T1, T2, ... that can reflect the periodic pattern. The variance of the asset interaction frequency within each base window and candidate window is then calculated. It is used to measure the volatility or regularity of the interaction frequency within the time window. The smaller the variance of the asset interaction frequency, the more significant the periodicity or stability of the interaction frequency within the window. F k represents the asset interaction frequency at the kth moment within the base window or candidate window W, μ W Represents the mean of the interaction frequency within the basic window or candidate window W. Finally, the T freq And the candidate window extracted from the cluster selects the first preset duration, that is,
[0034]
[0035] It's understandable that dynamically determining the first preset duration through Fourier transform and K-means clustering accurately captures the periodicity of asset interaction frequencies and avoids misjudging abnormal data caused by fixed window lengths. This method, based on periodicity and variance optimization, provides stable and regular input data for the subsequent establishment of the support vector model, thereby improving the model's accuracy in fitting normal data distributions and enhancing the accuracy and robustness of anomaly detection.
[0036] In one possible implementation, network data assets include servers, routers, switches, personal terminal devices, virtual machines, containers, IP addresses, MAC addresses, domain names, network interfaces, and protocol ports.
[0037] Establishing module 2, for combining the distribution center of historical network data, establishing a support vector description model based on historical network data, so as to determine the minimum hypersphere used to describe the normal operation status of network data assets.
[0038] The Support Vector Data Description (SVDD) model is a machine learning model used for anomaly detection. It constructs a minimal hypersphere in high-dimensional space, enclosing most normal data within it. Any data outside the sphere is considered anomalous. A minimal hypersphere is a geometric object in high-dimensional space, defined by its minimum radius to contain as many data points as possible, while allowing a small number of data points to deviate from the sphere and controlling for deviations through slack variables. This method can describe the normal distribution range of data and is used to detect anomalous data that deviates from this range.
[0039] By combining the distribution centers of historical network data and using the Support Vector Description (SVDD) model, a minimal hypersphere is constructed to accurately describe the normal operating range of network data assets, providing a benchmark for subsequent anomaly detection. This step dynamically adapts to data changes, improving the accuracy and practicality of anomaly detection while significantly reducing false positives and false negatives.
[0040] In a possible implementation, the establishment module 2 is specifically configured to:
[0041] Based on the asset interaction time and the asset interaction frequency at the asset interaction time, the historical network data state vector of each network data asset is established:
[0042]
[0043] Among them, t k represents the kth asset interaction moment of the network data asset, Indicates that the i-th network data asset is at t k Frequency of asset interaction, xi,t It represents the interaction state vector of the i-th network data asset at time t, and T1 and T2 represent the starting time and ending time of the preset duration respectively.
[0044] Among them, the historical network data state vector marks the asset interaction frequency of each network data asset at each asset interaction moment within a recording period. This data is used for subsequent data analysis to capture the periodic laws of each network data asset in time and the upper and lower limits of the interaction frequency in state.
[0045] Establish the smallest hypersphere that describes the normal operation status of network data assets:
[0046]
[0047] ||x i -c|| 2 ≤R 2 +ξ i ,ξ i ≥0
[0048] e≤ε
[0049]
[0050] Where R and c represent the radius and center of the hypersphere, respectively, i represents the slack variable for the i-th network data asset, i=1,2,…,N, N represents the total number of network data assets, C represents the penalty coefficient for controlling the degree of hypersphere relaxation, x i represents the interaction state vector of the i-th network data asset at any time, f represents the distribution center of historical network data, that is, the data center of mass of all historical network data, e represents the distance from the center of the hypersphere c to the data center of mass f, that is, the eccentricity, λ represents the penalty coefficient when the eccentricity e exceeds the dynamic eccentricity ε, L(R,c,ε,λ) represents the Lagrangian function with respect to R,c,ε,λ, and |||| represents the calculation of the Euclidean distance.
[0051] The dynamic eccentricity is determined based on the optimization conditions of Lagrangian duality theory, namely the KKT condition. The penalty coefficient for eccentricities exceeding the dynamic eccentricity ε acts as a Lagrangian multiplier and can be solved using the Lagrangian function during the minimum hypersphere solution. The data centroid is the geometric mean of all data points. When solving the minimum hypersphere, f is used as the eccentricity point to correct the center c of the minimum hypersphere, ensuring that the center does not deviate too much from the data center f, thereby better matching the actual data distribution of historical network data.
[0052] Using the Gaussian kernel function, we perform pairwise high-dimensional mapping on the historical network data state vectors of each network data asset:
[0053]
[0054] Among them, x j represents the interaction state vector of the jth network data asset at any time, K(x i ,x j ) represents the description of x i and x j The Gaussian kernel function value of similarity in high-dimensional space, exp represents the natural exponential function, σ i and σ j Respectively represent the description x i and x j Kernel width of local distribution characteristics.
[0055] In a possible implementation, the calculation formula for the kernel width is specifically:
[0056]
[0057] in, and Represents x i and x j The mth nearest neighbor point of , where M represents the preset neighborhood value.
[0058] Specifically, by selecting the M nearest neighbors of each point to calculate the kernel width, it is possible to adapt to the local data distribution characteristics, thereby preserving the local structure and similarity of the data during the high-dimensional mapping process. This dynamic kernel width method can effectively reduce the impact of distribution unevenness in high-dimensional data, avoid the overfitting or underfitting problems caused by a fixed kernel width, and improve the model's ability to handle complex data and the accuracy of high-dimensional mapping. By dynamically adjusting the kernel width and adapting to the local data distribution characteristics, a more accurate support vector description model can be established to more accurately identify anomalies in network data assets, effectively reducing false positives and false negatives, and improving the reliability and adaptability of security detection.
[0059] It should be noted that those skilled in the art can set the size of the preset neighborhood value according to actual needs, and the present invention does not limit this.
[0060] Solve the interaction state vector weight value through the Lagrangian optimization algorithm:
[0061]
[0062] Among them, α i Represents the interaction state vector weight value of the i-th network data asset, α j Represents the weight value of the interaction state vector of the j-th network data asset.
[0063] Calculate the radius and center of the hypersphere based on the obtained interaction state vector weight value:
[0064]
[0065] R 2 =||φ(x s )-c|| 2
[0066] Among them, x s Represents the interaction state vector of any network data asset, |||| 2 Calculates the square of the Euclidean norm.
[0067] Specifically, during the development of the support vector machine (SVM) model, a historical network data state vector is first constructed based on the time and frequency of asset interactions, identifying the periodic patterns and frequency limits of network data assets. Next, the SVM uses a Gaussian kernel function to map the state vector into a high-dimensional space, constructing a minimal hypersphere. The center and radius of the hypersphere describe the distribution range of normal data. During this process, an eccentricity e is introduced to correct the center c of the hypersphere, ensuring that c is closer to the center f of the historical data distribution. The introduction of the eccentricity prevents the center of the hypersphere from deviating from the actual data centroid, improving the robustness and fitting accuracy of the model, especially in the case of asymmetric data distribution or outliers. The Lagrangian optimization algorithm is used to calculate the weights of the interaction state vectors, ultimately determining the radius and center of the hypersphere. This allows the hypersphere to dynamically adapt to the data distribution, comprehensively and accurately describing the normal operation of network data assets and providing strong support for subsequent anomaly detection. The application of the eccentricity significantly improves the model's adaptability to complex data distributions and detection accuracy.
[0068] The first sending module 3 is used to send the support vector description model to each data processing node.
[0069] It should be noted that the first sending module distributes the support vector description model to each data processing node, enabling it to independently perform real-time data anomaly detection, thereby improving the system's distributed processing capabilities and detection efficiency.
[0070] The first collection module 4 is used to perform approximate collection of the real-time network data of each network data asset in combination with Chebyshev polynomials.
[0071] Chebyshev polynomials are a special type of orthogonal polynomial widely used in function approximation, numerical optimization, and error minimization. They possess excellent numerical stability and rapid convergence, enabling them to approximate target values in high-dimensional complex functions using a low-order form, thereby reducing computational effort and improving efficiency. The first acquisition module uses Chebyshev polynomials to estimate real-time network data. Leveraging their efficient approximation and error control properties, they can quickly and accurately acquire real-time network data even in the presence of highly concurrent data streams and partial data loss, significantly reducing latency and ensuring stable and reliable data acquisition.
[0072] In a possible implementation, the first acquisition module 4 is specifically configured to:
[0073] Count the total number of network data assets and record the asset interaction time of each network data asset.
[0074] The number of asset interactions of each network data asset is counted using Chebyshev polynomials:
[0075]
[0076] Among them, N(A i ) represents network data asset A i The total number of interactions, S represents the interaction subset containing network data assets, |S| represents the interaction subset containing network data assets A i The number of interaction subsets, k represents the upper limit of interaction subsets related to m, and m represents the total number of network data assets. Represents other network data assets A in S j The total number of interactions, represents the Chebyshev polynomial coefficients used to adjust the influence weight of the interaction subset S, and ! represents the factorial operation.
[0077] The number of asset interactions is used as the asset interaction frequency, and the real-time network data is obtained by combining the asset interaction moments.
[0078] in, When the number of subsets is large, the weight is reduced by combining the proportions. Control when the data volume is large and fast convergence is required. Symmetry is guaranteed, and the influence ratio of low-dimensional subsets and high-dimensional subsets is strictly distributed based on the number of combinations, which can avoid excessive participation of certain subsets and lead to estimation bias. This combined effect significantly reduces the influence of high-dimensional subsets while also achieving rapid convergence on errors. By adjusting the ratio of k and m, the influence of low- and high-dimensional subsets can be balanced, adapting to data scenarios of varying scales. Maintaining symmetry while rapidly converging high-dimensional subset weights through exponential decay makes this approach well-suited for large-scale network asset interaction analysis.
[0079] It's important to note that collecting real-time network data through a combination of Chebyshev polynomials and weight adjustment effectively leverages information from a subset of interactions to rapidly estimate the total number of asset interactions, avoiding the need to scan the entire data set and improving data processing efficiency. Furthermore, by dynamically adjusting the ratio of subset k to the total number of assets m, this method balances the influence of low- and high-dimensional subsets, reducing the impact of high-dimensional subsets on the results and ensuring statistical symmetry and accuracy. Even in scenarios with large data volumes, missing data, or high concurrency, this method achieves rapid convergence, reduces latency, and ensures efficient collection of real-time network data.
[0080] The second acquisition module 5 is used to collect load data of each data processing node.
[0081] In a possible implementation, the load data includes CPU usage, memory occupancy, network bandwidth utilization, and the length of a queue of pending tasks.
[0082] It should be noted that load data comprehensively evaluates the operating status of each node by monitoring CPU usage, memory occupancy, network bandwidth utilization and the length of the pending task queue, providing a real-time basis for task allocation and ensuring load balancing and system efficiency.
[0083] The second sending module 6 is used to send the real-time network data to each data processing node based on the load data and with the goal of minimizing the network data processing delay.
[0084] It can be understood that real-time network data is dynamically allocated based on the load data of the data processing node (including CPU usage, memory occupancy, network bandwidth utilization, etc.) to minimize network data processing delays, ensure system load balancing, and improve data processing efficiency and real-time performance.
[0085] In a possible implementation, the second sending module 6 is specifically configured to:
[0086] Determine the data delivery constraint function with the goal of minimizing network data processing latency:
[0087]
[0088] Among them, Minimize means taking the minimum value, R q 、B q and L q They represent the processing rate, network bandwidth, and network delay of data processing node q, respectively. q=1,2,…,Q, where Q represents the total number of data processing nodes, and D q Indicates the amount of network data sent to node q, Represents for any data processing node q.
[0089] Determine the node priority of each data processing node:
[0090]
[0091] Among them, P q Indicates the node priority of data processing node q.
[0092] Under the condition that the data allocation constraint function is satisfied, the real-time network data is sent to each data processing node according to the node priority.
[0093] It should be noted that the second dispatch module first optimizes the data allocation strategy by constructing a data dispatch constraint function aimed at minimizing network data processing latency. This function comprehensively considers the processing rate, network bandwidth, and latency of each data processing node. It then calculates node priority based on each node's processing rate, bandwidth, and latency, with high-priority nodes receiving tasks first. Ultimately, while satisfying the data dispatch constraint, real-time network data is distributed to each data processing node based on priority. This process significantly improves the efficiency and fairness of data processing, ensures system load balancing and minimizes latency, and adapts to the complex demands of dynamic network environments.
[0094] The identification module 7 is used to perform security identification on the received real-time network data through the support vector description model in each data processing node, and output abnormal real-time network data, that is, output real-time network data outside the minimum hypersphere.
[0095] In a possible implementation, the identification module is specifically configured to:
[0096] Establish a screening formula for abnormal real-time network data:
[0097] ||φ(x)-c|| 2 >R 2
[0098]
[0099] Among them, x represents the input real-time network data, φ(x) represents the high-dimensional mapping value of x, K(x,x) represents the inner product of the high-dimensional mapping value of x, and K(x,x) represents the inner product of the high-dimensional mapping value of x. i ) represents x and x i The kernel function value in high-dimensional space is similarity.
[0100] The real-time network data that meets the screening formula is output as abnormal real-time network data.
[0101] Specifically, during this process, the identification module maps the input real-time network data into a high-dimensional space using a high-dimensional mapping function and uses a kernel function to calculate the similarity between data points and support vectors. Subsequently, based on the center and radius of the minimum hypersphere, data points in the high-dimensional space that lie outside the sphere are identified as anomalies. By using the kernel function to capture the nonlinear characteristics of the data, high-precision anomaly detection is achieved, while effectively reducing false positives and missed negatives, improving the accuracy and reliability of network data asset security identification.
[0102] In a possible implementation, the method further includes:
[0103] The support vector description model is updated at intervals of a second preset time length.
[0104] It is understandable that by updating the support vector description model at intervals of the second preset time, it is possible to dynamically adapt to changes in network data assets, ensure the real-time performance and detection accuracy of the model, and improve the robustness and adaptability of the system.
[0105] It should be noted that those skilled in the art can set the second preset duration according to actual needs, and the present invention does not limit this.
[0106] In a possible implementation, the method further includes:
[0107] Issue early warnings in the presence of abnormal real-time network data.
[0108] It should be noted that when abnormal real-time network data is detected, the system will issue an early warning notification in a timely manner to facilitate rapid response and processing, thereby improving the initiative and practicality of security monitoring.
[0109] In actual application, the network data asset security discovery system, based on distributed machine learning, acquires historical network data from data center nodes and combines Fourier transforms and K-means clustering to determine a first preset duration. It then uses the support vector description model (SVDD) to establish a minimum hypersphere to describe normal asset status. After distributing the model to each data processing node, it incorporates Chebyshev polynomials to efficiently collect real-time network data and dynamically optimizes its distribution based on node load. Ultimately, each node uses the SVDD model to detect anomalies in real-time data, quickly identifying and outputting abnormal data. This system features low latency, high accuracy, and strong adaptability.
[0110] Compared with the prior art, the present invention has at least the following beneficial technical effects:
[0111] In the present invention, the data center node is used as the data allocation center, and data processing tasks are issued according to the load status of each data processing node, effectively avoiding the problems of overlong task queues and excessive processing delays, ensuring load balancing, improving data processing efficiency and real-time performance, making the security discovery of network data assets more accurate and efficient, and suitable for large-scale high-concurrency scenarios. In the process of secure identification of network data assets, a support vector description model based on historical network data is established in combination with the distribution center of historical network data to determine the minimum hypersphere used to describe the normal operating status of network data assets. This can not only adapt to dynamically changing network environments, but also more accurately distinguish normal from abnormal data, greatly improving the accuracy of anomaly detection. In the process of real-time data collection, Chebyshev polynomials are combined to estimate the approximate collection of real-time network data of each network data asset. This can effectively address the problems of data acquisition difficulties and excessive delays caused by data missing and high-concurrency data streams. It can quickly acquire network data in high-concurrency and data missing environments, significantly reducing delays and computational complexity, ensuring the real-time and reliability of data collection, and improving the efficiency and practicality of network data asset security discovery. Finally, at each data processing node, the established support vector description model is used to securely and quickly identify the collected data, achieving low latency and high accuracy, capable of handling large-scale data streams that meet real-world requirements. Throughout the entire process, from data collection to data identification, the system effectively reduces asset identification latency and accurately and promptly detects abnormal data states based on historical data patterns, resulting in a low false alarm rate and enhanced practicality.
[0112] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0113] The above embodiments merely illustrate several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.
Claims
1. A network data asset security discovery system based on distributed machine learning, characterized by: The network data asset security discovery system includes a data center node and a plurality of data processing nodes connected to the data center node; The network data asset security discovery system further includes: an acquisition module, configured to acquire historical network data through the data center node at a first preset time period, wherein the historical network data includes asset interaction times of various network data assets and asset interaction frequencies at the asset interaction times; Establishing a module for establishing a support vector description model based on the historical network data in combination with the distribution center of the historical network data to determine a minimum hypersphere for describing the normal operation status of the network data asset; A first sending module, configured to send the support vector description model to each data processing node; The first acquisition module is used to perform approximate acquisition of real-time network data of each network data asset in combination with Chebyshev polynomials; The second acquisition module is used to collect load data of each data processing node; A second sending module is used to send the real-time network data to each data processing node based on the load data and with the goal of minimizing network data processing delay; an identification module, configured to securely identify received real-time network data using a support vector description model in each data processing node, and output abnormal real-time network data, i.e., real-time network data outside the minimum hypersphere; The first acquisition module is specifically used for: Count the total number of network data assets and record the asset interaction time of each network data asset; The number of asset interactions of each network data asset is counted using Chebyshev polynomials: Among them, N(A i ) represents network data asset A i The total number of interactions, S represents the interaction subset containing network data assets, |S| represents the interaction subset containing network data assets A i The number of interaction subsets, k represents the upper limit of interaction subsets related to m, and m represents the total number of network data assets. Represents other network data assets A in S j The total number of interactions, represents the Chebyshev polynomial coefficients used to adjust the influence weight of the interaction subset S, and ! represents the factorial operation; The number of asset interactions is used as the asset interaction frequency, and the real-time network data is obtained by combining the asset interaction times.
2. The network data asset security discovery system based on distributed machine learning according to claim 1 is characterized in that: The network data assets include servers, routers, switches, personal terminal devices, virtual machines, containers, IP addresses, MAC addresses, domain names, network interfaces and protocol ports.
3. The network data asset security discovery system based on distributed machine learning according to claim 1 is characterized in that: The establishment module is specifically used for: Based on the asset interaction time and the asset interaction frequency at the asset interaction time, a historical network data state vector of each network data asset is established: Among them, t k represents the kth asset interaction moment of the network data asset, Indicates that the i-th network data asset is at t k Frequency of asset interaction, x i,t represents the interaction state vector of the i-th network data asset at time t, where T1 and T2 represent the start and end times of the preset duration respectively; Establish the smallest hypersphere that describes the normal operation status of network data assets: ||x i -c|| 2 ≤R 2 +ξ i ,x i ≥0 e≤ε Where R and c represent the radius and center of the hypersphere, respectively, i represents the slack variable for the i-th network data asset, i=1,2,…,N, N represents the total number of network data assets, C represents the penalty coefficient for controlling the degree of hypersphere relaxation, x i represents the interaction state vector of the i-th network data asset at any time, f represents the distribution center of historical network data, i.e., the data centroid of all historical network data, e represents the distance from the center of the hypersphere c to the data centroid f, i.e., the eccentricity, λ represents the penalty coefficient when the eccentricity e exceeds the dynamic eccentricity ε, L(R,c,ε,λ) represents the Lagrangian function with respect to R,c,ε,λ, and || || represents the calculation of the Euclidean distance; Using the Gaussian kernel function, we perform pairwise high-dimensional mapping on the historical network data state vectors of each network data asset: Among them, x j represents the interaction state vector of the jth network data asset at any time, K(x i ,x j ) represents the description of x i and x j The Gaussian kernel function value of similarity in high-dimensional space, exp represents the natural exponential function, σ i and σ j Respectively represent the description x i and x j kernel width of local distribution characteristics; Solve the interaction state vector weight value through the Lagrangian optimization algorithm: Among them, α i Represents the interaction state vector weight value of the i-th network data asset, α j Represents the interaction state vector weight value of the j-th network data asset; Calculate the radius and center of the hypersphere based on the obtained interaction state vector weight value: R 2 =‖φ(x s )-c‖ 2 Among them, x s Represents the interaction state vector of any network data asset, || || 2 Calculates the square of the Euclidean norm.
4. The network data asset security discovery system based on distributed machine learning according to claim 3 is characterized in that: The calculation formula of the kernel width is specifically: in, and Represents x i and x j The mth nearest neighbor point of , where M represents the preset neighborhood value.
5. The network data asset security discovery system based on distributed machine learning according to claim 1 is characterized in that: The load data includes CPU usage, memory occupancy, network bandwidth utilization and the length of the pending task queue.
6. The network data asset security discovery system based on distributed machine learning according to claim 1 is characterized in that: The second issuing module is specifically configured to: Determine a data delivery amount constraint function with the goal of minimizing the network data processing delay: Among them, Minimize means taking the minimum value, R q 、B q and L q They represent the processing rate, network bandwidth, and network delay of data processing node q, respectively. q=1,2,…,Q, where Q represents the total number of data processing nodes, and D q Indicates the amount of network data sent to node q, Represents any data processing node q; Determine the node priority of each data processing node: Among them, P q Indicates the node priority of data processing node q; When the data allocation constraint function is satisfied, the real-time network data is sent to each data processing node according to the node priority.
7. The network data asset security discovery system based on distributed machine learning according to claim 3 is characterized in that: The identification module is specifically used for: Establish a screening formula for abnormal real-time network data: Among them, x represents the input real-time network data, φ(x) represents the high-dimensional mapping value of x, K(x,x) represents the inner product of the high-dimensional mapping value of x, and K(x,x) represents the inner product of the high-dimensional mapping value of x. i ) represents x and x i The kernel function value in high-dimensional space is similarity; The real-time network data that meets the screening formula is output as abnormal real-time network data.
8. The network data asset security discovery system based on distributed machine learning according to claim 1 is characterized in that: Also includes: The support vector description model is updated at intervals of a second preset time length.
9. The network data asset security discovery system based on distributed machine learning according to claim 1 is characterized in that: Also includes: In the presence of the abnormal real-time network data, an early warning is issued.
Citation Information
Patent Citations
Task unloading method and device
CN113918240A
High-precision temperature prediction method of low-temperature sensor based on deep learning
CN114444662A
Production process monitoring system based on industrial internet
CN118915649A