A data processing method for cloud computing based on artificial intelligence
Through the cloud computing data processing method based on artificial intelligence, the problems of high cost, high complexity and poor scalability of traditional monitoring methods are solved, and efficient and sensitive network monitoring and abnormal detection are achieved.
Patent Information
- Application Number
- CN202411574639.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-11-06
AI Technical Summary
Traditional monitoring methods require the purchase of expensive dedicated hardware equipment or software tools, increasing costs and complexity while limiting monitoring scope and scalability.
Using a cloud computing data processing method based on artificial intelligence, the number of data packets transmitted by the port is obtained by setting calibration cycles and judgment cycles, the average value and real-time traffic are calculated, and the Gaussian density graph and convolutional neural network are used to determine whether there are abnormalities on the port.
It reduces the computing resource consumption when judging whether there is an abnormality on all ports, significantly reduces the overall computing demand, improves resource utilization, and improves the sensitivity and scalability of network monitoring.
Smart Images

Figure CN119449651B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a data processing method of cloud computing based on artificial intelligence. Background Art
[0002] Cloud computing is an Internet-based computing model that provides computing resources (such as servers, storage devices, databases, etc.) to users as a service to achieve on-demand acquisition and use of computing resources.
[0003] The port number is a digital identifier used to identify the logical endpoint of an application or service in network communication. Different applications or services usually use different port numbers for communication. A data packet is a unit of data transmitted in a computer network. It is the basic unit in network communication and can be used for network monitoring and analysis. By capturing and analyzing data packets, you can understand the usage, performance and security status of the network, help identify abnormal behavior, bottlenecks and failures in the network, and improve network security.
[0004] Traditional monitoring methods usually use dedicated hardware devices or software tools to capture and analyze data packets in the network. It is necessary to purchase expensive dedicated hardware devices or software tools, deploy and maintain them, which increases costs and complexity. At the same time, traditional monitoring methods usually require monitoring equipment to be deployed in specific locations, which limits the monitoring scope and scalability. Summary of the invention
[0005] The purpose of the present invention is to provide a data processing method for cloud computing based on artificial intelligence to solve the following technical problems:
[0006] Traditional monitoring methods usually use dedicated hardware devices or software tools to capture and analyze data packets in the network. It is necessary to purchase expensive dedicated hardware devices or software tools, deploy and maintain them, which increases costs and complexity. At the same time, traditional monitoring methods usually require monitoring equipment to be deployed in specific locations, which limits the monitoring scope and scalability.
[0007] The purpose of the present invention can be achieved through the following technical solutions:
[0008] A data processing method for cloud computing based on artificial intelligence, comprising the following steps:
[0009] S1: m calibration periods are set, where m is a first preset number, and time nodes are set at preset time intervals Δt within the calibration period, and the number of data packets transmitted by the port is obtained at the time nodes, and the average number of data packets transmitted by the port at the same time nodes is calculated to generate a first coordinate point (a, Ca), where Ca represents the average number of data packets transmitted by the port at time node a, and the first coordinate point is fitted to obtain a first curve f(t);
[0010] S2: Determine the m+1th calibration period, take it as the judgment period, divide the judgment period into n sub-periods of the same length, n is a second preset number, obtain the number of data packets transmitted by the port in real time within the sub-period, and draw a curve g(t) showing the change of the number of data packets over time;
[0011] S3: Calculate real-time traffic and compare traffic f'(t) represents the portion of the first curve within the sub-period, tsta and tend represent the start and end time points of the sub-period, respectively, and the positions of the real-time flow and the comparison flow on the x-axis and y-axis of the preset coordinate system F are taken as flow points;
[0012] Connect the flow points to obtain a flow line, take the angle between the flow line and the x-axis as the flow angle, and determine the radian value D of the flow angle;
[0013] The triangle formed by the flow point and the origin of the coordinate system F is used as a flow triangle, and the area S of the flow triangle is determined;
[0014] S4: Calculate the abnormal coefficient K=η*S / D, where η is a preset correction coefficient, set the abnormal coefficient threshold Kys, and when the abnormal coefficient K≥Kys, mark the corresponding port as a pending port to determine whether the pending port has an abnormality.
[0015] As a further solution of the present invention: in step S4, the process of determining whether the pending port has an abnormality specifically includes:
[0016] The number of data packets transmitted by the pending port in the sub-period is used as traffic data, and the characteristics of the traffic data of a single pending port are obtained based on the traffic sequence encoder, and the characteristics are used as the target characteristics;
[0017] The target features are fused and discretized based on the Gaussian density map, and the interaction features between the target features of different pending ports are determined through a mutually transposed convolutional network;
[0018] The pre-trained classification model is used to determine the pending ports with abnormalities.
[0019] As a further solution of the present invention: the process of training the classification model specifically includes:
[0020] Establishing a database, wherein the database stores interaction features labeled as abnormal;
[0021] A classification model is established based on the deep learning model, and is trained and verified through the database to obtain a pre-trained classification model.
[0022] As a further solution of the present invention: the abnormal coefficient threshold is determined based on the comparison flow, Kys=γ / C, γ is a preset coefficient.
[0023] As a further solution of the present invention: in the step S3, when the comparison flow and / or the real-time flow is 0, an early warning message is sent for prompting.
[0024] As a further solution of the present invention: in the step S1, when there is an abnormal pending port in the calibration period, a supplementary period is set to replace the calibration period with the abnormality.
[0025] As a further solution of the present invention: in the step S1, the process of calculating the average number of data packets also includes the following steps:
[0026] When the difference between the number of data packets in calibration period b and the corresponding average number of data packets is greater than the preset difference threshold, the number of data packets in calibration period b is removed and the average is calculated again, and the above steps are repeated until there is no difference greater than the preset difference threshold.
[0027] As a further solution of the present invention: taking the average number of data packets when there is no difference greater than a preset difference threshold as the target number, determining the number SL of calibration cycles corresponding to the target number, and sending an early warning message for prompting when the number SL is less than 0.8m.
[0028] Beneficial effects of the present invention: In this scheme, a calibration period is set and the number of transmitted data packets of the port is obtained at fixed time intervals within the period, and the average value is calculated to generate a first curve. The goal of this stage is to establish a historical benchmark (i.e., the first curve) that can reflect the typical change trend of the port transmission data packets under normal circumstances. Subsequently, the real-time data and the reference model (i.e., the first curve) can be compared to facilitate the detection of anomalies; a new judgment period is determined, and a plurality of sub-periods are divided in the period to obtain the number of real-time data packets, thereby generating a real-time curve g(t). In this stage, the division of the sub-periods is to capture data fluctuations in a shorter period of time, making the judgment more sensitive and detailed, and the sub-periods are divided into a plurality of sub-periods to obtain the number of real-time data packets, thereby generating a real-time curve g(t). The division of cycles can facilitate timely alarm after the abnormality is determined, so as to avoid the abnormality not being handled and affecting network security; then, the judgment parameters are determined according to the real-time flow and the comparative flow, that is, the area of the flow triangle and the radian value of the flow angle; it is worth noting that the flow point corresponding to the comparative flow is on the y-axis. When the comparative flow is constant, the larger the real-time flow, the larger the area of the flow triangle will be. At the same time, the flow angle will decrease, and the radian value of the flow angle will also decrease. Therefore, the larger the area of the flow triangle and / or the smaller the radian value of the flow angle, the larger the abnormal coefficient should be, that is, the greater the possibility of the existence of an abnormality; finally, after determining the pending port, it is judged whether there is an abnormality in the pending port. It is worth noting that the process of judging whether there is an abnormality in the pending port requires the use of a variety of complex algorithms and models, such as Gaussian density maps, convolutional neural networks, etc. However, although these models and algorithms can accurately identify and analyze abnormal patterns between ports, they all consume high computing resources. For example, the main function of the Gaussian density map is to fuse and discretize target features to help judge the regularity and deviation of data distribution. However: 1. Density estimation of multi-dimensional features: When performing Gaussian density estimation on multi-dimensional features, it is necessary to calculate the probability distribution of each data point in multi-dimensional space. Especially when the feature dimension is high, the amount of calculation for each data point will increase exponentially; 2. Application of complex kernel functions: Through Usually Gaussian density estimation will use the kernel density method to achieve data smoothing by selecting different kernel functions, which will further increase the computing requirements. The higher the data volume and dimension, the greater the computational overhead of density estimation; 3. Data discretization: Discretization maps the continuous Gaussian density feature distribution to different intervals. This process not only requires a lot of calculations, but also requires repeated discrete judgments, which further consumes computing resources; therefore, in this solution, first screen the pending ports and determine whether there are abnormalities in the pending ports, reduce the computing resource consumption when judging whether there are abnormalities for all ports, and concentrate computing resources on the pending ports, significantly reducing the overall computing requirements and improving resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The present invention will be further described below in conjunction with the accompanying drawings.
[0030] Figure 1 It is a flow chart of a data processing method of cloud computing based on artificial intelligence of the present invention. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0032] See also Figure 1 As shown, the present invention is a data processing method of cloud computing based on artificial intelligence, comprising the following steps:
[0033] S1: m calibration periods are set, where m is a first preset number, and time nodes are set at preset time intervals Δt within the calibration period, and the number of data packets transmitted by the port is obtained at the time nodes, and the average number of data packets transmitted by the port at the same time nodes is calculated to generate a first coordinate point (a, Ca), where Ca represents the average number of data packets transmitted by the port at time node a, and the first coordinate point is fitted to obtain a first curve f(t);
[0034] S2: Determine the m+1th calibration period, take it as the judgment period, divide the judgment period into n sub-periods of the same length, n is a second preset number, obtain the number of data packets transmitted by the port in real time within the sub-period, and draw a curve g(t) showing the change of the number of data packets over time;
[0035] S3: Calculate real-time traffic and compare traffic f'(t) represents the portion of the first curve within the sub-period, tsta and tend represent the start and end time points of the sub-period, respectively, and the positions of the real-time flow and the comparison flow on the x-axis and y-axis of the preset coordinate system F are taken as flow points;
[0036] Connect the flow points to obtain a flow line, take the angle between the flow line and the x-axis as the flow angle, and determine the radian value D of the flow angle;
[0037] The triangle formed by the flow point and the origin of the coordinate system F is used as a flow triangle, and the area S of the flow triangle is determined;
[0038] S4: Calculate the abnormal coefficient K=η*S / D, where η is a preset correction coefficient, set the abnormal coefficient threshold Kys, and when the abnormal coefficient K≥Kys, mark the corresponding port as a pending port to determine whether the pending port has an abnormality.
[0039] It should be noted that a calibration period is set and the number of transmitted data packets of the port is obtained at fixed time intervals within the period, and the average value is calculated to generate the first curve. The goal of this stage is to establish a historical benchmark (i.e., the first curve) that can reflect the typical change trend of the port transmission data packets under normal circumstances. The real-time data and the reference model (i.e., the first curve) can be compared later to facilitate the detection of anomalies; a new judgment period is determined, and multiple sub-periods are divided in the period to obtain the number of real-time data packets, thereby generating a real-time curve g(t). In this stage, the division of sub-periods is to capture data fluctuations in a shorter period of time, making the judgment more sensitive and detailed, and the division of sub-periods can be In order to facilitate timely alarm after the abnormality is determined, to avoid the abnormality not being handled and affecting network security; then, determine the judgment parameters according to the real-time flow and the comparison flow, that is, the area of the flow triangle and the radian value of the flow angle; it is worth noting that the flow point corresponding to the comparison flow is on the y-axis. When the comparison flow is constant, the larger the real-time flow, the larger the area of the flow triangle will be. At the same time, the flow angle will decrease, and the radian value of the flow angle will also decrease. Therefore, the larger the area of the flow triangle and / or the smaller the radian value of the flow angle, the larger the abnormal coefficient should be, that is, the greater the possibility of abnormality; finally, after determining the pending port, judge whether there is an abnormality on the pending port; it is worth noting that It should be noted that the process of judging whether there is an abnormality in the pending port requires the use of a variety of complex algorithms and models, such as Gaussian density maps, convolutional neural networks, etc. However, although these models and algorithms can accurately identify and analyze abnormal patterns between ports, they all consume high computing resources. For example, the main function of the Gaussian density map is to fuse and discretize target features to help judge the regularity and deviation of data distribution. However: 1. Density estimation of multi-dimensional features: When performing Gaussian density estimation on multi-dimensional features, it is necessary to calculate the probability distribution of each data point in multi-dimensional space. Especially when the feature dimension is high, the amount of calculation for each data point will increase exponentially; 2. Application of complex kernel functions: Usually high Gaussian density estimation uses the kernel density method to achieve data smoothing by selecting different kernel functions, which will further increase the computing requirements. The higher the data volume and dimension, the greater the computing overhead of density estimation. 3. Data discretization: Discretization maps the continuous Gaussian density feature distribution to different intervals. This process not only requires a lot of calculations, but also requires repeated discrete judgments, which further consumes computing resources. Therefore, in this solution, the pending ports are first screened and it is determined whether there are any abnormalities in the pending ports, which reduces the computing resource consumption when judging whether there are any abnormalities in all ports, and concentrates computing resources on the pending ports, significantly reducing the overall computing requirements and improving resource utilization.
[0040] In another preferred embodiment of the present invention, in step S4, the process of determining whether the pending port has an abnormality specifically includes:
[0041] The number of data packets transmitted by the pending port in the sub-period is used as traffic data, and the characteristics of the traffic data of a single pending port are obtained based on the traffic sequence encoder, and the characteristics are used as the target characteristics;
[0042] The target features are fused and discretized based on the Gaussian density map, and the interaction features between the target features of different pending ports are determined through a mutually transposed convolutional network;
[0043] The pre-trained classification model is used to determine the pending ports with abnormalities.
[0044] It is worth noting that, first, the number of data packets transmitted by the pending port in each sub-period is used as the flow data input, and the flow sequence encoder is used to extract the features of the pending port. This feature encoding process can compress the data so that it represents the core change pattern of the pending port flow and generates the target features. Then, the target features are processed by the Gaussian density map. First, they are fused, the relationship between the features is smoothed, and further discretized for subsequent analysis. The Gaussian density map helps capture the distribution law and degree of change of the flow characteristics of the pending port here, ensuring that the target features are more suitable for detecting abnormal patterns. On this basis, the feature interaction relationship between different pending ports is identified through mutually transposed convolutional networks. The convolutional network structure of this step is designed as a transposed network in order to explore the relationship between different ports. Through this network, the system can identify the potential commonalities or differences between flow characteristics, thereby providing more powerful feature support for subsequent classification judgments. Finally, the obtained interactive features are input into the pre-trained classification model to determine whether there is an anomaly. The model can effectively identify abnormal patterns through training with a large amount of data, thereby automatically screening out abnormal pending ports.
[0045] In another preferred embodiment of the present invention, the process of training the classification model specifically includes:
[0046] Establishing a database, wherein the database stores interaction features labeled as abnormal;
[0047] A classification model is established based on the deep learning model, and is trained and verified through the database to obtain a pre-trained classification model.
[0048] In another preferred embodiment of the present invention, the abnormal coefficient threshold is determined based on the comparative flow rate, Kys=γ / C, γ is a preset coefficient.
[0049] It should be noted that the flow point corresponding to the comparison flow is on the y-axis. When the flow point on the x-axis remains unchanged, the larger the comparison flow, the larger the area of the flow triangle. Therefore, it is necessary to adjust the abnormal coefficient threshold. The preset coefficient γ can be obtained through multiple tests.
[0050] In another preferred embodiment of the present invention, in the step S3, when the comparison flow and / or the real-time flow is 0, an early warning message is sent for prompting.
[0051] It should be noted that if the comparison traffic or real-time traffic is 0, it may indicate that the port has no data packet transmission during this period. This may mean that the port has a fault, connection problem, or configuration abnormality, and further investigation is required to ensure the normal operation of the network.
[0052] In another preferred embodiment of the present invention, in the step S1, when there is an abnormal pending port within the calibration period, a supplementary period is set to replace the calibration period with the abnormality.
[0053] It is understandable that the main purpose of the calibration cycle is to establish a historical benchmark to describe the typical changing trend of port data packet transmission under normal circumstances. If an abnormal situation occurs during the calibration cycle, the collected data will no longer represent the normal traffic pattern, affecting the accuracy of the benchmark; abnormal data will interfere with the calculation of the average value of the calibration cycle, thereby deviating from the actual normal traffic level. A supplementary cycle is set to re-collect normal traffic data to avoid interference of abnormal data on the benchmark model.
[0054] In another preferred embodiment of the present invention, in step S1, the process of calculating the average number of data packets further includes the following steps:
[0055] When the difference between the number of packets in the calibration period b and the corresponding average number of packets is greater than a preset difference threshold, the number of packets in the calibration period b is removed and the average is calculated again.
[0056] In another preferred embodiment of the present invention, the average number of data packets when there is no difference greater than a preset difference threshold is taken as the target number, and the number SL of calibration cycles corresponding to the target number is determined. When the number SL is less than 0.8m, an early warning message is sent for prompting.
[0057] It is worth noting that when calculating the average number of data packets, the data with a large difference between the number of data packets in the calibration period and the corresponding average value (abnormal data with a difference greater than the preset threshold) are removed, and the average value of the remaining number of data packets that does not exceed the preset difference threshold is used as the "target number". The target number represents a stable and expected traffic level; when the number of calibration cycles SL is less than 80% of the total number of calibration cycles (that is, SL < 0.8m), an early warning message will be sent. If the number of valid calibration cycles (that is, normal cycles that meet expectations) is insufficient, the representativeness of the benchmark model may not be sufficient to support accurate anomaly detection. At this time, an early warning is sent to remind system administrators or operators to pay attention to the reliability of the benchmark model so that it can be adjusted and processed in time.
[0058] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered to limit the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.
Claims
1. A data processing method for cloud computing based on artificial intelligence, characterized in that: The following steps are involved: S1: m calibration periods are set, where m is a first preset number, and time nodes are set at preset time intervals Δt within the calibration period, and the number of data packets transmitted by the port is obtained at the time nodes, and the average number of data packets transmitted by the port at the same time nodes is calculated to generate a first coordinate point (a, Ca), where Ca represents the average number of data packets transmitted by the port at time node a, and the first coordinate point is fitted to obtain a first curve f(t); S2: Determine the m+1th calibration period, take it as the judgment period, divide the judgment period into n sub-periods of the same length, n is a second preset number, obtain the number of data packets transmitted by the port in real time within the sub-period, and draw a curve g(t) showing the change of the number of data packets over time; S3: Calculate real-time traffic and compare traffic f'(t) represents the portion of the first curve within the sub-period, tsta and tend represent the start and end time points of the sub-period, respectively, and the positions of the real-time flow and the comparison flow on the x-axis and y-axis of the preset coordinate system F are taken as flow points; Connect the flow points to obtain a flow line, take the angle between the flow line and the x-axis as the flow angle, and determine the radian value D of the flow angle; The triangle formed by the flow point and the origin of the coordinate system F is used as a flow triangle, and the area S of the flow triangle is determined; S4: Calculate the abnormal coefficient K=η*S / D, where η is a preset correction coefficient, set the abnormal coefficient threshold Kys, and when the abnormal coefficient K≥Kys, mark the corresponding port as a pending port to determine whether the pending port has an abnormality; The process of determining whether the pending port is abnormal specifically includes: The number of data packets transmitted by the pending port in the sub-period is used as traffic data, and the characteristics of the traffic data of a single pending port are obtained based on the traffic sequence encoder, and the characteristics are used as the target characteristics; The target features are fused and discretized based on the Gaussian density map, and the interaction features between the target features of different pending ports are determined through a mutually transposed convolutional network; The pre-trained classification model is used to determine the pending ports with abnormalities.
2. The data processing method based on cloud computing based on artificial intelligence according to claim 1, characterized in that: The process of training the classification model specifically includes: Establishing a database, wherein the database stores interaction features labeled as abnormal; A classification model is established based on the deep learning model, and is trained and verified through the database to obtain a pre-trained classification model.
3. The data processing method based on cloud computing based on artificial intelligence according to claim 1, characterized in that: The abnormal coefficient threshold is determined based on the comparison flow rate, Kys=γ / C, γ is a preset coefficient.
4. The data processing method based on artificial intelligence cloud computing according to claim 1, characterized in that: In the step S3, when the comparison flow and / or the real-time flow is 0, an early warning message is sent for prompting.
5. The data processing method based on cloud computing based on artificial intelligence according to claim 1, characterized in that: In the step S1, when there is an abnormal pending port in the calibration period, a supplementary period is set to replace the calibration period with the abnormality.
6. The data processing method based on cloud computing based on artificial intelligence according to claim 1, characterized in that: In the step S1, the process of calculating the average number of data packets also includes the following steps: When the difference between the number of data packets in calibration period b and the corresponding average number of data packets is greater than the preset difference threshold, the number of data packets in calibration period b is removed and the average is calculated again, and the above steps are repeated until there is no difference greater than the preset difference threshold.
7. The data processing method based on cloud computing based on artificial intelligence according to claim 6, characterized in that: The average number of data packets when there is no difference greater than the preset difference threshold is taken as the target number, and the number SL of calibration cycles corresponding to the target number is determined. When the number SL is less than 0.8m, an early warning message is sent for prompting.
Citation Information
Patent Citations
Abnormal network traffic monitoring method, device and equipment based on deep learning and readable storage medium
CN118041661A
Real-time three-dimensional model data traffic anomaly detection method, device and equipment
CN118260699A