Monitoring data processing method based on big data analysis and electronic device
By constructing a monitoring transmission topology network and setting up interception and deployment processing terminals, the problem of data feature loss in large-scale distributed monitoring is solved by collecting and comparing monitoring data sources, thus achieving efficient and reliable transmission and processing of monitoring data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHONGNENG DIGITAL (TIANJIN) TECHNOLOGY CO LTD
- Filing Date
- 2025-08-01
- Publication Date
- 2026-07-03
AI Technical Summary
Existing technologies are ill-suited to the loss of data characteristics during the transmission of monitoring data in large-scale distributed monitoring, resulting in poor consistency and reliability of monitoring data.
Construct a monitoring transmission topology network, set up a first interception deployment and processing terminal and a second interception deployment and processing terminal, collect monitoring data sources, establish a feature vector space template to compare data feature vector loss, and optimize data feature loss indicators to improve data consistency and reliability.
It reduces the feature loss rate of monitoring data, improves the consistency and reliability of monitoring data, and ensures the integrity and accuracy of data during transmission.
Smart Images

Figure CN120856581B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to monitoring data processing methods and electronic devices based on big data analysis. Background Technology
[0002] Large-scale distributed monitoring is widely used in fields such as the Industrial Internet of Things (IIoT), smart cities, and cloud computing. It typically consists of a main monitoring processing terminal node, sub-monitoring processing terminal nodes, and a large number of distributed monitoring nodes, forming a complex multi-layered data transmission network. However, during the acquisition, transmission, and processing of monitoring data, issues such as network latency, data loss, and device heterogeneity make it difficult to guarantee data integrity and consistency, affecting the accuracy and real-time performance of data analysis. Traditional monitoring data processing methods struggle to effectively handle the high-efficiency acquisition and processing needs of massive amounts of monitoring data. Furthermore, because monitoring data undergoes multiple forwarding and aggregation processes during transmission, the temporal sequence, integrity, and correlation characteristics of the data may be lost or distorted to varying degrees, affecting the usability of the monitoring data and leading to deviations in the final monitoring analysis results.
[0003] Current technologies suffer from several technical problems, including difficulty in adapting to large-scale distributed monitoring, loss of data characteristics during data transmission, and poor consistency and reliability of monitoring data. Summary of the Invention
[0004] This application provides a monitoring data processing method and electronic device based on big data analysis, which solves the technical problems in the prior art that make it difficult to adapt to large-scale distributed monitoring and data feature loss during monitoring data transmission, resulting in poor consistency and reliability of monitoring data. It achieves the technical effect of reducing the monitoring data feature loss rate and improving the consistency and reliability of monitoring data.
[0005] This application provides a monitoring data processing method based on big data analysis. The method includes: constructing a monitoring transmission topology network based on a main monitoring processing terminal node, sub-monitoring processing terminal nodes, distributed monitoring nodes, and the data transmission relationships of each node; setting a first interception deployment processing terminal and a second interception deployment processing terminal in the monitoring transmission topology network, wherein the first interception deployment processing terminal is located near the distributed monitoring nodes, and the second interception deployment processing terminal is located near the main monitoring processing terminal node; collecting a first monitoring data source based on the first interception deployment processing terminal and collecting a second monitoring data source based on the second interception deployment processing terminal; establishing a feature vector space template; comparing the data feature vector loss of the first monitoring data source and the second monitoring data source based on the feature vector space template, outputting a data feature loss index, and optimizing the second monitoring data source according to the data feature loss index before entering the next-level monitoring processing terminal for processing.
[0006] In one possible implementation, the monitoring data processing method based on big data analysis is further used to perform the following processing: the sub-monitoring processing terminal node includes multi-level sub-monitoring processing terminal nodes, and the upper-level sub-monitoring processing terminal node is tree-connected with multiple lower-level sub-monitoring processing terminal nodes; wherein, the first interception deployment processing terminal is located near the distributed monitoring node and is communicatively connected to multiple sub-monitoring processing terminal nodes at the selected level, and is used to receive the first monitoring data source from the multiple sub-monitoring processing terminal nodes at the selected level; the second interception deployment processing terminal is located near the main monitoring processing terminal node and is communicatively connected to multiple sub-monitoring processing terminal nodes at the selected level, and is used to receive the second monitoring data source from the multiple sub-monitoring processing terminal nodes at the selected level.
[0007] In one possible implementation, the monitoring data processing method based on big data analysis is further used to perform the following processing: the data transmission relationship includes the transmission relationship between the main monitoring processing terminal node and the sub-monitoring processing terminal node, the transmission relationship between the upper-level sub-monitoring processing terminal node and the next-level sub-monitoring processing terminal node, and the transmission relationship between the sub-monitoring processing terminal node and the distributed monitoring node, and a monitoring transmission topology network is constructed based on the data transmission relationship.
[0008] In one possible implementation, the monitoring data processing method based on big data analysis is further configured to perform the following processing: obtaining the data feature vector of the second monitoring data source; establishing a feature vector space template based on the data feature vector of the second monitoring data source as a template; mapping the feature vector of the first monitoring data source according to the feature vector space template, and outputting the data feature vector of the first monitoring data source; comparing the loss of the data feature vector of the second monitoring data source and the data feature vector of the first monitoring data source, and outputting a data feature loss index.
[0009] In one possible implementation, the monitoring data processing method based on big data analysis is further configured to perform the following processing: identifying the interval levels between the first interception deployment processing terminal and the second interception deployment processing terminal; when the interval levels are greater than a preset interval levels, establishing a feature vector space template, wherein the feature vector space template is a feature vector space between the feature dimensions of the first monitoring data source and the feature dimensions of the second monitoring data source; performing high-dimensional feature vector mapping on the first monitoring data source according to the feature vector space template, and outputting the data feature vector of the first monitoring data source; performing low-dimensional feature vector mapping on the second monitoring data source according to the feature vector space template, and outputting the data feature vector of the second monitoring data source; performing a loss comparison between the data feature vector of the second monitoring data source and the data feature vector of the first monitoring data source, and outputting a data feature loss index.
[0010] In one possible implementation, the monitoring data processing method based on big data analysis is further configured to perform the following processing: calculating a similarity metric between the data feature vector of the second monitoring data source and the data feature vector of the first monitoring data source, wherein the similarity metric is obtained through KL divergence calculation; calculating a key feature similarity metric between the data feature vector of the second monitoring data source and the data feature vector of the first monitoring data source, wherein the key feature similarity metric is obtained through principal component loss rate calculation; and performing weighted calculations using the similarity metric and the key feature similarity metric to output a data feature loss index.
[0011] In one possible implementation, the monitoring data processing method based on big data analysis is further configured to perform the following processing: when the data feature loss index is greater than a preset data feature loss index, the data feature vectors of the first monitoring data source and the second monitoring data source are fused and repaired to obtain a third monitoring data source; the third monitoring data source is compared with the core fields, key features, and time-series trajectories of the first monitoring data source and the second monitoring data source, and a consistency verification result is output; if the consistency verification result passes, the third monitoring data source enters the next-level monitoring processing terminal for processing; if the consistency verification result fails, a re-fused and repaired third monitoring data source is obtained.
[0012] This application also provides an electronic device, including: a memory for storing executable instructions; and a processor for implementing a monitoring data processing method based on big data analysis when executing the executable instructions stored in the memory.
[0013] This application proposes a monitoring data processing method and electronic device based on big data analysis. The method constructs a monitoring transmission topology network based on the main monitoring processing terminal node, sub-monitoring processing terminal nodes, distributed monitoring nodes, and the data transmission relationships between these nodes. It sets up a first interception deployment processing terminal and a second interception deployment processing terminal; collects data from the first and second monitoring data sources; establishes a feature vector space template and compares data feature vector losses, outputting a data feature loss index; and optimizes the second monitoring data source according to the data feature loss index before sending it to the next-level monitoring processing terminal for processing. This solves the technical problems in existing technologies, such as difficulty in adapting to large-scale distributed monitoring and data feature loss during monitoring data transmission, leading to poor consistency and reliability of monitoring data. It achieves the technical effect of reducing the monitoring data feature loss rate and improving the consistency and reliability of monitoring data. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly described below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0015] Figure 1 This is a flowchart illustrating the monitoring data processing method based on big data analysis provided in an embodiment of this application.
[0016] Figure 2 The deployment tree diagram of the monitoring transmission topology network in the monitoring data processing method based on big data analysis provided in the embodiments of this application.
[0017] Figure 3 This is a schematic diagram illustrating the process of comparing data feature vector loss in the monitoring data processing method based on big data analysis provided in the embodiments of this application.
[0018] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0019] Explanation of reference numerals in the attached drawings: Input device 401, processor 402, memory 403, output device 404. Detailed Implementation
[0020] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below.
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of this application will be provided in conjunction with the accompanying drawings. The described embodiments should not be considered as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] In the following description, references to "some embodiments" describe a subset of all possible embodiments. However, it is understood that "some embodiments" can be the same or different subsets of all possible embodiments and can be combined with each other without conflict. The terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. The terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, product, or server that includes a series of steps is not necessarily limited to those steps explicitly listed, but may include other steps not explicitly listed or inherent to such processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only.
[0023] This application provides a monitoring data processing method based on big data analysis, such as... Figure 1 As shown, the method includes:
[0024] Step A10: Construct a monitoring transmission topology network based on the main monitoring processing terminal node, the sub-monitoring processing terminal nodes, the distributed monitoring nodes, and the data transmission relationships of each node.
[0025] Preferably, based on the data transmission relationship of distributed monitoring, the main monitoring processing terminal node, sub-monitoring processing terminal nodes, distributed monitoring nodes, and the data transmission paths between them are structurally modeled to form a hierarchical and manageable monitoring transmission topology network. This network describes the collection, transmission, aggregation, and processing flow of monitoring data, ensuring that data flows efficiently and reliably to the target analysis node. Distributed monitoring nodes, as the data acquisition layer, are located at the bottom layer of the monitoring transmission topology network and are typically terminal devices deployed at the monitoring site, such as sensors, cameras, and IoT devices, responsible for collecting and uploading raw data. Sub-monitoring processing terminal nodes, as the data aggregation layer, are located in the middle layer of the monitoring transmission topology network and are responsible for receiving data from multiple distributed monitoring nodes, performing preliminary cleaning, aggregation, or edge computing, and then transmitting it upwards. The main monitoring processing terminal node, as the core processing layer, is located at the top layer of the monitoring transmission topology network and is responsible for receiving data from multiple sub-monitoring processing terminals for global analysis, storage, or decision-making.
[0026] Preferably, the data transmission relationship between each node includes a hierarchical transmission relationship and a monitoring data flow direction. Monitoring data is uploaded level by level from distributed monitoring nodes to sub-monitoring processing terminal nodes to the main monitoring processing terminal node, forming a tree or mesh structure. Each sub-monitoring processing terminal node connects to multiple distributed monitoring nodes, and the main monitoring processing terminal node connects to multiple sub-monitoring processing terminal nodes. This monitoring transmission topology network clearly displays the source, aggregation path, and processing logic of the monitoring data. Interception points can also be deployed at key locations near the data source or core processing terminal for data verification and optimization, supporting efficient and reliable data acquisition, transmission, and processing.
[0027] Step A20: Set up a first interception deployment processing terminal and a second interception deployment processing terminal in the monitoring transmission topology network, wherein the first interception deployment processing terminal is set up near the distributed monitoring node, and the second interception deployment processing terminal is set up near the main monitoring processing terminal node.
[0028] Preferably, a two-tiered interception deployment processing end is set up in the monitoring transmission topology network, including a first interception deployment processing end and a second interception deployment processing end. The interception deployment processing end is a key functional component for data interception, feature extraction, quality verification, and transmission optimization. Its main functions include monitoring data acquisition, feature extraction and analysis, data optimization, and anomaly detection. Setting up two levels of interception deployment processing ends allows for quality control at different stages of the monitoring data lifecycle, forming a layered defense mechanism to prevent data features from gradually deteriorating during transmission. Specifically, the first interception deployment processing end is located between the distributed monitoring nodes and the sub-monitoring processing terminals, typically embedded in edge gateways or edge computing devices. It directly collects raw monitoring data from devices such as sensors and cameras, extracts key features, performs data cleaning such as noise filtering, missing value imputation, and data standardization, and then performs low-latency interception and feedback. For example, if data anomalies are detected, a local alarm or re-acquisition mechanism can be immediately triggered, reducing the backend processing burden. The second interceptor deployment processing terminal is located between the sub-monitoring processing terminal and the main monitoring processing terminal. It is usually deployed at the core network entry point or the front end of the cloud computing platform. It is used to collect the second monitoring data source after it has been aggregated by the sub-monitoring processing terminal, compare it with the original features of the first interceptor terminal, calculate the data feature loss index, and then dynamically adjust the data aggregation strategy according to the feature loss, such as weighted averaging and removing outliers. If the data quality meets the standard, it is allowed to pass to the main monitoring processing terminal node; otherwise, a retransmission or repair process is triggered to ensure that the monitoring data of different partitions meets the global consistency, thereby improving the data reliability and analysis efficiency of distributed monitoring.
[0029] Step A30: Collect a first monitoring data source based on the first interception deployment processing terminal, and collect a second monitoring data source based on the second interception deployment processing terminal.
[0030] Preferably, the first interception deployment processing end directly collects data from distributed monitoring nodes, such as the data output of IoT devices, industrial sensors, cameras, or the nearest transmission jump point, ensuring that it obtains the original information that has not been modified by intermediate nodes, forming the first monitoring data source, including the original sensor readings, video stream data, device status signals, etc.; the second interception deployment processing end monitors the output ports of distributed monitoring processing terminal nodes, such as edge servers and regional gateways, ensuring that it obtains the structured data after intermediate layer processing, forming the second monitoring data source, including regional statistical reports, event summaries, feature vector sets, etc.
[0031] Furthermore, such as Figure 2 As shown, step A30 further includes step A31, where the sub-monitoring processing terminal node includes multi-level sub-monitoring processing terminal nodes, and the upper-level sub-monitoring processing terminal node is tree-connected to multiple lower-level sub-monitoring processing terminal nodes; step A32, wherein the first interception deployment processing terminal is located near the distributed monitoring node and is communicatively connected to multiple sub-monitoring processing terminal nodes at the selected level, for receiving the first monitoring data source from the multiple sub-monitoring processing terminal nodes at the selected level; step A33, the second interception deployment processing terminal is located near the main monitoring processing terminal node and is communicatively connected to multiple sub-monitoring processing terminal nodes at the selected level, for receiving the second monitoring data source from the multiple sub-monitoring processing terminal nodes at the selected level.
[0032] Preferably, the monitoring transmission topology network has a hierarchical tree structure, with the root node being the main monitoring processing terminal, and the intermediate nodes being multiple sub-monitoring processing terminals. Each sub-monitoring processing terminal node includes multiple levels of sub-monitoring processing terminal nodes, with each higher-level sub-monitoring processing terminal node connected in a tree structure to multiple lower-level sub-monitoring processing terminal nodes. The leaf nodes are distributed monitoring nodes. The first interception deployment processing terminal is located between the distributed monitoring node and the lowest-level sub-monitoring processing terminal, close to the distributed monitoring node, and communicates with multiple sub-monitoring processing terminal nodes at the selected level, while simultaneously receiving first monitoring data sources from multiple sub-monitoring processing terminal nodes at the same level. The second interception deployment processing terminal is located between the top-level sub-monitoring processing terminal node and the main monitoring processing terminal node, close to the main monitoring processing terminal node, and communicates with multiple sub-monitoring processing terminal nodes at the selected level, receiving second monitoring data sources from multiple sub-monitoring processing terminal nodes at the same level. Through hierarchical quality inspection and cross-layer feedback in a multi-level interception deployment, data traceability and optimized allocation of processing resources are achieved.
[0033] Furthermore, step A30 also includes the data transmission relationship including the transmission relationship between the main monitoring processing terminal node and the sub-monitoring processing terminal node, the transmission relationship between the upper-level sub-monitoring processing terminal node and the next-level sub-monitoring processing terminal node, and the transmission relationship between the sub-monitoring processing terminal node and the distributed monitoring node, and constructing a monitoring transmission topology network based on the data transmission relationship.
[0034] Preferably, the data transmission relationships include the transmission relationships between the main monitoring and processing terminal node and the sub-monitoring and processing terminal nodes, the transmission relationships between the upper-level sub-monitoring and processing terminal nodes and the next-level sub-monitoring and processing terminal nodes, and the transmission relationships between sub-monitoring and processing terminal nodes and distributed monitoring nodes. Specifically, a structured monitoring transmission topology network is constructed based on the data flow and interaction rules between nodes at different levels. This clearly defines the complete transmission path of data from the data acquisition end to the core processing end and ensures efficient and reliable data flow at each stage. Specifically, the transmission relationship between the main monitoring and processing terminal node and the sub-monitoring and processing terminal nodes defines the bidirectional data flow between the core decision-making layer and the regional aggregation layer, such as command issuance and statistical report uploading; the transmission relationship between the upper-level sub-monitoring and processing terminal node and the next-level sub-monitoring and processing terminal node, i.e., the cascaded transmission link between sub-monitoring terminals, is used to establish data aggregation pipelines between multi-level edge computing nodes; and the transmission relationship between sub-monitoring and processing terminal nodes and distributed monitoring nodes is used to standardize the transmission method of raw data from sensors or monitoring equipment to the nearest processing node. This enables efficient governance of monitoring data in complex environments.
[0035] Step A40: Establish a feature vector space template, compare the data feature vector loss of the first monitoring data source and the second monitoring data source based on the feature vector space template, output a data feature loss index, optimize the second monitoring data source according to the data feature loss index, and then enter the next-level monitoring processing terminal for processing.
[0036] Furthermore, such as Figure 3 As shown, step A40 further includes step A41, obtaining the data feature vector of the second monitoring data source; step A42, establishing a feature vector space template based on the data feature vector of the second monitoring data source as a template; step A43, performing feature vector mapping on the first monitoring data source according to the feature vector space template, and outputting the data feature vector of the first monitoring data source; step A44, performing a loss comparison between the data feature vector of the second monitoring data source and the data feature vector of the first monitoring data source, and outputting a data feature loss index.
[0037] Preferably, a feature dimension is randomly selected from the second monitoring data source and combined to construct a data feature vector to establish a feature vector space template. This template is used to define a reasonable range of data features and serve as a comparison benchmark. Specifically, key feature dimensions, such as time-domain statistics, frequency-domain energy, and spatial distribution, are selected according to business needs. These dimensions are then standardized and converted into vector form, which serves as the feature vector space template. An allowable deviation threshold is set for each feature dimension. Then, the first monitoring data source is mapped to the feature vector space according to the feature vector space template. This means that the first monitoring data source is mapped to the feature vector space according to the same feature extraction rules, and the feature vector of the first monitoring data source is output.
[0038] Preferably, the loss of data feature vectors of the first monitoring data source and the second monitoring data source is compared based on the feature vector space template. That is, the loss of data feature vectors of the first monitoring data source and the second monitoring data source is compared. The deviation of the two feature vectors is calculated by KL divergence evaluation. It is checked whether the deviation of each feature exceeds the threshold, and then a comprehensive data feature loss index is output to quantify the degree of information loss of data in the transmission and aggregation process, such as a loss rate of 12%. Finally, the second monitoring data source is optimized according to the data feature loss index before entering the next-level monitoring processing terminal for processing. This involves dynamically adjusting the data flow strategy to ensure that the aggregated data entering the main monitoring processing terminal node simultaneously satisfies both monitoring information integrity and data processing efficiency. Specifically, the second monitoring data source is optimized based on the data feature loss index. This includes extracting corresponding feature values from the first monitoring data source and directly injecting them into the second data source when certain features are severely lost; triggering the sub-monitoring processing terminal nodes to re-execute monitoring data aggregation using more conservative parameters, such as shortening the time window and reducing the dimensionality reduction, to process the raw monitoring data collected by the distributed monitoring nodes; simultaneously enabling high-quality data to be directly transmitted to the main monitoring processing terminal node via a fast channel, while transferring low-quality data to a verification queue; thereby achieving hierarchical processing of monitoring data, ensuring low-latency processing of monitoring data, effectively balancing data quality and efficiency, and guaranteeing the consistency and reliability of monitoring data.
[0039] Furthermore, step A40 also includes step A45, identifying the interval levels of the first interception deployment processing terminal and the second interception deployment processing terminal; step A46, when the interval level is greater than a preset interval level, establishing a feature vector space template, wherein the feature vector space template is a feature vector space between the feature dimensions of the first monitoring data source and the feature dimensions of the second monitoring data source; step A47, performing high-dimensional feature vector mapping on the first monitoring data source according to the feature vector space template, and outputting the data feature vector of the first monitoring data source; step A48, performing low-dimensional feature vector mapping on the second monitoring data source according to the feature vector space template, and outputting the data feature vector of the second monitoring data source; step A49, performing a loss comparison between the data feature vector of the second monitoring data source and the data feature vector of the first monitoring data source, and outputting a data feature loss index.
[0040] Preferably, the number of interval levels between the first and second interception deployment processing terminals is automatically calculated and identified through topology analysis. The number of interval levels refers to the number of data processing layers between the first and second interception terminals. If the number of interval levels is greater than the preset number of interval levels, it indicates that the data has undergone multiple transformations, and a feature space template needs to be established to prevent semantic distortion. The feature vector space template is a feature vector space located between the feature dimensions of the first and second monitoring data sources, and is used to establish a mathematical mapping relationship between the high-dimensional features of the original data and the low-dimensional features of the aggregated data.
[0041] Preferably, the first monitoring data source undergoes high-dimensional feature vector mapping according to the feature vector space template. This involves extracting multi-scale features from the first monitoring data source using a feature pyramid algorithm and applying dimensionality compression rules specified by the template, such as retaining the first 10 dimensional components. The resulting simplified feature vector containing key information is then output as the data feature vector of the first monitoring data source. Similarly, the second monitoring data source undergoes low-dimensional feature vector mapping according to the feature vector space template. This involves recovering the implicit dimensional features of the second monitoring data source based on the feature expansion rules of the feature vector space template, reconstructing missing features using interpolation or generative models, and outputting an enhanced, comparable feature vector as the data feature vector of the second monitoring data source. Finally, the data feature vectors of the second and first monitoring data sources are compared for loss. This includes calculating the numerical deviation of features with the same dimensionality and checking whether the relative relationships of the features remain consistent. A data feature loss index is then output to ensure a precise balance between data availability and transmission efficiency regardless of how many layers the monitoring transmission topology network expands.
[0042] Furthermore, step A49 also includes step A491, calculating a similarity metric between the data feature vector of the second monitoring data source and the data feature vector of the first monitoring data source, wherein the similarity metric is obtained through KL divergence calculation; step A492, calculating a key feature similarity metric between the data feature vector of the second monitoring data source and the data feature vector of the first monitoring data source, wherein the key feature similarity metric is obtained through principal component loss rate calculation; and step A493, performing weighted calculations using the similarity metric and the key feature similarity metric, and outputting a data feature loss index.
[0043] Preferably, the similarity metric between the data feature vectors of the second monitoring data source and the data feature vectors of the first monitoring data source is obtained by KL divergence calculation to measure the difference in probability distribution between the two data feature vectors and reflect the similarity of the overall feature structure. Specifically, the data feature vectors of the first monitoring data source are converted into probability distributions by softmax normalization, and the data feature vectors of the second monitoring data source are converted into probability distributions in the same way. Then, the relative entropy of the probability distribution of the first monitoring data source to the probability distribution of the second monitoring data source is calculated by KL divergence as a similarity metric. If the relative entropy is 0, it means that the two data feature vectors are completely identical. Otherwise, there is information loss, and the larger the relative entropy, the more significant the difference. Finally, the global distribution difference between the features is obtained.
[0044] Preferably, principal component analysis (PCA) is used to identify key feature dimensions in the data and quantify the degree to which the aggregation process retains the main features. Specifically, PCA is performed on the first monitoring data source to determine the top N principal components with the highest contribution rates and record the variance contribution rate of each principal component, where N is a positive integer representing the number of key feature dimensions, for example, the top 3 components carry 80% of the information. Then, the second monitoring data source is projected onto the same principal component space, and the rate of change of the variance of each principal component is calculated. Finally, the key feature similarity metric is determined. Then, the similarity metric and the key feature similarity metric are weighted and a data feature loss index is output. The similarity metric reflects the overall data fidelity requirement and has a default weight of 0.4. The key feature similarity metric reflects the importance of key features and has a default weight of 0.6. When anomalies in key features are detected, the weight of the key feature similarity metric is automatically increased.
[0045] Furthermore, step A40 also includes step A410, where, when the data feature loss index is greater than a preset data feature loss index, the data feature vectors of the first monitoring data source and the second monitoring data source are fused and repaired to obtain a third monitoring data source; step A411, the third monitoring data source is compared with the core fields, key features, and time-series trajectories of the first monitoring data source and the second monitoring data source, and a consistency verification result is output; step A412, if the consistency verification result passes, the third monitoring data source enters the next-level monitoring processing terminal for processing; step A413, if the consistency verification result fails, a re-fused and repaired third monitoring data source is obtained.
[0046] Preferably, when the detected data feature loss index exceeds the preset data feature loss index, such as when the data feature loss index is greater than the preset threshold, it indicates that the information in the second monitoring data source is seriously distorted due to transmission or processing. At this point, the fusion repair process is initiated, which involves fusing and repairing the data feature vectors of the first monitoring data source and the second monitoring data source. Specifically, the lost feature dimensions in the second data source are extracted from the first monitoring data source and injected into the second data source proportionally to obtain the third monitoring data source. Then, the core fields, key features, and time series trajectories of the third monitoring data source are compared with those of the first and second monitoring data sources. Core field verification refers to checking whether the basic statistics are logically consistent, including whether the mean / extreme value after repair is between the original and aggregated data, and whether the numerical range conforms to physical laws. Key feature comparison refers to using principal component analysis for verification, that is, verifying whether the principal component variance recovery rate of the repaired third monitoring data source is ≥90%. Time series trajectory verification refers to calculating and determining whether the path deviation between the original time series and the repaired time series is less than the threshold and whether the mutation points are aligned through dynamic time warping algorithm.
[0047] Preferably, if the scores of all dimensions meet the preset thresholds, such as compliance of core fields and consistency of key features, the consistency verification result is passed; otherwise, it is failed. If the consistency verification result is passed, the third monitoring data source enters the next-level monitoring processing terminal for processing; if the consistency verification result is failed, the failed items are identified based on the verification result, then the precision of the time-series alignment sliding window is increased and the weight ratio of the original data in the fusion is increased, the fusion is re-performed and repaired, the third monitoring data source is generated, and then local verification is performed on the failed items to reduce computational overhead. This ensures a reduction in the feature loss rate of monitoring data and improves the consistency and reliability of monitoring data.
[0048] The big data analysis-based monitoring data processing system provided in this embodiment of the invention can execute the big data analysis-based monitoring data processing method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0049] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, showing a block diagram of an exemplary electronic device suitable for implementing the embodiments of the present invention. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present invention. This electronic device is in the form of a general-purpose computing device, and its components may include, but are not limited to, an input device 401, a processor 402, a memory 403, and an output device 404. The processor 402 may be one or more; the memory 403 may include a computer-readable medium and at least one program product having a set (at least one) of program modules configured to perform the functions of the embodiments of this application.
[0050] The memory 403 shown in this embodiment of the invention can be any combination of one or more computer-readable media. The computer-readable storage medium can be, but is not limited to, infrared, semiconductor systems, devices or components, or any combination thereof, for storing software programs, computer-executable programs and modules, such as the program instructions / modules corresponding to the big data analysis-based monitoring data processing method in this embodiment of the invention. The processor 402 executes various functional applications and data processing of the computer device by running the software programs, instructions and modules stored in the memory 403, thereby realizing the above-mentioned big data analysis-based monitoring data processing method.
[0051] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A monitoring data processing method based on big data analysis, characterized in that, The method includes: Based on the main monitoring processing terminal node, the sub-monitoring processing terminal nodes, the distributed monitoring nodes, and the data transmission relationships of each node, a monitoring transmission topology network is constructed. A first interception deployment processing terminal and a second interception deployment processing terminal are set in the monitoring transmission topology network, wherein the first interception deployment processing terminal is set near the distributed monitoring node, and the second interception deployment processing terminal is set near the main monitoring processing terminal node. The first monitoring data source is collected based on the first interception deployment processing terminal, and the second monitoring data source is collected based on the second interception deployment processing terminal; A feature vector space template is established. Based on the feature vector space template, the data feature vector loss of the first monitoring data source and the second monitoring data source is compared. A data feature loss index is output. The second monitoring data source is optimized according to the data feature loss index and then entered into the next-level monitoring processing terminal for processing. When the data feature loss index is greater than the preset data feature loss index, the data feature vector of the first monitoring data source and the second monitoring data source are fused and repaired to obtain the third monitoring data source. The third monitoring data source is compared with the core fields, key features and time-series trajectories of the first and second monitoring data sources, and the consistency verification result is output. If the consistency verification result passes, the third monitoring data source enters the next-level monitoring processing terminal for processing. If the consistency verification result fails, a third monitoring data source that has been re-integrated and repaired is obtained.
2. The monitoring data processing method based on big data analysis as described in claim 1, characterized in that, The sub-monitoring and processing terminal node includes multi-level sub-monitoring and processing terminal nodes, and the upper-level sub-monitoring and processing terminal node is connected to multiple lower-level sub-monitoring and processing terminal nodes in a tree structure. The first interception deployment processing terminal is located near the distributed monitoring node and is connected to multiple sub-monitoring processing terminal nodes under the selected level to receive the first monitoring data source from the multiple sub-monitoring processing terminal nodes under the selected level. The second interception deployment processing terminal is located near the main monitoring processing terminal node and is communicatively connected to multiple sub-monitoring processing terminal nodes under the selected level, for receiving the second monitoring data source from the multiple sub-monitoring processing terminal nodes under the selected level.
3. The monitoring data processing method based on big data analysis as described in claim 2, characterized in that, The data transmission relationships include the transmission relationship between the main monitoring and processing terminal node and the sub-monitoring and processing terminal nodes, the transmission relationship between the upper-level sub-monitoring and processing terminal node and the next-level sub-monitoring and processing terminal node, and the transmission relationship between the sub-monitoring and processing terminal node and the distributed monitoring node. A monitoring transmission topology network is constructed based on the data transmission relationships.
4. The monitoring data processing method based on big data analysis as described in claim 1, characterized in that, The method involves comparing the data feature vector loss of the first monitoring data source and the second monitoring data source based on the feature vector space template, including: Obtain the data feature vector of the second monitoring data source; A feature vector space template is established based on the data feature vector of the second monitoring data source. The first monitoring data source is mapped to a feature vector according to the feature vector space template, and the data feature vector of the first monitoring data source is output. The loss of the data feature vectors from the second monitoring data source and the first monitoring data source is compared, and the data feature loss index is output.
5. The monitoring data processing method based on big data analysis as described in claim 1, characterized in that, The method further includes comparing the data feature vector loss of the first monitoring data source and the second monitoring data source based on the feature vector space template: Identify the interval level between the first interception deployment processing terminal and the second interception deployment processing terminal; When the interval level is greater than the preset interval level, a feature vector space template is established. The feature vector space template is a feature vector space between the feature dimensions of the first monitoring data source and the feature dimensions of the second monitoring data source. The first monitoring data source is mapped to a high-dimensional feature vector according to the feature vector space template, and the data feature vector of the first monitoring data source is output. The second monitoring data source is mapped to a low-dimensional feature vector according to the feature vector space template, and the data feature vector of the second monitoring data source is output. The loss of the data feature vectors from the second monitoring data source and the first monitoring data source is compared, and the data feature loss index is output.
6. The monitoring data processing method based on big data analysis as described in claim 4, characterized in that, The method involves comparing the loss of the data feature vectors from the second monitoring data source and the first monitoring data source, and outputting a data feature loss index. Calculate the similarity metric between the data feature vector of the second monitoring data source and the data feature vector of the first monitoring data source. The similarity metric is obtained by KL divergence calculation. Calculate the key feature similarity measure between the data feature vector of the second monitoring data source and the data feature vector of the first monitoring data source. The key feature similarity measure is obtained by calculating the principal component loss rate. The similarity metric and the key feature similarity metric are weighted and the data feature loss index is output.
7. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable instructions; The processor, when executing executable instructions stored in the memory, implements the monitoring data processing method based on big data analysis as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Cloud network traffic monitoring system based on two-stage architecture
CN111683097A
Computer information security processing method and system based on big data
CN113626807A