Data analysis method, device and equipment

By creating a distributed cluster and file system, combined with a load balancing strategy, we solved the high missed acquisition rate and uneven load issues on northbound servers, and achieved an efficient and reliable data parsing process.

CN120751008APending Publication Date: 2025-10-03INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510675575.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In the existing technology, the northbound server adopts a decentralized data collection method, resulting in a high missed collection rate and difficulty in locating missed collections. It also lacks an effective load balancing mechanism, leading to uneven load and low resource utilization.

Method used

Create a distributed cluster, mount the nodes to the distributed file system, store data files through cyclic redundancy check, execute parsing tasks in parallel, dynamically allocate tasks based on load balancing strategies, and perform data continuity and cell coverage verification.

Benefits of technology

It realizes the centralized storage and management of massive data files, reduces the missed acquisition rate, simplifies problem location, dynamically balances node loads, improves resource utilization, and ensures data integrity and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751008A_ABST
    Figure CN120751008A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and provides a data analysis method, device and equipment, and the method comprises the steps: creating a distributed cluster, and mounting nodes in the distributed cluster to a distributed file system; data files downloaded from network management servers of all manufacturers are stored in the distributed file system, and the data files pass cyclic redundancy check; reading and analyzing the data file corresponding to the network management server of each manufacturer from the distributed file system based on the nodes in the distributed cluster to obtain analysis result data corresponding to the network management server of each manufacturer; and storing the analysis result data corresponding to the network management server of each manufacturer to a local database, and performing data continuity verification and cell coverage rate link ratio verification on the analysis result data. The whole data analysis process is more efficient and reliable through the distributed file system and the distributed cluster in combination with a three-level data integrity verification mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a data analysis method, device and equipment. Background Art

[0002] In the operation and optimization of modern communication networks, data analysis is a key component for network performance monitoring, fault diagnosis, and optimization. Northbound servers are a critical component in modern communication networks, responsible for providing network device data and management interfaces to higher-level systems (such as big data platforms and network optimization systems). Operators' network optimization services handle a wide variety of data files. These files are generated by base station equipment from numerous manufacturers and distributed across numerous network management servers. The sheer volume of data leads to high storage costs, poor data consistency, and complex management. With the continuous expansion of networks, the diversification of base station equipment manufacturers, and the explosive growth of data volumes, existing data analysis faces numerous challenges.

[0003] The current data collection and analysis process typically relies on northbound servers to first download data files from these network management servers, then decompress the data files and finally parse them. Because northbound servers utilize a decentralized data collection approach, missed data collection rates are high and difficult to locate. Furthermore, the lack of an effective load balancing mechanism can lead to overloaded servers while idle resources on other servers, resulting in uneven load and low resource utilization. Summary of the Invention

[0004] The present invention provides a data analysis method, device and equipment to solve the technical problems in the prior art that the northbound server adopts a decentralized data collection method, resulting in a high missed collection rate and difficulty in locating missed collections, and lacks an effective load balancing mechanism, resulting in uneven load.

[0005] The present invention provides a data analysis method, comprising the following steps: Creating a distributed cluster and mounting the nodes in the distributed cluster to a distributed file system; wherein the distributed file system stores data files downloaded from network management servers of all manufacturers, and the data files pass a cyclic redundancy check; Calling the nodes in the distributed cluster to execute in parallel the parsing task of the data files corresponding to the network management servers of each manufacturer in the distributed file system to obtain the parsing result data corresponding to the network management servers of each manufacturer; The analysis result data corresponding to each manufacturer's network management server is stored in a local database, and data continuity verification and cell coverage ratio ring verification are performed on the analysis result data.

[0006] According to a data parsing method provided by the present invention, calling the nodes in the distributed cluster to execute in parallel the parsing tasks of the data files corresponding to the network management servers of each manufacturer in the distributed file system, includes: When a file write close event is detected under a file path in a distributed file system, a file integrity check is performed on the data file corresponding to the file write close event; When the data file corresponding to the file write close event passes the file integrity check, a parsing task for the data file corresponding to the file write close event is created.

[0007] According to a data parsing method provided by the present invention, before monitoring a file write close event under a file path in a distributed file system, the method further includes: Define a file monitor class, which is used to monitor file write close events in a distributed file system; Initialize the file monitor class to create an observer object and an event handler, and associate the event handler with a file path in a distributed file system.

[0008] According to a data parsing method provided by the present invention, calling the nodes in the distributed cluster to execute in parallel the parsing tasks of the data files corresponding to the network management servers of each manufacturer in the distributed file system, includes: After creating a parsing task for a corresponding data file, determining a target node corresponding to the parsing task based on the load status of the nodes in the distributed cluster; The parsing task is assigned to the thread pool of the target node to process the parsing task based on the thread in the thread pool of the target node, and after detecting that the parsing task is completed, the task status corresponding to the parsing task is updated to a completed status.

[0009] According to a data parsing method provided by the present invention, the load status includes the length of the pending task queue and the number of parsing tasks completed per unit time; determining the target node corresponding to the parsing task based on the load status of the nodes in the distributed cluster includes: Determine the length of the pending task queue of the nodes in the current distributed cluster and the number of parsing tasks completed per unit time; Based on the length of the pending task queue and the number of completed parsing tasks, a weighted scoring strategy is adopted to determine the load status score of the nodes in the distributed cluster; the weight coefficient corresponding to the length of the pending task queue is positive, and the weight coefficient corresponding to the number of completed parsing tasks is negative; The node with the smallest load status score is determined as the target node corresponding to the parsing task.

[0010] According to a data parsing method provided by the present invention, the load status includes the length of a queue of pending tasks; after determining a target node corresponding to the parsing task based on the load status of the nodes in the distributed cluster, the method further includes: When it is detected that the length of the pending task queue of the target node is greater than a first task queue length threshold, determining the expansion quantity corresponding to each node in the distributed cluster based on the length of the pending task queue of each node in the distributed cluster; Create a corresponding number of expanded Pod instances for each node in the distributed cluster, and allocate the created Pod instances to the corresponding nodes in the distributed cluster; When it is detected that the length of the pending task queue of the target node is less than the second task queue length threshold, determining all nodes to be shrunk in the distributed cluster; the length of the pending task queue of the node to be shrunk is less than the second task queue length threshold; Based on the length of the pending task queue of each node to be shrunk, the corresponding shrinkage quantity of each node to be shrunk is determined, and the corresponding shrinkage quantity of Pod instances is released for each node to be shrunk.

[0011] According to a data analysis method provided by the present invention, data continuity verification is performed on the analysis result data, including: For each manufacturer's gateway server, detecting at every preset time interval whether the local database has stored the parsing result data corresponding to the manufacturer's gateway server within the preset time interval; When it is detected that the analysis result data corresponding to the gateway server of the manufacturer is stored in the local database within the preset time period, it is determined that the data continuity check has passed.

[0012] According to a data analysis method provided by the present invention, performing a cell coverage ratio check on the analysis result data includes: After the analysis result data corresponding to the network management servers of all manufacturers are stored every day, determining the cell coverage rate of the analysis result data corresponding to the network management servers of all manufacturers on that day and the analysis result data corresponding to the network management servers of all manufacturers on the previous day; When the cell coverage ratio verification is within a preset cell coverage ratio verification range, it is determined that the cell coverage ratio verification is passed.

[0013] The present invention also provides a data analysis device, comprising: A first data parsing module is configured to create a distributed cluster and mount the nodes in the distributed cluster to a distributed file system; wherein the distributed file system stores data files downloaded from network management servers of all manufacturers, and the data files pass a cyclic redundancy check; The second data parsing module is used to call the nodes in the distributed cluster to execute the parsing task of the data files corresponding to the network management servers of each manufacturer in the distributed file system in parallel, and obtain the parsing result data corresponding to the network management servers of each manufacturer; The third data analysis module is used to store the analysis result data corresponding to each manufacturer's network management server in a local database, and perform data continuity verification and cell coverage ratio ring verification on the analysis result data.

[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the data parsing method as described above is implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-described data parsing methods when executed by a processor.

[0016] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned data parsing methods.

[0017] The data parsing method, device and equipment provided by the present invention realize the centralized storage and management of massive data files by creating a distributed cluster and mounting the nodes to the distributed file system, and ensure the data integrity by using cyclic redundancy check. Furthermore, the distributed cluster nodes read the parsed data from the distributed file system in parallel, replacing the traditional decentralized collection mode, reducing the missed acquisition rate and simplifying the problem location. In addition, the distributed cluster can dynamically allocate parsing tasks, balance the node load, improve resource utilization, and the parsed data is again subjected to data continuity verification and cell coverage verification. The present invention uses a distributed file system and a distributed cluster, combined with a three-level data integrity verification mechanism, to make the entire data parsing process more efficient and reliable. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 This is one of the flow charts of the data analysis method provided by the present invention.

[0020] Figure 2 This is the second flow chart of the data analysis method provided by the present invention.

[0021] Figure 3 It is a structural diagram of the data analysis device provided by the present invention.

[0022] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0023] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0024] The data parsing method of the embodiment of the present invention is as follows: Figure 1 As shown, the process includes step 110 , step 120 and step 130 .

[0025] Step 110: Create a distributed cluster and mount the nodes in the distributed cluster to a distributed file system; wherein the distributed file system stores data files downloaded from network management servers of all manufacturers, and the data files pass a cyclic redundancy check.

[0026] In this embodiment, a distributed cluster serves as the parsing orchestration platform, responsible for scheduling and managing parsing tasks, and a distributed file system (Ceph File System, CephFS) serves as the file storage container. In a distributed cluster, each node (such as a server) is connected to CephFS via a network, and CephFS is accessed and managed as a unified storage resource.

[0027] It should be understood here that each manufacturer's network management server is manufactured by a different manufacturer, and therefore the base station data types involved in the data files of each manufacturer's network management server are also different.

[0028] In practice, you can create a Kubernetes cluster, which consists of a Master node (the control center of the Kubernetes cluster, responsible for managing the entire cluster) and Worker nodes (nodes that run application containers). The nodes are connected via a high-speed network, enabling message transmission and data exchange. Specifically, you can mount the Worker nodes in a Kubernetes cluster to CephFS.

[0029] Mounting is an operating system operation that attaches a file system (such as CephFS) to a directory (mount point), making the files and directories in that directory accessible. In a distributed cluster, each node attaches CephFS to a directory on its local file system through a mount operation.

[0030] For example, nodes in a distributed cluster can execute the "mount -t ceph : / / mnt / cephfs -oname=admin,secret=..." command to mount each node to CephFS, ensuring that all nodes can directly access unified storage resources.

[0031] It should be understood that CephFS is a component of the Ceph storage cluster. It is built on Ceph's distributed object storage (RADOS) and provides a POSIX-compliant file system interface. Data stored in CephFS is distributed across multiple nodes, supporting horizontal scalability and allowing multiple clients to mount and access the same file system simultaneously.

[0032] In this embodiment, a corresponding CephFS volume is pre-created in the Ceph cluster for each manufacturer's network management server. The CephFS volume can be mounted on the network management server to store data files. Furthermore, the CephFS volume can also be mounted on nodes (servers) in the distributed cluster to allow the nodes in the distributed cluster to access the stored data files. After the corresponding CephFS volume is created, the storage pool to be used is specified.

[0033] Based on the above configuration, data files are downloaded periodically from all manufacturers' network management servers and stored in CephFS. It should be noted that once the downloaded data files are stored in the corresponding CephFS directory, the file mapping information can be viewed using CephFS tools to determine the file path and storage directory of the data files.

[0034] Furthermore, to ensure data storage integrity and consistency, a cyclic redundancy check (CRC) is performed when data files are stored in CephFS. CRC is an error detection technique used to detect errors during data transmission or storage. For example, a CRC32 checksum can be used. When data files are written from the network management server to CephFS, the CRC32 checksum verifies whether data blocks have been corrupted during transmission or storage.

[0035] Step 120 : Calling the nodes in the distributed cluster to execute the parsing task of the data files corresponding to the network management servers of each manufacturer in CEPHFS in parallel, and obtaining the parsing result data corresponding to the network management servers of each manufacturer.

[0036] Here, we continue to use the K8S cluster as an example to explain. Specifically, the Master node in the K8S cluster can monitor the load status of each Worker node in real time. Based on the load status of each Worker node, it calls the Worker node in the K8S cluster to access and resolve the file paths and data files in the storage directory mounted by the Worker node.

[0037] In this embodiment, the data analysis method can be flexibly adjusted based on actual data analysis requirements and is not limited thereto.

[0038] Step 130: Store the analysis result data corresponding to each manufacturer's network management server in a local database, and perform data continuity verification and cell coverage ratio ring verification on the analysis result data.

[0039] After obtaining the parsing result data corresponding to each manufacturer's network management server, the parsing result data is loaded into a local database, such as a StarRocks MPP database or a Doris MPP database.

[0040] For easier understanding, refer to Figure 2 , Figure 2 This is a second flow diagram of the data parsing method provided by an embodiment of the present invention. As shown, a unified collection program is used to download data files from network management servers of different manufacturers and store them in CephFS. The parsing servers (i.e., Worker nodes) in the Kubernetes cluster mount a unified CephFS directory. The parsing servers directly parse the data files stored in CephFS using a no-download process and store the parsed data in a local analytical database.

[0041] Furthermore, after the data is stored in the local database, a double verification mechanism is adopted for the data in the local database to reduce the data leakage rate.

[0042] Data continuity verification refers to checking the timestamps or sampling points in data records to ensure data continuity. Cell coverage verification involves performing daily comparisons of cell coverage or performance indicators. This macroscopic analysis detects any unusual changes in coverage or performance indicators, and further verifies whether data has been missed.

[0043] The data parsing method of this embodiment realizes the centralized storage and management of massive data files by creating a distributed cluster and mounting the nodes to a distributed file system, and ensures data integrity by using cyclic redundancy check. Furthermore, the distributed cluster nodes read the parsed data from the distributed file system in parallel, replacing the traditional decentralized collection mode, reducing the missed acquisition rate and simplifying problem location. In addition, the distributed cluster can dynamically allocate parsing tasks, balance the node load, improve resource utilization, and the parsed data is again subjected to data continuity verification and cell coverage verification. The present invention uses a distributed file system and a distributed cluster, combined with a three-level data integrity verification mechanism, to make the entire data parsing process more efficient and reliable.

[0044] It should be noted that each implementation method of the present application can be freely combined, the order can be changed, or it can be executed separately, and does not need to rely on or depend on a fixed execution order.

[0045] In some embodiments, performing data continuity verification on the parsing result data includes: For each manufacturer's gateway server, detecting at every preset time interval whether the local database has stored the parsing result data corresponding to the manufacturer's gateway server within the preset time interval; When it is detected that the analysis result data corresponding to the gateway server of the manufacturer is stored in the local database within the preset time period, it is determined that the data continuity check has passed.

[0046] Specifically, the continuity and integrity of the timing data of each manufacturer's network management server are verified.

[0047] It should be understood that the data of the network management server is continuously generated, so in this embodiment, data files are continuously downloaded from each manufacturer's network management server and stored in CephFS every preset period. Based on this, in this embodiment, data continuity verification is performed on the data parsed by the nodes in the distributed cluster.

[0048] For example, every 15 minutes, the local database is checked to see if the parsing result data corresponding to the manufacturer's gateway server has been stored again. If the local database detects that the parsing result data corresponding to the manufacturer's gateway server has been stored again within the 15-minute period, then the parsing result data corresponding to the manufacturer's gateway server is determined to be present. Otherwise, data is marked as missing during this period.

[0049] The data analysis method of this embodiment verifies the continuity and integrity of the time series data of each manufacturer's network management server, thereby promptly discovering possible interruptions or anomalies in the data collection process and accurately locating data loss.

[0050] In some embodiments, performing a cell coverage comparison check on the parsed result data includes: After the analysis result data corresponding to the network management servers of all manufacturers are stored every day, determining the cell coverage rate of the analysis result data corresponding to the network management servers of all manufacturers on that day and the analysis result data corresponding to the network management servers of all manufacturers on the previous day; When the cell coverage ratio verification is within a preset cell coverage ratio verification range, it is determined that the cell coverage ratio verification is passed.

[0051] Specifically, at the community business logic level, the data integrity of each manufacturer's network management server is verified once a day.

[0052] It should be understood that by parsing the raw data reported by the network management server, the cell coverage rate can be generated. The cell coverage rate is a key performance indicator in communication networks (such as 4G / 5G) and is used to measure the signal coverage quality of a specific base station (cell).

[0053] In a normally operating mobile network, cell coverage typically fluctuates slowly rather than experiencing dramatic jumps. Therefore, in this embodiment, after the parsing result data corresponding to the network management servers of all manufacturers is stored daily, a month-on-month verification is performed based on the cell coverage in the parsing result data. If the month-on-month cell coverage is within a preset month-on-month cell coverage range, then the current data collection is determined to be normal, with no missed data. If the month-on-month cell coverage exceeds the preset month-on-month cell coverage range, for example, if the minimum month-on-month coverage of the preset month-on-month cell coverage range is 90%, and the month-on-month cell coverage on that day is 70%, then the data collection is determined to be abnormal (e.g., missed data).

[0054] In some embodiments, calling the nodes in the distributed cluster to perform in parallel the parsing task of the data files corresponding to the network management servers of each manufacturer in the distributed file system includes: When a file write close event is detected under a file path in a distributed file system, a file integrity check is performed on the data file corresponding to the file write close event; When the data file corresponding to the file write close event passes the file integrity check, a parsing task for the data file corresponding to the file write close event is created.

[0055] In this embodiment, the file monitoring mechanism and the file integrity check can be combined to call the nodes in the distributed cluster to perform the parsing task.

[0056] In this step, when a file write close event occurs under the file path in CephFS, the kernel will add the file write close event to the queue and call the nodes in the distributed cluster to execute the file write close event in the queue in parallel.

[0057] It should be understood that in CephFS, a file path refers to the full path of a file in CephFS. The file path starts from the mount point and points to a specific file or directory in CephFS. A file write close event is triggered when a file is closed after a write operation is completed.

[0058] In CephFS, you can use the inotify mechanism to monitor file write close events under a file path. Inotify is a file system event monitoring mechanism provided by the Linux kernel, allowing programs to monitor various file system events. Inotify maintains an event queue in the kernel. When a specified file system event occurs, the kernel adds the event information to the queue.

[0059] After detecting a file write close event, the system first performs a file integrity check on the corresponding data file. Specifically, it verifies the integrity and consistency of the file to ensure that the file was not damaged during the write process. For example, the MD5 (Message-Digest Algorithm 5) algorithm is used to verify data integrity. If the MD5 check passes, the data file is intact and undamaged.

[0060] After the data file corresponding to the file write close event passes the file integrity check, the download-free parsing process is triggered, and a parsing task for the data file corresponding to this file write close event is created. The file is directly parsed without downloading the file first.

[0061] In some embodiments, before monitoring a file write close event under a file path in a distributed file system, the method further includes: Define a file monitor class, which is used to monitor file write close events in a distributed file system; Initialize the file monitor class to create an observer object and an event handler, and associate the event handler with a file path in a distributed file system.

[0062] Specifically, first customize a file monitor class to encapsulate the file monitoring logic based on inotify, such as customizing a FileWatcher class.

[0063] Next, initialize the file monitor class. In the file monitor class's initialization method, create an Observer object to monitor file system changes. Then, create a CephEventHandler event handler to define the specific event processing logic. Finally, use the Observer object's schedule method to associate the CephEventHandler with a specified file path in CephFS. Here, the specified CephFS file path refers to the file path of the data file to be monitored.

[0064] In this embodiment, the above configuration enables calling the start method of the FileWatcher class, starting the Observer object and monitoring file system changes under the specified file path. When a file is closed and a write operation has occurred before, a file write close event is triggered. The Observer object captures the event and notifies the CephEventHandler. In the CephEventHandler's on_close_write method, an MD5 checksum is performed on the data file corresponding to the file write close event. If the MD5 checksum passes, the download-free parsing process is triggered, and the file is directly parsed.

[0065] The data parsing method of this embodiment uses a file monitoring mechanism to promptly detect changes in data files and performs a file integrity check on the data file after the file is written to ensure that the data file to be subsequently parsed is complete and undamaged. Finally, after the data file passes the file integrity check, the download-free parsing process is triggered, improving file processing efficiency and reducing unnecessary file downloads.

[0066] In some embodiments, calling the nodes in the distributed cluster to perform in parallel the parsing task of the data files corresponding to the network management servers of each manufacturer in the distributed file system includes: After creating a parsing task for a corresponding data file, determining a target node corresponding to the parsing task based on the load status of the nodes in the distributed cluster; The parsing task is assigned to the thread pool of the target node to process the parsing task based on the thread in the thread pool of the target node, and after detecting that the parsing task is completed, the task status corresponding to the parsing task is updated to a completed status.

[0067] In this embodiment, after creating a parsing task for a corresponding data file, a target node corresponding to the parsing task may be determined based on the load status of the nodes in the distributed cluster using a pre-configured load balancing strategy.

[0068] After the target node is determined, the parsing task is assigned to the thread pool of the target node, and the target node processes the parsing task through the threads in the thread pool.

[0069] Here, you can use Executors.newFixedThreadPool(custom number of threads) to create a thread pool with a dynamically adjustable number of threads via the configuration file. This ensures sufficient thread resources for parsing tasks and avoids the overhead of frequent thread creation. When a new data file is detected under the file path in CephFS, a parsing task is automatically generated and submitted to the target node's thread pool. Each thread in the thread pool is responsible for parsing a data file. Through thread pool management, multiple data files can be parsed simultaneously without blocking each other.

[0070] Furthermore, after the data file is parsed, its parsing task status will be marked as processed, so that subsequent repeated processing requests can be ignored or skipped directly.

[0071] The data parsing method of this embodiment first determines the target node corresponding to the parsing task based on the load status of the nodes in the distributed cluster to ensure the load balance of the system, and then distributes the parsing task to the thread pool of the target node to achieve efficient parallel processing of tasks. Finally, after the file parsing is completed, the task status corresponding to this parsing task is updated to a completed status to ensure the idempotence of the task and avoid data inconsistency problems caused by repeated processing.

[0072] In some embodiments, the load status includes the length of the pending task queue and the number of parsing tasks completed per unit time; determining the target node corresponding to the parsing task based on the load status of the nodes in the distributed cluster includes: Determine the length of the pending task queue of the nodes in the current distributed cluster and the number of parsing tasks completed per unit time; Based on the length of the pending task queue and the number of completed parsing tasks, a weighted scoring strategy is adopted to determine the load status score of the nodes in the distributed cluster; the weight coefficient corresponding to the length of the pending task queue is positive, and the weight coefficient corresponding to the number of completed parsing tasks is negative; The node with the smallest load status score is determined as the target node corresponding to the parsing task.

[0073] In this embodiment, load balancing tasks are allocated by considering two dimensional parameters: the length of the pending task queue and the number of parsing tasks completed per unit time.

[0074] The length of the pending task queue reflects the number of tasks a node currently needs to process and is an important indicator of node load. The longer the pending task queue, the heavier the node's task load.

[0075] The number of parsing tasks completed per unit time reflects the processing efficiency of a node and is another important indicator for measuring node load. A higher number of parsing tasks completed per unit time indicates a stronger node's processing capability or more abundant resources.

[0076] Specifically, a node's load status score is comprehensively quantified by the length of the pending task queue, which positively reflects the load pressure, and the number of completed parsing tasks, which negatively offsets the load pressure. By assigning a positive weight to the length of the pending task queue and a negative weight to the number of completed parsing tasks, an increase in the length of the pending task queue (positive weight) increases the score, indicating an increase in the current load, while the number of completed parsing tasks (negative weight) decreases the score, reflecting a strong processing capability. The higher the final score, the heavier the node load. Therefore, in this embodiment, the node with the lowest load status score is selected as the target node for the current parsing task.

[0077] The data parsing method of this embodiment accurately identifies the node load status by combining the length of the pending task queue that positively reflects the load pressure and the number of completed parsing tasks that negatively offsets the load pressure, and then preferentially allocates parsing tasks to target nodes with light loads, thereby achieving dynamic balance of the system load.

[0078] In some embodiments, the load status includes the length of a pending task queue; after determining the target node corresponding to the parsing task based on the load status of the nodes in the distributed cluster, the method further includes: When it is detected that the length of the pending task queue of the target node is greater than a first task queue length threshold, determining the expansion quantity corresponding to each node in the distributed cluster based on the length of the pending task queue of each node in the distributed cluster; Create a corresponding number of expanded Pod instances for each node in the distributed cluster, and allocate the created Pod instances to the corresponding nodes in the distributed cluster; When it is detected that the length of the pending task queue of the target node is less than the second task queue length threshold, determining all nodes to be shrunk in the distributed cluster; the length of the pending task queue of the node to be shrunk is less than the second task queue length threshold; Based on the length of the pending task queue of each node to be shrunk, the corresponding shrinkage quantity of each node to be shrunk is determined, and the corresponding shrinkage quantity of Pod instances is released for each node to be shrunk.

[0079] In this embodiment, the load status of each node is monitored in real time, and dynamic scaling is performed based on the load status. Here, after assigning the parsing task to the target node with the least load, the target node is used as a reference to determine whether the entire cluster needs dynamic scaling.

[0080] Specifically, when it is detected that the length of the pending task queue of the target node is greater than the first task queue length threshold, since the target node is the node with the smallest load in the current distributed cluster, it indicates that all nodes in the current distributed cluster have insufficient resources. Then, based on the length of the pending task queue of each node in the current distributed cluster, the current required capacity expansion of each node is determined, and then the corresponding number of Pod instances are created for each node, and the created Pod instances are allocated to the corresponding nodes for their use.

[0081] Furthermore, when it is detected that the target node's pending task queue length is less than a second task queue length threshold, it indicates that at least one node in the current distributed cluster has surplus resources. All nodes to be scaled down whose pending task queue lengths are less than the second task queue length threshold are then screened from the current distributed cluster. Based on the pending task queue lengths of each node to be scaled down, the corresponding scaling quantity is determined for each node to be scaled down, and the corresponding number of Pod instances are released for each node to achieve optimal resource allocation.

[0082] The data parsing method of this embodiment ensures efficient operation of the system and rational use of resources by dynamically scaling the capacity according to the length of the pending task queue of the node.

[0083] The data analysis device provided by the present invention is described below. The data analysis device described below and the data analysis method described above can be referenced to each other.

[0084] The data analysis device of the embodiment of the present invention is as follows: Figure 3 As shown, the following modules are included: a first data analysis module 310 , a second data analysis module 320 and a third data analysis module 330 .

[0085] The first data analysis module 310 is used to create a distributed cluster and mount the nodes in the distributed cluster to a distributed file system; wherein the distributed file system stores data files downloaded from network management servers of all manufacturers, and the data files pass cyclic redundancy check.

[0086] The second data parsing module 320 is used to call the nodes in the distributed cluster to execute in parallel the parsing tasks of the data files corresponding to the network management servers of each manufacturer in the distributed file system, and obtain the parsing result data corresponding to the network management servers of each manufacturer.

[0087] The third data analysis module 330 is used to store the analysis result data corresponding to each manufacturer's network management server in a local database, and perform data continuity verification and cell coverage ratio ring verification on the analysis result data.

[0088] The data parsing device of this embodiment realizes the centralized storage and management of massive data files by creating a distributed cluster and mounting the nodes to the distributed file system, and ensures the data integrity by using cyclic redundancy check. Furthermore, the distributed cluster nodes read the parsed data from the distributed file system in parallel, replacing the traditional decentralized collection mode, reducing the missed acquisition rate and simplifying the problem location. In addition, the distributed cluster can dynamically allocate parsing tasks, balance the node load, improve resource utilization, and the parsed data is again subjected to data continuity verification and cell coverage verification. The present invention uses a distributed file system and a distributed cluster, combined with a three-level data integrity verification mechanism, to make the entire data parsing process more efficient and reliable.

[0089] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the data parsing method, which includes: Creating a distributed cluster and mounting the nodes in the distributed cluster to a distributed file system; wherein the distributed file system stores data files downloaded from network management servers of all manufacturers, and the data files pass a cyclic redundancy check; Calling the nodes in the distributed cluster to execute in parallel the parsing task of the data files corresponding to the network management servers of each manufacturer in the distributed file system to obtain the parsing result data corresponding to the network management servers of each manufacturer; The analysis result data corresponding to each manufacturer's network management server is stored in a local database, and data continuity verification and cell coverage ratio ring verification are performed on the analysis result data.

[0090] Furthermore, the logic instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes any medium capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0091] On the other hand, the present invention further provides a computer program product, comprising a computer program, which may be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the data parsing method provided in each of the above methods, the method comprising: Creating a distributed cluster and mounting the nodes in the distributed cluster to a distributed file system; wherein the distributed file system stores data files downloaded from network management servers of all manufacturers, and the data files pass a cyclic redundancy check; Calling the nodes in the distributed cluster to execute in parallel the parsing task of the data files corresponding to the network management servers of each manufacturer in the distributed file system to obtain the parsing result data corresponding to the network management servers of each manufacturer; The analysis result data corresponding to each manufacturer's network management server is stored in a local database, and data continuity verification and cell coverage ratio ring verification are performed on the analysis result data.

[0092] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the data parsing method provided in each of the above methods, the method comprising: Creating a distributed cluster and mounting the nodes in the distributed cluster to a distributed file system; wherein the distributed file system stores data files downloaded from network management servers of all manufacturers, and the data files pass a cyclic redundancy check; Calling the nodes in the distributed cluster to execute in parallel the parsing task of the data files corresponding to the network management servers of each manufacturer in the distributed file system to obtain the parsing result data corresponding to the network management servers of each manufacturer; The analysis result data corresponding to each manufacturer's network management server is stored in a local database, and data continuity verification and cell coverage ratio ring verification are performed on the analysis result data.

[0093] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0094] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or alternatively, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes instructions for enabling a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or portions thereof.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in each of the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A data analysis method, characterized in that: include: Creating a distributed cluster and mounting the nodes in the distributed cluster to a distributed file system; wherein the distributed file system stores data files downloaded from network management servers of all manufacturers, and the data files pass a cyclic redundancy check; Calling the nodes in the distributed cluster to execute in parallel the parsing task of the data files corresponding to the network management servers of each manufacturer in the distributed file system to obtain the parsing result data corresponding to the network management servers of each manufacturer; The analysis result data corresponding to each manufacturer's network management server is stored in a local database, and data continuity verification and cell coverage ratio ring verification are performed on the analysis result data.

2. The data analysis method according to claim 1, characterized in that: Calling the nodes in the distributed cluster to execute in parallel the parsing task of the data files corresponding to the network management servers of each manufacturer in the distributed file system includes: When a file write close event is detected under a file path in a distributed file system, a file integrity check is performed on the data file corresponding to the file write close event; When the data file corresponding to the file write close event passes the file integrity check, a parsing task for the data file corresponding to the file write close event is created.

3. The data analysis method according to claim 2, characterized in that: Before monitoring the file write close event under the file path in the distributed file system, the method further includes: Define a file monitor class, which is used to monitor file write close events in a distributed file system; Initialize the file monitor class to create an observer object and an event handler, and associate the event handler with a file path in a distributed file system.

4. The data analysis method according to claim 1, wherein: Calling the nodes in the distributed cluster to execute in parallel the parsing task of the data files corresponding to the network management servers of each manufacturer in the distributed file system includes: After creating a parsing task for a corresponding data file, determining a target node corresponding to the parsing task based on the load status of the nodes in the distributed cluster; The parsing task is assigned to the thread pool of the target node to process the parsing task based on the thread in the thread pool of the target node, and after detecting that the parsing task is completed, the task status corresponding to the parsing task is updated to a completed status.

5. The data analysis method according to claim 4, characterized in that: The load status includes the length of the pending task queue and the number of parsing tasks completed per unit time; Determining a target node corresponding to the parsing task based on a load status of a node in the distributed cluster includes: Determine the length of the pending task queue of the nodes in the current distributed cluster and the number of parsing tasks completed per unit time; Based on the length of the pending task queue and the number of completed parsing tasks, a weighted scoring strategy is adopted to determine the load status score of the nodes in the distributed cluster; the weight coefficient corresponding to the length of the pending task queue is positive, and the weight coefficient corresponding to the number of completed parsing tasks is negative; The node with the smallest load status score is determined as the target node corresponding to the parsing task.

6. The data analysis method according to claim 4, characterized in that: The load status includes the length of the pending task queue; after determining the target node corresponding to the parsing task based on the load status of the nodes in the distributed cluster, the method further includes: When it is detected that the length of the pending task queue of the target node is greater than a first task queue length threshold, determining the expansion quantity corresponding to each node in the distributed cluster based on the length of the pending task queue of each node in the distributed cluster; Create a corresponding number of expanded Pod instances for each node in the distributed cluster, and allocate the created Pod instances to the corresponding nodes in the distributed cluster; When it is detected that the length of the pending task queue of the target node is less than the second task queue length threshold, determining all nodes to be shrunk in the distributed cluster; the length of the pending task queue of the node to be shrunk is less than the second task queue length threshold; Based on the length of the pending task queue of each node to be shrunk, the corresponding shrinkage quantity of each node to be shrunk is determined, and the corresponding shrinkage quantity of Pod instances is released for each node to be shrunk.

7. The data analysis method according to claim 1, characterized in that: Performing data continuity check on the analysis result data, including: For each manufacturer's gateway server, detecting at every preset time interval whether the local database has stored the parsing result data corresponding to the manufacturer's gateway server within the preset time interval; When it is detected that the analysis result data corresponding to the gateway server of the manufacturer is stored in the local database within the preset time period, it is determined that the data continuity check has passed.

8. The data analysis method according to claim 1, characterized in that: Performing a cell coverage comparison check on the analysis result data, including: After the analysis result data corresponding to the network management servers of all manufacturers are stored every day, determining the cell coverage rate of the analysis result data corresponding to the network management servers of all manufacturers on that day and the analysis result data corresponding to the network management servers of all manufacturers on the previous day; When the cell coverage ratio verification is within a preset cell coverage ratio verification range, it is determined that the cell coverage ratio verification is passed.

9. A data analysis device, characterized in that: include: A first data parsing module is configured to create a distributed cluster and mount the nodes in the distributed cluster to a distributed file system; wherein the distributed file system stores data files downloaded from network management servers of all manufacturers, and the data files pass a cyclic redundancy check; The second data parsing module is used to call the nodes in the distributed cluster to execute the parsing task of the data files corresponding to the network management servers of each manufacturer in the distributed file system in parallel, and obtain the parsing result data corresponding to the network management servers of each manufacturer; The third data analysis module is used to store the analysis result data corresponding to each manufacturer's network management server in a local database, and perform data continuity verification and cell coverage ratio ring verification on the analysis result data.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the data analysis method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Data storage method and device

    CN108304396A

  • Acquisition unit detection device and method based on self-adaptive comparative analysis technology

    CN110456300A

  • Report data verification method and device, computer equipment and storage medium

    CN110502531A

  • Data verification method and device, equipment and storage medium

    CN110503567A

  • Data acquisition method for multi-source science and technology innovation resources

    CN113918793A