Data real-time analysis output method and system
By using message middleware and stream processing framework in the wireless network operation and maintenance environment, performance data files are consumed and processed in real time, and the delay problem of real-time data analysis output in the prior art is solved, and efficient and accurate data processing and metric updates are achieved.
Patent Information
- Application Number
- CN202510570829.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The prior art is difficult to realize real-time data analysis and output in wireless network operation and maintenance environment, resulting in lengthy processing flow, high response delay, and many repeated calculations, making it difficult to support high-frequency indicator updates and rapid fault locations.
The trigger message is generated by scanning the preset directory of the manufacturer's server and writing it to the message middleware. The data source components in the stream processing framework consume these messages in real time, read the data file according to the file path, form the original data stream, and perform real-time conversion processing in the stream processing framework to generate an metric data stream, and finally write it to the target storage medium and/or message middleware.
It realizes the timeliness and accuracy of real-time data analysis output, avoids processing delays and resource waste in traditional batch mode, and supports high-frequency metric updates and rapid fault location.
Smart Images

Figure CN120086256A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method and system for real-time parsing and output of data. Background Art
[0002] In the current wireless network operation and maintenance environment, network elements such as base stations and cells periodically generate a large number of performance data files. There are manufacturer differences in data formats, and the file structure often uses compression and encapsulation and is written to a specified directory asynchronously.
[0003] To analyze this data, existing solutions generally adopt a batch processing architecture, that is, by scheduling tasks to periodically scan the data directory, decompress the files, and store the original data in the counter, and then perform index calculations in the database. Although this solution has a certain processing ability in the offline statistics scenario, due to its high dependence on disk I / O and multiple intermediate storages, and the need for manual adaptation processing for heterogeneous file formats, the processing process is long, the response delay is high, and there are many repeated calculations, making it difficult to support high-frequency index updates and rapid fault location.
[0004] Therefore, how to improve the timeliness and accuracy of real-time parsing and output of data has become a technical problem that needs to be solved urgently. Summary of the Invention
[0005] The present invention provides a method, system, electronic device and storage medium for real-time parsing and output of data, to solve the defects in the prior art and achieve the improvement of the timeliness and accuracy of real-time parsing and output of data.
[0006] The present invention provides a method for real-time parsing and output of data, including the following steps: Scan a preset directory of the manufacturer's server, and for each detected data file, generate a trigger message for characterizing the data file, and write the trigger message into the message middleware; In the stream processing framework, the data source component consumes the trigger message from the message middleware in real time, reads the corresponding data file according to the file path indicated in the trigger message, and forms an original data stream; In the stream processing framework, perform real-time conversion processing on the original data stream to generate an index data stream for characterizing the performance indicators of the target object; Write the index data stream into the target storage medium and / or the message middleware.
[0007] According to the method for real-time parsing and output of data provided by the present invention, in the stream processing framework, the data source component consumes the trigger message from the message middleware in real time, reads the corresponding data file according to the file path indicated in the trigger message, and forms an original data stream, including: The data source component subscribes to the topic corresponding to the trigger message in the message middleware, and pulls the trigger message in the order of the partition number and offset carried by the trigger message; Parsing the data file path and file size information carried in each of the trigger messages, and sequentially reading the data file content in a block manner through a distributed file system interface; Attach the current processing timestamp to each read block and encapsulate it as a data element; Output each of the data elements in sequence to obtain the original data stream.
[0008] According to a data real-time parsing and output method provided by the present invention, in the stream processing framework, real-time conversion processing is performed on the original data stream to generate an indicator data stream for characterizing the performance indicator of the target object, including: Performing a validity check on the original data stream, removing file records that do not meet preset rules, and obtaining a target file stream; In the stream processing framework, the target file stream is decompressed to obtain a file entity object including a manufacturer identification and file format information; According to the manufacturer identification and file format in the file entity object, a matching parser is selected to parse the file content, extract the underlying counter data and append tag information to form a marked counter data stream; Splitting the marker counter data stream according to the label information to generate a plurality of logical data streams including at least a cell-level data stream and a base station-level data stream; Convert the underlying counter data in each logical data stream into row records and register the corresponding table structure according to the row records to form a tabular data stream; Aggregate calculations are performed on the tabular data stream within a preset time window, and the dimension table is temporarily associated to obtain the indicator data stream.
[0009] According to a real-time data parsing and output method provided by the present invention, the target file stream is decompressed within the stream processing framework to obtain a file entity object containing a manufacturer identification and file format information, including: Determine, according to the file suffix of each file record in the target file stream, whether its compression format is any one of the preset suffixes; Dynamically calling a decompression plug-in corresponding to the compression format in the stream processing framework, decoding byte by byte in a streaming manner without storing on a disk to generate a decompressed file; During the decoding process, the file header or preset identification field of the decompressed file is parsed to extract the manufacturer identification and data file format information; The decompressed file content, the manufacturer identification and the file format information are encapsulated into the file entity object.
[0010] A method for real-time data parsing and output provided by the present invention, which selects a matching parser according to the manufacturer identifier and file format in the file entity object to parse the file content, extracts the underlying counter data and attaches tag information to form a marked counter data stream, including: Maintain a parser routing table in the stream processing framework in advance, and the parser routing table maps each combination key of the manufacturer identifier and file format to the corresponding parser class; When a file entity object is received, query the parser routing table and dynamically load the parser class that matches the combination key; When the data file format is a hierarchical markup file format, call an XML parser based on SAX, and when the data file format is a fixed-length text or comma-separated format, call a CSV parser based on row-column parsing; Use the parser to parse the file content to obtain the original counter data containing counter name-value pairs; Append tag information to each counter data, and the tag information at least includes the manufacturer identifier, data sampling time, and network element identifier or cell identifier parsed from the file name; Encapsulate the counter data after attaching the tag information into the marked counter data stream.
[0011] A method for real-time data parsing and output provided by the present invention, which splits the marked counter data stream according to the tag information to generate multiple logical data streams including at least a cell-level data stream and a base station-level data stream, including: Set a side output routing unit in the stream processing framework, and the routing unit matches according to the network element level tag attached to each counter data; When the network element level tag contains both a cell identifier and a base station identifier, allocate the marked counter data stream to the cell-level data stream; When the network element level tag only contains a base station identifier and does not contain a cell identifier, allocate the marked counter data stream to the base station-level data stream.
[0012] A method for real-time data parsing and output provided by the present invention, which converts the underlying counter data in each logical data stream into row records and registers the corresponding table structure according to the row records to form a tabular data stream, including: For any one of the logical data streams, parse the set of counter names therein, and automatically generate field description metadata based on three fields: sampling time, network element identifier, and counter name; Call the table registration interface of the stream processing framework to dynamically create a temporary table according to the field description metadata, and the temporary table uses the combined field of sampling time and network element identifier as the primary key; Concatenate each piece of counter data in the logical data stream with the corresponding tag field to form a row record, and insert it into the temporary table in real time to obtain a tabular data stream carrying the watermark timestamp.
[0013] According to a data real-time parsing and output method provided by the present invention, the aggregating calculation of the tabular data stream within a preset time window and temporarily associating with a dimension table to obtain the metric data stream includes: Configure a rolling time window for the tabular data stream, where the length of the rolling time window is a preset duration, and the step size is the same as the length; Within each rolling time window, group the row records according to the network element identifier and counter name, and calculate the sum, maximum value, and minimum value of the counter values in each group to obtain the window aggregation result; Before outputting the window aggregation result, according to the watermark timestamp at the end time of the window, temporarily associate the window aggregation result with a dimension table containing city, district, and manufacturer information, and supplement the regional dimension and manufacturer dimension fields; Use the associated window aggregation result as the metric data stream.
[0014] The present invention also provides a data real-time parsing and output system, including the following modules: A first processing module for scanning a preset directory of a manufacturer server, generating a trigger message for each detected data file, and writing the trigger message into a message middleware; A second processing module for, in a stream processing framework, consuming the trigger message from the message middleware in real time through a data source component, reading the corresponding data file according to the file path indicated in the trigger message, and forming an original data stream; A third processing module for, in the stream processing framework, performing real-time conversion processing on the original data stream to generate a metric data stream for characterizing the performance metrics of a target object; A fourth processing module for writing the metric data stream into a target storage medium and / or the message middleware.
[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the data real-time parsing and output method as described in any one of the above when executing the program.
[0016] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data real-time parsing and output method as described in any one of the above.
[0017] The present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the data real-time parsing and output method as described in any one of the above.
[0018] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: By scanning a preset directory of the manufacturer's server, a trigger message for characterizing each detected data file is generated, and the trigger message is written into the message middleware, thereby realizing real-time perception and asynchronous driving of performance data generation events, effectively avoiding the processing delay and resource waste caused by the traditional timed polling mode. On this basis, in the stream processing framework, the trigger message is consumed in real time from the message middleware through the data source component, and the corresponding data file is read according to the file path indicated in the trigger message to form an original data stream, thereby realizing the full-process automated processing from message-driven to data loading. This path ensures the accuracy and orderliness of the file reading process by binding the trigger message to the specific file path. Then, real-time conversion processing is performed on the original data stream in the stream processing framework to generate an index data stream for characterizing the performance indicators of the target object, thereby realizing the rapid conversion from the original file content to structured and computable performance indicators. Finally, the index data stream is written into at least one of the target storage medium and the message middleware for subsequent systems to consume in real time, thereby effectively improving the timeliness and accuracy of data real-time parsing and output. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 is one of the flow diagrams of the data real-time parsing and output method provided by the present invention.
[0021] Figure 2 is another flow diagram of the data real-time parsing and output method provided by the present invention.
[0022] Figure 3 is yet another flow diagram of the data real-time parsing and output method provided by the present invention.
[0023] Figure 4 is still another flow diagram of the data real-time parsing and output method provided by the present invention.
[0024] Figure 5It is the fifth flowchart of the data real-time parsing and output method provided by the present invention.
[0025] Figure 6 It is the sixth flowchart of the data real-time parsing and output method provided by the present invention.
[0026] Figure 7 It is the seventh flowchart of the data real-time parsing and output method provided by the present invention.
[0027] Figure 8 It is the eighth flowchart of the data real-time parsing and output method provided by the present invention.
[0028] Figure 9 It is the structural schematic diagram of the data real-time parsing and output system provided by the present invention.
[0029] Figure 10 It is the structural schematic diagram of the electronic device provided by the present invention. Detailed implementation manners
[0030] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0031] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, the element defined by the phrase "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element. The orientation or positional relationship indicated by the terms "upper", "lower", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the system or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0032] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0033] The following will describe the data real-time parsing and output method, system, electronic device, and storage medium provided by the present invention in conjunction with Figures 1 - 10 the following description.
[0034] Figure 1 is one of the flow diagrams of the data real-time parsing and output method provided by the present invention. As shown in Figure 1 the figure, it includes steps 101 to 104: Step 101: Scan the preset directory of the manufacturer's server. For each detected data file, generate a trigger message for characterizing the data file, and write the trigger message into the message middleware.
[0035] In a preferred embodiment of the present invention, step 101 involves scanning the preset directory of the manufacturer's server, generating a trigger message for characterizing the data file when a data file is detected, and writing the trigger message into the message middleware. This step, as the starting action of the data real-time parsing and output process of the present invention, mainly aims to achieve an incremental discovery and timely notification mechanism for the performance data files on the manufacturer side, so as to support the event-driven data reading and processing operations in the subsequent stream processing framework, thereby overcoming the disadvantages of the traditional batch processing method in terms of time granularity and processing delay.
[0036] In the specific implementation process, the system monitors the preset manufacturer's server directory in a periodic or event-driven manner. This directory is usually the unified storage path for the manufacturer's equipment to upload performance data files according to specific rules. When the system detects a newly appeared data file in this directory (for example, the file name is not marked by historical records or the creation time is after the last scan), it is considered that the data file is new data to be processed. At this time, the system generates a trigger message based on the data file, where the trigger message at least carries keyword fields such as the path information, creation time, file size, and manufacturer identification of the data file, so as to accurately identify the file and provide necessary context information for subsequent processing.
[0037] The generated trigger message is then written to a preset message middleware, which can be a streaming message system such as Kafka or RocketMQ with high throughput and distributed subscription capabilities. By writing the trigger message to the message middleware, a data-driven decoupled processing mode can be achieved, that is, the detection and parsing processes of data files are logically isolated, the constraint of time synchronization is removed, which is conducive to supporting the construction of an elastic and scalable real-time processing pipeline. In addition, the partition and offset mechanisms provided by the message middleware also provide a good foundation for achieving accurate consumption, fault tolerance recovery, and data idempotency processing in subsequent stream processing systems.
[0038] Step 102: In the stream processing framework, the data source component consumes the trigger message from the message middleware in real time, and reads the corresponding data file according to the file path indicated in the trigger message to form an original data stream.
[0039] In this embodiment, step 102 aims to drive the stream processing framework to realize the automatic reading of data files and the generation of the original data stream based on the trigger message written to the message middleware in step 101. The design intention of this step is that the parsing of traditional performance files is mostly to actively pull directory files by a scheduling script. This "timed polling + batch loading" mode faces significant delays and resource waste problems in scenarios of multi-vendor and large-scale heterogeneous data access. By constructing a "message-driven data acquisition mechanism", it is possible to realize the on-demand reading of files based on trigger events, effectively improving the response speed from files to data streams and resource utilization efficiency.
[0040] In a possible implementation manner, Figure 2 is the second flow schematic diagram of the data real-time parsing and output method provided by the present invention. As Figure 2 shown, step 102 specifically includes the following steps: Step 201: The data source component subscribes to the topic corresponding to the trigger message in the message middleware, and pulls the trigger message in the order of the partition number and offset carried by the trigger message.
[0041] Step 202: Parse the data file path and file size information carried in each trigger message, and sequentially read the content of the data file in a block manner through the distributed file system interface.
[0042] Step 203: Append the current processing timestamp to each read block and encapsulate it as a data element.
[0043] Step 204: Output each data element in sequence to obtain the original data stream.
[0044] In this embodiment, steps 201 to 204 are used to specifically implement how the data source component pulls trigger messages from the message middleware and efficiently reads data files based on the information contained in the trigger messages, and finally constructs an original data stream with time stamps. As the specific implementation of step 102, the key of this sub-process is to ensure the controllability, sequentiality, and data structural integrity of file reading, so that the subsequent stream processing link can directly complete real-time metric extraction without introducing intermediate disk landing.
[0045] First, in step 201, the data source component is configured to subscribe to a specified topic in the message middleware (such as Kafka), which is used to carry the trigger messages generated in step 101. To ensure consistent and traceable reading order, the data source component pulls trigger messages in order based on the partition number and offset fields in each trigger message, and records the consumption position. When the system restarts or fails and recovers, breakpoint resume reading can be achieved based on this position to avoid duplicate reading or omission of files.
[0046] Next, in step 202, the data source component parses the trigger message and extracts fields such as the data file path and file size carried therein. According to the path field, the component initiates a sequential block-by-block reading operation through a pre-configured distributed file system interface (such as HDFS or OSS). The system disassembles large files into several consecutive data blocks using a fixed block size (such as 64KB or 128KB) and reads them sequentially in a streaming manner, avoiding loading the entire file into memory and causing resource blockage.
[0047] Subsequently, in step 203, to enable time-driven downstream processing and support time series operator operations such as window aggregation, the system immediately attaches the current processing timestamp after each read block is generated, and encapsulates the timestamp and the read block together into a structured data element. This timestamp can not only represent the actual file processing time, but also cooperate with the business sampling time in the file content to form multi-granularity time characteristics.
[0048] Finally, in step 204, the system outputs the above structured data elements in sequence to form the constituent units of the original data stream. This original data stream has orderliness, structural integrity, and time attributes, and can be directly received and processed by the downstream legality verification, decompression, and metric parsing modules without intermediate caching or data landing on disk.
[0049] In summary, steps 201 to 204 technically achieve an efficient bridge from trigger messages to a structured original data stream. Firstly, it ensures the pulling order and recoverability through position control. Secondly, it guarantees system resource load balance through block-level reading. Thirdly, it improves the time series processing ability of the data stream by attaching timestamps, significantly enhancing the real-time performance, stability, and scalability of performance data collection as a whole.
[0050] Step 103: In the stream processing framework, real-time conversion processing is performed on the original data stream to generate an indicator data stream for characterizing the performance indicators of the target object.
[0051] In this embodiment, step 103 aims to convert the original data stream formed in step 102 into an indicator data stream with analytical significance. This step is the core conversion link in the entire data processing link. Its design purpose is to complete the whole process of parsing, filtering, labeling, structuring and calculating from file data to performance indicators through the stream processing framework without interrupting the data path, so as to ensure that the final result has real-time and integrity.
[0052] In one possible implementation, Figure 3 FIG. 3 is a flow chart of the real-time data analysis and output method provided by the present invention. Figure 3 As shown, step 103 specifically includes steps 301-306: Step 301: Perform a validity check on the original data stream, remove file records that do not meet preset rules, and obtain a target file stream.
[0053] In this embodiment, the role of step 301 is to first complete the legitimacy check of the data file before the original data stream formally enters the decompression, parsing and other computationally intensive processes. The original intention of designing this step is that the performance data files generated by the manufacturer may have abnormal situations during the upload or generation process, such as file truncation, naming errors, inconsistent sampling time, format damage and other problems. If they directly enter the parsing link, it may not only cause the program to terminate abnormally, but also waste a lot of computing resources and even cause indicators to be misreported. Therefore, establishing a "legality filtering" link to ensure data availability at the source is a key link to ensure the stability and accuracy of the overall streaming link.
[0054] In a specific implementation method, the legality check executes the following content verification logic for each file record in the original data stream. First, the system verifies whether the file name meets the preset naming specifications based on the configured manufacturer rules, such as whether it contains valid manufacturer identification, timestamp, area code and other fields to ensure that the file can be uniquely identified and attributed. Secondly, verify whether the sampling time recorded in the file header is in the valid time window before the current system time to prevent abnormal data in the future from entering the indicator calculation logic. Thirdly, read the file size field recorded in the file header, and compare it with the actual number of bytes read to determine whether there is a problem of file content loss or truncation. Finally, for data files with a checksum field, the system recalculates the check value based on the content, and compares it with the checksum given in the file header to verify the content integrity.
[0055] The legality verification process is implemented as a streaming function in the stream processing framework, featuring non-blocking, parallelizable, and state-maintainable characteristics. If a file record fails to pass any of the above verification conditions, the record is filtered out without triggering an exception and will no longer participate in subsequent processing steps, thus ensuring the stability of the system processing link.
[0056] By introducing this step, significant technical effects are achieved: First, the quality of input data is greatly improved, avoiding link interruptions or parsing failures caused by abnormal files; Second, pre-screening is completed in a parallel filtering manner in the stream processing framework, making full use of cluster resources to achieve both high performance and high robustness; Third, the file pre-screening ability in the traditional batch processing process is migrated to the real-time data channel, providing a pre-guarantee mechanism for the stable access of large-scale heterogeneous data from manufacturers. As the first sub-step of step 103, this implementation method provides a clean and stable data basis for subsequent operations such as decompression, parsing, and table creation.
[0057] Step 302: Within the stream processing framework, decompress the target file stream to obtain a file entity object containing manufacturer identification and file format information.
[0058] In this embodiment, step 302 is used to perform an online decompression operation on the target file stream after legality verification to obtain a file entity object containing manufacturer identification and file format information. The reason for implementing this step is that performance data files generated by communication manufacturers are usually stored and transmitted in a compressed format to reduce disk occupancy and network transmission burden. However, in the traditional batch processing architecture, compressed files are often decompressed into intermediate files before being parsed, which not only increases the I / O cost but also introduces additional burdens of disk writing and cleaning, and is not suitable for real-time processing scenarios with extremely high requirements for low latency and high throughput. Therefore, this step introduces a streaming decompression mechanism in the stream processing framework, aiming to decompress files on-demand in memory and immediately convert them into parseable objects, thereby opening up the real-time path from the original compressed data to structured metrics.
[0059] In a possible implementation manner, Figure 4 is the fourth flowchart of the data real-time parsing and output method provided by the present invention. As Figure 4 shown, step 302 specifically includes the following steps: Step 401: Determine whether the compression format of each file record in the target file stream is any of the preset suffixes according to the file suffix.
[0060] Step 402: Dynamically call the decompression plugin corresponding to the compression format within the stream processing framework, and decode byte by byte in a streaming manner without writing to disk to generate the decompressed file.
[0061] Step 403: During the decoding process, parse the file header or the preset identification field of the decompressed file, and extract the manufacturer identification and the data file format information.
[0062] Step 404: Package the content of the decompressed file together with the manufacturer identification and the file format information into a file entity object.
[0063] In this embodiment, steps 401 to 404 are specific sub-steps of step 302, which describe in detail the online decompression process of the target file stream, including how to identify the file compression format, perform the decompression operation, extract key information, and construct a file entity object. This process not only realizes the real-time conversion of compressed data into a parsable object, but also provides a structured input carrier for the subsequent selection of parsers and index extraction, with clear intermediate processing objectives and technical advantages.
[0064] First, in step 401, by analyzing the file name suffix field of the target file, it is determined whether its compression format belongs to the preset format set supported by the system. This format set may include, but is not limited to, compression formats such as ".tar.gz", ".zip", ".gz", and ".tar", and the preset support range can be dynamically extended according to the platform running environment and the actual access of manufacturers. The method of the system matching the compression format based on the file suffix name has the advantages of strong versatility, non-intrusiveness, and low implementation cost, and can quickly complete the compression type identification, serving as the scheduling entry for the decompression operation.
[0065] In step 402, the system calls the decompression plug-in corresponding to the identified compression format type within the stream processing framework. The plug-in decompresses the file content byte by byte (or in chunks) based on the streaming decoding mechanism, and does not write the decompression result to disk, but directly passes it in the pipeline through the memory buffer method. This processing method effectively avoids the disk I / O bottleneck, ensuring the low latency and high throughput characteristics of the decompression process, and is especially suitable for scenarios with high requirements for computing and response speed in real-time performance data processing. The decompression plug-in is designed according to the standard interface, facilitating the isolation and extension of the decoding logic for different compression formats.
[0066] In step 403, the system performs a structured parsing on the beginning part of the decompressed file, and extracts the manufacturer identification field and the data file format field from it. These key information may appear in specific identification segments, configuration areas, or naming fields of the file header. The system automatically extracts and standardizes them according to the preset parsing rules. For example, the manufacturer identification can be standardized as "HUAWEI", "ZTE", "ERICSSON", etc., while the data format is uniformly identified as "XML", "CSV", "TXT", etc. The extracted meta-information is used to guide the subsequent selection of parsers and data structure identification.
[0067] Finally, in step 404, the system encapsulates the content of the file obtained after decompression above, together with the extracted manufacturer identifier and data file format information, into a file entity object. This object contains fields such as the original content reference, compression format identifier, source path, decompression timestamp, etc., and has good structural and serializable characteristics, and serves as the direct input unit for subsequent parsing steps (such as step 303).
[0068] By implementing steps 401 to 404, the present invention realizes the unified recognition and efficient decompression of multi-format compressed performance files, does not rely on the external storage disk process, and has extremely strong real-time performance and resource utilization efficiency. At the same time, by completing the extraction and encapsulation of meta-information during the decompression stage, the subsequent parser selection and processing logic are made clearer and more automated, improving the modularity and extensibility of the parsing link, and is an important technical guarantee for realizing the real-time processing ability of cross-manufacturer and multi-format performance data.
[0069] Step 303: According to the manufacturer identifier and file format in the file entity object, select a matching parser to parse the file content, extract the underlying counter data and attach label information to form a labeled counter data stream.
[0070] In this embodiment, step 303 is used to perform a parsing operation on the file entity object obtained in step 302, aiming to extract the underlying counter data in the data file and attach label information with semantic indication to each counter data, and finally form a labeled counter data stream for downstream streaming calculation. The purpose of implementing this step is to establish a mapping path from the file entity to the structured indicator source data, and solve the key problems such as complex parser adaptation, semantic loss, and inconsistent structure existing in the real-time parsing process of multi-manufacturer and multi-format data.
[0071] In a possible implementation manner Figure 5 is the fifth flow diagram of the data real-time parsing output method provided by the present invention, as Figure 5 shown, step 303 specifically includes the following steps Step 501: Maintain a parser routing table in the stream processing framework in advance. The parser routing table maps each combination key of the manufacturer identifier and file format to the corresponding parser class.
[0072] Step 502: When receiving the file entity object, query the parser routing table and dynamically load the parser class that matches this combination key.
[0073] Step 503: When the data file format is a hierarchical markup file format, call the XML parser based on SAX, and when the data file format is a fixed-length text or comma-separated format, call the CSV parser based on row-column parsing.
[0074] Step 504: Parse the file content using a parser to obtain the original counter data containing counter name-value pairs.
[0075] Step 505: Append tag information to each counter data. The tag information includes at least the manufacturer identifier, data sampling time, and network element identifier or cell identifier parsed from the file name.
[0076] Step 506: Package the counter data with appended tag information into a tagged counter data stream.
[0077] In this embodiment, steps 501 to 506 are specific sub-steps of step 303, aiming to refine the parsing process of the file entity object and complete the whole process from parser matching to tagged counter data stream generation. This process solves the heterogeneous parsing problem of multi-vendor and multi-format data in real-time stream processing, ensuring that each output data not only has numerical significance but also carries complete service tag information, with good identifiability and processability, and is the basis for subsequent splitting and metric calculation.
[0078] First, in step 501, the system pre-maintains a parser routing table inside the stream processing framework. This routing table uses the "manufacturer identifier + file format" as the combined key and maps it to the corresponding parser class. For example, the combination of the manufacturer identifier "HUAWEI" and the format "XML" will point to the SAX parser class dedicated to parsing the Huawei XML structure; if the format is "CSV", it will be mapped to the CSV parser class that parses according to the field position. This routing table can be dynamically updated by the system configuration module, supports manufacturer expansion and format evolution, and has good generality and maintainability.
[0079] Next, in step 502, when the system receives the file entity object, it will automatically query the parser routing table according to the manufacturer identifier and file format fields carried by it, obtain the corresponding parser class and dynamically load it in the runtime environment. After loading, the system initializes the parser according to the decompressed file content and inputs the data stream into the parser to execute the parsing task.
[0080] In step 503, the system selects an appropriate parsing method according to the format of the data file. If the file format is a hierarchical markup language (such as XML), the parser adopts an event-driven parsing logic based on the SAX model, continuously reads tag nodes and outputs the parsing results in real time, which is suitable for scenarios with complex structures but sensitive to memory; if the format is fixed-length text or comma-separated format (such as CSV), the method of reading by line and mapping field names and counter values according to field positions is used for parsing. Through the classification parsing method, it effectively adapts to the organizational differences of performance data from different manufacturers.
[0081] Subsequently, in step 504, the parser traverses the file content to extract the original counter data, which usually exists in the form of "counter name - value pair". These original data represent the original observations of the basic network metrics and are the underlying basis for constructing subsequent complex performance metrics. The system stores each piece of counter data extracted in a standard structure and prepares to append context label information.
[0082] In step 505, the system appends label information to each piece of original counter data. The label includes but is not limited to manufacturer identification, data sampling time (parsed from the file name or internal fields of the file), network element identification (such as eNodeB ID, gNB ID), and cell identification (such as Cell ID). The process of appending label information can be automatically completed through methods such as rule matching, field extraction, and context reasoning to ensure that all data has a complete semantic background, facilitating subsequent flow splitting and aggregation operations by label.
[0083] Finally, in step 506, the system encapsulates each piece of counter data with appended labels into a tagged counter data unit with a unified structure and forms a tagged counter data stream as the output. This data stream has a high degree of structuring and stable fields and already has the ability to directly participate in stream processing, window calculation, and dimensional association, and can be used as the input source for subsequent logical data stream construction and tabular processing.
[0084] Through the implementation of steps 501 to 506, the system realizes the end-to-end automatic conversion from the file entity object to the tagged counter data stream. Its technical effects are reflected in the following aspects: First, through the dual-factor parser matching mechanism, a unified parsing entry for multiple manufacturers and multiple formats is realized; second, the parsing process runs in a streaming manner throughout, supporting the real-time computing requirements of high concurrency and high throughput; third, the output data has a complete label system, laying a structural foundation for multi-dimensional analysis and dynamic routing, and significantly improving the automation and scalability of the overall metric output link.
[0085] Step 304: Split the tagged counter data stream according to the label information to generate multiple logical data streams including at least cell-level data stream and base station-level data stream.
[0086] In this embodiment, step 304 is used to further logically split the marked counter data stream generated in step 303 to form multiple logical data streams with clear semantic levels, including at least cell-level data streams and base station-level data streams. The purpose of this step is to cope with the multi-level network element structure naturally existing in wireless network performance data, such as cells, gNodeBs, regions, etc. If classification is not performed according to the label dimension at the data parsing stage, problems such as data mixing, high processing complexity, and semantic conflicts will be faced in subsequent stages such as metric modeling and aggregation calculation. Therefore, performing logical splitting at the initial stage of data stream formation to generate multiple logical data streams is a necessary condition for realizing efficient stream computing and accurate metric output.
[0087] In a possible implementation manner, Figure 6 is the sixth flowchart of the data real-time parsing and output method provided by the present invention. As Figure 6 shown, step 304 specifically includes the following steps: Step 601: Set a side output routing unit in the stream processing framework, and the routing unit matches according to the network element level label attached to each counter data.
[0088] Step 602: When the network element level label contains both a cell identifier and a base station identifier, allocate the marked counter data stream to the cell-level data stream.
[0089] Step 603: When the network element level label only contains a base station identifier and does not contain a cell identifier, allocate the marked counter data stream to the base station-level data stream.
[0090] In this embodiment, steps 601 to 603 specifically refine the splitting logic of the marked counter data stream in step 304, and describe how to perform shunt operations on data at different network element levels by setting a routing mechanism, and finally generate cell-level data streams and base station-level data streams with clear structures and explicit semantics. The core design of this process is that wireless performance data may mix metrics with different granularities during transmission and parsing. If reasonable differentiation is not performed at an early stage, subsequent table building, aggregation, and metric generation will not only have redundant calculations, but may also result in semantic misjudgments. Therefore, ensuring that data has a clear hierarchical attribution before entering different service processing channels through a label-driven automatic shunt mechanism is a key step in improving stream processing efficiency and metric accuracy.
[0091] First, in step 601, the system configures a side output routing unit in the stream processing framework. This unit is implemented through a custom stream processing operator and is used to parse the tag information carried in each piece of marker counter data and perform real-time data shunting based on the "network element level" field in the tag. This field is extracted or calculated during the file parsing process and represents the network structure level to which the current counter data belongs, such as whether it contains a cell identifier (e.g., Cell ID) or only a base station identifier (e.g., gNodeB ID).
[0092] Subsequently, entering step 602, the system performs conditional matching based on the tag information: when it is found that the tag of a certain piece of counter data contains both a cell identifier and a base station identifier, the system determines that its granularity is at the cell level and has a clear cell affiliation relationship. Therefore, this data is routed to the cell-level data stream. This type of data usually involves key performance indicators of the radio access link, such as uplink PRB utilization rate, RRC connection establishment success rate, etc., and needs to be aggregated and modeled at the cell level.
[0093] In step 603, when the system detects that the tag of a certain piece of counter data contains only a base station identifier but does not contain a cell identifier, it determines that the data is at the base station level granularity. For example, some board performance indicators or base station power status information provided by certain manufacturers do not involve specific cell division. Such data is routed by the system to the base station-level data stream for generating resource status, health monitoring, and other indicators at the site level.
[0094] The entire side output routing mechanism is implemented inside the stream processing framework, with high-throughput and low-latency processing characteristics. It does not introduce data replication or broadcast processes, ensuring the optimal resource utilization efficiency. At the same time, the routing logic can be dynamically adjusted through a configuration file, supporting the introduction of more levels (such as regional level, neighboring cell level) or adjusting the tag field matching rules, with good flexibility and expansion capabilities.
[0095] By implementing steps 601 to 603, the system realizes an automated multi-level data splitting mechanism driven by tags. The main technical effects brought about include: First, ensuring that the same type of metric data is processed under the same semantic dimension, avoiding incorrect aggregation or metric conflicts; Second, migrating the splitting logic from static rules to real-time streams to complete, improving the system response efficiency and the ability to adapt to heterogeneous data formats; Third, providing a clear data stream branch structure for subsequent logical table registration and metric calculation, reducing the complexity of the calculation link, which is a key technical link for realizing the decoupling and efficient processing of large-scale wireless performance data metrics.
[0096] Step 305: Convert the underlying counter data in each logical data stream into row records and register the corresponding table structures according to the row records to form tabular data streams.
[0097] In this embodiment, step 305 is used to further convert each logical data stream obtained by splitting in step 304 into a tabular data stream with structured semantics. The reason for implementing this step is that traditional performance data mostly exists in text or semi-structured form, lacks a standardized structure, and is not conducive to the unified scheduling and calculation of operators in the stream processing framework. In streaming processing platforms such as Flink, by converting data into a table structure and registering it as a temporary view, the data can be uniformly modeled, windowed, and dimensionally associated with SQL or Table API, greatly improving the expressiveness and computing efficiency of indicator construction. Therefore, completing the table structure registration before the logical data flow enters the computing layer to form a tabular data flow is a necessary link to open up the "stream data-indicator model-real-time calculation" link.
[0098] In one possible implementation, Figure 7 FIG. 7 is a flow chart of the real-time data analysis and output method provided by the present invention. Figure 7 As shown, step 305 specifically includes the following steps: Step 701: for any logical data flow, parse the counter name set therein, and automatically generate field description metadata based on the three fields of sampling time, network element identifier, and counter name.
[0099] Step 702: Call the table registration interface of the stream processing framework to dynamically create a temporary table based on the field description metadata. The temporary table uses the combined field of sampling time and network element identifier as the primary key.
[0100] Step 703: Each counter data in the logical data stream is concatenated with the corresponding tag field into a row record, and inserted into a temporary table in real time to obtain a tabular data stream carrying a watermark timestamp.
[0101] In this embodiment, steps 701 to 703 specifically implement the internal process of "converting each logical data stream into a tabular data stream" in step 305, and detail how to dynamically build a field structure based on the logical data stream, generate a table structure, and implement data writing, so as to support subsequent aggregation calculations, SQL queries, and dimension association operations in the stream processing framework. The core of the setting of this process is that the original performance data source is highly heterogeneous, and the counter fields in different manufacturers or formats vary significantly. If a static field definition method is used, the system's adaptability and processing flexibility will be severely limited. Therefore, automatically parsing fields and dynamically registering table structures during operation is the key path to achieving high versatility and high real-time indicator processing.
[0102] First, in step 701, for any split logical data stream (such as cell level or base station level), the system automatically identifies and parses the set of counter names contained in the data stream. This set of counters is derived from the original metric fields extracted in the aforementioned parsing step and forms a complete set of fields in combination with metadata fields such as sampling time and network element identifier. The system constructs a field description metadata structure with "sampling time", "network element identifier", and "counter name" as the core dimensions. This structure records attributes such as field name, field type, and whether it is a primary key field, forming the basis for the table structure definition.
[0103] Subsequently, in step 702, the system dynamically creates a temporary table structure based on the field description metadata generated in step 701 by calling the table registration interface provided by the stream processing framework (such as Flink Table API). The temporary table does not depend on a physical database and exists only as a logical view in the stream processing runtime environment, but has a complete field definition and primary key constraint. Among them, the system uses the combination of the "sampling time" and "network element identifier" fields as the logical primary key to ensure the unique identification of counter records for the same time and network object, and supports window aggregation and grouping operations based on these primary key fields. After the table structure is registered, the system completes the dynamic transformation from the "original data field set" to the "standard table structure", providing structural support for subsequent streaming calculations based on SQL or window functions.
[0104] Next, in step 703, each marked counter data in the logical data stream is mapped into a standardized row record according to the field description and inserted into the registered temporary table in real time. The row record not only contains various counter value fields but also label fields such as manufacturer identifier, cell ID, sampling time, etc. To meet the requirements of time series calculation, the system also attaches an event timestamp to each row record and maintains the maximum acceptable delay range of the system through the watermark mechanism. Finally, all structured row records are continuously generated driven by timestamps, forming a tabular data stream with table structure, time attributes, and queryability.
[0105] Through the implementation of the above steps 701 to 703, the system completes the structural reconstruction of the logical data stream into a tabular data stream, and the technical effects include: First, by using the automatic field parsing and dynamic table structure registration mechanism, the system's adaptability to various data formats and different metric dimensions is greatly improved; Second, through the logical primary key and structure mapping, the data semantic expression form is unified, facilitating subsequent indicator development and maintenance; Third, through the embedding of timestamps and watermarks, the time consistency guarantee in window processing and latency tolerance of streaming data is achieved. Overall, this process effectively supports the structured access and efficient stream calculation of massive heterogeneous performance data and is one of the core steps in building a large-scale indicator output system.
[0106] Step 306: Aggregate and calculate the tabular data stream within a preset time window, and temporarily associate with the dimension table to obtain the metric data stream.
[0107] In this embodiment, step 306 is used to perform window aggregation calculation on the tabular data stream generated in step 305, and temporarily associate with the dimension table on the basis of aggregation, and finally generate a metric data stream for characterizing the performance status of the target object. As the end calculation unit of the entire real-time data parsing link, the core objective of this step is to convert a large amount of underlying counter data into key performance indicators with business semantics (such as connection rate, disconnection rate, resource utilization rate, etc.), and supplement context fields such as geography and manufacturer by associating dimension information, so that the output result has business usability and system consumability.
[0108] In a possible implementation manner, Figure 8 is the eighth schematic diagram of the process of the real-time data parsing and output method provided by the present invention. As Figure 8 shown, step 306 specifically includes the following steps: Step 801: Configure a rolling time window for the tabular data stream. The length of the rolling time window is a preset duration, and the step size is the same as the length.
[0109] Step 802: Within each rolling time window, group the row records according to the network element identifier and counter name, and calculate the sum, maximum value, and minimum value of the counter values of each group to obtain the window aggregation result.
[0110] Step 803: Before outputting the window aggregation result, according to the watermark timestamp of the window end time, temporarily associate the window aggregation result with the dimension table containing city, district, and manufacturer information, and supplement the regional dimension and manufacturer dimension fields.
[0111] Step 804: Use the associated window aggregation result as the metric data stream.
[0112] In this embodiment, steps 801 to 804 further refine the key calculation process in step 306, specifically describe how to complete window aggregation based on the time-driven mechanism, and implement temporary association with the dimension table in combination with the watermark strategy of event time, and finally output a metric data stream with complete semantics. This sub-process not only realizes the real-time conversion of counter data into performance indicators, but also improves the business interpretability and analysis value of the result data by introducing a dynamic dimension extension mechanism. It is the most comprehensive and output-oriented calculation link in the wireless performance data calculation process.
[0113] First, in step 801, the system configures a rolling time window for the tabular data stream. This window has a fixed length and a fixed step size. The window length is set to fifteen minutes, and the step size is the same as the length, ensuring that the system can output the complete index calculation results every fifteen minutes. The window is driven by event time and runs in combination with the timestamp field carried in the aforementioned tabular data stream. The system uses a watermark strategy in conjunction to tolerate the delays caused by out-of-order data and network jitter, ensuring that data can still be accurately assigned to the appropriate window within the maximum acceptable delay range, guaranteeing the timeliness and accuracy of the calculation.
[0114] In step 802, when each rolling window closes, the system performs an aggregation calculation on all the data within the window. The aggregation operation uses the combination of the "network element identifier + counter name" fields as the grouping basis, and executes a preset aggregation function for each group of data, including but not limited to: summation (such as the cumulative number of RRC connection requests), maximum value (such as the maximum PRB utilization rate within 15 minutes), minimum value (such as the minimum CQI quality value), etc. These statistical results represent the change range of the underlying counter within a cycle window and are the core basic data for generating derived performance indicators.
[0115] Subsequently, in step 803, to enhance the semantic integrity and business value of the indicators, the system performs a dimension table association operation on the aggregation results. This association uses the "temporary dimension table" strategy and is matched through the event-time-driven Temporal Join method, ensuring that each aggregation result uses the version of the dimension table snapshot that is consistent with its timestamp when associating dimension information. The dimension table can include fields such as city, district, manufacturer type, and site ownership relationship. Its data comes from an external operation and maintenance system or a configuration center. The system loads the dimension table data into the runtime environment through a broadcast mechanism or a side input mechanism to achieve low-latency access and matching.
[0116] During the matching process, the system uses the network element identifier as the connection primary key field, finds the corresponding regional dimension and manufacturer dimension information based on this identifier, and supplements it to the current aggregation result to form a complete-structured and richly-labeled index data record. Through this operation, each index result not only has an index name and a value but also has a clear time, space, and operation attribution background.
[0117] Finally, in step 804, the system encapsulates the above aggregation results with completed dimension supplementation into an index data stream in a unified format and outputs it as the final result to the target storage medium or real-time message channel. The index data stream has unified field definitions, time series identifiers, dimension information, and performance values, facilitating subsequent system operations such as real-time display, historical archiving, and intelligent alarm.
[0118] In summary, steps 801 to 804 achieve the last-hop conversion from raw data to consumable metrics by constructing a time window, performing grouped aggregation, associating dimension information, and outputting structured metric data.
[0119] Step 104: Write the metric data stream to a target storage medium and / or message middleware.
[0120] In this embodiment, step 104 is used to write the metric data stream generated in step 306 (especially its sub-step 804) to at least one of a target storage medium and a message middleware for subsequent systems to perform real-time consumption, display, archiving, or alarm processing. The motivation for this step is that although the previous steps have completed the parsing of performance data and metric calculation, if the metric results cannot be transmitted to downstream systems in a timely manner or persistently stored, a complete closed-loop data service cannot be formed. Therefore, the final output and distribution of metric data are not only a necessary end point of the entire process but also a key technical guarantee for building a highly available, highly visible, and highly responsive real-time metric platform.
[0121] In the specific implementation process, the system classifies the output targets of the metric data stream into two categories according to preset configurations: one is a target storage medium, such as relational databases (such as Greenplum, StarRocks), distributed data warehouses (such as Hive, Hudi), etc.; the other is a real-time message middleware, such as Kafka, Pulsar, etc., for immediate subscription and use by real-time alarm modules, data buses, or visualization platforms. The metric data stream is serialized into a corresponding format according to the target type before output. For example, Kafka uses the Avro or JSON structure, and the database uses batch write or streaming write interfaces.
[0122] The system writes the metric data to the target end in an asynchronous and non-blocking manner by configuring Sink operators in the stream processing framework. For message middleware, the system sets the Topic or partition key according to dimension fields such as the region or manufacturer to which the metric belongs to achieve orderly distribution and parallel consumption of data within the message channel. For database storage, the system supports automatic table creation, field mapping, and primary key deduplication mechanisms to ensure the consistency and standardization of data writing.
[0123] Meanwhile, the system supports a multi-target output mode, that is, the metric data can be written to multiple target ends simultaneously. For example, it can be written to Kafka for real-time consumption and to a data warehouse for subsequent historical analysis, improving the business compatibility and data reusability of the system. The output path can be dynamically selected and switched through configuration, with good scalability and disaster tolerance capabilities.
[0124] It should be noted that after step 104, the method can further perform anomaly detection based on the metric data stream and generate and output structured alarm events when detecting anomaly events.
[0125] In this embodiment, after completing the writing operation of the metric data stream in step 104, the system further introduces an extended processing flow for metric anomaly detection and alarm linkage to achieve real-time identification of abnormal fluctuations in key performance indicators and trigger automated alarm linkage operations based on the identification results. The purpose of this step is to make up for the deficiencies of traditional wireless performance monitoring systems in terms of real-time performance and response mechanisms. Especially in high-density network deployment environments, where performance data is updated frequently and faults change rapidly, if abnormal indicators cannot be quickly detected and linked for disposal at the initial stage of the anomaly, it often leads to problems such as fault amplification, positioning delay, and service loss. Therefore, the present invention introduces a flow-embedded anomaly detection module in the metric flow processing link and constructs a closed-loop system from metric flow to operation and maintenance response in combination with a flexible alarm trigger mechanism.
[0126] In the specific implementation process, based on the flow processing framework, the system deploys a set of anomaly detection operators beside the output path of the metric data stream. These detection operators directly receive the metric data stream as input and have low-latency and concurrently scalable processing capabilities. The system supports multiple detection strategies to run in parallel, including static threshold rules, dynamic change rate monitoring, and sliding window comparison. For example, for key metrics such as call connection rate and disconnection rate, the system can preset upper and lower threshold values. When the value of a certain metric data exceeds the threshold range, an anomaly flag is triggered. At the same time, the system can also maintain historical window data, calculate the deviation degree between the current value and the past average value, and set a deviation percentage threshold. If the current value changes too quickly, an anomaly identification can also be triggered. In addition, for multi-dimensional metric joint anomaly scenarios, the system supports associating and modeling multiple metrics. When there is a structural change in the combined relationship (such as a sharp increase in the number of requests while the success rate drops), it can also be determined as an abnormal event.
[0127] Once the detection operator identifies an abnormal event, the system enters the linkage processing flow, encapsulates the abnormal metric record as a structured alarm event object, which includes multiple fields such as abnormal metric items, metric values, trigger time, network element identifier, geographical information, and the manufacturer. The system supports sending alarm events in multiple ways, including pushing them to a dedicated alarm channel of the message middleware for real-time subscription by the operation and maintenance platform, or integrating with SMS, email, or instant messaging platforms through calling the alarm management API for notification, and even linking with the automated operation and maintenance system to execute temporary disposal actions such as configuration distribution and fault isolation.
[0128] Refer to Figure 9 , Figure 9 is a schematic structural diagram of the data real-time parsing and output system provided by the present invention. The system includes: The first processing module is used to scan a preset directory of the manufacturer's server. For each detected data file, it generates a trigger message for characterizing the data file and writes the trigger message into the message middleware. The second processing module is used to, in the stream processing framework, consume the trigger message from the message middleware in real time through the data source component, read the corresponding data file according to the file path indicated in the trigger message, and form a raw data stream. The third processing module is used to, in the stream processing framework, perform real-time conversion processing on the raw data stream to generate an index data stream for characterizing the performance metrics of the target object. The fourth processing module is used to write the index data stream into the target storage medium and / or the message middleware.
[0129] In a possible implementation manner, the second processing module is further used to: Have the data source component subscribe to the topic corresponding to the trigger message in the message middleware and pull the trigger message in the order of the partition number and offset carried by the trigger message; Parse the data file path and file size information carried in each trigger message, and sequentially read the data file content in a chunked manner through the distributed file system interface; During the reading process, append the current processing timestamp to each read block and encapsulate it as a data element; Output each data element in sequence to obtain the raw data stream.
[0130] In a possible implementation manner, the third processing module is further used to: Perform a legality check on the raw data stream, eliminate file records that do not meet the preset rules, and obtain a target file stream; In the stream processing framework, decompress the target file stream to obtain a file entity object containing the manufacturer identifier and file format information; According to the manufacturer identifier and file format in the file entity object, select a matching parser to parse the file content, extract the underlying counter data and append label information to form a labeled counter data stream; Split the labeled counter data stream according to the label information to generate multiple logical data streams including at least a cell-level data stream and a base station-level data stream; Convert the underlying counter data in each logical data stream into row records and register the corresponding table structure according to the row records to form a tabular data stream; Perform aggregation calculation on the tabular data stream within a preset time window and temporarily associate with the dimension table to obtain the index data stream.
[0131] In a possible implementation manner, the third processing module is further used to: Determine whether the compression format of each file record in the target file stream is any of the preset suffixes according to the file suffix; Dynamically call the decompression plugin corresponding to the compression format within the stream processing framework, and decode byte by byte in a streaming manner without disk writing to generate the decompressed file; Parse the file header or preset identification field of the decompressed file during the decoding process, and extract the manufacturer identification and data file format information; Encapsulate the decompressed file content, manufacturer identification, and file format information into a file entity object.
[0132] In a possible implementation manner, the third processing module is further configured to: Maintain a parser routing table in the stream processing framework in advance. The parser routing table maps each combination key of manufacturer identification and file format to the corresponding parser class; When receiving the file entity object, query the parser routing table and dynamically load the parser class that matches the combination key; When the data file format is the hierarchical markup file format, call the XML parser based on SAX, and when the data file format is the fixed-length text or comma-separated format, call the CSV parser based on row-column parsing; Use the parser to parse the file content to obtain the original counter data containing the counter name-value pairs; Append label information to each counter data. The label information includes at least the manufacturer identification, data sampling time, and network element identification or cell identification parsed from the file name; Encapsulate the counter data with the appended label information into a marked counter data stream.
[0133] In a possible implementation manner, the third processing module is further configured to: Set a side output routing unit in the stream processing framework. The routing unit performs matching according to the network element level label attached to each counter data; When the network element level label contains both the cell identification and the base station identification, allocate the marked counter data stream to the cell-level data stream; When the network element level label only contains the base station identification and does not contain the cell identification, allocate the marked counter data stream to the base station-level data stream.
[0134] In a possible implementation manner, the third processing module is further configured to: For any logical data stream, parse the set of counter names therein, and automatically generate field description metadata based on three fields: sampling time, network element identification, and counter name; Call the table registration interface of the stream processing framework, and dynamically create a temporary table according to the field description metadata. The temporary table uses the combined field of sampling time and network element identification as the primary key; Concatenate each counter data in the logical data stream with the corresponding label field to form a row record, and insert it into the temporary table in real time to obtain a tabular data stream carrying the watermark timestamp.
[0135] In a possible implementation manner, the third processing module is further configured to: Configure a rolling time window for the tabular data stream, where the length of the rolling time window is a preset duration, and the step size is the same as the length; Within each rolling time window, group the row records according to the network element identifier and the counter name, and calculate the sum, maximum value, and minimum value of the counter values in each group to obtain the window aggregation result; Before outputting the window aggregation result, temporarily associate the window aggregation result with a dimension table containing city, district, and manufacturer information according to the watermark timestamp of the window end time, and supplement the regional dimension and manufacturer dimension fields; Use the associated window aggregation result as the metric data stream.
[0136] It should be noted that the data real-time parsing and output system provided by the present invention can execute the data real-time parsing and output method of any of the above embodiments during specific operation, and this embodiment will not be elaborated herein.
[0137] Figure 10 is a schematic structural diagram of an electronic device provided by the present invention, as Figure 10 shown, the electronic device may include: a processor 1010 (processor), a communication interface 1020 (Communications Interface), a memory 1030 (memory), and a communication bus 1040. Among them, the processor 1010, the communication interface 1020, and the memory 1030 communicate with each other through the communication bus 1040. The processor 1010 can call the logical instructions in the memory 1030 to execute the data real-time parsing and output method, and the method includes: scanning a preset directory of the manufacturer server, for each detected data file, generating a trigger message for characterizing the data file, and writing the trigger message into the message middleware; in the stream processing framework, the data source component consumes the trigger message from the message middleware in real time, reads the corresponding data file according to the file path indicated in the trigger message to form a raw data stream; in the stream processing framework, perform real-time conversion processing on the raw data stream to generate a metric data stream for characterizing the performance metrics of the target object; write the metric data stream into the target storage medium and / or the message middleware.
[0138] In addition, when the logical instructions in the above-mentioned memory 1030 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, external hard drives, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0139] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the data real-time parsing and output method provided in the above-mentioned various embodiments. The method includes: scanning a preset directory of the manufacturer's server, for each detected data file, generating a trigger message for characterizing the data file, and writing the trigger message into the message middleware; in the stream processing framework, the data source component consumes the trigger message from the message middleware in real time, reads the corresponding data file according to the file path indicated in the trigger message, and forms a raw data stream; in the stream processing framework, performing real-time conversion processing on the raw data stream to generate an index data stream for characterizing the performance indicators of the target object; writing the index data stream into the target storage medium and / or the message middleware.
[0140] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the data real-time parsing and output method provided in the above-mentioned various embodiments. The method includes: scanning a preset directory of the manufacturer's server, for each detected data file, generating a trigger message for characterizing the data file, and writing the trigger message into the message middleware; in the stream processing framework, the data source component consumes the trigger message from the message middleware in real time, reads the corresponding data file according to the file path indicated in the trigger message, and forms a raw data stream; in the stream processing framework, performing real-time conversion processing on the raw data stream to generate an index data stream for characterizing the performance indicators of the target object; writing the index data stream into the target storage medium and / or the message middleware.
[0141] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0142] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A method for real-time data analysis and output, characterized in that: include: Scan the preset directory of the manufacturer's server, generate a trigger message for characterizing each data file detected, and write the trigger message into the message middleware; In the stream processing framework, the trigger message is consumed in real time from the message middleware through the data source component, and the corresponding data file is read according to the file path indicated in the trigger message to form an original data stream; In the stream processing framework, real-time conversion processing is performed on the original data stream to generate an indicator data stream for characterizing the performance indicator of the target object; The indicator data stream is written into a target storage medium and / or the message middleware.
2. The method for real-time data analysis and output according to claim 1, characterized in that: In the stream processing framework, the trigger message is consumed in real time from the message middleware through the data source component, and the corresponding data file is read according to the file path indicated in the trigger message to form an original data stream, including: The data source component subscribes to the topic corresponding to the trigger message in the message middleware, and pulls the trigger message in the order of the partition number and offset carried by the trigger message; Parsing the data file path and file size information carried in each of the trigger messages, and sequentially reading the data file content in a block manner through a distributed file system interface; Attach the current processing timestamp to each read block and encapsulate it as a data element; Output each of the data elements in sequence to obtain the original data stream.
3. The data real-time analysis and output method according to claim 1, characterized in that: In the stream processing framework, the original data stream is converted in real time to generate an indicator data stream for characterizing the performance indicator of the target object, including: Performing a validity check on the original data stream, removing file records that do not meet preset rules, and obtaining a target file stream; In the stream processing framework, the target file stream is decompressed to obtain a file entity object including a manufacturer identification and file format information; According to the manufacturer identification and file format in the file entity object, a matching parser is selected to parse the file content, extract the underlying counter data and append tag information to form a marked counter data stream; Splitting the marker counter data stream according to the label information to generate a plurality of logical data streams including at least a cell-level data stream and a base station-level data stream; Convert the underlying counter data in each logical data stream into row records and register the corresponding table structure according to the row records to form a tabular data stream; Aggregate calculations are performed on the tabular data stream within a preset time window, and the dimension table is temporarily associated to obtain the indicator data stream.
4. The method for real-time data analysis and output according to claim 3, characterized in that: In the stream processing framework, the target file stream is decompressed to obtain a file entity object containing a manufacturer identification and file format information, including: Determine, according to the file suffix of each file record in the target file stream, whether its compression format is any one of the preset suffixes; Dynamically calling a decompression plug-in corresponding to the compression format in the stream processing framework, decoding byte by byte in a streaming manner without storing on a disk to generate a decompressed file; During the decoding process, the file header or preset identification field of the decompressed file is parsed to extract the manufacturer identification and data file format information; The decompressed file content, the manufacturer identification and the file format information are encapsulated into the file entity object.
5. The method for real-time data analysis and output according to claim 3, characterized in that: The method of selecting a matching parser to parse the file content according to the manufacturer identifier and the file format in the file entity object, extracting the underlying counter data and appending tag information to form a marked counter data stream includes: Maintaining a parser routing table in advance in the stream processing framework, the parser routing table mapping each combination key of the manufacturer identifier and the file format to a corresponding parser class; When a file entity object is received, query the parser routing table and dynamically load the parser class matching the composite key; When the data file format is a hierarchical markup file format, an XML parser based on SAX is called, and when the data file format is a fixed-length text or comma-delimited format, a CSV parser based on row and column parsing is called; Parsing the file content using the parser to obtain raw counter data including counter name-value pairs; Adding tag information to each counter data, wherein the tag information includes at least a manufacturer identifier, a data sampling time, and a network element identifier or a cell identifier obtained by parsing the file name; The counter data to which the tag information is attached is encapsulated into the marked counter data stream.
6. The method for real-time data analysis and output according to claim 3, characterized in that: The step of splitting the marker counter data stream according to the label information to generate a plurality of logical data streams including at least a cell-level data stream and a base station-level data stream comprises: A side output routing unit is provided in the flow processing framework, and the routing unit performs matching according to a network element level label attached to each counter data; When the network element level label includes both a cell identifier and a base station identifier, allocating the marking counter data flow to the cell level data flow; When the network element level label only includes a base station identifier but not a cell identifier, the marking counter data flow is allocated to a base station level data flow.
7. The method for real-time data analysis and output according to claim 3, characterized in that: The step of converting the underlying counter data in each logical data stream into row records and registering a corresponding table structure according to the row records to form a tabular data stream includes: For any of the logical data flows, a set of counter names are parsed, and field description metadata is automatically generated based on three fields: sampling time, network element identifier, and counter name; Calling the table registration interface of the stream processing framework, dynamically creating a temporary table according to the field description metadata, wherein the temporary table uses the combined field of sampling time and network element identifier as the primary key; Each counter data in the logical data stream is spliced with the corresponding label field into a row record, and inserted into the temporary table in real time to obtain a tabular data stream carrying a watermark timestamp.
8. The method for real-time data analysis and output according to claim 3, characterized in that: The step of performing aggregation calculation on the tabular data stream within a preset time window and temporarily associating the dimension table to obtain the indicator data stream includes: Configure a rolling time window for the tabular data stream, wherein the length of the rolling time window is a preset time length, and the step length is consistent with the length; In each of the rolling time windows, the row records are grouped according to the network element identifier and the counter name, and the sum, maximum value and minimum value of the counter values of each group are calculated to obtain the window aggregation result; Before outputting the window aggregation result, the window aggregation result is temporarily associated with the dimension table containing the city, district, county and manufacturer information according to the watermark timestamp of the window end time, and the regional dimension and manufacturer dimension fields are supplemented; The associated window aggregation result is used as the indicator data stream.
9. A real-time data analysis and output system, characterized in that: include: A first processing module is used to scan a preset directory of a manufacturer's server, generate a trigger message for characterizing each data file detected, and write the trigger message into a message middleware; A second processing module is used to consume the trigger message from the message middleware in real time through a data source component in a stream processing framework, read the corresponding data file according to the file path indicated in the trigger message, and form an original data stream; A third processing module is used to perform real-time conversion processing on the original data stream in the stream processing framework to generate an indicator data stream for characterizing the performance indicator of the target object; The fourth processing module is used to write the indicator data stream into a target storage medium and / or the message middleware.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the real-time data parsing and output method as described in any one of claims 1 to 8 is implemented.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the real-time data parsing and output method as described in any one of claims 1 to 8 is implemented.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the real-time data parsing and output method as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Hbase-based index monitoring method and device, equipment and storage medium
CN113468019A
Data object analysis method and device and electronic equipment
CN113568677A
General parallel analysis method and device for airborne binary files and electronic equipment
CN113742298A
Data synchronization method and device under distributed architecture and data synchronization system
CN116107985A
Concurrent traffic monitoring method and device, terminal equipment and storage medium
CN116743558A
Cited By
Standardized processing method, system and equipment for original engineering parameters of flight simulator
CN120596443A
A method, system and device for standardizing raw engineering parameters of a flight simulator
CN120596443B
Method and system for transmitting data to Greenplum database
CN121051081A