Data routing method and device, electronic equipment and storage medium
By extracting business fields and timestamps from the Flume component and writing the data directly to HDFS, the problem of high data routing latency is solved, achieving more efficient data processing and storage.
Patent Information
- Application Number
- CN202510684572.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-10-31
AI Technical Summary
The high latency of data routing in existing technologies, especially when using Kafka as a relay strategy, causes data to have to go through multiple levels of transfer sequentially.
By extracting the business fields and timestamps of the target data from the Flume component, determining the table storage path using the Flume configuration file, and generating the partition table path based on the timestamps, the data is directly written to HDFS without needing to go through Kafka as an intermediary.
It reduces data routing latency, lowers hardware costs, saves storage space, improves data processing efficiency, and reduces data loss rate and analysis job time.
Smart Images

Figure CN120880978A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data routing technology, and in particular to a data routing method, apparatus, electronic device, and storage medium. Background Technology
[0002] As a core distributed log collection tool in the big data field, Apache Flume's Source-Channel-Sink architecture supports real-time data collection from various data sources such as Kafka and writing to target storage such as Hadoop Distributed File System (HDFS), file storage, databases, and message queues.
[0003] Traditional data routing solutions often employ a Kafka relay strategy. This involves parsing business fields using custom interceptors and distributing the data to different topics via KafkaSink. Downstream devices consume the topic data and write it to the corresponding HDFS path. While this solution achieves multi-target routing, the high latency results from the data needing to pass through multiple levels of Kafka, Flume, Kafka, and HDFS sequentially. Summary of the Invention
[0004] This invention provides a data routing method, apparatus, electronic device, and storage medium to solve the technical problem of high data routing latency in the prior art.
[0005] This invention provides a data routing method, comprising: Obtain the target data from the data source and encapsulate the target data into an Event object in Flume; Extract the business fields and timestamp of the target data, and write the business fields and timestamp into the Header of the Event; Locate the table storage path corresponding to the business field in the Flume configuration file, write the table storage path into the Header, and store the Event into the Flume Channel; The Event is retrieved from the Channel, and a partition table path is generated based on the table storage path in the Header and the timestamp. The Event is then written into the partition table corresponding to the partition table path.
[0006] According to a data routing method provided by the present invention, the step of extracting the business field includes: If the target data is structured data, the business fields are extracted using regular expressions or JSON Path. If the target data is unstructured, the business fields are extracted using an entity recognition model.
[0007] According to a data routing method provided by the present invention, the step of writing the timestamp into the header includes: The extracted timestamps are then subjected to time zone-independent conversion and format standardization in sequence; The timestamp, after being formatted and standardized, is written into the Header.
[0008] According to a data routing method provided by the present invention, the extracted timestamps include a custom format.
[0009] According to a data routing method provided by the present invention, after extracting the business fields and timestamps of the target data and writing the business fields and timestamps into the header of the Event, the method further includes: Verify the integrity of the Event's data structure; The event that failed verification is flagged, and the flagged event is routed to a separate, isolated directory.
[0010] According to a data routing method provided by the present invention, the service field includes service type and device identifier; The verification of the integrity of the Event's data structure includes: Verify whether the service type is in the whitelist and whether the device identifier is empty; If the service type is outside the whitelist or the device identifier is empty, it indicates that the verification has failed.
[0011] According to a data routing method provided by the present invention, after extracting the business fields and timestamps of the target data and writing the business fields and timestamps into the header of the Event, the method further includes: The Event is filtered for null values, duplicates, and exceeding thresholds using regular expressions and a business rules engine.
[0012] According to a data routing method provided by the present invention, writing the Event into the partition table corresponding to the partition table path includes: Set the Flume's Sink type to hdfs, set hdfs.fileType to SequenceFile, and enable Snappy compression; Set both hdfs.rollInterval and hdfs.rollSize to 0 to disable the time- and file-size-based scrolling mechanism, retain only the scrolling by record count, and set the threshold corresponding to hdfs.rollCount; The Events pulled from the Channel in batches are written to an HDFS temporary file according to the configured serialization rules, and the number of Events in the HDFS temporary file is obtained in real time. If the number of Events in the HDFS temporary file reaches the threshold corresponding to hdfs.rollCount, then the HDFS temporary file is closed and renamed to the official SequenceFile; During the process of writing the Event to the HDFS temporary file, the offset of the HDFS temporary file is recorded by the CheckpointManager and stored in the specified checkpoint directory; After Flume restarts, the offset in the checkpoint directory is read, and the Event is written to the HDFS temporary file starting from the offset.
[0013] The present invention also provides a data routing device, comprising: The encapsulation module is used to obtain the target data from the data source and encapsulate the target data into an Event object in Flume. The interception module is used to extract the business fields and timestamp of the target data, and write the business fields and timestamp into the header of the Event; Locate the table storage path corresponding to the business field in the Flume configuration file, write the table storage path into the Header, and store the Event into the Flume Channel; The routing module is used to retrieve the Event from the Channel, generate a partition table path based on the table storage path in the Header and the timestamp, and write the Event into the partition table corresponding to the partition table path.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the data routing method as described above.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data routing method as described above.
[0016] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the data routing methods described above.
[0017] The data routing method, apparatus, electronic device, and storage medium provided by this invention extract the business fields and timestamps of the target data through the Flume component, determine the table storage path corresponding to the business fields in the configuration file, improve the table storage path according to the timestamp to obtain the final partition table path, and finally write the data according to the partition table path. This is equivalent to realizing direct data writing based on the Flume component, without the need for data distribution through Kafka, thus reducing the latency of data routing. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the data routing method provided by the present invention.
[0020] Figure 2 This is a schematic diagram of the data routing device provided by the present invention.
[0021] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] It should be noted that in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The terms "upper," "lower," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Unless otherwise expressly specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly, for example, as a fixed connection, a detachable connection, or an integral connection; a mechanical connection or an electrical connection; a direct connection or an indirect connection through an intermediate medium; or a connection within two elements. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0024] The terms "first," "second," etc., used in this invention are used to distinguish similar objects, not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0025] The following is combined Figures 1-3 This invention describes the data routing method, apparatus, electronic device, and storage medium provided by the present invention.
[0026] Flume employs a three-layer architecture: Source, responsible for receiving or fetching data; Channel, acting as a buffer connecting Source and Sink; and Sink, responsible for writing data to the target system. The Flume data stream processing is as follows: Source obtains raw data from the data source, encapsulates it into Event objects, processes the Events using interceptors, stores the Events in the Channel, Sink retrieves the Events from the Channel, and writes them to the target system. Interceptors are a crucial component of Flume's data stream processing, allowing users to intercept and process Events before data is written to the Channel.
[0027] Data is transmitted between Flume components in units of Events. An Event consists of a Header and a Body. The Header mainly includes fields such as metadata, timestamp, and hostname.
[0028] like Figure 1 As shown, the data routing method provided by the present invention includes steps S1-S4, and is applied to a data routing device. The data routing device includes a Flume component. Step S1 is completed by the Source, steps S2-S3 are completed by the interceptor, and step S4 is completed by the Sink.
[0029] Step S1: Obtain the target data from the data source and encapsulate the target data into an Event object in Flume.
[0030] The data source can include various servers, such as IT cloud, network cloud, mobile cloud, etc. The target data can be Kafka data, which can include structured data and unstructured data.
[0031] Step S2: Extract the business fields and timestamp of the target data, and write the business fields and timestamp into the Header of the Event.
[0032] The service field can include service type (event_type) and service tag. The service tag can further include device identifier (device_id) and error code (error_code). event_type includes FireWall, LoadBalance, Switch, etc.
[0033] This allows you to write the timestamp field in the header.
[0034] Step S3: Locate the table storage path corresponding to the business field in the Flume configuration file, write the table storage path into the Header, and store the Event into the Flume Channel.
[0035] The Flume configuration file predefines the table storage path corresponding to each business field, allowing the data write path to be determined based on the extracted business field. For example, `event_type=FireWall` corresponds to ` / user / hive / warehouse / ods.db / resops_perf_o_firewall_delta_o`, meaning the table storage path for the business field `FireWall` is ` / user / hive / warehouse / ods.db / resops_perf_o_firewall_delta_o`. Here, ` / user / hive / warehouse` is the default table storage root directory for Hive, `ods.db` is the Hive database name, and `resops_perf_o_firewall_delta_o` is the table name (firewall table). Overall, it means that firewall data will be written to the `ods.resops_perf_o_firewall_delta_o` table in Hive.
[0036] Specifically, the table storage path can be written into the target_path field in the Header for dynamic calling by the Sink.
[0037] Of course, events need to be written according to time partitions, so the table storage path supports multi-level path nesting, for example: / resops_perf_o_firewall_delta_o / partitionday=%{year}%{month}%{day} / partitionhour=%{hour} / partitionmin=%{minute}); `partitionday=%{year}%{month}%{day}` means partitioning by day, `partitionhour=%{hour}` means partitioning by hour, and `partitionmin=%{minute}` means partitioning by minute.
[0038] Step S4: Retrieve the Event from the Channel, generate a partition table path based on the table storage path and timestamp in the Header, and write the Event into the partition table corresponding to the partition table path.
[0039] The partition table path is generated based on the table storage path and timestamp in the Header. Specifically, %{year}, %{month}, %{day}, %{hour}, and %{minute} in the table storage path are automatically replaced with the actual time in the timestamp field in the Header, avoiding hard-coded paths.
[0040] Writing an Event to the partition table corresponding to the partition table path can be done by writing it to an HDFS table. The complete HDFS path and filename can be rendered by combining the Event's context variables (such as hostname and timestamp), allowing you to infer the Flume node's write activity based on the file information in HDFS. For example, an Event encapsulated with firewall data can be written to the HDFS firewall table, an Event encapsulated with port data can be written to the HDFS port table, an Event encapsulated with CPU data can be written to the HDFS CPU table, an Event encapsulated with router data can be written to the HDFS router table, and an Event encapsulated with switch data can be written to the HDFS switch table.
[0041] As described above, the data routing method of this invention extracts the business fields and timestamps of the target data through the Flume component, determines the table storage path corresponding to the business fields in the configuration file, improves the table storage path based on the timestamp to obtain the final partition table path, and finally writes the data according to the partition table path. This is equivalent to realizing direct data writing based on the Flume component, without the need for data distribution through Kafka, thus reducing the latency of data routing.
[0042] In some implementations, step S2, the step of extracting business fields, may further include: If the target data is structured data, extract the business fields using regular expressions or JSON Path. If the target data is unstructured, business fields are extracted using an entity recognition model.
[0043] The interceptor of this invention supports the extraction of structured data in various formats such as JSON, CSV, and Syslog, as well as the extraction of unstructured data (such as log text). For structured data, the business field can be event_type; for unstructured data, the business field can be business tags (including device_id and error_code).
[0044] Specifically, `event_type` can be extracted using regular expressions (such as `Pattern.compile("\\{\"event_type\":\"(.*?)\"\\}")`) or JSON Path (such as `JsonPath.read(message,"$.event_type")`), and business tags can be automatically extracted using entity recognition models integrated with the Stanford CoreNLP library. Regular expressions (Regex or RegExp) are tools for matching, finding, and replacing text patterns. JSON Path is a query language for locating and extracting specific nodes in JSON data, specifically designed for JSON structures. Stanford CoreNLP is an open-source library developed by the Stanford NLP Group at Stanford University that enables named entity recognition.
[0045] This allows for the extraction of business fields from both structured and unstructured data.
[0046] In some implementations, step S2, writing the timestamp into the Header, may further include: The extracted timestamps are sequentially subjected to timezone-independent conversion and format standardization. Write the standardized timestamp into the header.
[0047] The interceptor of this invention has 15 built-in time format parsers, which can parse ISO 8601 (the international standard for date and time representation developed by the International Organization for Standardization (ISO)), Unix timestamps, and custom formats such as dd / MMM / yyyy:HH:mm:ss Z and other time formats.
[0048] Specifically, timezone-independent conversion of timestamps can be achieved using DateTimeFormatter (a package in the java.time.format package used to parse and format date and time objects in the java.time package). The timestamp format can be standardized to yyyy-MM-dd HH:mm:ss format, where yyyy, MM, dd, HH, mm, and ss represent year, month, day, hour, minute, and second, respectively.
[0049] By performing timezone-independent conversion and format standardization on timestamps, the timezone and format of timestamps in various Events can be made the same, facilitating unified processing.
[0050] In some implementations, the format of the extracted timestamp in step S2 may include a custom format.
[0051] The interceptor of this invention also supports custom time formats, allowing independent fields of year, month, day, hour, minute, and second to be written into the header to adapt to storage requirements partitioned by time.
[0052] This allows for timezone-independent conversion and format standardization of custom-formatted timestamps, supporting user-defined timestamp formats.
[0053] In some embodiments, after step S2, the data routing method of the present invention may further include: Verify the integrity of the Event's data structure; Events that fail validation are flagged and routed to separate, isolated directories.
[0054] Specifically, the integrity of the data structure can be verified using JsonSchemaValidator (a tool or library for validating whether JSON data conforms to a specific structure and constraints).
[0055] Events that fail to be validated can be marked as status=invalid, indicating that the data validation has failed. The marked events are then routed to a separate HDFS isolated directory, which is equivalent to writing the event separately to the invalid_data directory of the data table, thus avoiding any impact on the main data.
[0056] By verifying and marking data structure integrity, events with incomplete data structures can be routed to independent, isolated directories, thus avoiding impact on the main data.
[0057] In some embodiments of the present invention, verifying the integrity of the Event's data structure may further include: Verify whether the service type is in the whitelist and whether the device identifier is empty; If the business type is outside the whitelist or the device identifier is empty, it means that the verification failed.
[0058] The default whitelist records trusted service types. If a service type is not in the whitelist, it means that the service type is not trusted and needs to be filtered out. An empty device identifier means that the device identifier is invalid and also needs to be filtered out.
[0059] Specifically, when the business field is `event_type`, it verifies whether `event_type` is in the whitelist. If `event_type` is not in the whitelist, the data structure integrity verification of the Event fails. When the business field is `device_id`, it verifies whether `device_id` is empty. If `device_id` is empty, the data structure integrity verification of the Event fails.
[0060] This can filter invalid data that is not in the whitelist for the service type and that has an empty device identifier, thus avoiding invalid data routing and reducing data routing efficiency.
[0061] In some embodiments, after step S2, the data routing method of the present invention may further include: Filter events with null values, duplicates, and exceeding thresholds using regular expressions and a business rules engine.
[0062] Specifically, regular expressions (such as ^\\s*$) and the integrated Drools business rule engine (used to separate business rules from application code) can be used to filter out null values, duplicates, and data exceeding thresholds (such as CPU utilization > 100%). Drools is an open-source business rule management system developed by Red Hat that allows business logic to be separated from application code and managed and executed in the form of declarative rules. Regular expressions can generally filter null values and simple duplicate data, but they cannot filter complex duplicate data or data exceeding thresholds. The integrated Drools business rule engine, however, can filter null values, duplicates, and data exceeding thresholds.
[0063] This can filter out null values, duplicates, and abnormal data that exceed the threshold, avoiding routing abnormal data and reducing data routing efficiency.
[0064] In addition, this invention can also monitor changes in data filtering rules and hot-load them into memory through ZooKeeper (an open-source distributed coordination service in the Hadoop ecosystem, used to manage configuration information, naming services, and provide distributed synchronization and group services in distributed systems), and use Keepalived (an open-source solution for implementing Linux server failover and load balancing) to achieve service-free restart updates of the interceptor.
[0065] When writing the Event to the partition table in step S4, the HDFS API can be used to check if the partition table exists before writing. If it does not exist, it will be created automatically in recursive mode. During the writing process, subdirectories automatically inherit the Access Control List Policy (ACL) of the parent directory to ensure data security. An HTTP / 2 connection pool can be built using HttpClient 5.x, supporting multiplexing (parallel transmission of multiple file blocks over a single connection), improving throughput compared to HTTP / 1.1. Header compression (HPACK algorithm) can be enabled to reduce network overhead when uploading small files.
[0066] In some implementations, step S4, writing the Event to the partition table corresponding to the partition table path, may further include: Set Flume's Sink type to hdfs, set hdfs.fileType to SequenceFile, and enable Snappy compression; Set both hdfs.rollInterval and hdfs.rollSize to 0 to disable the time- and file-size-based scrolling mechanism, retain only the scrolling by record count, and set the threshold corresponding to hdfs.rollCount; Events pulled from the Channel in batches are written to the HDFS temporary file according to the configured serialization rules, and the number of events in the HDFS temporary file is obtained in real time. If the number of Events in the HDFS temporary file reaches the threshold corresponding to hdfs.rollCount, then the HDFS temporary file is closed and renamed to the official SequenceFile. During the process of writing the event to the HDFS temporary file, the offset of the HDFS temporary file is recorded by the CheckpointManager and stored in the specified checkpoint directory. After Flume restarts, it reads the offset in the checkpoint directory and continues writing events to the HDFS temporary file starting from that offset.
[0067] Among them, hdfs.fileType is used to specify the data file format written to HDFS, SequenceFile is the binary key-value pair file format built into Hadoop, hdfs.rollInterval is the time interval for rolling new files, hdfs.rollSize is the size threshold for rolling new files, and hdfs.rollCount is the threshold for the number of records in rolling new files.
[0068] Setting both `hdfs.rollInterval` and `hdfs.rollSize` to 0 disables the time- and file-size-based scrolling mechanism, retaining only the record-count-based scrolling. Setting `hdfs.codeC=snappy` enables Snappy compression. The threshold corresponding to `hdfs.rollCount` defaults to 10000, meaning Flume will only generate a new file after accumulating 10000 events.
[0069] When writing data to the HDFS temporary file, Flume periodically (e.g., every 1000 events processed) records the current file offset through the CheckpointManager (a core component in distributed computing frameworks used for fault tolerance and state recovery; it periodically creates consistent snapshots of the application state to ensure that jobs can recover from the most recent checkpoint in case of node failure or task failure, avoiding data loss and ensuring computational accuracy). This offset is stored in a designated checkpoint directory (e.g., / flume / checkpoint in HDFS). Upon restarting after a failure (e.g., Flume process crashes or HDFS becomes temporarily unavailable), the Sink first reads the latest checkpoint directory to obtain the last written offset and then continues writing data from that position, rather than starting from the beginning.
[0070] By merging scattered events into a single SequenceFile that conforms to HDFS storage optimization goals, the metadata burden on the NameNode (the core component of HDFS, responsible for managing the file system's namespace and client access to files) can be significantly reduced, improving the efficiency of subsequent data processing. Resuming interrupted downloads can prevent duplicate data writes, further improving data write efficiency.
[0071] Furthermore, this invention can integrate the Flume transaction interface, committing the transaction only after HDFS confirms successful file writing; otherwise, it rolls back and retryes. The retry interval can be dynamically adjusted based on the error type (e.g., network timeout, NameNode congestion) (initially 100ms, increasing exponentially to 10s); continuously failing data is transferred to a local disk queue for asynchronous retry by an independent thread, avoiding main process blocking.
[0072] The data routing method of this invention achieves the following advantages: It eliminates the Kafka relay layer, reducing the number of Kafka Broker nodes by 5 (original cluster size 10 nodes → optimized to 5 nodes), thus reducing hardware costs by 50%; it saves 24% of HDFS storage space through data compression and invalid data filtering; it primarily reduces the number of connections due to HTTP / 2 multiplexing; Flume node CPU utilization drops from 85% to 45%; thanks to data pre-partitioning by business fields, downstream full table scan overhead is reduced, shortening Spark analysis job runtime by 30%; the latency from data collection to queryability drops from 12 minutes to 5 minutes; and the data loss rate drops from 0.3% to 0.08% (guaranteed by transactional writes).
[0073] like Figure 2 As shown, the data routing device provided by the present invention includes: The encapsulation module is used to obtain the target data from the data source and encapsulate the target data into an Event object in Flume; The interception module is used to extract the business fields and timestamps of the target data and write them into the Header of the Event. Find the table storage path corresponding to the business field in the Flume configuration file, write the table storage path into the Header, and store the Event into the Flume Channel; The routing module is used to retrieve Events from the Channel, generate partition table paths based on the table storage path and timestamp in the Header, and write the Events into the partition table corresponding to the partition table path.
[0074] It should be noted that the data routing device provided by the present invention can execute the data routing method of any of the above embodiments during specific operation, which will not be elaborated in this embodiment.
[0075] In some implementations, the interception module can also be used for: If the target data is structured data, extract the business fields using regular expressions or JSON Path. If the target data is unstructured, business fields are extracted using an entity recognition model.
[0076] In some implementations, the interception module can also be used for: The extracted timestamps are sequentially subjected to timezone-independent conversion and format standardization. Write the standardized timestamp into the header.
[0077] In some implementations, the extracted timestamp may be in a custom format.
[0078] In some implementations, the interception module can also be used for: Verify the integrity of the Event's data structure; Events that fail validation are flagged and routed to separate, isolated directories.
[0079] In some implementations, the service field may include service type and device identifier; The interception module can also be used for: Verify whether the service type is in the whitelist and whether the device identifier is empty; If the business type is outside the whitelist or the device identifier is empty, it means that the verification failed.
[0080] In some implementations, the interception module can also be used for: Filter events with null values, duplicates, and exceeding thresholds using regular expressions and a business rules engine.
[0081] In some implementations, the routing module can also be used for: Set Flume's Sink type to hdfs, set hdfs.fileType to SequenceFile, and enable Snappy compression; Set both hdfs.rollInterval and hdfs.rollSize to 0 to disable the time- and file-size-based scrolling mechanism, retain only the scrolling by record count, and set the threshold corresponding to hdfs.rollCount; Events pulled from the Channel in batches are written to the HDFS temporary file according to the configured serialization rules, and the number of events in the HDFS temporary file is obtained in real time. If the number of Events in the HDFS temporary file reaches the threshold corresponding to hdfs.rollCount, then the HDFS temporary file is closed and renamed to the official SequenceFile. During the process of writing the event to the HDFS temporary file, the offset of the HDFS temporary file is recorded by the CheckpointManager and stored in the specified checkpoint directory. After Flume restarts, it reads the offset in the checkpoint directory and continues writing events to the HDFS temporary file starting from that offset.
[0082] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 3As shown, the electronic device may include a processor, a communications interface, memory, and a communication bus. The processor, communications interface, and memory communicate with each other via the communication bus. The processor can invoke logical instructions in the memory to execute a data routing method. This method includes: acquiring target data from the data source and encapsulating the target data into an Event object in Flume; extracting the business fields and timestamps of the target data and writing them into the Event's Header; finding the table storage path corresponding to the business fields in the Flume configuration file, writing the table storage path into the Header, and storing the Event in Flume's Channel; retrieving the Event from the Channel, generating a partition table path based on the table storage path and timestamp in the Header, and writing the Event into the partition table corresponding to the partition table path.
[0083] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0084] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, when the program instructions are executed by a computer, the computer is able to execute the data routing method provided in the above embodiments, the method including: obtaining target data from a data source and encapsulating the target data into an Event object in Flume; extracting the business fields and timestamps of the target data and writing the business fields and timestamps into the Header of the Event; finding the table storage path corresponding to the business fields in the Flume configuration file, writing the table storage path into the Header, and storing the Event into the Channel of Flume; retrieving the Event from the Channel, generating a partition table path according to the table storage path and timestamp in the Header, and writing the Event into the partition table corresponding to the partition table path.
[0085] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the data routing method provided in the above embodiments. The method includes: obtaining target data from a data source and encapsulating the target data into an Event object in Flume; extracting the business fields and timestamps of the target data and writing the business fields and timestamps into the Header of the Event; finding the table storage path corresponding to the business fields in the Flume configuration file, writing the table storage path into the Header, and storing the Event into the Channel of Flume; retrieving the Event from the Channel, generating a partition table path based on the table storage path and timestamp in the Header, and writing the Event into the partition table corresponding to the partition table path.
[0086] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0087] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data routing method, characterized in that, include: Obtain the target data from the data source and encapsulate the target data into an Event object in Flume; Extract the business fields and timestamp of the target data, and write the business fields and timestamp into the Header of the Event; Locate the table storage path corresponding to the business field in the Flume configuration file, write the table storage path into the Header, and store the Event into the Flume Channel; The Event is retrieved from the Channel, and a partition table path is generated based on the table storage path in the Header and the timestamp. The Event is then written into the partition table corresponding to the partition table path.
2. The data routing method according to claim 1, characterized in that, The steps for extracting the business fields include: If the target data is structured data, the business fields are extracted using regular expressions or JSON Path. If the target data is unstructured, the business fields are extracted using an entity recognition model.
3. The data routing method according to claim 1, characterized in that, The step of writing the timestamp into the header includes: The extracted timestamps are then subjected to time zone-independent conversion and format standardization in sequence; The timestamp, after being formatted and standardized, is written into the Header.
4. The data routing method according to claim 3, characterized in that, The extracted timestamps can be in custom formats.
5. The data routing method according to claim 1, characterized in that, After extracting the business fields and timestamp of the target data and writing the business fields and timestamp into the Header of the Event, the method further includes: Verify the integrity of the Event's data structure; The event that failed verification is flagged, and the flagged event is routed to a separate, isolated directory.
6. The data routing method according to claim 5, characterized in that, The business fields include business type and device identifier; The verification of the integrity of the Event's data structure includes: Verify whether the service type is in the whitelist and whether the device identifier is empty; If the service type is outside the whitelist or the device identifier is empty, it indicates that the verification has failed.
7. The data routing method according to claim 1, characterized in that, After extracting the business fields and timestamp of the target data and writing the business fields and timestamp into the Header of the Event, the method further includes: The Event is filtered for null values, duplicates, and exceeding thresholds using regular expressions and a business rules engine.
8. The data routing method according to claim 1, characterized in that, The step of writing the Event into the partition table corresponding to the partition table path includes: Set the Flume's Sink type to hdfs, set hdfs.fileType to SequenceFile, and enable Snappy compression; Set both hdfs.rollInterval and hdfs.rollSize to 0 to disable the time- and file-size-based scrolling mechanism, retain only the scrolling by record count, and set the threshold corresponding to hdfs.rollCount; The Events pulled from the Channel in batches are written to an HDFS temporary file according to the configured serialization rules, and the number of Events in the HDFS temporary file is obtained in real time. If the number of Events in the HDFS temporary file reaches the threshold corresponding to hdfs.rollCount, then the HDFS temporary file is closed and renamed to the official SequenceFile; During the process of writing the Event to the HDFS temporary file, the offset of the HDFS temporary file is recorded by the CheckpointManager and stored in the specified checkpoint directory; After Flume restarts, the offset in the checkpoint directory is read, and the Event is written to the HDFS temporary file starting from the offset.
9. A data routing device, characterized in that, include: The encapsulation module is used to obtain the target data from the data source and encapsulate the target data into an Event object in Flume. The interception module is used to extract the business fields and timestamp of the target data, and write the business fields and timestamp into the header of the Event; Locate the table storage path corresponding to the business field in the Flume configuration file, write the table storage path into the Header, and store the Event into the Flume Channel; The routing module is used to retrieve the Event from the Channel, generate a partition table path based on the table storage path in the Header and the timestamp, and write the Event into the partition table corresponding to the partition table path.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the data routing method as described in any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the data routing method as described in any one of claims 1 to 8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the data routing method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Log collection method and system
CN106776715A
Flume data verification statistical method and device based on Hive
CN112131209A
Flume metadata information analysis and extraction method and related components
CN112685364A
Data writing method, device and equipment and readable storage medium
CN115221116A
Flume heterogeneous data-based acquisition and standardization method
CN116521715A