A method for real-time situation data acquisition and processing

By creating dynamic flow tasks in real-time situational data processing and performing many-to-one dictionary table matching, the problem of inability to perform many-to-one matching in the existing technology is solved, and fast and accurate data processing and high-quality data distribution are achieved.

CN113901090BActive Publication Date: 2025-06-03BEIJING INST OF COMP TECH & APPL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111266120.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-28
Publication Date
2025-06-03
Estimated Expiration
2041-10-28

AI Technical Summary

Technical Problem

The existing technology cannot effectively match many-to-one dictionary tables, resulting in inefficient processing of real-time situation data.

Method used

By creating a dynamic flow task, use the target name in the flow data to compare and judge with the target name or keyword in the mapped dictionary table. If it is successfully matched, use the machine hull number in the dictionary table to replace the machine hull number of the flow data and unify the machine hull number.

Benefits of technology

It realizes fast and accurate processing of real-time situation data, can dynamically match many-to-one dictionary tables, and improves data quality and processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113901090B_ABST
    Figure CN113901090B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for real-time situation data acquisition and processing, belonging to the field of radar data processing. The method of the present invention includes: acquisition of real-time data; preprocessing of data, including non-empty judgment and Chinese-English conversion; unification of data aircraft numbers, supplementation of missing attributes, and correlation matching of dynamic streaming data with static dictionary data, and filling the matched static dictionary information back into the real-time situation stream; data storage and push: one copy of the matched real-time data is persistently stored, and the other copy is simply processed and pushed into the message queue for application systems to call; dynamic maintenance and refresh of static dictionaries. The method for real-time data acquisition and processing proposed by the present invention is convenient to access, flexibly expands the data processing function by dynamically adding streaming task sql, and can obtain fast and accurate processing results. It has important application value in scenarios where real-time data needs to be cleaned, real-time dictionary matching, and streaming data storage and push.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of radar data processing, and particularly relates to a method for real-time situation data access and processing. Background Art

[0002] With the advent of the big data era, the data formats generated in daily life are diverse, and the efficient storage and fast and accurate processing of data have become increasingly important. Among them, the processing of real-time data is an important part of big data processing. Facing a large amount of streaming data, stability and reliability have become problems that the streaming data processing system needs to face directly.

[0003] The source of real-time situation data is collected by different radar signal sources, and the information is automatically judged by the program and supplemented by manual intervention. Therefore, there may be some naming differences such as target names and ship numbers for the same target. For users, one is to process real-time data and display it without delay on the map, and the other is to improve the quality of data in real-time processing and send down the data with higher quality. The existing technology generally uses programming to access real-time streams and uses a simple dictionary for one-to-one matching, and cannot perform many-to-one dictionary table matching. Summary of the Invention

[0004] (1) Technical Problems to be Solved

[0005] The technical problem to be solved by the present invention is how to provide a method for real-time situation data access and processing to solve the problem that the existing technology generally uses programming to access real-time streams, uses a simple dictionary for one-to-one matching, and cannot perform many-to-one dictionary table matching.

[0006] (2) Technical Solutions

[0007] To solve the above technical problems, the present invention proposes a method for real-time situation data access and processing, and the method includes the following steps:

[0008] S1. Real-time stream access step: Create an input stream, connect to the data source, and configure the connection parameters of the message queue;

[0009] S2. Data preprocessing step: Create a derivative stream to process the input stream created in S1, including data filtering and Chinese-English conversion;

[0010] S3. Create a dynamic stream task to unify the key attributes of the derivative stream according to the mapping dictionary table: Compare and judge the target name in the stream data with the target name or keyword in the mapping dictionary table. If the match is successful, replace the ship number in the stream data with the ship number in the dictionary table to unify the ship numbers;

[0011] S4. Data Persistence and Push: Store the streaming data created in S3 into the HBase table, and send the processed streaming data through the message queue for use.

[0012] Further, when selecting the message queue in step S1, the supported message queues include Kafka and RabbitMQ.

[0013] Further, step S1 specifically includes: Using Hive SQL to configure the ZooKeeper and Kafka nodes, where Kafka is the received message queue, and ZooKeeper is the distributed scheduling platform on which the message queue depends, storing some Kafka metadata information; configuring the serialization encoding information for receiving character streams and the Kafka topic, and splitting the received character stream into paragraph texts according to the line break character.

[0014] Further, step S2 specifically includes: Using Hive SQL to create a derived stream, using the JSON parsing function json_tuple() supported by Hive to parse the Chinese character fields in the input stream in S1, and converting them to English encoding using the as keyword to complete the Chinese-English encoding conversion; the source of the streaming data is the input stream created in S1, and the filtering condition for the streaming data is that the JSON string is not empty and the target name of the input stream is not empty.

[0015] Further, creating the dynamic stream task in step S3 specifically includes: Creating a dynamic stream task, with the from source being the name of the derived stream in the previous step.

[0016] Further, in step S3, use the target name in the streaming data to compare and judge with the target name or keyword in the mapping dictionary table. If the match is successful, replace the hull number in the streaming data with the hull number in the dictionary table to unify the hull numbers, which specifically includes: Associating the dynamic stream with the static mapping dictionary table through the left join association method, and judging whether the dictionary value matches by whether the target name in the streaming data is equal to the target name in the mapping dictionary table; if the match is not successful at this time, then use whether the target name in the streaming data is included in the keyword in the mapping dictionary table, and the keyword maintains all alias fields of the same target; if the match is successful, replace the hull number in the streaming data with the hull number in the dictionary table to unify the hull numbers.

[0017] Further, step S3 also includes: If the match fails, it means that there is no corresponding dictionary in the dictionary table data, or there is no alias that can be matched for the target, and SQL is used for marking, and later manual selection of the dictionary is performed for historical data matching.

[0018] Further, the step S4 specifically includes: First, use hive sql to create a streaming task. After "from", follow the stream name created by S3, map all fields in the stream to a pre-established hbase table for persistent storage; continue to create a streaming task. After "from", follow the stream name created by S3, fetch data from the stream in the previous step, restore the originally English fields in the stream back to Chinese fields, concatenate them into a json string, and use the streaming task to publish the json string.

[0019] Further, after the step S4, there is also a step S5: the online maintenance step of the dictionary. If target data needs to be added to the dictionary, select the relevant dictionary tables in the basic dictionary library, and map the same attribute fields in different dictionary tables to the mapping dictionary table:

[0020] Further, the step S5 specifically includes: Click the new mapping dictionary target in the situation fetching system. The system selects multiple isomorphic tables from the basic dictionary table for tree-shaped target display. These tables have a tree structure and similar field contents. Map the fields of these similar heterogeneous tables to the mapping dictionary table uniformly. At the same time, expand the keywords in the dictionary table; for the keywords in the expansion table, when adding a new dictionary target, copy the keyword information in the original dictionary table to the new table, and support manual modification of the keywords in the dictionary table. Append the target alias to be added to the mapping dictionary keyword.

[0021] (III) Beneficial effects

[0022] The present invention proposes a method for real-time situation data fetching and processing. The present invention uses the relatively stable hive sql at the present stage as the real-time fetching and processing process, which is highly flexible and configurable. It uses multiple dictionary tables to map to a unified dictionary table, and flexibly maintains the keywords in the dictionary table, so that the dictionary table can be timely matched with the real-time stream after modification, without the need for complex business restart operations. The method for real-time data fetching and processing proposed by the present invention is convenient to access. The function of data processing can be flexibly expanded by dynamically adding the stream task sql. And it can obtain fast and accurate processing results. It has important application value in scenarios where real-time data needs to be data-cleaned, real-time dictionary matched, and stream data stored and pushed. Description of the drawings

[0023] Figure 1 It is a flowchart of the real-time fetching method of the present invention. Detailed implementation manners

[0024] To make the objectives, contents, and advantages of the present invention clearer, the following further describes in detail the specific implementation manners of the present invention with reference to the drawings and embodiments.

[0025] The present invention discloses a method for real-time data acquisition and processing, which includes: (1) Real-time data acquisition. (2) Data preprocessing, including non-empty judgment and Chinese-English conversion. (3) Unifying the data aircraft numbers and supplementing the missing attributes. The dynamic streaming data needs to be associated and matched with the static dictionary data, and the matched static dictionary information is filled back into the real-time situation stream. (4) Data storage and push. One copy of the matched real-time data is persistently stored, and the other copy is simply processed and pushed into the message queue for the application system to call. (5) Dynamic maintenance and refresh of the static dictionary. The data in multiple standard dictionary tables are uniformly collected into a mapping dictionary, and then the real-time stream is automatically refreshed so that the association and matching operation in the second step can be matched with the data in the latest dictionary table. The method for real-time data acquisition and processing proposed by the present invention is convenient to access, and the data processing function can be flexibly expanded by dynamically adding the stream task sql. And fast and accurate processing results can be obtained. It has important application value in scenarios where real-time data needs to be cleaned, real-time dictionary matching, and streaming data storage and push.

[0026] The real-time data processing of the present invention is divided into real-time data acquisition, data governance, and data usage. The real-time acquisition provides general and delay-free data access. The data governance provides flexible governance methods and customized stream processing rules with fast configuration. The data usage provides PB-level massive data storage services and low-latency data push services.

[0027] The purpose of the present invention is to provide a method for real-time acquisition of situation data to meet the requirements of fast, accurate, and high-quality data processing for real-time data.

[0028] To achieve the above purpose, the present invention proposes a method for real-time data acquisition and processing, which includes:

[0029] Real-time stream acquisition step.

[0030] Data preprocessing step.

[0031] Unifying the data aircraft numbers and supplementing the attributes.

[0032] Data persistence and push.

[0033] Online maintenance of the dictionary.

[0034] Figure 1 Is a flowchart of a real-time acquisition method of the present invention. As Figure 1 shown, the method includes:

[0035] S1. Real-time stream acquisition step. Support mainstream message queues such as kafka, rabbit mq, etc. Create an input stream, connect to the data source, and configure the connection parameters of the message queue.

[0036] In specific implementation, hive sql is used to configure message queue information, such as zookeeper and kafka nodes. Among them, kafka is the received message queue, and zookeeper is the distributed scheduling platform on which the message queue depends, storing some kafka metadata information. Configure the serialization encoding information for receiving character streams and the kafka topic. Split the received character stream into paragraph texts according to line breaks.

[0037] S2. Data preprocessing step. Create a derived stream to process the input stream created in S1, including data filtering and Chinese-English conversion.

[0038] In specific implementation, hive sql is used to create a derived stream. The json_tuple() function supported by hive is used to parse the Chinese character fields in the input stream in S1 and convert them to English encoding using the as keyword to complete the Chinese-English encoding conversion.

[0039] Furthermore, the source of the stream data is the input stream created in S1, and the filtering condition of the stream data is that the json string is not empty and the target name of the input stream is not empty.

[0040] S3. Create a dynamic stream task to unify the key attributes of the derived stream according to the mapping dictionary table: Use the target name in the stream data to compare and judge with the target name or keyword in the mapping dictionary table. If the match is successful, use the hull number in the dictionary table to replace the hull number in the stream data to unify the hull number.

[0041] In specific implementation, in the primary key consistency judgment stage, create a dynamic stream task, and the from source is the derived stream name in the previous step.

[0042] Among them, the dynamic stream is associated with the static mapping dictionary table through the left join association method, and the equality of the target name in the stream data and the target name in the mapping dictionary table is used to judge whether the dictionary value matches;

[0043] At this time, if the match is not successful, then check whether the target name in the stream data is included in the keywords in the mapping dictionary table. The maintenance of the mapping dictionary table is in S5, and the keywords maintain all alias fields of the same target.

[0044] If the match is successful, use the hull number in the dictionary table to replace the hull number in the stream data to unify the hull number.

[0045] On the contrary, if the match fails, it means that there is no corresponding dictionary in the dictionary table data, or there is no alias that can be matched for the target, and sql makes a mark, and later manually selects a dictionary for matching for historical data.

[0046] S4. Data Persistence and Push. Store the stream data created in S3 into the hbase table, and send the processed stream data through the message queue for use.

[0047] In specific implementation, first use hive sql to create a stream task. After "from", follow the stream name created in S3, map all fields in the stream to the pre-established hbase table for persistent storage.

[0048] Continue to create a stream task. After "from", follow the stream name created in S3, fetch data from the stream in the previous step, restore the originally English fields in the stream back to Chinese fields, concatenate them into a json string, and use the stream task to publish the json string.

[0049] S5. Online Maintenance Steps of the Dictionary. If you need to add target data to the dictionary, select the relevant dictionary tables in the basic dictionary library, and map the same attribute fields in different dictionary tables to the mapping dictionary table.

[0050] In specific implementation, click the new mapping dictionary target in the situation fetching system. The system selects multiple isomorphic tables (such as the general equipment table) from the basic dictionary table for tree-shaped target display. These tables have a tree structure and similar field contents. Map the fields of these similar heterogeneous tables to the mapping dictionary table uniformly, and at the same time, extend the keywords in the dictionary table.

[0051] Furthermore, for the keywords in the extended table, when adding a new dictionary target, copy the keyword information in the original dictionary table to the new table, and support manual modification of the keywords in the dictionary table. Append the target aliases to be added to the mapping dictionary keywords.

[0052] The present invention provides a method for fetching and online processing real-time situation data, including:

[0053] (1) Support mainstream message queues.

[0054] (2) Use sql to flexibly process stream data.

[0055] (3) Use dynamic stream data to perform real-time matching with static dictionaries.

[0056] (4) Use the method of real-time maintaining the dictionary to dynamically change the stream matching status.

[0057] Further, the support for mainstream message queues includes kafka, rabbit mq, rocket mq, etc.

[0058] Further, in the use of sql to flexibly process stream data, without affecting performance, the data in the stream data can be flexibly processed for convenient viewing and modification.

[0059] Furthermore, in the real-time matching of the dynamic flow data and the static dictionary, the hive sql method is used to flexibly match the field values in the dynamic flow and the static dictionary table.

[0060] Furthermore, in the method of using the real-time maintenance dictionary to dynamically change the flow matching status, the static dictionary table can be modified externally, updated to the real-time flow, and the flow can be automatically refreshed to make the static dictionary matched by the flow take effect.

[0061] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.

Claims

1. A method for real-time situation data extraction and processing, Characterized in that, The method includes the following steps: S1. Real-time stream extraction step, create an input stream, connect to the data source, and configure the connection parameters of the message queue; S2. Data preprocessing step, create a derived stream, process the input stream created in S1, including data filtering and Chinese-English conversion; S3. Create a dynamic stream task, unify the key attributes of the derived stream according to the mapping dictionary table: use the target name in the stream data to compare and judge with the target name or keyword in the mapping dictionary table. If the match is successful, replace the vessel number in the stream data with the vessel number in the dictionary table to unify the vessel number; S4. Data persistence and push, store the stream data created in S3 into the hbase table, and send the processed stream data through the message queue for use.

2. The method for real-time situation data extraction and processing according to claim 1, Characterized in that, When selecting the message queue in step S1, the supported message queues include kafka and rabbit mq.

3. The method for real-time situation data extraction and processing according to claim 1, Characterized in that, Step S1 specifically includes: using hive sql, configure the zookeeper and kafka nodes, where kafka is the received message queue, and zookeeper is the distributed scheduling platform relied on by the message queue, storing some kafka metadata information; configure the serialization encoding information for receiving character streams and the kafka topic, and split the received character stream into paragraph texts according to the line break character.

4. The method for real-time situation data extraction and processing according to any one of claims 1-3, Characterized in that, Step S2 specifically includes: using hive sql to create a derived stream, using the json parsing function json_tuple() supported by hive to parse the Chinese character fields in the input stream in S1, and converting them to English encoding using the as keyword to complete the Chinese-English encoding conversion; the source of the stream data is the input stream created in S1, and the filtering condition of the stream data is that the json string is not empty and the target name of the input stream is not empty.

5. The method for real-time situation data extraction and processing according to claim 1, Characterized in that, Step S3 to create a dynamic stream task specifically includes: create a dynamic stream task, and the from source is the name of the derived stream in the previous step.

6. The method for real-time situation data extraction and processing according to claim 5, Characterized in that, In step S3, the target name in the streaming data is compared with the target name or keyword in the mapping dictionary table. If a successful match is found, the vessel number in the dictionary table is used to replace the vessel number in the streaming data to unify the vessel numbers. Specifically, it includes: associating the dynamic stream with the static mapping dictionary table through a left join, and judging whether the dictionary value matches by checking whether the target name in the streaming data is equal to the target name in the mapping dictionary table. If the match is not successful at this time, then check whether the target name in the streaming data is included in the keywords in the mapping dictionary table, where the keywords maintain all alias fields of the same target. If a successful match is found, the vessel number in the dictionary table is used to replace the vessel number in the streaming data to unify the vessel numbers.

7. The method for real-time situation data ingestion and processing according to claim 6, characterized in that, step S3 further includes: if the match is not successful, it means that there is no corresponding dictionary in the dictionary table data, or there is no alias that can be matched for the target, and sql is used for marking, and later manual selection of the dictionary is performed for historical data matching.

8. The method for real-time situation data ingestion and processing according to any one of claims 5-7, characterized in that, step S4 specifically includes: first, use hive sql to create a stream task, with the stream name created by S3 following from, map all fields in the stream to a pre-established hbase table for persistent storage; continue to create a stream task, with the stream name created by S3 following from, ingest data from the stream in the previous step, restore the originally English fields in the stream to Chinese fields again, concatenate them into a json string, and use the stream task to publish the json string.

9. The method for real-time situation data ingestion and processing according to claim 8, characterized in that, after step S4, there is further step S5: the online maintenance step of the dictionary. If target data needs to be added to the dictionary, select the relevant dictionary tables in the basic dictionary library, and map the same attribute fields in different dictionary tables to the mapping dictionary table.

10. The method for real-time situation data ingestion and processing according to claim 9, characterized in that, step S5 specifically includes: click to add a mapping dictionary target in the situation ingestion system, the system selects multiple isomorphic tables from the basic dictionary table for tree-shaped target display, these tables have a tree structure and similar field contents, map the fields of these similar heterogeneous tables to the mapping dictionary table uniformly, and at the same time, expand the keywords in the dictionary table; for the keywords in the extended table, when adding a new dictionary target, copy the keyword information in the original dictionary table to the new table, and support manual modification of the keywords in the dictionary table, and append the target alias to be added to the mapping dictionary keywords.

Citation Information

Patent Citations

  • Data connection and visualization method and system

    CN110569298A

  • Real-time data alarm method based on stream processing engine and rule engine

    CN111444291A