Text format meteorological data configuration general decoding storage method and device

Through configured general decoding programs and data mapping technology, the problem of inefficient decoding of meteorological data is solved, rapid decoding and database entry are achieved, development difficulty is reduced, and meteorological data processing efficiency is improved.

CN120066522APending Publication Date: 2025-05-30STATE QIXIANG INFORMATION CENT
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510470117.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the decoding processing of meteorological data is inefficient, and the code needs to be frequently modified in the face of diversified meteorological data format changes, resulting in high development costs and low efficiency.

Method used

The configuration-based general decoding program is adopted, combined with RabbitMQ message middleware and directory polling methods to obtain meteorological data files. Through file-level, report-level, group-level and feature-level inspections, it uses configuration rules to decode and map it to database fields to achieve rapid library entry.

Benefits of technology

提高了气象数据的解码处理效率,减少了开发难度,普通运维人员也能操作,快速响应数据格式变化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066522A_ABST
    Figure CN120066522A_ABST
Patent Text Reader

Abstract

The invention discloses a text format meteorological data configuration general decoding storage method and device, and relates to the technical field of meteorological information processing, and the method comprises the steps: obtaining a to-be-decoded meteorological data file based on a configured general decoding program through employing a real-time or directory polling data access mode; checking the format of the meteorological data file, and decoding the meteorological data file based on a preset configuration analysis rule to obtain a report so as to obtain elements; and performing calculation conversion mapping on the obtained elements based on configuration to form a storage statement, and calling a storage interface to realize data storage. According to the invention, the decoding processing efficiency of the meteorological data can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of meteorological information processing, and particularly relates to a method and device for configuring and general decoding and warehousing of text-format meteorological data. Background Art

[0002] For the diverse data generated by meteorological observations, including raw observation data such as ground, upper air, ocean, radiation, and atmospheric composition, as well as various dataset products, many of these meteorological data are stored in a structured manner. In actual use, it is necessary to parse the original meteorological data file to obtain information such as station number, time, and each element value, and then store the element data in a relational database for users to retrieve, providing services for meteorological forecasting and early warning and other services.

[0003] The encoding rules of different data types are different, and the included elements are also different. Therefore, the decoding of the original meteorological data files all adopts a dedicated development method, developing corresponding programs for specific data formats. Even a slight change in the data format requires modifying the code and recompiling and packaging again. At the same time, with the increase in the number of data types, this development method is extremely time-consuming and laborious when dealing with hundreds or thousands of types of meteorological data files. Summary of the Invention

[0004] This application provides a method and device for configuring and general decoding and warehousing of text-format meteorological data, which can effectively improve the decoding and processing efficiency of meteorological data.

[0005] In a first aspect, an embodiment of this application provides a method for configuring and general decoding and warehousing of text-format meteorological data. The method for configuring and general decoding and warehousing of text-format meteorological data includes: Based on the configured general decoding program, and using a real-time or directory polling data access method to obtain the meteorological data file to be decoded; Check the format of the meteorological data file, and decode the meteorological data file based on a preset configuration parsing rule to obtain a report and then obtain elements; Perform calculation conversion mapping on the obtained elements based on the configuration to form a warehousing statement, and call the warehousing interface to implement data warehousing.

[0006] Combined with the first aspect, in an implementation manner, The real-time data access method is in the form of combining a notification message and a data file. The notification message is implemented using the RabbitMQ message middleware, and each notification message stores the data type and path information of the meteorological data file; The directory polling data access method is based on the directory to be polled and the data type encoding input in the configuration file, and performs polling processing in the corresponding directory to obtain the absolute paths of all meteorological data files in the corresponding directory.

[0007] In one implementation manner in combination with the first aspect, The format check of the meteorological data file includes file-level check, report-level check, group-level check, and element-level check; The file-level check includes whether the file name is empty, whether the file exists, whether it is a file, and whether the file is empty; The report-level check is to check whether the report length is within the set length range, and to judge whether the number of elements obtained after splitting the report by the delimiter is the same as the specified value; The group-level check includes group length check and group code check; The element-level check includes element code check, element length check, latitude check, longitude check, time format check, and time limit check.

[0008] In one implementation manner in combination with the first aspect, decoding the meteorological data file according to the preset configuration parsing rule to obtain a report and then obtain elements, specifically including: Traverse the meteorological data file line by line, and obtain the data traversed this time to form a report when reaching the report delimiter. The report delimiter is default to the line break character, and the report delimiter can be modified through configuration; Parse the report based on the element delimiter to obtain multiple strings, and number the strings. Then each numbered string is an element obtained by decoding. The element delimiter is default to a space or a tab character, and the element delimiter can be modified through configuration.

[0009] In one implementation manner in combination with the first aspect, Form the storage statement based on the basic configuration and the storage field mapping configuration; The basic configuration is used to obtain the queue information, data encoding, program behavior control, delimiter setting, and station longitude and latitude of the task. And the basic configuration includes public configuration and private configuration. The formats of the public configuration and the private configuration are the same, and they are loaded in the order of loading the public configuration first and then the private configuration; The storage field mapping configuration is used to establish a mapping relationship between the decoded elements and each field in the data table of the database, and realize the generation of the configured storage statement.

[0010] In one implementation manner in combination with the first aspect, The key configuration items of the basic configuration include report delimiter, first element delimiter, second element delimiter, maximum number of lines to read, location of the station information configuration file, update strategy, and data storage time check range; The report delimiter is used to split the meteorological data file to obtain a report; The first element delimiter is used to split the report to obtain the strings of each element; The second element delimiter is used to perform secondary splitting on the elements obtained by splitting with the first element delimiter; The station information configuration file is used to store the longitude, latitude, and altitude information of the station, so that the longitude, latitude, and altitude information of the corresponding station can be obtained from the station information configuration file based on the configured station number; The station information configuration file is in text format and stores information using key-value pairs and value values. The content of the key-value includes the station number and the station network code, and the content of the value value includes the station number, longitude, altitude, administrative division code, station type, switching level, spare field, and district association code; The update strategy is a processing strategy for duplicate data.

[0011] In combination with the first aspect, in an implementation manner, The mapping configuration of the fields for storage in the database is configured in XML format. The attribute index is used to represent the number of elements obtained by splitting the meteorological data file with the report delimiter and then splitting the report with the element delimiter. The label value is used to represent the name of the field for storage in the database, and built-in functions are used to implement expression calculation processing; The key configuration items of the mapping configuration of the fields for storage in the database include whether to perform group number verification, the number of group number verifications, the type of data for storage in the database, the value taking method, the position of the report for value taking, the default default value, the time zone of the data time, the landing time for obtaining the meteorological data file, date formatting, and the expression for configuring the field value of the current row.

[0012] In combination with the first aspect, in an implementation manner, for the mapping between the elements and the fields in the data table of the database, the processing process includes: converting the data type of the value of the element, performing mathematical calculations on the value of the element, taking the value of the element and performing truncation, string concatenation, converting Beijing time to world time, converting world time to Beijing time, obtaining station information, and processing specific fields; Among them, when the name of the data table changes, modify the name of the data table in the configuration. When a new field is added to the data table, add a row of configuration for the value taking of the new field. When a field is reduced in the data table, delete the configuration of the corresponding field.

[0013] In a second aspect, an embodiment of the present application provides a text format meteorological data configuration-based general decoding and storage device. The text format meteorological data configuration-based general decoding and storage device includes: An acquisition unit, which is used to obtain the meteorological data file to be decoded based on the configured general decoding program and using a real-time or directory polling data access method; A decoding unit for checking the meteorological data file format and decoding the meteorological data file based on a preset configuration parsing rule to obtain a report and then obtain elements; A calculation and conversion unit for performing calculation conversion mapping on the obtained elements based on the configuration to form a storage statement; A storage unit for calling a storage interface for the formed storage statement to implement data storage.

[0014] Combined with the second aspect, in an implementation manner, The real-time data access method is in the form of combining a notification message and a data file. The notification message is implemented using the RabbitMQ message middleware, and each notification message stores the data type and path information of the meteorological data file; The directory polling data access method is based on the directory to be polled and the data type code input in the configuration file, and polling processing is performed in the corresponding directory to obtain the absolute paths of all meteorological data files in the corresponding directory.

[0015] The beneficial effects brought by the technical solution provided by the embodiments of the present application include: It can realize the rapid decoding of text format meteorological data, complete the mapping between the decoded elements and the database fields according to the configuration, realize the rapid storage of meteorological data, reduce the development difficulty of meteorological data decoding, and at the same time reduce the technical difficulty brought by the addition of new meteorological data and format changes, enabling ordinary operation and maintenance personnel to perform corresponding operations, and effectively improving the decoding processing efficiency of meteorological data. Description of the Drawings

[0016] Figure 1 It is a schematic flowchart of the configuration-based general decoding and storage method for text format meteorological data of the present application; Figure 2 It is an example report; Figure 3 It is a schematic diagram of the functional modules of the configuration-based general decoding and storage device for text format meteorological data of the present application. Detailed Embodiments

[0017] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0018] To make the purpose, technical solution and advantages of the present application clearer, the embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings.

[0019] In a first aspect, an embodiment of the present application provides a method for general decoding and warehousing of text format meteorological data with configuration.

[0020] In one embodiment, referring to Figure 1 , Figure 1 is a schematic flow diagram of the method for general decoding and warehousing of text format meteorological data in the present application. As Figure 1 shown, the method for general decoding and warehousing of text format meteorological data includes: S1: Based on the configured general decoding program, and using real-time or directory polling data access mode to obtain the meteorological data file to be decoded; S2: Check the format of the meteorological data file, and decode the meteorological data file based on the preset configuration parsing rules to obtain a report and then obtain elements; S3: Perform calculation conversion mapping on the obtained elements based on the configuration to form a warehousing statement, and call the warehousing interface to implement data warehousing.

[0021] For the decoding and warehousing in the present application, the specific process is as follows: Split the meteorological data file to obtain a report, then parse the report to obtain elements, then perform calculation conversion on the elements to adapt to the data table fields in the database, combine the calculated and converted elements with some preset variables to form a warehousing statement, and finally call the warehousing interface to warehouse the warehousing statement.

[0022] In the present application, a combination of RabbitMQ (message-oriented middleware) message notification and directory polling is used to obtain the meteorological data file to be processed. First, check the format of the meteorological data file, then extract the corresponding report, parse the report to obtain elements, and then according to the conversion rules, splice the elements into a warehousing statement, and finally call the warehousing interface to implement data warehousing.

[0023] It should be noted that the configured general decoding program supports two data access modes: real-time and directory polling. The real-time data access mode is in the form of combining a notification message and a data file. The notification message is implemented using the RabbitMQ message middleware, and each notification message stores the data type and path information of the meteorological data file.

[0024] The specific format of the notification message is "meteorological data file data type: full path of meteorological data file", for example: A.0060.0001.R001: / space / dpc / work / data / A / A.0060.0001.R001 / 202412 / 2024121809 / 01 / Z_SEVP_C_BEHB_P_20241218092153_ARTSTN_202412180900.csv Among them, A.0060.0001.R001 represents the data type of the meteorological data file, which is a data type code designed internally by the meteorological department. Generally, the same type of data has a similar file name rule; the full path of the meteorological data file is after the colon. After receiving the notification message from the specified queue in rabbitMQ, the processing framework reads the meteorological data file from the shared file system according to the full path, and calls the pre-configured processing logic according to the data type to implement the full process processing of format checking, decoding, and warehousing of the meteorological data file.

[0025] The data access method of directory polling is based on the directories to be polled and data type codes input in the configuration file, and polling is performed in the corresponding directories to obtain the absolute paths of all meteorological data files in the corresponding directories. Specifically, for the data access method of directory polling, it is necessary to input the directories to be polled and data type codes in the corresponding configuration file. The processing framework will perform polling in the specified directories to obtain the absolute paths of all meteorological data files in these directories, and call the pre-configured processing logic according to the data type of the meteorological data file to implement the full process processing of format checking, decoding, and warehousing of the meteorological data file.

[0026] Furthermore, in one embodiment, after receiving the meteorological data file, the processing framework first calls the format checking module to check the format of the meteorological data file, and performs the next step according to the check result. The format check of the meteorological data file includes file-level check, report-level check, group-level check, and element-level check.

[0027] It should be noted that a meteorological data file contains one or more reports, and each report ends or starts with a fixed format. A report consists of one or more lines of data to form the smallest complete meteorological data coding unit; usually, a report generally only contains one line. Therefore, usually, one line of data is a report. Each string generated by splitting the report by the element separator is an element. The key to data decoding is to effectively extract these elements and write them into the database table according to the designed warehousing rules.

[0028] The format checking module performs four types of checks: file-level, report-level, group-level, and element-level, according to the data type of the obtained meteorological data file, based on the coding rules of various materials, and combined with the errors that are likely to occur in the actual decoding process.

[0029] Specifically, the file-level check includes whether the file name is empty, whether the file exists, whether it is a file, and whether the file is empty. For the check of whether the file name is empty, that is, to determine whether the file name is empty, if it is empty, an error code is returned, error log information is output, and whether the corresponding EI code (error code) is sent according to the configuration, and then the next meteorological data file is judged; for the check of whether the file exists, that is, to determine whether the file exists, if it does not exist, an error code is returned, error log information is output, and whether the corresponding EI code is sent according to the configuration, and then the next meteorological data file is judged; for the judgment of whether it is a file, that is, to determine whether the meteorological data file is in file form, if not (for example, in directory form), an error code is returned, error log information is output, and whether the corresponding EI code is sent according to the configuration, and then the next meteorological data file is judged; for the check of whether the file is empty, that is, to determine whether the size of the meteorological data file is 0, if it is 0, an error code is returned, error log information is output, and whether the corresponding EI code is sent according to the configuration, and then the next meteorological data file is judged.

[0030] Specifically, the group-level check includes group length check and group coding check. For the group length check, it is to check whether the whole group of data is within the specified length range, otherwise it returns an error code, outputs error log information, and decides whether to send the corresponding EI code according to the configuration, and the elements contained in the current group are all treated as missing, and then continue to process the next group of data; for the group coding check, it is to check whether the character coding range of the whole group of data is within the specified range, otherwise it returns an error code, outputs error log information, and decides whether to send the corresponding EI code according to the configuration, and the elements contained in the current group are all treated as missing, and then continue to process the next group of data. Group coding is implemented through configuration according to the data type.

[0031] Specifically, the report-level check is to check whether the report length is within the set length range, and to determine whether the number of elements obtained after the report is split by a delimiter is the same as the specified value. For checking whether the report length is within the set length range, it is to determine whether the report length is too long. If so, an error code is returned, error log information is output, and it is determined according to the configuration whether the corresponding EI code is sent, and then the current report is skipped, the number of error reports is increased by 1, and the next report is processed; and, it is determined whether the report length is less than the specified length. If so, an error code is returned, error log information is output, and it is determined according to the configuration whether the corresponding EI code is sent, the number of error reports is increased by 1, and then the current report is skipped and the next report is processed. The threshold of the report length is implemented through configuration according to the data type.

[0032] Specifically, the feature-level check includes feature code check, feature length check, latitude check, longitude check, time format check, and time limit check.

[0033] For element coding check, to check whether the placeholder code of the element coding is within the specified range, otherwise return an error code, output error log information, and decide whether to send the corresponding EI code according to the configuration. Except for special elements, the current element is treated as missing measurement.

[0034] For element length check, to check whether the length of the segmented element placeholder is within the specified range, otherwise return an error code, output error log information, and decide whether to send the corresponding EI code according to the configuration. Except for special elements, the current element is treated as missing measurement.

[0035] For latitude check, the general range of latitude is [-90, 90]. Check whether the decoded latitude is within the specified range, otherwise return an error code, output an error log, and decide the subsequent processing strategy according to the configuration.

[0036] For longitude check, the general range of longitude is [-180, 180]. Longitudes greater than 180 need to be converted. The conversion rule is the decoded value - 360. Output a log for explanation when the conversion occurs. Check whether the converted longitude is within the specified range, otherwise return an error code, output an error log, and decide the subsequent processing strategy according to the configuration.

[0037] For time format check, check whether the decoded time is a valid time format, otherwise return an error code, output error log information, and decide whether to send the corresponding EI code according to the configuration. Determine that the current element is a special element and skip the processing of the current report.

[0038] For time limit check, check the time difference between the decoded data observation time and the file landing time. If it is within the specified range, process it normally, otherwise return an error code, output error log information, and decide whether to send the corresponding EI code according to the configuration. Determine that the current element is a special element and skip the processing of the current report.

[0039] The time limit check is implemented through configuration according to the data type, and there is a separate control option to determine whether to perform it. Calculate the difference between the file landing time and the data observation time, with the unit of minutes. A negative value indicates that the file arrived before the specified observation time, and a positive value indicates that it arrived after the specified observation time.

[0040] Further, in one embodiment, parsing the meteorological data file based on a preset configuration parsing rule to obtain a report and then obtain elements, specifically including: S201: Traverse the meteorological data file line by line, and obtain the data traversed this time to form a report when reaching the report separator. The report separator is defaulted to the line break character, and the report separator can be modified through configuration; S202: Parse the report based on the element separator to obtain multiple character strings, and number the character strings. Each numbered character string is an element obtained by decoding. The element separator is a space or a tab by default, and the element separator can be modified through configuration.

[0041] In one possible case, the elements include station, observation time, precipitation, temperature value, wind speed value, wind direction value, maximum wind speed value, relative humidity value, and air pressure value. In practical applications, the actual content of the elements is determined by the type of meteorological data file.

[0042] That is, after the format check is passed, the processing framework calls the feature parsing module to perform the corresponding feature parsing. When parsing the meteorological data file, the data is read line by line. The line starting with the comment character will be skipped. When the report delimiter is encountered, the read data will be formed into a whole report. The report delimiter defaults to a newline character, which can be modified through the configuration file. After obtaining the report, the report is parsed according to the feature delimiter. The feature delimiter defaults to a space or a tab character, and consecutive delimiters are regarded as one. The parsed strings are numbered from 0 to n. Figure 2 As shown in the figure, the meteorological data file is a report per line (i.e., the line break is the report separator), the line starting with # is a comment line, and the comma is the element separator. After parsing, each report has 9 elements, which are one of the source data for the storage conversion configuration.

[0043] In order to perform more refined processing, a secondary segmentation can be performed after the primary segmentation. For example, the value of the element numbered 1 is "2024-07-12 08:00:00". After segmentation by [-:space], information such as year, month, and day can be easily extracted.

[0044] Further, in one embodiment, after the element parsing is completed, the storage conversion configuration module is called according to the data type to realize the conversion from the parsed element to the database storage record. The present application realizes the formation of the storage statement based on the basic configuration and the storage field mapping configuration.

[0045] The basic configuration is used to obtain the task's queue information, data encoding, program behavior control, separator settings, and site longitude and latitude, and the basic configuration includes public configuration and private configuration, and the format of the public configuration and private configuration is the same. When the processing framework is started, the public configuration file is loaded first, and then the private configuration file is loaded. When there are duplicates, the private configuration shall prevail. The final configuration of the general decoding program is determined by the combination of public + private.

[0046] For the public and private configurations of the basic configuration, examples are as follows: Public configuration file name: default dpc_global_config.properties Private configuration file: data_code_name.properties Where data_code_name is defined differently according to the data type.

[0047] In this application, the key configuration items of the basic configuration include report separator, first element separator, second element separator, maximum number of lines to read, location of the station information configuration file, update strategy, and data storage time check range; the report separator is used to split the meteorological data file to obtain reports, and the default is the line break character; the first element separator is used to split the report to obtain the strings of each element; the second element separator is used to perform secondary splitting on the elements obtained by splitting with the first element separator, and the default is empty and only takes effect when configured; the maximum number of lines to read is the upper limit of the number of lines read at one time for large files, to avoid problems such as memory overflow caused by processing too many files at one time.

[0048] The station information configuration file is used to save information such as the longitude, latitude, and altitude of the station, so that the longitude, latitude, altitude, etc. of the corresponding station can be obtained from the station information configuration file based on the configured station number. The station information configuration file is in text format and stores information using key-value pairs. The content of the key includes the station number and the station network code, and the content of the value includes the station number, longitude, altitude, administrative division code, station type, exchange level, spare field, and regional association code.

[0049] Meteorological data is sensitive to longitude and latitude, but often the longitude and latitude information of the station is stored separately from the observation data, that is, the original meteorological data file does not contain longitude and latitude information. Therefore, a special configuration file is designed to save the longitude and latitude information of the station. The general decoding program can obtain the corresponding longitude and latitude information from the configuration file according to the configured station number and generate the storage statement. The format of this configuration file is text format, with one record per line. The key and value are separated by a space, and the key cannot be missing and is the station number + station network code. The value contains different attribute information, separated by commas, and are the sub-station number, longitude, altitude, administrative division code, station type, exchange level, spare field, and regional association code in sequence. Different attributes in the value can be missing, and the default value is null. For example: 54511+01 null,116.4667,39.8,31.3,2250,110115,012,101,null,2.

[0050] The update strategy is the processing strategy for duplicate data. For the processing strategy of duplicate data, generally, meteorological data has its own correction sequence, from CCA to CCX. However, there are also many data that do not contain a correction sequence. These data generally follow the principle of first-in-first-served or later-in-first-served, and need to be set according to different data.

[0051] For the inspection range of data storage time, during the real-time decoding process of meteorological data, there are abnormal data such as some historical data and data beyond the current time range. It is necessary to filter according to the specific data type to avoid abnormal data from being stored in the database.

[0052] The following gives examples of the configuration items of the basic configuration.

[0053] The general public configuration includes the following: Configuration description: dpc_global_config.properties Configure the connection information of the message middleware Configuration example: # IP address configuration of the RabbitMQ message middleware server RabbitMQ.host=10.40.120.49 # Username configuration of the RabbitMQ message middleware server RabbitMQ.user=******* # Password configuration of the RabbitMQ message middleware server user RabbitMQ.passWord=******* # Storage program access port number configuration of the RabbitMQ message middleware server RabbitMQ.port=5676 # Whether DI records logs, the default is DI.WRITE.LOG=0 # Inspection range of data storage time, unit: hour D_DATETIME_BEFORE_DAY=24 D_DATETIME_AFTER_DAY=672 # Configure the global DI and EI sending control switches di.option=1 ei.option=1.

[0054] The general private configuration includes the following: Configuration example: data_code_name.properties # RabbitMQ message queue name RabbitMQ.queueName=UPAR_ORI_B.0011.0008.R002_001 # Number of sending processing threads, default is 1 (greater than or equal to 1) dpc.control.process.ThreadCount=1 # Number of DI sending processing threads, default is 1 (greater than or equal to 1) dpc.control.di.ThreadCoun=1 # Whether to send DI log information '', '' default = 1 (=1 for normal sending, =0 means not sending DI) dpc.control.di.option=1 # Whether to send EI, default = 1 (=1 for normal sending, =0 means not sending EI) dpc.control.ei.option=1 # Data processing mode, default is 0 '', '' (0 File path notification message; 1 Directory polling mode; 2 Data message processing mode) dpc.control.processmode.option=0 # Polling flag, whether to exit after polling, default is 1 (1 Exit after polling, 0 Do not exit after polling) dpc.control.fileloop.quitflag=0 # If polling is empty, the waiting time for the next polling, unit is millisecond '', '' default is 10000 dpc.control.fileloop.sleepTime=10000 # Table name for storage dpc.process.conf.tables=UPAR_WEA_CHN_MUL_FTM_DROP_TAB # CTS encoding dpc.process.conf.ctsCode=B.0011.0008.R002 # SOD encoding dpc.process.conf.sodCodes=B.0011.0008.S002 # BDMAIN: Big data platform main process; BDBAK : Big data platform backup process monitor.di.option.data_flow=BDMAIN # Data source directory, only effective during directory polling dpc.control.fileloop.sourcedir= / space / dpc / work / data / B / B.0011.0008.R002 / in ## Path to move to after data processing is completed, only effective during directory polling dpc.control.fileloop.targetdir= / space / dpc / work / data / B / B.0011.0008.R002 / out # Directory + file name of the data processing log dpc.log.message.url=.. / logs / %d{yyyy-MM-dd} / dpc-UPAR_WEA_CHN_MUL_FTM_DROP_TAB-message-{port}-%d{yyyy-MM-dd}.log # Directory + file name of the data processing log dpc.log.process.url=.. / logs / %d{yyyy-MM-dd} / dpc-UPAR_WEA_CHN_MUL_FTM_DROP_TAB-process-{port}-%d{yyyy-MM-dd}.log # Report separator dpc.check.report.splitregex1=\n # Element separator 1 dpc.check.element.splitregex1=\\t # Element separator 2 dpc.check.element.splitregex2=\\t #dpc.check.report.splitregex=, # Starting line number for report processing dpc.check.report.startline=1 # Ending line number for report processing dpc.check.report.endline= # Table name for storage, stored according to the configuration in db.properties dpc.custom.storage.dbname=rdb # Large report parsing, number of lines read each time, default is 1000 lines dpc.custom.decode.count=10000 # Station information configuration file format: station number, longitude and latitude, etc. Note: The station number should be in the first column dpc.custom.station.excel.path= / space / cmadaas / dpc / CMADAAS_DPC_DECODE / bin / config / stationconfig.txt # Whether to update using the correction report, '', default is 1 (1 uses the BBB item to update, 0 ignores the BBB item, first come first served, 2, later arrivals prevail) dpc.check.update.useBBB=0

[0055] In this application, the storage field mapping configuration is used to establish a mapping relationship between the elements obtained by decoding and each field in the data table of the database, so as to generate a configurable storage statement. That is, the storage field mapping configuration module realizes the establishment of a mapping relationship between the parsed elements and each field in the data table of the database, so as to generate a configurable storage statement. The storage field mapping configuration is configured in XML format, using the attribute index to represent the number of the report obtained by splitting the meteorological data file with the report separator and the elements obtained by splitting the report with the element separator, using the tag value to represent the storage field name, and the built-in function to implement the expression calculation process. That is, the storage field mapping configuration is configured in XML format, using the index function to represent the element sequence number parsed from the report, using the tag value to represent the storage field name, and implementing special processing such as expression calculation through the built-in function.

[0056] In this application, the database is a structured database such as ORACLE, Xugu, MYSQL, etc. A data table in the database may contain multiple records, each record contains one or more fields, and each field has its corresponding application meaning. The processing framework can obtain all or part of the records of the data table through the database access interface and calculate one or more fields to obtain the required information.

[0057] The table names and field names for configured parsing and warehousing can be modified according to the corresponding configuration files, so as to realize the warehousing of data in different formats into different data tables. When adding a new field to a data table, only one line of configuration needs to be added to correspond to the value of the field; when reducing a field in a data table, only the configuration of the corresponding field needs to be deleted in the corresponding configuration file. Specifically, when the data table name changes, only the data table name in the configuration needs to be modified; when a new field is added to the data table, only one line of configuration needs to be added to correspond to the value of the new field; when a field in the data table is reduced, only the configuration of the corresponding field needs to be deleted in the corresponding configuration file. The data tables and related fields for configured warehousing need to be created in the database before the general decoding program is started, otherwise an error will be reported when calling the warehousing interface.

[0058] The key configuration items for the mapping configuration of warehousing fields include whether to perform group number verification, the quantity of group number verification (only passing when the number of elements after splitting is equal to this value), the data type for warehousing, the value-taking method (default to take the default value, report to take the reported value), the reported value-taking position (the subscript of the corresponding value after splitting according to the report delimiter, starting from 0), the default default value, the time zone of the data time (GMT+8 represents Beijing time), the landing time for obtaining meteorological data files (the default value), date formatting, and the expression for configuring the field value of the current line.

[0059] Regarding the expression, the explanations are as follows: (1) The data time takes the subscript of the corresponding value in the report, starting with [ and ending with ], and multiple values are separated by commas, e.g., [4,5,6,7]; (2) The addition, subtraction, multiplication, and division calculations of data start with the calculation operator, e.g., -273.2; (3) Data truncation, [start,end] configures the start and end positions, indicating taking the value from the start position to the end position of the report intercepted at the corresponding index position (reported value-taking position) in the report (String type); (4) Filling in data directly means taking the current data (in string format), e.g., 08; if it is longitude and latitude, the value taken is the station network number of the station.

[0060] In this application, for the mapping between elements and fields in the data tables of the database, the processing process includes: taking the value of the element for data type conversion, taking the value of the element for mathematical calculations (i.e., performing addition, subtraction, multiplication, and division calculations on the value of the element), taking the value of the element and performing truncation, string concatenation, converting Beijing time to world time, converting world time to Beijing time, obtaining station information, and processing specific fields.

[0061] Regarding taking the value of the element for data type conversion, the examples are as follows: dataType="double" content="report" value="999999" index="1" Indicates the second set of values of the element, converted to double type, with a default value of 999999.

[0062] For mathematical calculations on the values of the element, examples are as follows: dataType="double" content="report" expression="-273.2" index="1" Indicates the second set of values of the element "-273.2".

[0063] For truncating the values of the element, examples are as follows: content="report" index="10" expression="[17:19]" Indicates the second set of values of the element and truncates the 17th to 19th digits.

[0064] For string concatenation, for example, the data time is concatenated from V04001, V04002... and the time field format is specified. The example is as follows: dataType="datetime" content="report" expression="[4,5,6,7]" format="yyyyMMddHH" Indicates the 5th, 6th, 7th, and 8th sets of data of the element (i.e., V04001\V04002\V04003\V04004) are concatenated and converted to a time in the yyyyMMddHH format.

[0065] For converting Beijing time to Coordinated Universal Time (UTC), examples are as follows: dataType="datetime" content="report" index="10" timeZone="GMT+8"format="yyyy-MM-dd hh:mm:ss" Indicates the 11th set of data of the element (data time) is converted to UTC, with the format yyyy-MM-dd hh:mm:ss.

[0066] For obtaining station information, examples are as follows: dataType="double" content="default" value="V06001" expression="08" It represents retrieving the longitude of the station with a battle net of 08 in the lua file.

[0067] For the processing of specific fields, examples are as follows: Warehousing time, update time: content = "default" value = 'SYSDATE' It represents retrieving the current system time in the format of yyyy-MM-dd hh:mm:ss; File reception time: content = "default" value = "FILE_LAST_MODIF" It represents retrieving the file landing time in the format of yyyy-MM-dd hh:mm:ss; Data identifier: content = "default" value = "B.0011.0008.S002" It represents assigning the value of B.0011.0008.S002.

[0068] Furthermore, it also includes a multi-variable calculation module, which is used to support complex expression calculations for multiple variables and expand the data mapping ability. The multi-variable calculation module represents the decoded elements in a variable manner. The multi-variable calculation module forms the final warehousing statement through automatic identification, calculation, and conversion of variables. For example: The first segment after splitting ${0} The first segment after splitting ${0:1:2}, taking 2 characters starting from 1 The second segment after the secondary splitting of the 6th element of ${6.1}.

[0069] The first segment after secondary splitting of ${6.1:0:2}, taking 2 characters starting from 0 The last segment after splitting ${-1:0:2}, taking 2 characters starting from 0 [from left to right].

[0070] For example, an example of a multi-variable mixed calculation: ($ {0}-1000) / 10 + $ {6.1:0:2}.

[0071] Furthermore, in one embodiment, after the warehousing field mapping configuration, the general decoding program calls the database warehousing interface to warehouse the generated sql statement. The warehousing interface supports different databases such as Xugu, MYSQL, and ORACLE, and can be modified through configuration. To improve the warehousing speed during warehousing, the batch submission mode can be adopted, and the batch submission quantity can be quickly modified through configuration to adapt to different data. At the end of the file, regardless of whether the remaining records meet the batch requirements, they are immediately submitted for warehousing. After the file is processed, the processing framework will delete or retain it according to the configuration.

[0072] When the text - format meteorological data configuration - based general decoding and warehousing method of the present application is actually applied, the general decoding program is deployed on multiple servers, and the servers can form an effect of both load balancing and mutual backup, thereby improving the speed and reliability of data processing. The general decoding program obtains file - path messages from the queue, reads meteorological data files from the file system according to the paths, performs operations such as format checking, data decoding, field mapping, and variable calculation according to the data type, automatically generates warehousing statements according to the configuration, and finally calls the warehousing interface to batch - warehouse the statements. After warehousing, the original meteorological data files can be deleted or retained according to the configuration.

[0073] Due to the rapid development of meteorological technology, various new observation and exploration devices are constantly emerging, generating a variety of new data. At the same time, a large amount of data is obtained through exchanges with external departments. Many of these meteorological data are encoded in text format. Although their forms and elements are different, there are also commonalities. Therefore, in view of the characteristics of similar formats and different elements in text - format meteorological data, the present application designs a configuration - based general decoding algorithm, which realizes the core links such as format checking of meteorological data files, element parsing, and generation of warehousing statements through configuration technology. At the same time, it integrates interfaces such as task acquisition, data warehousing, and monitoring information sending to realize the full - process decoding processing of text - format meteorological data files, improve the decoding and development efficiency of new data, and can quickly respond to changes in data formats.

[0074] The text - format meteorological data configuration - based general decoding and warehousing method of the embodiments of the present application can realize the rapid decoding of text - format meteorological data, complete the mapping between decoding elements and database fields according to the configuration, and realize the rapid warehousing of meteorological data. In actual applications, nearly 80 - plus types of text - format data on the Tianqing, the core data platform of the meteorological department, are decoded and warehoused through the method described in the present application, reducing the difficulty of meteorological data decoding development. At the same time, it reduces the technical difficulties brought by new meteorological data and format changes, enabling ordinary operation and maintenance personnel to perform corresponding operations, and effectively improving the decoding and processing efficiency of meteorological data.

[0075] In a second aspect, the embodiments of the present application also provide a text - format meteorological data configuration - based general decoding and warehousing device.

[0076] In one embodiment, referring to Figure 3 , Figure 3 is a schematic diagram of the functional modules of the text - format meteorological data configuration - based general decoding and warehousing device of the present application. As Figure 3 shown, the text - format meteorological data configuration - based general decoding and warehousing device includes: an acquisition unit, a decoding unit, a calculation and conversion unit, and a warehousing unit.

[0077] The acquisition unit is used to obtain the meteorological data file to be decoded based on the configured general decoding program, and adopts real-time or directory polling data access method; the decoding unit is used to check the format of the meteorological data file, and decode the meteorological data file based on the preset configuration parsing rules to obtain a report and then obtain the elements; the calculation conversion module is used to perform calculation conversion mapping on the obtained elements based on the configuration to form a storage statement; the storage unit is used to call the storage interface for the formed storage statement to realize data storage.

[0078] In the present application, the real-time data access method is a combination of notification messages and data files. The notification messages are implemented using RabbitMQ message middleware, and each notification message stores the data type and path information of the meteorological data file; the directory polling data access method is based on the directory and data type encoding to be polled entered in the configuration file, and polling is performed in the corresponding directory to obtain the absolute path of all meteorological data files in the corresponding directory.

[0079] The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices. The terms "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit "first", "second" and "third" to different types.

[0080] In the description of the embodiments of the present application, "exemplary", "for example" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary", "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary", "for example" or "for example" is intended to present related concepts in a specific way.

[0081] In the description of the embodiments of the present application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; the “and / or” in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, “multiple” refers to two or more than two.

[0082] In some of the processes described in the embodiments of the present application, multiple operations or steps appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of the present application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.

[0083] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device to execute the methods described in the various embodiments of the present application.

[0084] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A method for configuring and storing meteorological data in text format in a universal manner, characterized in that: The text format meteorological data configuration universal decoding storage method comprises: Based on the configured universal decoding program, the meteorological data files to be decoded are obtained by real-time or directory polling data access; Check the format of the meteorological data file, and decode the meteorological data file based on a preset configuration parsing rule to obtain a report and then obtain elements; The obtained elements are calculated, transformed and mapped based on the configuration to form storage statements, and the storage interface is called to implement data storage.

2. A text format meteorological data configuration universal decoding and warehousing method as claimed in claim 1, characterized in that: The real-time data access method is a combination of notification messages and data files. The notification messages are implemented using RabbitMQ message middleware, and each notification message stores the data type and path information of the meteorological data file. The data access method of the directory polling is to perform polling processing in the corresponding directory based on the directory to be polled and the data type code input in the configuration file, and obtain the absolute path of all meteorological data files in the corresponding directory.

3. A text format meteorological data configuration universal decoding and warehousing method as claimed in claim 1, characterized in that: The format check of meteorological data files includes file-level check, report-level check, group-level check and element-level check; The file-level check includes whether the file name is empty, whether the file exists, whether it is a file, and whether the file is empty; The report-level check is to check whether the report length is within the set length range, and to determine whether the number of elements obtained after the report is divided by a delimiter is the same as the specified value; The group level check includes a group length check and a group coding check; The element-level check includes element coding check, element length check, latitude check, longitude check, time format check, and time limit check.

4. A text format meteorological data configuration universal decoding and warehousing method as claimed in claim 1, characterized in that: The step of decoding the meteorological data file based on the preset configuration parsing rules to obtain a report and then obtain elements specifically includes: Traverse the meteorological data file line by line, and obtain the traversed data to form a report when traversing to the report delimiter. The report delimiter defaults to a line break, and the report delimiter can be modified through configuration; The report is parsed based on the element separator to obtain multiple strings, and the strings are numbered. Each numbered string corresponds to an element obtained by decoding. The element separator defaults to a space or a tab, and the element separator can be modified through configuration.

5. A text format meteorological data configuration universal decoding and warehousing method as claimed in claim 1, characterized in that: The formation of the storage statement is realized based on the basic configuration and storage field mapping configuration; The basic configuration is used to obtain the queue information, data encoding, program behavior control, separator setting, and site longitude and latitude of the task, and the basic configuration includes public configuration and private configuration, and the format of the public configuration and the private configuration are the same, and are loaded in the order of public configuration first and private configuration later; The storage entry field mapping configuration is used to establish a mapping relationship between the decoded elements and each field in the data table of the database, so as to realize the generation of configured storage entry statements.

6. A text format meteorological data configuration universal decoding and warehousing method as claimed in claim 5, characterized in that: The key configuration items of the basic configuration include report delimiter, first element delimiter, second element delimiter, maximum number of read rows, station information configuration file location, update strategy, and data entry time check range; The report separator is used to separate the meteorological data file to obtain a report; The first element separator is used to segment the report to obtain character strings of each element; The second element separator is used to perform secondary segmentation on the elements obtained by segmentation by the first element separator; The station information configuration file is used to store the latitude, longitude and altitude information of the station, so that the latitude, longitude and altitude information of the corresponding station can be obtained from the station information configuration file based on the configured station number; The station information configuration file is in text format and uses key values ​​and value values ​​to store information. The key value includes the station number and the station network code, and the value includes the station number, longitude, altitude, administrative division code, station type, exchange level, spare field, and regional association code. The update strategy is a strategy for processing duplicate data.

7. A text format meteorological data configuration universal decoding and warehousing method as claimed in claim 5, characterized in that: The input field mapping configuration is configured in XML format, the attribute index is used to indicate that the meteorological data file is segmented by the report separator to obtain the report and the element separator is used to segment the report to obtain the element number, the tag value is used to indicate the input field name, and the built-in function is used to implement the expression calculation processing; The key configuration items of the input field mapping configuration include whether to check the number of groups, the number of group checks, the type of data input, the value method, the report value location, the default value, the time zone of the data time, the landing time of the meteorological data file, the date formatting, and the expression for configuring the current row field value.

8. A text format meteorological data configuration universal decoding and warehousing method as claimed in claim 7, characterized in that: For the mapping between elements and fields in the data table of the database, the processing process includes: taking the value of the element for data type conversion, taking the value of the element for mathematical calculation, taking the value of the element and intercepting it, string concatenation, converting Beijing time to world time, converting world time to Beijing time, obtaining station information, and processing specific fields; When the name of the data table changes, modify the data table name in the configuration; when a new field is added to the data table, add a row to configure the value of the corresponding new field; when the number of fields in the data table decreases, delete the configuration of the corresponding field.

9. A text format meteorological data configuration universal decoding storage device, characterized in that: The text format meteorological data configuration universal decoding storage device comprises: An acquisition unit, which is used to acquire the meteorological data file to be decoded based on the configured universal decoding program and adopts a real-time or directory polling data access method; A decoding unit, which is used to check the format of the meteorological data file, and decode the meteorological data file based on a preset configuration parsing rule to obtain a report and then obtain elements; A calculation conversion unit, which is used to perform calculation conversion mapping on the obtained elements based on the configuration to form a storage statement; The storage unit is used to call the storage interface for the formed storage statement to realize data storage.

10. A text format meteorological data configuration universal decoding storage device as claimed in claim 9, characterized in that: The real-time data access method is a combination of notification messages and data files. The notification messages are implemented using RabbitMQ message middleware, and each notification message stores the data type and path information of the meteorological data file. The data access method of the directory polling is to perform polling processing in the corresponding directory based on the directory to be polled and the data type code input in the configuration file, and obtain the absolute path of all meteorological data files in the corresponding directory.

Citation Information

Patent Citations

  • Decoding method of aviation weather report

    CN101853248A

  • Meteorological data processing method and system, electronic device and storage medium

    CN111103635A

  • Meteorological mode data decoding processing method

    CN116010525A

  • Universal configurable unstructured meteorological data processing method and device

    CN116089366A

  • Method and device for processing mass meteorological observation data based on Storm

    CN116126552A