A data acquisition format protocol conversion method and device

CN116303717BActive Publication Date: 2026-09-08CHINA TELECOM CLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310246430.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-10
Publication Date
2026-09-08
Estimated Expiration
2043-03-10

AI Technical Summary

Technical Problem

[0004]另一方面,随着生产设备和相关技术的智能化升级,以及全球市场无时无刻都在变化的需求,行业内实时数据的采集与计算相关标准已经提升到了秒级要求,当前的批处理数据架构难以应对,需要构建新一代的实时数据架构体系以实现“换挡加速”

Benefits of technology

[0053] The advantages of this application's embodiments are as follows: The data acquisition format protocol based on data synchronization provided in this application can be converted using maxwell-json, cannel-json, and debezium-json protocols, thus ensuring compatibility with the acquisition synchronization framework protocol and improving data synchronization transmission efficiency; the data acquisition format protocol based on data acquisition synchronization solves the problems of excessively high data network bandwidth consumption, high disk read/write rates, and high disk storage ratios in related technologies' data synchronization acquisition systems; furthermore, through the provided Kafka serialization and deserialization methods based on the data acquisition synchronization data format protocol, disk space can be further compressed and data synchronization transmission efficiency can be improved through data serialization processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116303717B_ABST
    Figure CN116303717B_ABST
Patent Text Reader

Abstract

The application relates to a data collection format protocol conversion method and device, electronic equipment and a medium, the method comprising: converting a protocol of a first collection format into a protocol of a DATA-CVS format by constructing serialization and deserialization of Kafka and constructing the DATA-CVS format of Flink; wherein the serialization and deserialization of Kafka comprises: constructing serialization and deserialization classes of Kafka data streams by constructing a serialization module; obtaining a created custom data format by creating a POJO serialization module, converting the custom data format into a POJO class corresponding to a field name of the custom data format; and converting the POJO class into the DATA-CVS format by a conversion module; and the protocol of the first collection format at least comprises protocols of canal_json, maxwel_json and debezium_json. The application can improve the disk utilization efficiency of data collection storage and the efficiency of development work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data protocol technology, and in particular to a data acquisition format protocol conversion method, apparatus, electronic device, and medium. Background Technology

[0002] Regardless of the stage of a company's digital transformation, data collection and synchronization is the most practical and frequent requirement for businesses.

[0003] On the one hand, the demand for real-time data from enterprises' refined operations is constantly expanding. Real-time data can help enterprises collect data from sensors such as machine speed, temperature, pressure, and flow in the industrial field, as well as stock quotes, server logs, traditional databases, and even Hadoop systems at the fastest speed. Mining valuable information in real-time or near-real-time is of great significance for enterprises to make rapid decisions.

[0004] On the other hand, with the intelligent upgrading of production equipment and related technologies, and the ever-changing demands of the global market, the industry's standards for real-time data collection and computation have been raised to the second level. The current batch processing data architecture is insufficient to cope with this, and a new generation of real-time data architecture system needs to be built to achieve "gear shift and acceleration".

[0005] In real-time data acquisition and synchronization solutions, there are scenarios involving message middleware for peak shaving and valley filling, as well as many-to-many data synchronization. Currently, there are three data acquisition formats on the market: canal_json, maxwel_json, and debezium_json. However, they all have a drawback: excessive data redundancy, which occupies unnecessary disk space. How to effectively simplify disk read and write, reduce disk I / O, reduce memory consumption, and thus reduce the overall resource pool consumption of the server is a problem that needs to be solved. Summary of the Invention

[0006] In view of the above problems, this application provides a data acquisition format protocol conversion method, apparatus, electronic device and medium.

[0007] In a first aspect, embodiments of this application provide a data acquisition format protocol conversion method, including:

[0008] The protocol of the first acquisition format is converted into a DATA-CSV format protocol by constructing a DATA-CSV format, which includes Kafka serialization and deserialization and Flink format.

[0009] The construction includes Kafka serialization and deserialization, including:

[0010] By building a serialization module, we can construct serialization and deserialization classes for Kafka data streams; by creating a POJO serialization module, we can obtain the created custom data format and convert the custom data format into POJO classes with corresponding field names; and by using a conversion module, we can convert the POJO classes into -DATA-CSV format.

[0011] The first type of acquisition format protocol includes at least the following protocols: canal_json, maxwel_json, and debezium_json.

[0012] Furthermore, in the aforementioned data acquisition format protocol conversion method, the construction of Flink's DATA-CSV format includes:

[0013] Flink's table-api format and connector are used to extract data from the protocol of the first acquisition format.

[0014] The data in the protocol of the first acquisition format is converted into -DATA-CSV data using Flink SQL;

[0015] Perform SQL processing on -DATA-CSV data;

[0016] In the -DATA-CSV format protocol, -DATA-CSV is a delimited data format.

[0017] Furthermore, in the aforementioned data acquisition format protocol conversion method, the rules for the -DATA-CSV format protocol include at least the following seven:

[0018] 1. Do not leave blanks at the beginning; use lines as units.

[0019] Second: Column names may or may not be included; if column names are included, they should be enclosed in parentheses.

[0020] Three: Data in a single line does not span multiple lines and there are no blank lines;

[0021] 4. Use invisible characters as delimiters; even if a column is empty, its existence must be indicated.

[0022] 5. If ASCII characters exist in the column content, replace them with escape characters and enclose the field value in half-width quotation marks;

[0023] Six: When reading and writing files, the ASCII code operation rules are inverses.

[0024] 7. No restrictions on internal encoding format.

[0025] Furthermore, in the aforementioned data acquisition format protocol conversion method, the escaping requirements in the -DATA-CSV format protocol must include at least the following three:

[0026] 1. For fields containing the ASCII code corresponding to the type, the ASCII code corresponding to the key, and the newline character, add an escape character before the ASCII code corresponding to the type, the ASCII code corresponding to the key, and the newline character.

[0027] 2. The ASCII codes corresponding to the types inside the fields and the ASCII codes corresponding to the keys are converted by adding an escape character in front of them to achieve the encoding of text quotation marks;

[0028] Third, for specific fields required for synchronization, use the corresponding ASCII codes to map them one by one.

[0029] Furthermore, in the aforementioned data acquisition format protocol conversion method, specific fields required for synchronization are mapped one-to-one using corresponding ASCII codes, including:

[0030] The data source corresponds to (ACII code / LF = 0x0B);

[0031] The log collection time ts_ms corresponds to (ACII code / LF = 0x0C);

[0032] The operation type op corresponds to (ACII code / LF = 0x0D);

[0033] The metadata schema corresponds to (ACII code / LF = 0x0F).

[0034] Furthermore, in the above-mentioned data acquisition format protocol conversion method, the file in the -DATA-CSV format protocol is a text file separated by newline characters;

[0035] Text files store tabular data in plain text format;

[0036] A text file is a sequence of characters;

[0037] A text file consists of multiple records, separated by a newline character; each record consists of fields, separated by characters or strings; multiple records contain the exact same sequence of fields.

[0038] Open the text file with WordPad or Notepad to record the text;

[0039] Each record in a text file is a line-terminating newline character (ASCII code / LF = 0x0A) or a carriage return and newline character (ASCII code / CRLF = 0x0D0A);

[0040] In C#, 0x0A represents the character '\n', and 0x0D0A represents the string "\r\n".

[0041] Furthermore, in the above-mentioned data acquisition format protocol conversion method, the field values ​​in the -DATA-CSV format protocol contain multiple types, and each type is assigned a corresponding ASCII code;

[0042] - Field values ​​in the DATA-CSV format protocol are enclosed in brackets corresponding to the ASCII code of the type. An empty field in a line is enclosed in brackets corresponding to the ASCII code of the type.

[0043] - In the DATA-CSV format protocol, if the field packet contains the ASCII code corresponding to the type, it should be enclosed in the ASCII code corresponding to the type, and the ASCII code corresponding to the type should be escaped.

[0044] In a DATA-CSV format protocol, if the field value contains the ASCII code corresponding to the type, double-write the ASCII code corresponding to the type.

[0045] Secondly, embodiments of this application also provide a data acquisition format protocol conversion device, including: a construction module,

[0046] The building module is used to build Kafka serialization and deserialization and to build Flink's DATA_DATA_CSV format, converting the first acquisition format protocol to the DATA_CSV format protocol;

[0047] The building module is used to construct Kafka's serialization and deserialization mechanisms, including:

[0048] The module for building serialization is used to construct serialization and deserialization classes for Kafka data streams; the module for creating POJO serialization is used to obtain the created custom data format and convert the custom data format into POJO classes with corresponding field names; the module for conversion is used to convert POJO classes into DATA-CSV format.

[0049] The first type of acquisition format protocol includes at least the following protocols: canal_json, maxwel_json, and debezium_json.

[0050] Thirdly, embodiments of the present invention also provide an electronic device, including: a processor and a memory;

[0051] The processor executes a data acquisition format protocol conversion method as described above by calling the program or instructions stored in the memory.

[0052] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a program or instructions that cause a computer to execute a data acquisition format protocol conversion method as described in any of the preceding claims.

[0053] The advantages of this application's embodiments are as follows: The data acquisition format protocol based on data synchronization provided in this application can be converted using maxwell-json, cannel-json, and debezium-json protocols, thus ensuring compatibility with the acquisition synchronization framework protocol and improving data synchronization transmission efficiency; the data acquisition format protocol based on data acquisition synchronization solves the problems of excessively high data network bandwidth consumption, high disk read / write rates, and high disk storage ratios in related technologies' data synchronization acquisition systems; furthermore, through the provided Kafka serialization and deserialization methods based on the data acquisition synchronization data format protocol, disk space can be further compressed and data synchronization transmission efficiency can be improved through data serialization processing. Attached Figure Description

[0054] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 A schematic diagram of a data acquisition format protocol conversion method provided in this application embodiment. Figure 1 ;

[0056] Figure 2 A schematic diagram of a data acquisition format protocol conversion method provided in this application embodiment. Figure 2 ;

[0057] Figure 3 A schematic diagram of a data acquisition format protocol conversion device provided in an embodiment of this application;

[0058] Figure 4 This is a schematic block diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0059] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0061] The consistent hashing algorithm involved in this application will be introduced first below.

[0062] Figure 1 A schematic diagram of a data acquisition format protocol conversion method provided in this application embodiment. Figure 1 .

[0063] In a first aspect, embodiments of this application provide a data acquisition format protocol conversion method, including:

[0064] The protocol of the first acquisition format is converted into a DATA-CSV format protocol by constructing a DATA-CSV format, which includes Kafka serialization and deserialization and Flink format.

[0065] Specifically, in this embodiment, the construction of the DATA-CSV format, which includes Kafka serialization and deserialization and Flink format, is an optimized storage format based on data acquisition and synchronization message middleware. It will be open-sourced to the Apache community in the future.

[0066] Figure 1 A schematic diagram of a data acquisition format protocol conversion method provided in this application embodiment. Figure 1 .

[0067] This includes building Kafka's serialization and deserialization processes, combined with... Figure 1 ,include:

[0068] S101: By constructing a serialization module, construct serialization and deserialization classes for Kafka data streams;

[0069] S102: By creating a POJO serialization module, obtain the created custom data format, and convert the custom data format into a POJO class with the corresponding field names of the custom data format;

[0070] S103: Convert the POJO class to -DATA-CSV format using the conversion module;

[0071] The first type of acquisition format protocol includes at least the following protocols: canal_json, maxwel_json, and debezium_json.

[0072] Specifically, in this embodiment, the construction includes Kafka serialization and deserialization. By serializing the data, disk space can be compressed and the transmission efficiency of data synchronization can be improved.

[0073] Figure 2 A schematic diagram of a data acquisition format protocol conversion method provided in this application embodiment. Figure 2 .

[0074] Furthermore, in the aforementioned data acquisition format protocol conversion method, the Flink format—DATA-CSV—is constructed, combined with... Figure 2 It includes three steps: S201 to S203.

[0075] S201: Flink's table-api format and connector, extracting data from the protocol of the first acquisition format;

[0076] S202: Convert the data in the protocol of the first acquisition format to obtain -DATA-CSV data through Flink SQL;

[0077] S203: Perform SQL processing on -DATA-CSV data;

[0078] In the -DATA-CSV format protocol, -DATA-CSV is a delimited data format.

[0079] Specifically, in this embodiment, the -DATA-CSV format is a delimited data format with field / column separators (ASCII code / LF = 0x05) and record / row separators (ASCII code / LF = 0x0A).

[0080] Furthermore, in the aforementioned data acquisition format protocol conversion method, the rules for the -DATA-CSV format protocol include at least the following seven:

[0081] 1. Do not leave blanks at the beginning; use lines as units.

[0082] Second: Column names may or may not be included; if column names are included, they should be enclosed in parentheses.

[0083] Three: Data in a single line does not span multiple lines and there are no blank lines;

[0084] 4. Use invisible characters as delimiters; even if a column is empty, its existence must be indicated.

[0085] 5. If ASCII characters exist in the column content, replace them with escape characters and enclose the field value in half-width quotation marks;

[0086] Six: When reading and writing files, the ASCII code operation rules are inverses.

[0087] 7. No restrictions on internal encoding format.

[0088] Specifically, in the embodiments of this application, the second rule includes column names enclosed in (ASCII code / LF = 0x0F); the fourth rule uses an invisible character (i.e., (ASCII code / LF = 0x05)) as a separator, and the column must be empty to indicate its existence; the fifth rule, if the column content contains (ASCII code), replaces it with the escape character \+ASCII code, i.e., (\ASCII code) escape, that is, encloses the field value in half-width quotation marks (i.e., ""); the seventh rule has no limit on the internal code format, and can be ASCII, Unicode, or others.

[0089] Furthermore, in the aforementioned data acquisition format protocol conversion method, the escaping requirements in the -DATA-CSV format protocol must include at least the following three:

[0090] 1. For fields containing the ASCII code corresponding to the type, the ASCII code corresponding to the key, and the newline character, add an escape character before the ASCII code corresponding to the type, the ASCII code corresponding to the key, and the newline character.

[0091] 2. The ASCII codes corresponding to the types inside the fields and the ASCII codes corresponding to the keys are converted by adding an escape character in front of them to achieve the encoding of text quotation marks;

[0092] Third, for specific fields required for synchronization, use the corresponding ASCII codes to map them one by one.

[0093] Specifically, in this embodiment, the first type of escaping requirement includes fields containing the ASCII code corresponding to the type, the ASCII code corresponding to the key, or newline characters, which must have an escape character added before them. The escape character can be a user-defined character used to escape special symbols. The second type of escaping requirement requires fields containing the ASCII code corresponding to the type and the ASCII code corresponding to the key, which must have an escape character added before them. The escape character can also be a user-defined character used to encode text quotation marks. For specific fields required for synchronization, such as the third type of escaping requirement (metadata schema, operation type op, data source source, log collection time ts_ms), corresponding ASCII codes are used to represent them one-to-one.

[0094] Furthermore, in the aforementioned data acquisition format protocol conversion method, specific fields required for synchronization are mapped one-to-one using corresponding ASCII codes, including:

[0095] The data source corresponds to (ACII code / LF = 0x0B);

[0096] The log collection time ts_ms corresponds to (ACII code / LF = 0x0C);

[0097] The operation type op corresponds to (ACII code / LF = 0x0D);

[0098] The metadata schema corresponds to (ACII code / LF = 0x0F).

[0099] Furthermore, in the above-mentioned data acquisition format protocol conversion method, the file in the -DATA-CSV format protocol is a text file separated by newline characters;

[0100] Text files store tabular data in plain text format;

[0101] A text file is a sequence of characters;

[0102] A text file consists of multiple records, separated by a newline character; each record consists of fields, separated by characters or strings; multiple records contain the exact same sequence of fields.

[0103] Open the text file with WordPad or Notepad to record the text;

[0104] Each record in a text file is a line-terminating newline character (ASCII code / LF = 0x0A) or a carriage return and newline character (ASCII code / CRLF = 0x0D0A);

[0105] In C#, 0x0A (newline character) represents the character '\n', and 0x0D0A (carriage return and line feed) represents the string "\r\n".

[0106] Specifically, in this embodiment, the -DATA-CSV format file is essentially a text file separated by newline characters (ASCII code / LF=0x05) and ((ASCII code / LF=0x0A)). -DATA-CSV ((ASCII code / LF=0x05) separated value file format) Comma-Separated Values ​​[the delimiter for each line must be (ASCII code / LF=0x05)], comma-separated values ​​(DATA_CSV, sometimes also called character-separated values, because the delimiter character can also be other than (ASCII code / LF=0x05)). The file stores tabular data in plain text form, and the content of the tabular data is numbers and text; the text file is a sequence of characters and does not contain data that must be interpreted like binary numbers; the delimiter between fields is a character or string, most commonly (ASCII code / LF=0x05).

[0107] Furthermore, in the above-mentioned data acquisition format protocol conversion method, the field values ​​in the -DATA-CSV format protocol contain multiple types, and each type is assigned a corresponding ASCII code;

[0108] - Field values ​​in the DATA-CSV format protocol are enclosed in brackets corresponding to the ASCII code of the type. An empty field in a line is enclosed in brackets corresponding to the ASCII code of the type.

[0109] - In the DATA-CSV format protocol, if the field packet contains the ASCII code corresponding to the type, it should be enclosed in the ASCII code corresponding to the type, and the ASCII code corresponding to the type should be escaped.

[0110] In a DATA-CSV format protocol, if the field value contains the ASCII code corresponding to the type, double-write the ASCII code corresponding to the type.

[0111] Specifically, in this embodiment, each type is assigned a corresponding ASCII code, such as the String type corresponding to (ASCII code / LF = 0x02). Users can customize the ASCII code for each type. If an item in a line is empty, it can be enclosed in the ASCII code corresponding to the type. If a field contains a special character (i.e., the ASCII code corresponding to the type), it must be enclosed in the ASCII code corresponding to the type, and the corresponding ASCII code must be escaped. When the value of a field contains a special character (i.e., the ASCII code corresponding to the type), the ASCII code corresponding to the type can also be doubled, just like using an ASCII code corresponding to the type as an escape character.

[0112] Figure 3 This is a schematic diagram of a data acquisition format protocol conversion device provided in an embodiment of this application.

[0113] Secondly, embodiments of this application also provide a data acquisition format protocol conversion device, including: a construction module 300.

[0114] Module 300 is used to build a protocol that includes Kafka serialization and deserialization and to build the Flink DATA_DATA_CSV format, converting the first acquisition format protocol into the DATA_CSV format protocol;

[0115] Module 300 is used to build Kafka serialization and deserialization systems, including:

[0116] The serialization module 301 is used to build serialization and deserialization classes including Kafka data streams; the POJO serialization module 302 is used to obtain the created custom data format and convert the custom data format into POJO classes with corresponding field names; the conversion module 303 is used to convert the POJO classes into -DATA-CSV format.

[0117] The first type of acquisition format protocol includes at least the following protocols: canal_json, maxwel_json, and debezium_json.

[0118] Thirdly, embodiments of the present invention also provide an electronic device, including: a processor and a memory;

[0119] The processor executes a data acquisition format protocol conversion method as described above by calling the program or instructions stored in the memory.

[0120] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a program or instructions that cause a computer to execute a data acquisition format protocol conversion method as described in any of the preceding claims.

[0121] Figure 4 This is a schematic block diagram of an electronic device provided in an embodiment of this disclosure.

[0122] like Figure 4 As shown, the electronic device includes at least one processor 401, at least one memory 403, and at least one communication interface 403. The various components of the electronic device are coupled together via a bus system 404. The communication interface 403 is used for information transmission with external devices. It is understood that the bus system 404 is used to implement communication between these components. In addition to a data bus, the bus system 404 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 4 The general designated all buses as Bus System 404.

[0123] It is understood that the memory 402 in this embodiment can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory.

[0124] In some implementations, memory 402 stores elements such as executable units or data structures, or subsets thereof, or extended sets thereof: operating systems and applications.

[0125] The operating system includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application programs include various applications, such as media players and browsers, used to implement various application functions. A program implementing any method in the data acquisition format protocol conversion method provided in this application embodiment can be included in the application programs.

[0126] In this embodiment, the processor 401 executes the steps of various embodiments of the data acquisition format protocol conversion method provided in this application by calling the program or instructions stored in the memory 402, specifically, the program or instructions stored in the application.

[0127] The protocol of the first acquisition format is converted into a DATA-CSV format protocol by constructing a DATA-CSV format, which includes Kafka serialization and deserialization and Flink format.

[0128] The construction includes Kafka serialization and deserialization, including:

[0129] By building a serialization module, we can construct serialization and deserialization classes for Kafka data streams; by creating a POJO serialization module, we can obtain the created custom data format and convert the custom data format into POJO classes with corresponding field names; and by using a conversion module, we can convert the POJO classes into -DATA-CSV format.

[0130] The first type of acquisition format protocol includes at least the following protocols: canal_json, maxwel_json, and debezium_json.

[0131] Any of the methods in the data acquisition format protocol conversion method provided in this application embodiment can be applied to or implemented by the processor 401. The processor 401 can be an integrated circuit chip with signal capabilities. During implementation, each step of the above method can be completed by the integrated logic circuits in the hardware of the processor 401 or by instructions in software form. The processor 401 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional device.

[0132] The steps of any method in the data acquisition format protocol conversion method provided in this application embodiment can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software units in the decoding processor. The software units can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 402. The processor 401 reads the information in memory 402 and, in conjunction with its hardware, completes the steps of the data acquisition format protocol conversion method.

[0133] Those skilled in the art will understand that although some embodiments described herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of this application and form different embodiments.

[0134] Those skilled in the art will understand that the descriptions of the various embodiments have different focuses, and for parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0135] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A data acquisition format protocol conversion method, characterized in that, include: The protocol of the first acquisition format is converted into a DATA-CSV format protocol by constructing a DATA-CSV format, which includes Kafka serialization and deserialization and Flink format. The construction includes Kafka serialization and deserialization, including: By constructing a serialization module, serialization and deserialization classes for Kafka data streams are built; by creating a POJO serialization module, the created custom first collection format is obtained, and the custom first collection format is converted into a POJO class with the corresponding field names of the custom first collection format; the conversion module converts the POJO class into -DATA-CSV format. The protocols for the first type of data acquisition format include at least the protocols for canal_json, maxwel_json, and debezium_json; The Flink-based DATA-CSV format converts the protocol of the first acquisition format into a DATA-CSV format protocol, including: Flink's table-api format and connector are used to extract data from the protocol of the first acquisition format. The data in the protocol of the first acquisition format is converted into -DATA-CSV data using Flink SQL; Perform SQL processing on the -DATA-CSV data; In the protocol described in the -DATA-CSV format, -DATA-CSV format is a delimited data format; The file in the -DATA-CSV format protocol is a text file separated by newline characters; The text file stores the table data in plain text format; The text file is a sequence of characters; The text file consists of multiple records, separated by a newline character; each record consists of fields, separated by characters or strings; multiple records contain the exact same sequence of fields. The text file is opened and recorded using WordPad or Notepad; Each record in the text file is a line-terminating newline character (ASCII code / LF = 0x0A) or a carriage return and newline character (ASCII code / CRLF = 0x0D0A); In C#, 0x0A represents the character '\n', and 0x0D0A represents the string "\r\n".

2. The data acquisition format protocol conversion method according to claim 1, characterized in that, The rules of the DATA-CSV format protocol include at least the following seven:

1. Do not leave blanks at the beginning; use lines as units. Second: Column names may or may not be included; if column names are included, they should be enclosed in parentheses. Three: Data in a single line does not span multiple lines and there are no blank lines; 4. Use invisible characters as delimiters; even if a column is empty, its existence must be indicated.

5. If ASCII characters exist in the column content, replace them with escape characters and enclose the field value in half-width quotation marks; Six: When reading and writing files, the ASCII code operation rules are inverses.

7. No restrictions on internal encoding format.

3. The data acquisition format protocol conversion method according to claim 1, characterized in that, The escaping requirements in the -DATA-CSV format protocol include at least the following three:

1. For fields containing the ASCII code corresponding to the type, the ASCII code corresponding to the key, and the newline character, add an escape character before the ASCII code corresponding to the type, the ASCII code corresponding to the key, and the newline character.

2. The ASCII codes corresponding to the types inside the fields and the ASCII codes corresponding to the keys are converted by adding an escape character in front of them to achieve the encoding of text quotation marks; Third, for specific fields required for synchronization, use the corresponding ASCII codes to map them one by one.

4. The data acquisition format protocol conversion method according to claim 3, characterized in that, The specific fields required for synchronization are mapped one-to-one using corresponding ASCII codes, including: The data source corresponds to (ACII code / LF = 0x0B); The log collection time ts_ms corresponds to (ACII code / LF = 0x0C); The operation type op corresponds to (ACII code / LF = 0x0D); The metadata schema corresponds to (ACII code / LF = 0x0F).

5. The data acquisition format protocol conversion method according to claim 1, characterized in that, The field values ​​in the -DATA-CSV format protocol contain multiple types, and each type is assigned a corresponding ASCII code; In the DATA-CSV format protocol, field values ​​are enclosed in brackets corresponding to the ASCII code of the type. If an item in a line is empty, it is enclosed in brackets corresponding to the ASCII code of the type. If the field packet in the -DATA-CSV format protocol contains a type-corresponding ASCII code, it shall be enclosed in the type-corresponding ASCII code and the type-corresponding ASCII code shall be escaped. If the field value in the -DATA-CSV format protocol contains the ASCII code corresponding to the type, then double-write the ASCII code corresponding to the type.

6. A data acquisition format protocol conversion device, characterized in that, include: Module building The building module is used to construct a protocol that includes Kafka serialization and deserialization and the Flink DATA_DATA_CSV format, converting the first acquisition format protocol into a DATA_CSV format protocol. The building module is used to construct serialization and deserialization mechanisms for Kafka, including: A serialization module is built to construct serialization and deserialization classes for Kafka data streams; a POJO serialization module is created to obtain the created custom first collection format and convert the custom first collection format into POJO classes with corresponding field names; a conversion module is used to convert the POJO classes into -DATA-CSV format. The protocols for the first type of data acquisition format include at least the protocols for canal_json, maxwel_json, and debezium_json; The Flink-based DATA-CSV format converts the protocol of the first acquisition format into a DATA-CSV format protocol, including: Flink's table-api format and connector are used to extract data from the protocol of the first acquisition format. The data in the protocol of the first acquisition format is converted into -DATA-CSV data using Flink SQL; Perform SQL processing on the -DATA-CSV data; In the protocol described in the -DATA-CSV format, -DATA-CSV format is a delimited data format; The file in the -DATA-CSV format protocol is a text file separated by newline characters; The text file stores the table data in plain text format; The text file is a sequence of characters; The text file consists of multiple records, separated by a newline character; each record consists of fields, separated by characters or strings; multiple records contain the exact same sequence of fields. The text file is opened and recorded using WordPad or Notepad; Each record in the text file is a line-terminating newline character (ASCII code / LF = 0x0A) or a carriage return and newline character (ASCII code / CRLF = 0x0D0A); In C#, 0x0A represents the character '\n', and 0x0D0A represents the string "\r\n".

7. An electronic device, characterized in that, include: Processor and memory; The processor executes a data acquisition format protocol conversion method as described in any one of claims 1 to 5 by calling the program or instructions stored in the memory.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions that cause a computer to perform a data acquisition format protocol conversion method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • CDN (Content Delivery Network) log statistical method and device and electronic equipment

    CN114443606A

  • Integrating logic in micro batch based event processing systems

    US20190394259A1