Custom structured log generator
By designing a custom structured log generator, structured logs are automatically generated based on user-defined data structures and constraints, solving the problems of inconsistent log formats and sensitive information leakage in the existing technology, and achieving the security, authenticity and efficiency of log data.
Patent Information
- Application Number
- CN202510183305.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-27
AI Technical Summary
It is difficult for the prior art to generate custom structured logs that meet the needs of different application scenarios, and real log data may contain sensitive information and is difficult to disclose.
Design a custom structured log generator, which adopts the data structure module, the data field constraint module, the algorithm module for generating logs, and the log module for outputting the required format, so as to automatically generate structured logs based on user-defined data structures and constraints, and supports multiple log file formats.
It realizes automatic generation of structured logs based on user needs, solves the problem of inconsistent log formats, ensures the security and authenticity of log data, and improves the analyticity of log data.
Smart Images

Figure CN120045536A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing and log generation, and particularly to a custom structured log generator. Background Art
[0002] With the rapid development and construction of IT information systems, different systems have logs in different formats, and the log formats are not unified; moreover, with the accumulation of time and the expansion of the system, enterprises generate a large amount of log data every year. Log data in production systems plays an increasingly important role in software development positioning, operation and maintenance monitoring of system operation, data analysis and compression, security against ransomware, etc. Log data usually has three types: structured logs, semi-structured logs, and unstructured logs. Unstructured logs and semi-structured logs bring great difficulties to the positioning, analysis, and processing of upper-layer applications due to unclear regulations. With the development of artificial intelligence and big data analysis technologies, the need for structuring log data by upper-layer processing and analysis software is becoming increasingly clear.
[0003] Structured log data has fixed fields and formats, which is very convenient for retrieval and analysis by upper-layer systems. At the same time, it is also conducive to compressing and storing files to reduce storage space. However, there are also many difficulties in obtaining structured log data during the operation of a real environment. The first point is that the log data generated during the operation of a real system may very likely contain sensitive personal information of enterprise customers, and it is difficult to make it public on the Internet. Once it is made public, even if desensitization operations are performed, there may still be a risk of privacy leakage. The second point is that the generation of real log data is restricted by the actual upper-layer business scenarios and it is difficult to cover all scenarios. Therefore, there will be fewer fault logs. The third point is that existing log generation tools usually can only generate logs in fixed formats, lacking flexibility and it is difficult to meet the specific requirements of users for log formats and contents in different application scenarios. Summary of the Invention
[0004] The object of the present invention is to provide a custom structured log generator in view of the deficiencies of the prior art. The custom structured log generator is composed of a data structure definition module, a data field constraint definition module, an algorithm module for generating logs, and a log module for outputting the required format, which can automatically generate structured logs according to the data structure, constraint conditions, and mutual exclusion conditions defined by the user, and support output and saving in multiple common log file formats (csv / txt / parquet). When generating logs, the system can consider the association and mutual exclusion relationships between fields, thereby improving the authenticity of the generated logs. As a general underlying tool, it is applicable to software design, development, operation and maintenance, big data analysis, testing and other links, and can help developers quickly generate a large amount of log data in the design, development, testing, and demonstration stages without relying on a real-deployed software system. It can generate real production log data in a real environment and also solve the problem of the sensitivity of log data. It is applicable to software design, development, operation and maintenance, big data analysis, testing and other links, and has the advantages of data security, automated high efficiency, and consistency of log formats, with good application prospects.
[0005] The object of the present invention is achieved as follows: A custom structured log generator, characterized in that it is composed of a data structure definition module, a data field constraint definition module, an algorithm module for generating logs, and a log module for outputting the required format, which can automatically generate structured logs according to the data structure, constraint conditions, and mutual exclusion conditions defined by the user. The data structure definition module is used to initialize and define the data structure of the logs required by the user, and the data structure contains basic information such as field names, data types, and nested relationships; after defining the log data structure, the data field constraint definition module selects options for defining the association and mutual exclusion relationships between fields, as well as the enumeration range and numerical value range of a certain field; the algorithm module for generating logs automatically fills the log data according to the log data structure and the number of logs specified by the user, using algorithms including: an algorithm for generating simulated time, an algorithm for generating simulated strings, an algorithm for generating simulated IP addresses, or an algorithm for randomly generating according to enumeration values; the log module for outputting the required format saves the automatically generated log data in the specified csv / parquet / txt file format.
[0006] The present invention has the following beneficial technical effects and remarkable technical progress compared with the prior art: 1) Data security: The generated log data does not contain the usage data of real users, avoiding the risk of sensitive data leakage and solving the problem that enterprises choose not to disclose log data sets due to the security sensitivity of production log data.
[0007] 2) High efficiency: Enabling the log data to be repeatedly and automatically generated, facilitating data analysis and processing of upper-layer applications.
[0008] 3) Log consistency: The generated simulated log data can be customized by customers as needed to meet the requirements of different systems, and has a high degree of consistency with the format of the original data, facilitating data analysis and processing by the upper-layer application platform. Brief Description of the Drawings
[0009] Figure 1 is a schematic structural diagram of the present invention; Figure 2 is a flowchart of the present invention. Detailed Embodiment
[0010] Refer to Figure 1 , the present invention includes: a data structure definition module, a data field constraint definition module, a log generation algorithm module, and a log output module in the required format. The data structure definition module is used to initialize and define the data structure of the log required by the user, and the data structure contains basic information such as field names, data types, and nested relationships; after defining the log data structure, the data field constraint definition module selects options for defining the association relationship and mutual exclusion relationship between fields, as well as the enumeration range and numerical value range of a certain field; the log generation algorithm module automatically fills the log data according to the log data structure and the number of logs specified by the user, using algorithms including: an algorithm for generating simulated time, an algorithm for generating simulated strings, an algorithm for generating simulated IP addresses, or an algorithm for randomly generating according to enumeration values; the log output module in the required format saves the automatically generated log data as a specified csv / parquet / txt file format.
[0011] The following further elaborates on the present invention through specific embodiments. Embodiment
[0012] Refer to Figure 2 , the present invention realizes the automatic generation of structured logs according to the data structure, constraint conditions, and mutual exclusion conditions defined by the user in the following steps: The first step: Input the data structure of the log data, including basic information such as field names, data types, and nested relationships. The following briefly explains these three data structure types.
[0013] The field name is required: The user must define the field names included in the log, and the content is user-defined, which can be "timestamp", "user ID name", "operation type", etc.
[0014] The data type is required: The user must specify the corresponding data type for each field, such as string, integer, float, boolean, timestamp, etc.
[0015] The nested relationship is not mandatory: The user can set the nested relationship for each field. For example, the first-level field of "User Basic Information" contains second-level fields such as "Name", "Gender", "Age", etc. The nested relationship helps in understanding the data, facilitating indexing and display.
[0016] Step 2: Identify whether the field needs to call the existing interface for simulating common data fields. If so, call the interface for simulating field data integrated in advance by the system. For example: Automatically generate common interface functions such as user name, gender, age, date of birth, email address, contact phone number, country, access date, ipv4 / ipv6 address, MAC address of the device, URL, etc. If it is not in the pre-integrated category, automatically generate data based on data similarity and constraints later.
[0017] Step 3: Identify the range of a single field. For numerical values, define the upper and lower limits, length, and format; for enumerated values, define the enumerated list; for strings, set the string length.
[0018] Step 4: Identify whether there are association relationships and mutual exclusion relationships between different fields. Association relationship: The user can specify the association relationship between different fields, such as the association relationship between "Birthday" and "ID number". Mutual exclusion relationship: The user can specify the mutual exclusion relationship between different fields to improve the authenticity of log data through constraints between fields.
[0019] Step 5: Select the log output format, such as csv, txt in row storage format or parquet in column storage format, for use in different log analysis platforms.
[0020] Step 6: Automatically fill the logs. According to the above definitions and constraints, fill a certain amount of log data. When filling the logs, use various algorithms to automatically fill the log data.
[0021] Step 7: Output the final required logs, which the user can use for viewing, display, development, analysis, or for log data compression and query algorithm testing.
[0022] The above is only a further explanation of the present invention and is not intended to limit the present invention. Equivalent implementations without departing from the spirit and scope of the present invention concept should be included within the scope of the claims of the present invention.
Claims
1. A custom structured log generator, characterized in that: A custom structured log generator consisting of a data structure definition module, a data field constraint definition module, a log generation algorithm module, and a log output module in the required format is used to automatically generate structured logs according to the data structure, constraint conditions, and mutually exclusive conditions defined by the user. The data structure definition module is used to initialize the data structure of the log required by the user, and the data structure contains basic information such as field name, data type, and nested relationship; after defining the log data structure, the data field constraint definition module selects whether to define the association relationship and mutually exclusive relationship between the fields, as well as the enumeration range and value size range of a given field; The algorithm module for generating logs automatically fills in log data according to the log data structure and the number of logs specified by the user, using algorithms including: simulated time generation algorithm, simulated string generation algorithm, simulated IP address generation algorithm, or random generation algorithm according to enumeration values; the log module for outputting the required format saves the automatically generated log data into the specified csv / parquet / txt file format.