File analysis method and device, electronic equipment and medium

By generating intermediate configuration files and using balanced binary tree and hash value to parse configuration files, the problem of low parsing efficiency under ten thousand-level configuration is solved, and efficient file parsing and resource management is achieved.

CN120492416APending Publication Date: 2025-08-15INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510638350.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When the configuration number reaches 10,000, the traditional full-scale parsing method leads to inefficient file parsing, large resource consumption, and weak concurrency support capabilities, making it impossible to achieve on-demand loading or incremental updates.

Method used

Generate an intermediate configuration file, including a configuration identification structure, parse the configuration file using a balanced binary tree and a target hash value, and generates a first configuration identification structure corresponding to the change operation, determines the second configuration identification structure based on the balanced binary tree and a target hash value, and finally determines the binary configuration file of the original configuration file based on the identification content and location information.

Benefits of technology

It improves file parsing efficiency in scenarios where the number of configurations reaches 10,000, reduces resource consumption, and supports efficient incremental updates and rapid retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492416A_ABST
    Figure CN120492416A_ABST
Patent Text Reader

Abstract

The invention discloses a file analysis method and device, electronic equipment and a medium, and relates to the technical field of computer software, the file analysis method comprises the steps that in response to a change operation for an original configuration file, an intermediate configuration file is generated, and the intermediate configuration file comprises a first configuration identification structure corresponding to the change operation; the first configuration identification structure at least comprises identification content, an identification type, a target hash value and position information; based on a balanced binary tree and a target hash value corresponding to the first configuration identification type, analyzing the intermediate configuration file and determining a second configuration identification structure from the first configuration identification structure; and in response to the completion of the analysis of the intermediate configuration file, determining the binary configuration file corresponding to the original configuration file based on the identification content and the position information in the second configuration identification structure, thereby solving the technical problem that the analysis efficiency is low when the full-amount analysis is carried out when the configuration number magnitude reaches a ten thousand level. The technical effect of improving the file analysis efficiency under the condition that the configuration number reaches the ten thousand level is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer software technology, and in particular to file parsing methods, devices, electronic devices, and media. Background Art

[0002] In related file parsing solutions, all configurations are usually placed in an independent configuration file. For scenarios involving additions, deletions, and modifications, the entire file needs to be loaded and parsed for all configurations before performing the corresponding operations and configuration changes. This makes little difference in scenarios where the number of configurations and files is relatively small. However, when the number of configurations reaches tens of thousands, full parsing is required, resulting in low file parsing efficiency. Summary of the Invention

[0003] The present application provides a file parsing method, device, electronic device and medium to at least solve the problem in the related art of low parsing efficiency when performing full parsing when the number of configurations reaches tens of thousands.

[0004] This application provides a file parsing method, including:

[0005] In response to a change operation on an original configuration file, an intermediate configuration file corresponding to the change operation is generated, the intermediate configuration file including a first configuration identification structure corresponding to the change operation, the first configuration identification structure including at least first configuration identification content, first configuration identification type, target hash value of the first configuration identification, and location information of the first configuration identification;

[0006] Parsing the intermediate configuration file and determining a second configuration identification structure from the first configuration identification structure based on a balanced binary tree corresponding to the first configuration identification type and a target hash value;

[0007] In response to the completion of parsing the intermediate configuration file, a binary configuration file corresponding to the original configuration file is determined based on the identification content and location information in the second configuration identification structure.

[0008] The present application also provides a file parsing device, comprising:

[0009] a generating unit, configured to generate, in response to a change operation on an original configuration file, an intermediate configuration file corresponding to the change operation, the intermediate configuration file including a first configuration identification structure corresponding to the change operation, the first configuration identification structure including at least first configuration identification content, first configuration identification type, a target hash value of the first configuration identification, and location information of the first configuration identification;

[0010] a first determining unit, configured to parse the intermediate configuration file and determine a second configuration identification structure from the first configuration identification structure based on a balanced binary tree corresponding to the first configuration identification type and a target hash value;

[0011] The second determining unit is configured to determine, in response to completion of parsing of the intermediate configuration file, a binary configuration file corresponding to the original configuration file based on the identification content and location information in the second configuration identification structure.

[0012] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned file parsing methods when executing the computer program.

[0013] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned file parsing methods are implemented.

[0014] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned file parsing methods when executed by a processor.

[0015] Through this application, in response to a change operation on an original configuration file, an intermediate configuration file corresponding to the change operation is generated, the intermediate configuration file includes a first configuration identification structure corresponding to the change operation, the first configuration identification structure including at least the first configuration identification content, the first configuration identification type, the target hash value of the first configuration identification, and the location information of the first configuration identification; based on the balanced binary tree corresponding to the first configuration identification type and the target hash value, the intermediate configuration file is parsed and a second configuration identification structure is determined from the first configuration identification structure; in response to the completion of the intermediate configuration file parsing, the binary configuration file corresponding to the original configuration file is determined based on the identification content and location information in the second configuration identification structure. This can solve the technical problem of low parsing efficiency when performing full parsing when the number of configurations reaches tens of thousands in related solutions, and achieve the technical effect of improving file parsing efficiency in scenarios where the number of configurations reaches tens of thousands. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 A flowchart of a file parsing method provided in an embodiment of the present application;

[0018] Figure 2 A flowchart of a newly added post-configuration file parsing method provided in an embodiment of the present application;

[0019] Figure 3A schematic diagram of the structure of a binary configuration file provided in an embodiment of the present application;

[0020] Figure 4 A schematic diagram of the structure of a file parsing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0023] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0024] The rapid development of modern information technology and the widespread adoption of emerging technologies such as cloud computing, microservices architecture, the Internet of Things, and artificial intelligence have placed higher demands on storage file system configuration management. Traditional configuration file parsing methods are no longer able to meet the current needs of large-scale, highly concurrent, and dynamically changing scenarios. Especially in enterprise-level systems, where the number and complexity of configuration files are growing exponentially, traditional methods have numerous shortcomings in terms of processing efficiency, resource utilization, and concurrency.

[0025] In typical modern storage systems, common configuration file formats such as JavaScript Object Notation (JavaScript Object Notation) and eXtensible Markup Language (XML) are widely used. As business scale expands, a single configuration file often contains tens of thousands or even millions of configuration items. The traditional approach is to store all configurations in a single text file. When the program starts or the configuration is updated, the entire configuration is loaded into memory at once for parsing and processing. This full loading and parsing method is not very effective when the amount of configuration data is small, but its performance bottlenecks become increasingly prominent when faced with large-scale configuration data. Specifically, the following major problems exist in traditional solutions:

[0026] Low parsing efficiency: Every time the configuration changes or the system restarts, the entire configuration file needs to be reloaded and parsed. On-demand loading or incremental updates cannot be achieved, resulting in a significant increase in response time.

[0027] High resource consumption: Fully loading the configuration will result in excessive memory usage, especially in scenarios where multiple instances are deployed or multiple services share the configuration, resulting in unnecessary resource waste.

[0028] Weak concurrency support: Most systems use a single-threaded approach for configuration parsing. Each time a configuration is parsed, an independent linked list is constructed to traverse and retrieve all configurations. Reloading the configuration is time-consuming and inefficient.

[0029] Inefficient change operations: Adding, modifying, or deleting configuration items usually requires re-parsing the entire configuration file before the change can be executed. This is particularly time-consuming when there are a large number of configurations, seriously affecting system operation efficiency.

[0030] In related file parsing solutions, all configurations are usually placed in independent configuration files. For scenarios involving additions, deletions, and modifications, the entire file needs to be loaded and parsed for all configurations before performing corresponding operations and configuration changes. This makes little difference in scenarios with a small number of configurations and files. However, when the number of configurations reaches tens of thousands, the parsing efficiency is low when performing full parsing.

[0031] Through this application, in response to a change operation on an original configuration file, an intermediate configuration file corresponding to the change operation is generated, the intermediate configuration file includes a first configuration identification structure corresponding to the change operation, the first configuration identification structure including at least the first configuration identification content, the first configuration identification type, the target hash value of the first configuration identification, and the location information of the first configuration identification; based on the balanced binary tree corresponding to the first configuration identification type and the target hash value, the intermediate configuration file is parsed and a second configuration identification structure is determined from the first configuration identification structure; in response to the completion of the intermediate configuration file parsing, the binary configuration file corresponding to the original configuration file is determined based on the identification content and location information in the second configuration identification structure. This can solve the technical problem of low parsing efficiency when performing full parsing when the number of configurations reaches tens of thousands in related solutions, and achieve the technical effect of improving file parsing efficiency in scenarios where the number of configurations reaches tens of thousands.

[0032] The file parsing method provided by the embodiment of the present disclosure can be executed by a server and a data center. The file parsing method can be applied to scenarios that require processing large-scale configuration files, frequently updating configurations, and requiring fast startup and response.

[0033] The embodiments of the present application provide a file parsing method, and the method is described in detail in conjunction with the execution process of the file parsing method.

[0034] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0035] Figure 1 A flowchart of a file parsing method provided by an embodiment of the present disclosure.

[0036] like Figure 1 As shown, the following steps are included:

[0037] Step 101: In response to a change operation on an original configuration file, an intermediate configuration file corresponding to the change operation is generated, where the intermediate configuration file includes a first configuration identification structure corresponding to the change operation, where the first configuration identification structure includes at least first configuration identification content, first configuration identification type, target hash value of the first configuration identification, and location information of the first configuration identification.

[0038] In some embodiments, a change operation refers to operations such as modifying, adding, or deleting configuration items in the original configuration file. The change operation is usually performed to adapt to the ever-changing needs during the operation of the file storage system, such as adjusting server parameters, updating service functions, or repairing incorrect configurations.

[0039] In some embodiments, for the change operation of the original configuration file, generating an intermediate configuration file corresponding to the change operation means generating intermediate configuration files corresponding to the new operation, modification operation and deletion operation respectively. The intermediate configuration file is used to record the above-mentioned various change operations. All pending change operations can be concentrated in one file for unified management and processing, significantly reducing parsing time and resource consumption and allowing local updates to be applied without reloading the entire configuration, thereby improving the flexibility and response speed of the system.

[0040] In some embodiments, the first configuration identification structure is used to store identification objects representing configuration fields in the intermediate configuration file. For example, the first configuration identification structure can be a token data structure, where a token is a type of marking symbol. The first configuration identification content refers to the specific content of the new configuration item. The first configuration identification type is used to indicate the type of this configuration item. The target hash value of the first configuration identification refers to the calculated hash value of the new configuration item, which is used for fast search and conflict avoidance. The location information of the first configuration identification refers to the location of the new configuration item in the total configuration file, including the starting address and length of the configuration item in the total configuration file. In addition, the first configuration identification structure can also include a hash tree node of the first configuration identification. The hash tree node refers to the specific location or reference of the configuration item in the hash tree to which it belongs, and usually contains pointers to the parent node, left child node, right child node, and information such as the hash value of the node. This can further optimize the parsing and retrieval efficiency of the configuration file. By organizing the configuration items into a hash tree (such as a balanced binary tree), the search speed can be significantly improved, conflicts can be reduced, and more efficient update operations can be supported.

[0041] Step 102: Parse the intermediate configuration file based on the balanced binary tree corresponding to the first configuration identification type and the target hash value and determine the second configuration identification structure from the first configuration identification structure;

[0042] In some embodiments, the balanced binary tree corresponding to the first configuration identification type refers to a balanced binary tree corresponding to each type in the first configuration identification type. The balanced binary tree can be a self-balancing binary search (AVL) tree or a red-black tree, which is used to efficiently store and retrieve configuration items. The type here can be a new type, a modified type, and a deleted type determined according to the change operation, or it can be a type determined according to the content (first configuration identification content) of the configuration item (first configuration identification structure).

[0043] In some embodiments, the content in the intermediate configuration file is read, and the corresponding node is searched in the balanced binary tree based on the target hash value of the configuration item, and the second configuration identification structure is determined from the first configuration identification structure based on the search result. The second configuration identification structure refers to the configuration identification structure that already exists in the balanced binary tree.

[0044] In some embodiments, each first configuration identification structure is read from the intermediate configuration file, and if the hash value has not been calculated, the hash value can be calculated based on the content of the configuration item.

[0045] Step 103: In response to the completion of parsing the intermediate configuration file, a binary configuration file corresponding to the original configuration file is determined based on the identification content and location information in the second configuration identification structure.

[0046] In some embodiments, a consistency check or event notification mechanism can be used to determine that the intermediate file parsing is complete. Specifically, the intermediate configuration file parsing is complete by comparing the change record in the intermediate configuration file with the actual updated data structure. An event notification mechanism can also be used to mark the status of parsing completion, and a callback function or other synchronization mechanism can be used to notify the thread that the parsing is complete.

[0047] In some embodiments, the location information refers to the starting address and length of the configuration in the configuration file.

[0048] In some embodiments, by determining the binary configuration file corresponding to the original configuration file, the loading efficiency can be quickly improved and the time can be shortened during the program restart process.

[0049] Through this application, in response to a change operation on an original configuration file, an intermediate configuration file corresponding to the change operation is generated, the intermediate configuration file includes a first configuration identification structure corresponding to the change operation, the first configuration identification structure including at least the first configuration identification content, the first configuration identification type, the target hash value of the first configuration identification, and the location information of the first configuration identification; based on the balanced binary tree corresponding to the first configuration identification type and the target hash value, the intermediate configuration file is parsed and a second configuration identification structure is determined from the first configuration identification structure; in response to the completion of the intermediate configuration file parsing, the binary configuration file corresponding to the original configuration file is determined based on the identification content and location information in the second configuration identification structure. This can solve the technical problem of low parsing efficiency when performing full parsing when the number of configurations reaches tens of thousands in related solutions, and achieve the technical effect of improving file parsing efficiency in scenarios where the number of configurations reaches tens of thousands.

[0050] In some embodiments, based on the balanced binary tree corresponding to the first configuration identification type and the target hash value, parsing the intermediate configuration file and determining the second configuration identification structure from the first configuration identification structure includes:

[0051] Obtaining a first hash value and a second hash value corresponding to each first configuration identifier, where the first hash value and the second hash value corresponding to each first configuration identifier are determined by different hash algorithms;

[0052] In some embodiments, the hash algorithms used by the first hash value and the second hash value corresponding to each first configuration identification content may be partially the same or completely the same, as long as different hash algorithms are used for the same configuration item. For example, for configuration item 1, the Secure Hash Algorithm (SHA) and the Message-Digest Algorithm 5 (MD5) may be used; for configuration item 2, the MD5 algorithm and the Cyclic Redundancy Check (CRC) algorithm may be used, or two hash algorithms that are exactly the same as those of configuration item 1 may be used, or two algorithms that are completely different from those of configuration item 1 may be used, the purpose of which is to reduce hash conflicts.

[0053] Determine a target hash value corresponding to each first configuration identification content based on the first hash value and the second hash value;

[0054] In some embodiments, the first hash value and the second hash value can be summed to obtain the target hash value, the first hash value and the second hash value can be averaged to obtain the target hash value, or the first hash value and the second hash value can be weighted to obtain the target hash value. This application does not limit the specific method of determining the target hash value based on the first hash value and the second hash value.

[0055] In some embodiments, during the configuration file parsing process, in order to improve hash collision detection capability and unique identification efficiency, a mechanism of double hash calculation of the same configuration item using multiple hash algorithms is used to determine the target hash value corresponding to each first configuration identification content.

[0056] Determining, based on a balanced binary tree corresponding to the first configuration identifier type, a node value in the balanced binary tree;

[0057] In some embodiments, each node in the balanced binary tree represents a configuration item, and the node value generally refers to the data stored in the node, such as the specific content of the configuration item or its hash value, which is generally the hash value of the previously inserted configuration item.

[0058] The intermediate configuration file is parsed, and if the node value in the balanced binary tree is the same as the target hash value, the second configuration identification structure is determined from the first configuration identification structure.

[0059] In some embodiments, if the node value in the balanced binary tree is the same as the target hash value, it means that the configuration item already exists and does not need to be inserted repeatedly.

[0060] In some embodiments, for each first configuration identification structure extracted from the intermediate configuration file, check whether its corresponding target hash value matches the data stored in the node in the balanced binary tree. If the match is successful, further judgment can be made based on the identification content in the first configuration identification structure to find the corresponding identification structure, thereby determining the second configuration identification structure from the first configuration identification structure.

[0061] In some embodiments, the second configuration identification structure is used to indicate the first configuration identification structure that already exists in the balanced binary tree.

[0062] In some embodiments, after determining the node value in the balanced binary tree based on the balanced binary tree corresponding to the first configuration identification type, the method further includes:

[0063] Compare the node value in the balanced binary tree with the target hash value. If the node value in the balanced binary tree and the target hash value are different, insert the target hash value into the balanced binary tree.

[0064] In some embodiments, if the node value in the balanced binary tree is different from the target hash value, it means that the configuration item is a new configuration item and needs to be inserted into the tree.

[0065] In some embodiments, by comparing the node value in the balanced binary tree with the target hash value, if the node value in the balanced binary tree and the target hash value are different, the target hash value is inserted into the balanced binary tree, so that the configuration item set can be efficiently managed and deduplication, incremental update and fast retrieval functions can be achieved.

[0066] In some embodiments, determining the binary configuration file corresponding to the original configuration file based on the identification content and location information in the second configuration identification structure includes:

[0067] Creating configuration item data based on the identification content and location information in the second configuration identification structure, where the configuration item data at least includes the configuration item name, pointers to each field of the configuration item, and the starting position and length of the configuration item in the overall configuration file;

[0068] In some embodiments, configuration item data refers to a data structure used for configuration items in a configuration file, and each configuration has a different configuration name. Specifically, the data structure includes but is not limited to the following key contents: configuration item name, pointers to each field of the configuration item, and the starting position and length of the configuration item in the total configuration file, wherein the configuration item name refers to the name or key name of the configuration item, the pointers to each field of the configuration item are pointers to each field inside the configuration item, the starting position of the configuration item in the total configuration file refers to the starting offset of the configuration item in the original configuration file, and the length refers to the length of the configuration item data.

[0069] In some embodiments, by creating configuration item data based on the identification content and location information in the second configuration identification structure, required information can be effectively extracted from the second configuration identification structure and the configuration item data can be created.

[0070] Based on the configuration item data, a binary configuration file corresponding to the original configuration file is determined, where the binary configuration file at least includes a serial number of the text configuration file and a serial number of the binary configuration file.

[0071] In some embodiments, based on the data structure of the created configuration item, the identification content, the starting address and length of the configuration in the file are filled into the data structure, and then the filled configuration item data is added to the binary configuration file, a binary configuration file corresponding to the original configuration file can be obtained.

[0072] In some embodiments, the serial number of the text configuration file is used to identify the version of the text configuration file, such as the modification timestamp, Git commit hash, etc.; the serial number of the binary configuration file is used to identify the version of the current binary configuration file, including which version of the text configuration file the binary file is generated based on. The serial number of the text configuration file and the serial number of the binary configuration file can help manage different versions of configuration files and improve the configuration loading efficiency in large-scale systems.

[0073] In some embodiments, the file parsing method further includes:

[0074] In response to the binary configuration file existing in the storage file system, based on the serial number of the text configuration file and the serial number of the binary configuration file, determining a file to be parsed mapped to the memory, where the file to be parsed is one of the text configuration file and the binary configuration file;

[0075] In some embodiments, during system startup or operation, whether a valid binary configuration file exists and its version consistency with the text configuration file determines whether the text configuration file or the binary configuration file should be mapped into memory for parsing and use.

[0076] In some embodiments, it can be determined whether a binary configuration file exists in the storage file system based on the suffix of the configuration file. The binary configuration file is usually converted from a text configuration file during the last compilation or hot update process.

[0077] In some embodiments, by comparing the serial numbers of the text configuration file and the binary configuration file to see if they are consistent, it can be determined whether the current binary configuration file is the latest version generated based on the current text configuration file. If not, it means that the text configuration has been modified and the binary configuration file has expired and needs to be regenerated or ignored. Specifically, if the binary configuration file does not exist, the text configuration file is loaded; if the binary configuration file exists and the serial numbers match, the binary configuration file is loaded, skipping the text parsing process and improving startup speed.

[0078] Parse the original configuration file based on the location information of the file to be parsed.

[0079] In some embodiments, the type of file to be parsed (one of a text configuration file and a binary configuration file) is first determined, and then the original configuration file is located and parsed based on the location information of the selected configuration file.

[0080] In some embodiments, the original configuration file is parsed by reading corresponding data segments according to the location information of the file to be parsed. This file parsing method is suitable for application scenarios that require efficient management and rapid loading of large amounts of configuration data.

[0081] In some embodiments, the intermediate configuration file includes a syntax header, a change configuration, and a syntax tail. The syntax header is used to verify the generation timestamp, configuration quantity, and header verification of the update operation. The change configuration is a new configuration, and the syntax tail is used to verify the integrity of the intermediate configuration file.

[0082] In some embodiments, the syntax header includes a magic number and a format version, and the syntax tail includes a total file size and an overall checksum, where the magic number is a specific byte sequence, usually located at the beginning of the file, used to identify the file type or format, which can help the program quickly identify whether the file is of the expected type without having to parse the entire file content; the format version indicates the specific version number that the file follows, which helps with compatibility management and backward compatibility. If the file format changes (such as adding new fields, changing the encoding method, etc.), the version number can be used to distinguish different versions of parsing logic; the total file size refers to the number of bytes of the entire file, including all data parts and possible metadata. The file is detected to be damaged or incomplete by comparing the declared file size with the actual data length read; the overall checksum refers to a hash value calculated based on all or part of the file content, which is used to verify file integrity and prevent data tampering or transmission errors. The recipient can recalculate the checksum of the file and compare it with the provided value. If they are consistent, the file is considered to be unmodified and complete. Otherwise, there may be errors or malicious changes.

[0083] In some embodiments, a standard configuration file syntax header is added to the first half of the intermediate file, a standard configuration file syntax tail is added to the second half of the intermediate file, and a new configuration that needs to be parsed and executed is added in the middle of the configuration file. It can also be a modified configuration, such as directly loading the intermediate configuration file to be parsed when the program executes the new configuration operation.

[0084] In some embodiments, the file parsing method further includes:

[0085] In response to the change operation being a modification operation or a deletion operation, generating an intermediate configuration file corresponding to the modification operation or the deletion operation;

[0086] In some embodiments, for a modification operation, the generated intermediate configuration file typically includes the name of the modified configuration item and its new value; while for a deletion operation, the generated intermediate configuration file only needs to identify the name of the configuration item to be deleted.

[0087] In response to the modification operation or the deletion operation being completed, if the binary configuration file exists in the storage file system, the original configuration file is parsed based on the location information of the binary configuration file.

[0088] In some embodiments, by responding to the completion of a modification operation or a deletion operation, if a binary configuration file exists in the storage file system, the original configuration file is parsed based on the location information of the binary configuration file, so as to effectively manage and respond to configuration change operations and ensure that the system configuration status can be correctly updated or queried after the operation is completed.

[0089] In some embodiments, as Figure 2 As shown, Figure 2A flow chart of a newly added configuration post-file parsing method provided in an embodiment of the present application, wherein, in step 201, a new configuration is added to the total configuration file (original configuration file); in step 202, the configuration to be added is filtered from the total configuration file before performing the configuration operation parsing; in step 203, a configuration intermediate file is created; in step 204, the identifiers in the configuration intermediate file are parsed in sequence; in step 205, a configuration identifier is constructed; in step 206, an identifier object is retrieved in the corresponding identifier type; in step 207, if the identifier object is not retrieved in the corresponding identifier type, it is added to the hash tree corresponding to the configuration identifier; in step 208, if the identifier object is retrieved in the corresponding identifier type, the configuration identifier is returned and the configuration item data is constructed; in step 209, the configuration item data is added to the binary configuration file; in step 210, the binary configuration file header and the binary configuration file serial number are updated; in step 211, whether the serial number of the text configuration file and the serial number of the binary configuration file are consistent are determined. If they are consistent, in step 212, the binary configuration file is preloaded at the early stage of program startup to implement file parsing. If they are inconsistent, as in step 213, the text configuration file is loaded to implement file parsing.

[0090] In some embodiments, when the file system configuration file is very large, adding, deleting, or modifying the configuration does not require a full search through all configuration items. Instead, an intermediate configuration file is generated and a balanced binary tree is used to construct a configuration identification structure, allowing for rapid configuration parsing and validation. For configuration items reaching tens or even millions, binary configuration files can be parsed to quickly complete configuration parsing, loading, and validation, improving loading efficiency and reducing program startup time.

[0091] In some embodiments, as Figure 3 As shown, Figure 3 A schematic diagram of the structure of a binary configuration file provided in an embodiment of the present application includes a syntax header, configuration items and a syntax tail, and the number of configuration items is determined by the number of configuration items corresponding to the change operation.

[0092] Through this application, in response to a change operation on an original configuration file, an intermediate configuration file corresponding to the change operation is generated, the intermediate configuration file includes a first configuration identification structure corresponding to the change operation, the first configuration identification structure including at least the first configuration identification content, the first configuration identification type, the target hash value of the first configuration identification, and the location information of the first configuration identification; based on the balanced binary tree corresponding to the first configuration identification type and the target hash value, the intermediate configuration file is parsed and a second configuration identification structure is determined from the first configuration identification structure; in response to the completion of the intermediate configuration file parsing, the binary configuration file corresponding to the original configuration file is determined based on the identification content and location information in the second configuration identification structure. This can solve the technical problem of low parsing efficiency when performing full parsing when the number of configurations reaches tens of thousands in related solutions, and achieve the technical effect of improving file parsing efficiency in scenarios where the number of configurations reaches tens of thousands.

[0093] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0094] The embodiment of the present application further provides a file parsing device 400, Figure 4 A schematic diagram of the structure of a file parsing device provided in an embodiment of the present disclosure is shown in FIG. Figure 4 Shown, including:

[0095] A generating unit 401 is configured to generate, in response to a change operation on an original configuration file, an intermediate configuration file corresponding to the change operation, the intermediate configuration file including a first configuration identification structure corresponding to the change operation, the first configuration identification structure including at least first configuration identification content, first configuration identification type, target hash value of the first configuration identification, and location information of the first configuration identification;

[0096] A first determining unit 402 is configured to parse the intermediate configuration file and determine a second configuration identification structure from the first configuration identification structure based on a balanced binary tree corresponding to the first configuration identification type and a target hash value;

[0097] The second determining unit 403 is configured to determine, in response to the completion of parsing the intermediate configuration file, a binary configuration file corresponding to the original configuration file based on the identification content and location information in the second configuration identification structure.

[0098] Furthermore, in a possible implementation of the embodiment of the present disclosure, the first determining unit 402 is configured to:

[0099] Obtaining a first hash value and a second hash value corresponding to each first configuration identifier, where the first hash value and the second hash value corresponding to each first configuration identifier are determined by different hash algorithms;

[0100] Determine a target hash value corresponding to each first configuration identification content based on the first hash value and the second hash value;

[0101] Determining, based on a balanced binary tree corresponding to the first configuration identifier type, a node value in the balanced binary tree;

[0102] The intermediate configuration file is parsed, and if the node value in the balanced binary tree is the same as the target hash value, the second configuration identification structure is determined from the first configuration identification structure.

[0103] Furthermore, in a possible implementation of the embodiment of the present disclosure, the file parsing device 400 further includes an inserting unit, which is configured to:

[0104] Compare the node value in the balanced binary tree with the target hash value. If the node value in the balanced binary tree and the target hash value are different, insert the target hash value into the balanced binary tree.

[0105] Furthermore, in a possible implementation of the embodiment of the present disclosure, the second determining unit 403 is configured to:

[0106] Creating configuration item data based on the identification content and location information in the second configuration identification structure, where the configuration item data at least includes the configuration item name, pointers to each field of the configuration item, and the starting position and length of the configuration item in the overall configuration file;

[0107] Based on the configuration item data, a binary configuration file corresponding to the original configuration file is determined, where the binary configuration file at least includes a serial number of the text configuration file and a serial number of the binary configuration file.

[0108] Furthermore, in a possible implementation of the embodiment of the present disclosure, the file parsing device 400 further includes a first parsing unit, which is configured to:

[0109] In response to the binary configuration file existing in the storage file system, based on the serial number of the text configuration file and the serial number of the binary configuration file, determining a file to be parsed mapped to the memory, where the file to be parsed is one of the text configuration file and the binary configuration file;

[0110] Parse the original configuration file based on the location information of the file to be parsed.

[0111] In a possible implementation of the embodiment of the present disclosure, the intermediate configuration file includes a syntax header, a change configuration, and a syntax tail. The syntax header is used to verify the generation timestamp, configuration quantity, and header verification of the update operation. The change configuration is a newly added configuration. The syntax tail is used to verify the integrity of the intermediate configuration file.

[0112] Furthermore, in a possible implementation of the embodiment of the present disclosure, the file parsing device 400 further includes a second parsing unit, which is configured to:

[0113] In response to the change operation being a modification operation or a deletion operation, generating an intermediate configuration file corresponding to the modification operation or the deletion operation;

[0114] In response to the modification operation or the deletion operation being completed, if the binary configuration file exists in the storage file system, the original configuration file is parsed based on the location information of the binary configuration file.

[0115] Through this application, in response to a change operation on an original configuration file, an intermediate configuration file corresponding to the change operation is generated, the intermediate configuration file includes a first configuration identification structure corresponding to the change operation, the first configuration identification structure including at least the first configuration identification content, the first configuration identification type, the target hash value of the first configuration identification, and the location information of the first configuration identification; based on the balanced binary tree corresponding to the first configuration identification type and the target hash value, the intermediate configuration file is parsed and a second configuration identification structure is determined from the first configuration identification structure; in response to the completion of the intermediate configuration file parsing, the binary configuration file corresponding to the original configuration file is determined based on the identification content and location information in the second configuration identification structure. This can solve the technical problem of low parsing efficiency when performing full parsing when the number of configurations reaches tens of thousands in related solutions, and achieve the technical effect of improving file parsing efficiency in scenarios where the number of configurations reaches tens of thousands.

[0116] For the description of the features in the embodiment corresponding to the file parsing device, reference can be made to the relevant description of the embodiment corresponding to the file parsing method, which will not be repeated here.

[0117] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned file parsing method embodiments.

[0118] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned file parsing method embodiments when running.

[0119] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0120] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned file parsing method embodiments are implemented.

[0121] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned file parsing method embodiments are implemented.

[0122] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0123] The above is a detailed introduction to a file parsing method, device, electronic device and medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A file parsing method, characterized in that: include: In response to a change operation on an original configuration file, generating an intermediate configuration file corresponding to the change operation, the intermediate configuration file including a first configuration identification structure corresponding to the change operation, the first configuration identification structure including at least first configuration identification content, first configuration identification type, target hash value of the first configuration identification, and location information of the first configuration identification; Parsing the intermediate configuration file and determining a second configuration identification structure from the first configuration identification structure based on a balanced binary tree corresponding to the first configuration identification type and the target hash value; In response to the completion of parsing the intermediate configuration file, a binary configuration file corresponding to the original configuration file is determined based on the identification content and location information in the second configuration identification structure.

2. The file parsing method according to claim 1, wherein: The step of parsing the intermediate configuration file and determining the second configuration identification structure from the first configuration identification structure based on the balanced binary tree corresponding to the first configuration identification type and the target hash value includes: Obtaining a first hash value and a second hash value corresponding to each first configuration identifier, where the first hash value and the second hash value corresponding to each first configuration identifier are determined by different hash algorithms; Determine a target hash value corresponding to each first configuration identification content based on the first hash value and the second hash value; Determining, based on a balanced binary tree corresponding to the first configuration identifier type, a node value in the balanced binary tree; The intermediate configuration file is parsed, and if the node value in the balanced binary tree is the same as the target hash value, the second configuration identification structure is determined from the first configuration identification structure.

3. The file parsing method according to claim 2, wherein: After determining the node values in the balanced binary tree based on the balanced binary tree corresponding to the first configuration identifier type, the method further includes: The node value in the balanced binary tree is compared with the target hash value. If the node value in the balanced binary tree is different from the target hash value, the target hash value is inserted into the balanced binary tree.

4. The file parsing method according to claim 1, wherein: The determining, based on the identification content and location information in the second configuration identification structure, the binary configuration file corresponding to the original configuration file includes: Creating configuration item data based on the identification content and location information in the second configuration identification structure, wherein the configuration item data at least includes a configuration item name, pointers to each field of the configuration item, and a starting position and length of the configuration item in the overall configuration file; Based on the configuration item data, a binary configuration file corresponding to the original configuration file is determined, where the binary configuration file at least includes a serial number of the text configuration file and a serial number of the binary configuration file.

5. The file parsing method according to claim 4, characterized in that: The method further comprises: In response to the binary configuration file existing in the storage file system, based on the serial number of the text configuration file and the serial number of the binary configuration file, determining a file to be parsed mapped to the memory, the file to be parsed being one of the text configuration file and the binary configuration file; The original configuration file is parsed according to the location information of the file to be parsed.

6. The file parsing method according to claim 1, wherein: The intermediate configuration file includes a syntax header, a change configuration and a syntax tail. The syntax header is used to verify the generation timestamp, configuration quantity and header verification of the update operation. The change configuration is a newly added configuration. The syntax tail is used to verify the integrity of the intermediate configuration file.

7. The file parsing method according to claim 1, wherein: The method further comprises: In response to the change operation being a modification operation or a deletion operation, generating an intermediate configuration file corresponding to the modification operation or the deletion operation; In response to the modification operation or the deletion operation being completed, if a binary configuration file exists in the storage file system, the original configuration file is parsed based on the location information of the binary configuration file.

8. A file parsing device, characterized in that: include: a generating unit, configured to generate, in response to a change operation on an original configuration file, an intermediate configuration file corresponding to the change operation, the intermediate configuration file including a first configuration identification structure corresponding to the change operation, the first configuration identification structure including at least first configuration identification content, first configuration identification type, a target hash value of the first configuration identification, and location information of the first configuration identification; a first determining unit, configured to parse the intermediate configuration file and determine a second configuration identification structure from the first configuration identification structure based on a balanced binary tree corresponding to the first configuration identification type and the target hash value; The second determining unit is configured to determine, in response to completion of parsing of the intermediate configuration file, a binary configuration file corresponding to the original configuration file based on identification content and location information in the second configuration identification structure.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the file parsing method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the file parsing method according to any one of claims 1 to 7 are implemented.