Large-scale machine data query method, system and equipment and storage medium
By storing and querying real-time data generation in the ion implanter in a structured storing and querying the real-time data generation of the ion implanter, the problem of low efficiency of large-scale electrical data query in the prior art is solved, efficient and stable data processing and query are achieved, and the overall performance of the ion implanter is improved.
Patent Information
- Application Number
- CN202510167594.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-17
AI Technical Summary
The prior art is difficult to efficiently process and query large-scale real-time electrical data generated in ion implanters, resulting in low reading performance and long time, affecting user experience and possibly causing software crashes and reducing production capacity.
By creating storage files in real time in the ion implanter and sorting the files by the creation timestamp, using a structured storage format (including file pointer position, number of row data bytes, type of machine data strings and line data), the files are filtered according to the query time range during query, and the query ignores the number of rows written by the line data to calculate the query ignores the number of rows, improving the data query efficiency.
It realizes standardized storage and efficient query of machine data, improves query efficiency of ion implanter, reduces time-consuming data query, enhances software stability and reduces capacity loss.
Smart Images

Figure CN120029985A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electrical data processing, and in particular to a method, system, device and storage medium for querying large-scale machine data. Background Art
[0002] In the semiconductor field, ion implanters are highly integrated, with complex and diverse electrical circuits. They usually work 24 hours a day, and most of the data in the machines is generated in real time. Therefore, the data generated by ion implanters is very large. These data are very important for maintaining the normal operation of ion implanters, and require reliable and efficient storage and reading solutions.
[0003] At present, the commonly used data reading and writing scheme is to store and read data by relying on SQL statements (a structured query language specifically used to operate relational databases) of relational databases (a database that stores data in the form of binary tables). The above scheme often involves data going through many intermediate steps such as SQL parsing and data compression during the reading and writing process. However, for the huge and real-time electrical data generated in the ion implanter, the above method cannot guarantee the reading performance under large data volumes, and is prone to serious reading time-consuming problems, affecting the user experience. In severe cases, it may also cause the crash of the ion implanter software and reduce the production capacity of the ion implanter. Summary of the invention
[0004] The purpose of the present invention is to provide a method, system, device and storage medium for querying large-scale machine data.
[0005] The technical solution of the present invention is as follows: A method for querying large-scale machine data includes the following operations: S1. In the process of real-time operation of the ion implanter to generate machine data, several files for storing machine data are created in sequence to obtain several storage files; the storage file name is the file creation timestamp, and the storage file is composed of structure 1, structure 2, structure 3 and structure 4; the content of structure 1 includes the file pointer position, the content of structure 2 includes the number of row data bytes, the content of structure 3 includes several types of machine data character strings, and the content of structure 4 includes several row data, each row data includes the corresponding row data write timestamp and several types of machine data codes; S2. Based on the query start time and the query end time, a number of storage files are screened to obtain a number of files to be queried; for the several files to be queried, the target data query and acquisition operations are performed in sequence from the earliest to the latest in terms of the file creation timestamps, to obtain the target query data of the several files, thus forming a target query data set; the process of executing the target data query and acquisition operation for the current file to be queried is specifically as follows: obtaining the position of the target type data in the structure 3 of the current file to be queried to obtain the target type data position; obtaining the timestamp difference between the write timestamp of the first row data and the timestamp of the second row data in the structure 4 of the current file to be queried, and taking the quotient of the query time interval and the timestamp difference as the number of query ignored rows; based on the query ignored rows, the row data in the structure 4 of the current file to be queried is screened to obtain a number of row data to be queried; obtaining the data corresponding to the position of the target type data in each row data to be queried to obtain the target query data of the current file.
[0006] The number of bytes of row data in structure 2 in S1 is the sum of the number of bytes occupied by one row data in structure 4; the number of bytes of each type of machine data encoding is equal to the number of bytes of the timestamp written in each row data.
[0007] When the operation of S2 is not executed, the file pointer position in the structure 1 of S1 is the position of the last byte in the storage file at the current opening time; if the difference between the file pointer position at the current opening time and the file pointer position at the previous opening time, and the remainder of the number of row data bytes are not 0, the machine data of the storage file at the current opening time is incomplete, and a certain number of 0s are added after the last byte of the storage file at the current opening time until the difference between the file pointer position at the current opening time and the file pointer position at the previous opening time after the update, and the remainder of the number of row data bytes is 0.
[0008] In S2, in the process of executing the target data query and acquisition operation on the first file to be queried, the row data write timestamps of structure 4 in the first file to be queried are traversed one by one. During the traversal process, if the previous row data write timestamp is less than the query start time, and the current row data write timestamp is greater than the query start time, the row data corresponding to the current row data write timestamp is written as the first row data, and the row data screening operation is performed.
[0009] The target category position in S2 is the target category data in the current file to be queried, and structure 3 corresponds to the position in the string array.
[0010] The operation of obtaining the target query data in the current row data to be queried in S2 is specifically as follows: after the file pointer is moved to the byte corresponding to the timestamp of the row data written to the current row data to be queried, the file pointer continues to move backward DataIndex1×a bytes and then pauses, and the code of a bytes is read after the position of the file pointer, and the file pointer moves backward to the last code of the a-byte code to obtain the first target type data; the file pointer continues to move backward (DataIndex2-DataIndex1-1)×a bytes and then pauses, and the code of a bytes is read after the position of the file pointer, and the file pointer moves backward to the last code of the a-byte code to obtain the second target type data; the file pointer continues to move backward (DataIndex3-DataIndex2-1)×a bytes and then pauses, and the code of a bytes is read after the position of the file pointer, and the file pointer moves backward to a The third target type data is obtained at the last code in the encoding of bytes; and so on, several target query data in the current row data to be queried are obtained; DataIndex1, DataIndex2, and DataIndex3 are the first target type data position, the second target type data position, and the third target type data position, respectively, and a is the number of bytes of a type of machine data.
[0011] In S2, if the file creation timestamp of the current stored file is less than the timestamp corresponding to the query start time, and the file creation timestamp of the next stored file is greater than the timestamp corresponding to the query start time, the current stored file is used as the file to be queried; if the file creation timestamp of the current stored file is less than the timestamp corresponding to the query start time, and the file creation timestamp of the next stored file is greater than the timestamp corresponding to the query end time, the current stored file is used as the file to be queried.
[0012] A large-scale machine data query system, used to implement the large-scale machine data query method, includes: A storage file generation module is used to sequentially create a number of files for storing machine data in the process of real-time operation of the ion implanter to generate machine data, thereby obtaining a number of storage files; the storage file name is a file creation timestamp, and the storage file is composed of structure 1, structure 2, structure 3 and structure 4; the content of structure 1 includes the file pointer position, the content of structure 2 includes the number of row data bytes, the content of structure 3 includes several types of machine data character strings, and the content of structure 4 includes several row data, each row data includes a corresponding row data write timestamp and several types of machine data codes; The data query module is used to screen a number of storage files based on the query start time and the query end time to obtain a number of files to be queried; the target data query and acquisition operations are performed on the files to be queried in order from the earliest to the latest according to the file creation timestamps, to obtain the target query data of the files, and to form a target query data set; the process of executing the target data query and acquisition operation on the current file to be queried is specifically as follows: the position of the target type data in the structure 3 of the current file to be queried is obtained to obtain the target type data position; the timestamp difference between the write timestamp of the first row data and the timestamp of the second row data in the structure 4 of the current file to be queried is obtained, and the quotient of the query time interval and the timestamp difference is used as the query ignored row number; based on the query ignored row number, the row data in the structure 4 of the current file to be queried is screened to obtain a number of row data to be queried; the data corresponding to the target type data position in each row data to be queried is obtained to obtain the target query data of the current file.
[0013] A large-scale machine data query device comprises a processor and a memory, wherein the processor implements the large-scale machine data query method when executing a computer program stored in the memory.
[0014] A computer-readable storage medium is used to store a computer program, wherein the computer program implements the above-mentioned large-scale machine data query method when executed by a processor.
[0015] The beneficial effects of the present invention are: The present invention provides a large-scale machine data query method. In the process of generating machine data by real-time operation of an ion implanter, several files for storing machine data are created in sequence to obtain several storage files. Each storage file is composed of a structure 1 including a file pointer position, a structure 2 including a row data byte number, a structure 3 including several types of machine data character strings, and a structure 4 including several row data, so as to realize the standardization and clarification of machine data storage, and facilitate subsequent inspection and rapid search of machine data. In addition, in the data query process, according to the query start time and query end time, the storage files are filtered to obtain files to be queried that meet the query time range, and according to the timestamps written between two adjacent row data and the query time interval, the row data in the structure of each file to be queried is filtered to improve the data query efficiency and obtain several row data to be queried. Finally, according to the position of the target type data in the structure 3 of the current file to be queried, the corresponding data in each row data to be queried is obtained to obtain the target query data. The method is applied to the large-scale machine data query of the ion implanter, and can improve the query efficiency, reduce the time consumption of large-scale machine data query, improve the stability of the ion implanter software, and reduce the loss of the ion implanter production capacity. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] By reading the detailed description of the preferred embodiment below, the scheme and advantages of the present application will become clear to those skilled in the art. The accompanying drawings are only for the purpose of illustrating the preferred embodiment and are not to be considered as limiting the present invention.
[0017] In the attached picture: Figure 1 Schematic diagram of the structure of the storage file in the embodiment. DETAILED DESCRIPTION
[0018] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings.
[0019] This embodiment provides a method for querying large-scale machine data, including the following operations: S1. In the process of real-time operation of the ion implanter to generate machine data, several files for storing machine data are created in sequence to obtain several storage files; the storage file name is the file creation timestamp, and the storage file is composed of structure 1, structure 2, structure 3 and structure 4; the content of structure 1 includes the file pointer position, the content of structure 2 includes the number of row data bytes, the content of structure 3 includes several types of machine data character strings, and the content of structure 4 includes several row data, each row data includes the corresponding row data write timestamp and several types of machine data codes; S2. Based on the query start time and the query end time, a number of stored files are screened to obtain a number of files to be queried; the target data query and acquisition operations are performed on the files to be queried in order from the earliest to the latest file creation timestamps to obtain the target query data of the files, thereby forming a target query data set.
[0020] S1. In the process of real-time operation of the ion implanter to generate machine data, a plurality of files for storing the machine data are sequentially created to obtain a plurality of storage files.
[0021] During the real-time operation of the ion implanter, a huge amount of machine data is generated. In order to accurately store the machine data, several files for storing the machine data are created in sequence during the process of generating the machine data to obtain several storage files.
[0022] The storage file name is the file creation timestamp, which can be the total number of seconds from 00:00:00 Greenwich time on January 1, 1970 (08:00:00 Beijing time on January 1, 1970) to the present, or the number of cycles from midnight on January 1, 0001 to the present with a cycle of 100ns. If the data size stored in the current storage file is equal to the standard file storage size, the next storage file is created. The later the storage file is created, the larger the corresponding file creation timestamp is. At the same time, the file creation timestamp is less than the machine data acquisition timestamp, that is, the file name time of the storage file is earlier than the machine data acquisition time in the corresponding storage file, that is, the storage file is created first, and then the machine data is written.
[0023] In order to standardize the storage of machine data and facilitate subsequent inspection and quick search of machine data, in this embodiment, see Figure 1 , the storage file consists of structure 1, structure 2, structure 3 and structure 4, that is, the content of the storage file consists of the contents of structure 1, structure 2, structure 3 and structure 4.
[0024] The content in structure 1 includes the file pointer position. When the file pointer position changes, the content of structure 1 will change. When the operation of S2 is not performed, the file pointer position in structure 1 of the storage file is the position of the last byte in the storage file, and the file pointer position will change as the file pointer moves. That is, when the storage file does not perform data query, the file pointer position is the position of the last byte in the storage file. When the storage file is performing the operation of S2, the file pointer will move according to the operation instruction, and the file pointer position will change as the file pointer moves.
[0025] The content in structure 2 includes the number of row data bytes, which is the sum of the number of bytes of all types of machine data to be stored in the storage file and the number of bytes of the timestamp of a row data, that is, the total number of bytes occupied by a row data. The number of bytes of each type of machine data (or type of machine data encoding) and the number of bytes of the timestamp of each row data are equal and fixed, both a, indicating the size of the number of bytes of a row data. For example, each row data in the storage file (structure 4) will store 20 types of machine data, so the number of row data bytes is (20+1)×a bytes.
[0026] The content in structure 3 includes several types of machine data character strings, which are arranged in sequence in the form of: machine data DataA character string, machine data DataB character string, machine data DataC character string...machine data DataM character string.
[0027] The content of structure 4 includes several rows of data, each of which includes the corresponding row of data write timestamp (the timestamp when the row of data starts to be created) and several types of machine data codes. The machine data code is in binary form, so the form of a row of data is: row data write timestamp, machine data DataA binary code stream, machine data DataB binary code stream, machine data DataC binary code stream... machine data DataM binary code stream. In other words, the number of row data bytes in structure 2 is the sum of the number of bytes occupied by a row of data in structure 4.
[0028] In addition, when all types of machine data in structure 3 are written into the current row of data in structure 4, it is considered that the writing of the current row of data is completed, and the writing of the next row of data is executed.
[0029] In addition, in order to prevent the row data in structure 4 from being incomplete during the actual use of the storage file, especially during repeated opening and closing, when the S2 operation is not executed, the file pointer position in structure 1 of the storage file corresponds to the position of the last byte in the storage file at the current opening moment.
[0030] If the difference between the file pointer position at the current opening time and the file pointer position at the previous opening time, and the remainder of the number of row data bytes are not 0, the machine data of the file stored at the current opening time is incomplete, and a number of 0s are added after the last byte of the file stored at the current opening time. At this time, the file pointer position automatically moves to the last byte position to realize automatic updating of the file pointer position until the difference between the file pointer position at the current opening time and the file pointer position at the previous opening time after the update, and the remainder of the number of row data bytes is 0, so that the total row data bytes of the storage file structure 4 are the same each time it is opened, reaching the preset standard total row data bytes.
[0031] For example, the file pointer position at the current opening time is 27, and the file pointer position at the last opening time is 30. The difference between the file pointer position at the current opening time and the file pointer position at the last opening time is 3, and the remainder with the number of row data bytes is not 0. In this case, the machine data of the file stored at the current opening time is incomplete, so 3 zeros are added after the last byte of the file stored at the current opening time.
[0032] S2. Based on the query start time and the query end time, a number of stored files are screened to obtain a number of files to be queried; the target data query and acquisition operations are performed on the files to be queried in order from the earliest to the latest file creation timestamps to obtain the target query data of the files, thereby forming a target query data set.
[0033] According to the query start time and query end time, the storage files are filtered to obtain the files to be queried that meet the query time range, and according to the timestamps written between two adjacent rows of data and the query time interval, the row data in the structure of each file to be queried is filtered to improve the data query efficiency and obtain several rows of data to be queried; finally, according to the position of the target type data in the structure 3 of the current file to be queried, the corresponding data in each row of data to be queried is obtained to obtain the target query data.
[0034] First, based on the query start time and the query end time, a number of storage files are screened to select storage files that meet the data query time range, thereby obtaining a number of files to be queried.
[0035] One screening method is (when the query start time and query end time are used for machine data writing time): if the file creation timestamp of the current storage file is less than the timestamp corresponding to the query start time, and the file creation timestamp of the next storage file is greater than the timestamp corresponding to the query start time, then the current storage file is used as the file to be queried. If the file creation timestamp of the current storage file is less than the timestamp corresponding to the query end time, and the file creation timestamp of the next storage file is greater than the timestamp corresponding to the query end time, then the current storage file is used as the file to be queried.
[0036] Another filtering method is (when the query start time and query end time are used for the storage file creation time): the storage files with file creation timestamps within the query start time and query end time range are taken as the files to be queried.
[0037] Then, the target data query acquisition operation is performed in sequence for several files to be queried according to the order of the file creation timestamps from earliest to latest, and the target query data of several files are obtained to form a target query data set. There are two target data query acquisition methods, which are as follows.
[0038] In the first method, in the target data query and acquisition method based on row data filtering, the process of executing the target data query and acquisition operation on the current file to be queried is specifically as follows.
[0039] Step 1: Obtain the location of the target category data in the structure 3 of the current file to be queried, and obtain the location of the target category data.
[0040] Specifically, according to the target type data name and the number of types to be queried at the same time, the structure 3 is converted into a string array through the function of operating string type data in the programming language, the target type data name is searched for its position in the string array, and each target type data position is recorded as DataIndex1, DataIndex2, DataIndex3...DataIndexN. That is, the target type data position is the position of the target type data in the current file to be queried, and the structure 3 corresponds to the position in the string array.
[0041] For example, structure 3 is machine data DataA string, machine data DataB string, machine data DataC string, machine data DataD string, machine data DataE string, and the corresponding string array positions are [0,1,2,3,4]. The target type data names are machine data DataB, machine data DataD, and machine data DataE. The number of types queried simultaneously is 3. The corresponding target type data positions are 1, 3, and 4, respectively, recorded as DataIndex1, DataIndex2, and DataIndex3.
[0042] Step 2: Obtain the timestamp difference between the write timestamp of the first row of data and the timestamp of the second row of data in structure 4 of the current file to be queried, and use the quotient of the query time interval and the timestamp difference as the number of rows to be ignored in the query; based on the number of rows to be ignored in the query, filter the row data in structure 4 of the current file to be queried to obtain a number of row data to be queried.
[0043] The method for obtaining the number of rows ignored by the query is as follows: after the file pointer moves to the byte corresponding to the first row data write timestamp in the structure 4 of the current file to be queried, the file pointer continues to move backward by the number of bytes of the difference between the number of bytes of the row data and the number of bytes corresponding to the first row data write timestamp, so as to jump to the byte corresponding to the second row data write timestamp; according to the timestamp difference between the first row data write timestamp and the second row data timestamp, the actual writing time of the first row data can be determined. Then the quotient of the preset query time interval and the timestamp difference (the actual writing time of the first row data) is the number of rows that need to be ignored in the data query, which is the number of rows ignored by the query and is recorded as Span.
[0044] Next, based on the query ignored row number Span, the row data in structure 4 of the current file to be queried is filtered to reduce the amount of row data to be queried, improve data query efficiency, and obtain a number of row data to be queried. For example, when the query ignored row number Span is 3, if the original first row data in structure 4 is the row data to be queried, the original fifth row data in structure 4 is used as the row data to be queried.
[0045] Step 3: Obtain the data corresponding to the target type data position in each row of data to be queried, and obtain the target query data of the current file.
[0046] The specific operation method for obtaining the target query data in the current row data to be queried is as follows.
[0047] One method is (reading data after the file pointer position): after the file pointer is moved to the row data corresponding to the timestamp of the current row data to be queried, the file pointer continues to move backward DataIndex1×a bytes and then pauses. At this time, the file pointer is located in front of the type data to be obtained, and then a bytes of code are read after the file pointer position. During the data reading process, the file pointer moves to the corresponding read data position. After reading a bytes of code, the file pointer also moves backward to the last code in the a bytes of code to obtain the first target type data; the file pointer continues to move backward (DataIndex2-DataIndex1-1)×a bytes and then pauses, and reads a bytes of code after the file pointer position. After the data is read, the file pointer also moves backward to the last code in the a bytes of code to obtain the second target type data; the file pointer continues to move backward (DataIndex3-DataIndex2-1)×a bytes and then stops, and reads a bytes of code after the file pointer position. After the data is read, the file pointer moves backward to the last code in the a-byte code to obtain the third target type data; and so on, the file pointer continues to move backward (DataIndexN-DataIndex(N-1)-1)×a bytes and then stops, and reads a-byte codes after the file pointer position to obtain the Nth target type data, thereby obtaining several target query data in the current row data to be queried; DataIndex1, DataIndex2, and DataIndex3 are the first target type data position, the second target type data position, the third target type data position, the N-1th target type data position, and the Nth target type data position, respectively, and a is the number of bytes of a type of machine data.
[0048] For example, when the target data type names are machine data DataB, machine data DataD, and machine data DataE, the corresponding target data type positions are DataIndex1, DataIndex2, and DataIndex3, respectively. 1, 3, and 4, the file pointer is moved to the row data corresponding to the timestamp of the row data to be queried, and the file pointer continues to move backward for 1×a bytes and then pauses. The file pointer reaches the machine data DataB, and then reads the a-byte code behind the file pointer to obtain the machine data DataB. After the machine data DataB is acquired, the file pointer moves to the last code in the binary code of the machine data DataB; the file pointer continues to move backward for (3-1-1)×a bytes and then pauses, and reads the a-byte code behind the file pointer to obtain the machine data DataD; after the machine data DataD is acquired, the file pointer moves to the last code in the binary code of the machine data DataD, which is also just before the machine data DataE, and the file pointer continues to move backward for (4-3-1)×a bytes and then pauses, that is, the file pointer does not move, and reads the a-byte code behind the file pointer to obtain the machine data DataE.
[0049] Another method is (reading data before the file pointer position): after the file pointer is moved to the row data corresponding to the timestamp of the current row data to be queried, the file pointer continues to move backward (DataIndex1+1)×a bytes and then pauses. After the file pointer reaches the position of the target type data to be obtained, read the a-byte code before the file pointer position to obtain the first target type data. After the data is obtained, the file pointer moves to the last code in the a-byte code; the file pointer continues to move backward (DataIndex2-DataIndex1-1)×a bytes and then pauses. Read the a-byte code before the file pointer position to obtain the second target type data. After the data is obtained, the file pointer moves to the last code in the a-byte code; the file pointer continues to move backward (DataIndex3-DataIndex2-1)×a bytes and then pauses. Read the a-byte code before the file pointer position to obtain the third target type data. After the data is obtained, the file pointer moves to a By analogy, the file pointer continues to move backward by (DataIndexN-DataIndex(N-1)-1)×a bytes and then pauses, and reads the a-byte code before the file pointer to obtain the Nth target type data, thereby obtaining several target query data in the current row data to be queried.
[0050] The second method, in the method for querying and obtaining target data without data filtering, that is, when Span = 0, the process of performing the target data query and acquisition operation on the currently to-be-query file is as follows: Obtain the position of the target type of data in Structure 3 of the currently to-be-query file to obtain the target type data position; obtain the data corresponding to the target type data position in each row of data to obtain the target query data of the current file. The methods for obtaining the target type data position and the target query data are the same as above.
[0051] In addition, during the process of performing the target data query and acquisition operation on the first to-be-query file, traverse the write timestamps of the row data in Structure 4 of the first to-be-query file one by one. During the traversal, if the write timestamp of the previous row of data is less than the query start time and the write timestamp of the current row of data is greater than the query start time, then write the row data corresponding to the write timestamp of the current row of data as the first row of data and perform the operation of filtering the row data.
[0052] This embodiment also provides a query system for large-scale machine data for implementing the above-mentioned method for querying large-scale machine data, including: A storage file generation module, configured to sequentially create several files for storing machine data during the process of the ion implanter running in real time to generate machine data, obtaining several storage files; the names of the storage files are the file creation timestamps, and the storage files are composed of Structure 1, Structure 2, Structure 3, and Structure 4; the content in Structure 1 includes the file pointer position, the content in Structure 2 includes the number of bytes of the row data, the content in Structure 3 includes several types of machine data strings, and the content in Structure 4 includes several rows of data, and each row of data includes the write timestamp of the corresponding row of data and several types of machine data codes; A data query module, configured to screen several storage files based on the query start time and the query end time to obtain several to-be-query files; the several to-be-query files sequentially perform the target data query and acquisition operation in the order from the earliest to the latest file creation timestamp to obtain the target query data of several files, forming a target query data set; the process of performing the target data query and acquisition operation on the currently to-be-query file is specifically as follows: Obtain the position of the target type of data in Structure 3 of the currently to-be-query file to obtain the target type data position; obtain the timestamp difference between the write timestamp of the first row of data and the second row of data timestamp in Structure 4 of the currently to-be-query file, and use the quotient of the query time interval and the timestamp difference as the number of rows to be ignored in the query; based on the number of rows to be ignored in the query, screen the row data in Structure 4 of the currently to-be-query file to obtain several to-be-query row data; obtain the data corresponding to the target type data position in each to-be-query row data to obtain the target query data of the current file.
[0053] This embodiment further provides a large-scale machine data query device, including a processor and a memory, wherein the processor implements the large-scale machine data query method when executing a computer program stored in the memory.
[0054] This embodiment also provides a computer-readable storage medium for storing a computer program, wherein the computer program implements the above-mentioned large-scale machine data query method when executed by a processor.
[0055] The present embodiment provides a large-scale machine data query method. In the process of generating machine data by real-time operation of an ion implanter, several files for storing machine data are created in sequence to obtain several storage files. Each storage file is composed of a structure 1 including a file pointer position, a structure 2 including a row data byte number, a structure 3 including several types of machine data character strings, and a structure 4 including several row data, so as to realize the standardization and clarification of machine data storage, and facilitate subsequent inspection and rapid search for machine data. In addition, in the data query process, according to the query start time and query end time, the storage files are filtered to obtain files to be queried that meet the query time range, and according to the timestamps written between two adjacent row data and the query time interval, the row data in the structure of each file to be queried is filtered to improve the data query efficiency and obtain several row data to be queried. Finally, according to the position of the target type data in the structure 3 of the current file to be queried, the corresponding data in each row data to be queried is obtained to obtain the target query data. The method is applied to the large-scale machine data query of the ion implanter, which can improve the query efficiency, reduce the time consumption of large-scale machine data query, improve the stability of the ion implanter software, and reduce the loss of the ion implanter production capacity.
Claims
1. A method for querying large-scale machine data, characterized in that: The following operations are included: S1. In the process of real-time operation of the ion implanter to generate machine data, a plurality of files for storing machine data are sequentially created to obtain a plurality of storage files; the storage file name is the file creation timestamp, and the storage file consists of structure 1, structure 2, structure 3 and structure 4; The content of structure 1 includes the file pointer position, the content of structure 2 includes the number of bytes of row data, the content of structure 3 includes several types of machine data strings, and the content of structure 4 includes several rows of data, each row of data includes the corresponding row data write timestamp and several types of machine data codes; S2. Based on the query start time and the query end time, a number of stored files are screened to obtain a number of files to be queried; the target data query acquisition operation is performed on the files to be queried in order from the earliest to the latest file creation timestamps to obtain the target query data of the files, thereby forming a target query data set; The specific process of executing the target data query and acquisition operation for the current file to be queried is as follows: Obtain the location of the target type data in structure 3 of the current file to be queried, and obtain the location of the target type data; Get the timestamp difference between the write timestamp of the first row of data and the timestamp of the second row of data in structure 4 of the current file to be queried, and use the quotient of the query time interval and the timestamp difference as the number of rows to be ignored in the query; Based on the number of ignored rows in the query, the row data in structure 4 of the current file to be queried is filtered to obtain a number of row data to be queried; The data corresponding to the target type data position in each row of data to be queried is obtained to obtain the target query data of the current file.
2. The method for querying large-scale machine data according to claim 1, characterized in that: The number of bytes of row data in structure 2 in S1 is the sum of the number of bytes occupied by one row of data in structure 4; The number of bytes used to encode each type of machine data is equal to the number of bytes used to write the timestamp for each row of data.
3. The method for querying large-scale machine data according to claim 1, characterized in that: When the operation of S2 is not performed, the file pointer position in the structure 1 of S1 is the position of the last byte in the storage file at the current opening time; If the difference between the file pointer position at the current opening time and the file pointer position at the previous opening time, and the remainder of the number of row data bytes is not 0, the machine data of the file stored at the current opening time is incomplete, and a certain number of 0s are added after the last byte of the file stored at the current opening time until the difference between the file pointer position at the current opening time and the file pointer position at the previous opening time after update, and the remainder of the number of row data bytes is 0.
4. The method for querying large-scale machine data according to claim 1, characterized in that: In S2, during the process of executing the target data query and acquisition operation on the first file to be queried, the row data write timestamps of structure 4 in the first file to be queried are traversed one by one. During the traversal process, if the previous row data write timestamp is less than the query start time, and the current row data write timestamp is greater than the query start time, the row data corresponding to the current row data write timestamp is written as the first row data, and the row data screening operation is performed.
5. The method for querying large-scale machine data according to claim 1, characterized in that: In S2, the target category position is the position of the target category data in the current file to be queried, and structure 3 corresponds to the position in the string array.
6. The method for querying large-scale machine data according to claim 1, characterized in that: In S2, the operation of obtaining the target query data in the current row data to be queried is specifically: After the file pointer is moved to the byte corresponding to the timestamp of the row data written in the current row data to be queried, the file pointer continues to move backward DataIndex1×a bytes and then pauses, and reads the code of a bytes after the position of the file pointer, and the file pointer moves backward to the last code in the code of a bytes to obtain the first target type data; the file pointer continues to move backward (DataIndex2-DataIndex1-1)×a bytes and then pauses, and reads the code of a bytes after the position of the file pointer, and the file pointer moves backward to the last code in the code of a bytes to obtain the second target type data; the file pointer continues to move backward (DataIndex3-DataIndex2-1)×a bytes and then pauses, and reads the code of a bytes after the position of the file pointer, and the file pointer moves backward to the last code in the code of a bytes to obtain the third target type data; and so on, several target query data in the current row data to be queried are obtained; DataIndex1, DataIndex2, DataIndex3 are the first target type data position, the second target type data position, and the third target type data position respectively, and a is the number of bytes of one type of machine data.
7. The method for querying large-scale machine data according to claim 1, characterized in that: In S2, if the file creation timestamp of the current storage file is less than the timestamp corresponding to the query start time, and the file creation timestamp of the next storage file is greater than the timestamp corresponding to the query start time, the current storage file is used as the file to be queried; If the file creation timestamp of the current storage file is less than the timestamp corresponding to the query start time, and the file creation timestamp of the next storage file is greater than the timestamp corresponding to the query end time, the current storage file is used as the file to be queried.
8. A large-scale machine data query system, used to implement the large-scale machine data query method according to claim 1, characterized in that: include: A storage file generation module is used to sequentially create a number of files for storing machine data in the process of real-time operation of the ion implanter to generate machine data, thereby obtaining a number of storage files; the storage file name is a file creation timestamp, and the storage file is composed of structure 1, structure 2, structure 3 and structure 4; the content of structure 1 includes the file pointer position, the content of structure 2 includes the number of row data bytes, the content of structure 3 includes several types of machine data character strings, and the content of structure 4 includes several row data, each row data includes a corresponding row data write timestamp and several types of machine data codes; The data query module is used to screen a number of storage files based on the query start time and the query end time to obtain a number of files to be queried; the target data query and acquisition operations are performed on the files to be queried in order from the earliest to the latest file creation timestamps to obtain the target query data of the files to form a target query data set; the process of executing the target data query and acquisition operation on the current file to be queried is specifically as follows: obtaining the position of the target type data in the structure 3 of the current file to be queried to obtain the position of the target type data; obtaining the timestamp difference between the first row data write timestamp and the second row data timestamp in the structure 4 of the current file to be queried, and taking the quotient of the query time interval and the timestamp difference as the number of rows to be ignored in the query; Based on the number of query ignored rows, the row data in structure 4 of the current file to be queried is screened to obtain a number of row data to be queried; the data corresponding to the target type data position in each row data to be queried is obtained to obtain the target query data of the current file.
9. A large-scale machine data query device, characterized in that: The method comprises a processor and a memory, wherein when the processor executes the computer program stored in the memory, the method for querying large-scale machine data as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the large-scale machine data query method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Aggregate query method and system for traffic data flows
CN104156524A
Oil and gas field time series data storage method and device, oil and gas field time series data query method and device and storage medium
CN112286867A
Generation method and device of slow query log, equipment and medium
CN117971887A
Point-in-time query method and system
US20070271242A1
Systems and methods for range keys to enable efficient bulk writes in log-structured merge tree storage
US20240126738A1