Log blacklist keyword detection method and system

Through concurrent processing and streaming processing technology, and using SIMD instruction set to perform fine-grained parallel processing within each thread, the problem of low efficiency in large-scale log file processing is solved, and efficient log analysis and extraction is achieved.

CN119961235AActive Publication Date: 2025-05-09POWERLEADER COMPUTER SYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510438274.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently process and analyze large-scale log files, resulting in high system resource utilization and low processing efficiency, which cannot meet the needs of fast and accurate analysis.

Method used

Concurrency processing technology and streaming processing methods are adopted to process log files in parallel through multi-threads or multiple processes, and fine-grained parallel processing is achieved through SIMD instruction set within each thread to improve keyword search efficiency.

Benefits of technology

It significantly improves the processing efficiency of large-scale log files, reduces the occupation of system resources, and realizes efficient extraction and analysis of server logs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961235A_ABST
    Figure CN119961235A_ABST
Patent Text Reader

Abstract

The invention relates to a log blacklist keyword detection method and system. The method comprises the following steps: initializing data, and determining the size of a log block to be processed by each thread; starting a main program, obtaining a log file path, opening a file, calculating the size of the file, and sequentially executing log blocking and thread starting; creating threads and task allocation, calculating a task range of each thread and a file offset range needing to be processed by each thread, creating an independent file pointer for each thread and transmitting the corresponding task range; the log blocks are read in a streaming mode and processed, in each thread, file content is read step by step, and then keyword matching processing of SIMD optimization is carried out on each data block; executing 1 to N log blocks through a thread; and clearing and releasing the memory. The server log file analysis method provided by the invention has the advantages of efficient keyword search, streaming processing and fine-grained parallelization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method and system for detecting log blacklist keywords, and belongs to the field of communications. Background Art

[0002] As the core equipment of the network, the server undertakes key tasks such as data storage, processing and transmission. In the daily operation of the server, a large amount of log information is generated. These log information records important information such as the server's operating status, user operations, system events, etc., which are extremely valuable for monitoring server performance, troubleshooting, ensuring network security, and optimizing system configuration.

[0003] However, as the amount of data processed by servers continues to increase and business complexity increases, the size of log files is also showing an explosive growth trend. In actual applications, even the log files generated by small servers may reach several megabytes or even tens of megabytes in size, while the log files generated by large servers or data centers may reach several gigabytes or even several terabytes in size.

[0004] At present, for small-scale log files, they can generally be opened and viewed through a simple text editor or log viewing tool, and searched by manually entering keywords or blacklists to obtain the required information. However, this method is powerless when faced with large-scale log files. When trying to open a large log file, the computer's memory and processing power often cannot meet the needs due to the large file content, causing the computer to run slowly or even freeze, greatly reducing work efficiency. In addition, the manual search method is not only time-consuming and labor-intensive, but also easy to miss important information, and cannot meet the needs of fast and accurate analysis of log data. Summary of the invention

[0005] The present invention provides a method and system for detecting a log blacklist keyword, aiming to solve at least one of the technical problems existing in the prior art.

[0006] The technical solution of the present invention relates to a method for detecting log blacklist keywords, and the method according to the present invention comprises the following steps: S100, initializing data and determining the log block size to be processed by each thread; S200, the main program starts, obtains the log file path and opens the file, calculates the file size, and performs log segmentation and thread startup in sequence; S300, creating threads and task allocation, calculating the task range of each thread and the file offset range that each thread needs to process, creating an independent file pointer for each thread and passing the corresponding task range; S400, streaming reading and processing of log blocks, in each thread, gradually reading the file content and then performing SIMD optimized keyword matching processing on each data block; S500, executing 1 to N log blocks through threads; S600, clean up and release memory.

[0007] Further, the step S100 includes: S110, load the required library and header files; S120, defining constants and global variables, including setting BLOCK_SIZE for calculating the size of the log block to be processed by each thread and a predefined constant NUM_THREADS for indicating the number of threads to be created; and defining a function of init_keyword_first_chars for initializing a variable keyword_first_chars, wherein the variable keyword_first_chars is used to store a mask of the first character of the keyword; S130, defining a blacklist keyword list; S140, define auxiliary functions and structures; among them, detailed_match is used to check whether the complete keyword is matched; thread_data_t is used to store the processing result of each thread; init_keyword_first_chars is used to initialize the SIMD vector to make it the first character of the keyword.

[0008] Further, the step S200 includes: S210, parsing and checking command line parameters to obtain a log file path; if the number of parameters is incorrect, outputting usage instructions and exiting the program in a failed state; S220, opening the log file, and attempting to open the file specified in the command line parameter in read-only binary mode; wherein, if the file opening fails, outputting an error message, and exiting the program in a failure state; S230, obtaining the file size; wherein, the file pointer is moved to the end of the file, the position of the current file pointer is obtained, and is stored in the log_data_size variable, and the file pointer is reset to the beginning of the file again for subsequent file reading operations.

[0009] Further, the step S300 includes: S310, calculating the task scope of each thread, and calculating the file offset range that each thread should process according to the file size and the number of threads; S320, create threads, create an independent file pointer for each thread, and pass the corresponding task scope.

[0010] Furthermore, in step S300, a for loop is used to initialize the data of each thread and create a thread, which includes the steps of: S301, in the loop, the data structure thread_data[i] of each thread is initialized; S302, opening the file pointer, and opening the file specified in the command line parameter in read-only binary mode; wherein, if the file opening fails, outputting an error message and exiting the program; S303, calculating the start and end offsets of each thread so that each thread processes a continuous block of the log file; S304. Use the pthread_create function to create a thread. The function executed by the thread is process_block. The parameter passed to the pthread_create function is the address of thread_data[i].

[0011] Further, the step S400 includes: S410, initializing the SIMD vector to the first character of the keyword; S420, using SIMD instructions to quickly filter possible keyword starting positions; S430, performing detailed string matching on the starting position of the keyword to confirm whether the keyword is actually matched; S440, handle the situation where the last block is less than BLOCK_SIZE, and ensure that all data are processed correctly.

[0012] Further, the step S500 includes: S510, using a for loop to initialize threads and thread data, wherein the loop variable i starts from 0 and goes up to NUM_THREADS; S520, inside the loop, the data structure thread_data[i] of each thread is initialized; S530, setting the start and end offsets of each thread; S540, setting the ID of each thread; wherein thread_id is set to the index i of the current loop; S550, then create a thread, wherein a pthread_create function is used to create a new thread, and the thread will execute the process_block function; &threads[i] is passed as the thread identifier, NULL as the thread attribute, process_block as the thread function, and &thread_data[i] as the parameter of the thread function.

[0013] Further, the step S600 includes: S610, use a for loop to traverse all threads, variable i starts from 0 until NUM_THREADS; S620, inside the loop, waiting for each thread to end; if the thread ends successfully, returning 0; if an error occurs, returning a non-zero value, indicating an error occurred while waiting for the thread, obtaining an error description, and exiting the program with a failed status; S630: After the loop ends, close the previously opened file.

[0014] The technical solution of the present invention also relates to a computer-readable storage medium on which program instructions are stored, and the above-mentioned method is implemented when the program instructions are executed by a processor.

[0015] The technical solution of the present invention also relates to a detection system for log blacklist keywords, the system comprising a computer device, and the computer device comprises the above-mentioned computer-readable storage medium.

[0016] The beneficial effects of the present invention are as follows: The log blacklist keyword detection method and system of the present invention significantly improve the processing efficiency of log files, reduce the occupation of system resources, and realize efficient extraction and analysis of server logs through efficient keyword search, concurrent processing, streaming processing and fine-grained parallel processing.

[0017] The present invention realizes vectorized processing of log data by introducing SIMD instruction set, so that the keyword search process can be executed in parallel on multiple data blocks. Concurrent processing technology is adopted to process log files in parallel by multi-threading or multi-process, and the log files are divided into multiple subtasks, and are assigned to multiple threads or processes for simultaneous processing. Considering that the log file may be very large, it is unrealistic to load the entire file into the memory at one time. The present invention adopts a streaming processing method, which reduces the memory usage by reading the log file block by block. In addition to thread-level parallelism, the present invention also realizes fine-grained parallelism by using SIMD instruction set within each thread, further tapping the potential of parallel processing, so that each thread can more efficiently utilize CPU resources when processing log data. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a basic flow chart of the method according to the present invention. DETAILED DESCRIPTION

[0019] The concept, specific structure and technical effects of the present invention will be clearly and completely described below in combination with the embodiments and drawings to fully understand the purpose, scheme and effect of the present invention.

[0020] It should be noted that, unless otherwise specified, when a feature is referred to as being "fixed" or "connected" to another feature, it may be directly fixed or connected to another feature, or it may be indirectly fixed or connected to another feature. The singular forms "a", "said" and "the" used herein are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art. The terms used in this specification are intended only to describe specific embodiments and are not intended to limit the invention. The term "and / or" used herein includes any combination of one or more of the related listed items.

[0021] It should be understood that, although the term first, second, third etc. may be adopted to describe various elements in the present disclosure, these elements should not be limited to these terms. These terms are only used to distinguish the same type of elements from each other. For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language ("for example", "such as" etc.) provided herein is only intended to better illustrate embodiments of the present invention, and unless otherwise required, will not impose limitations on the scope of the present invention.

[0022] Reference Figure 1 In some embodiments, the method for detecting a log blacklist keyword according to the present invention comprises at least the following steps: S100, initialize data and define configuration.

[0023] S200: The main program starts to open the log file to be read.

[0024] S300: After creating a thread and allocating tasks, start executing the thread.

[0025] S400, streaming reading and processing of log blocks.

[0026] S500 , executing 1 to N processing log blocks (process_block) through threads.

[0027] S600: After waiting for all threads to finish, clean up and release memory.

[0028] In view of the fact that current technologies lack analysis of large log files and cannot efficiently extract keywords and solve practical scenario problems, the present invention proposes an efficient and low-resource log processing method, which significantly improves the processing efficiency of large-scale log files, reduces the occupation of system resources, and realizes efficient extraction and analysis of server logs through efficient keyword search, concurrent processing, streaming processing and fine-grained parallel processing.

[0029] The present invention realizes vectorized processing of log data by introducing SIMD (Single Instruction Multiple Data) instruction set. In traditional log analysis methods, keyword search usually adopts byte-by-byte or character-by-character sequential scanning, which is inefficient when processing large-scale log files and has low utilization of CPU resources. In contrast, the SIMD instruction set used in the present invention can process multiple data units simultaneously in one instruction cycle, thereby significantly improving the efficiency of data processing.

[0030] The present invention adopts concurrent processing technology to process log files in parallel through multi-threading or multi-processing. Compared with the traditional single-thread processing method which is often limited by the processing power of a single CPU core, the present invention makes full use of the computing power of multi-core CPUs. Also, considering that the log file may be very large, it is unrealistic to load the entire file into the memory at one time. The present invention adopts streaming processing technology to reduce the memory usage by reading the log file block by block.

[0031] Considering that the log file may be very large, it is unrealistic to load the entire file into the memory at one time. The present invention adopts a streaming processing method, which reduces the memory usage by reading the log file block by block. The core idea of ​​streaming processing is to divide the log file into multiple small blocks, process only one data block at a time, release the memory occupied by the data block after the processing is completed, and then read the next data block to continue processing. The streaming processing method avoids the problem of insufficient memory caused by loading the entire file at one time, so that the present invention can efficiently process ultra-large-scale log files.

[0032] In addition to thread-level parallelism, the present invention also uses SIMD instruction sets to achieve fine-grained parallelism within each thread. In traditional parallel processing, thread-level parallelism mainly depends on the number of cores of a multi-core CPU, while the fine-grained parallelism of the present invention performs data parallel processing within each thread. For example, in one thread, multiple data units can be processed simultaneously through the SIMD instruction set, rather than the traditional one-by-one processing method, which further improves the efficiency of data processing.

[0033] In some embodiments, the initialization and preparation phase of the present invention first loads the necessary libraries and header files, including<immintrin.h> ,<stdio.h> ,<string.h> ,<pthread.h> C language libraries. Then define constants and global variables, specifically, set BLOCK_SIZE (64-byte alignment), BUFFER_SIZE (for example, 64KB) and NUM_THREADS (number of threads). Then, define the blacklist keyword list keywords, such as "error", "failure", "timeout", etc. Then, define auxiliary functions and structures, specifically, detailed_match is used to check whether the complete keyword is matched; thread_data_t is used to store the processing results of each thread; init_keyword_first_chars is used to initialize the SIMD vector to make it the first character of the keyword.

[0034] Specifically, first, a string array named blacklist is defined, which contains three strings, namely "error", "failure" and "timeout". Then, a structure thread_data_t is defined to store thread-related data. Among them, the structure hread_data_t contains the following members: FILE *file, which is a pointer to the file and may be used for file operations; off_t start_offset, which is the starting offset of the file; off_t end_offset, which is the ending offset of the file; int thread_id, which is the identifier of the thread. Then, a function named detailed_match is defined to check string matching. Its parameters include: a pointer block pointing to a character, an integer pos, and a pointer keyword pointing to a keyword. Furthermore, the function uses the strncmp function to compare whether the substring of block starting from position pos is the same as keyword. If they are the same, 0 is returned; otherwise, a non-zero value is returned.

[0035] Then, a function named init_keyword_first_chars is defined to initialize the variable keyword_first_chars, which is used to store the mask of the first character of the keyword. It first uses the _mm512_setzero_si512() function to initialize keyword_first_chars to all zeros; then it traverses the blacklist array, and for each non-empty string, checks whether its length is greater than 0. If the length is greater than 0, it uses the _mm512_mask_storeu_epi8 function to store the first character of the string in keyword_first_chars. Further, using the mask operation, only the lower 6 bits of the character (& 0x3F) are stored and stored in the corresponding position of keyword_first_chars; finally, the initialized keyword_first_chars is returned.

[0036] It should be noted that the present invention adopts a concurrent processing mechanism, which defines a thread array and a thread data array, as well as a block_size for calculating the log block size to be processed by each thread, wherein NUM_THREADS is a predefined constant indicating the number of threads to be created, so that each thread will process a subset of the file, thereby achieving parallel processing.

[0037] In some embodiments, when the main program of the present invention is started, the main function will parse the command line parameters at the beginning to obtain the log file path, and then try to open the file and calculate the file size, and then start to execute log segmentation and thread startup in sequence.

[0038] Specifically, it first checks whether the number of command line parameters is not equal to 2 (argc != 2). If so, it is considered that the number of parameters is incorrect and the program will output the usage (Usage: %s<log_file> ), where the symbol %s is replaced by the name of the program,<log_file> It is used to prompt the user to provide a path for a log file. The program exits with a failure status (exit(EXIT_FAILURE)). Then the log file is opened. The fopen function is used to try to open the file specified in the command line parameter (argv[1]) in read-only binary mode ("rb"). If the file opening fails (file is NULL), the perror function is used to output an error message ("Failed to open file") and the program exits with a failure status. Then the file size is obtained. The fseek function is used to move the file pointer to the end of the file (SEEK_END). The ftell function is used to obtain the current file pointer position, that is, the file size, and store it in the log_data_size variable. The fseek function is used again to reset the file pointer to the beginning of the file (SEEK_SET) for subsequent file reading operations.

[0039] It should be noted that this paper adopts streaming processing, which effectively reduces the memory usage. It determines the size of the log block to be processed by each thread through the block_size variable, so that the file can be evenly distributed to all threads for processing.

[0040] It should be noted that the present invention adopts a concurrent processing mechanism, which assigns tasks to each thread and creates threads through a loop. In this loop, the file pointer, start and end offsets and thread ID are initialized for each thread, and then the pthread_create function is used to create the thread, so that each thread will execute the process_block function, which is responsible for processing the log block assigned to the thread.

[0041] In some embodiments, in the thread creation and task allocation of the present invention, first, the task scope of each thread is calculated, and the file offset range that each thread should process is calculated according to the file size and the number of threads. Then, the thread is created, and an independent file pointer is created for each thread using pthread_create, and the corresponding task scope is passed.

[0042] Specifically, first, declare an array threads of type pthread_t to store thread identifiers, and the array size is NUM_THREADS, which is a predefined constant indicating the number of threads to be created. Also, declare an array thread_data of type thread_data_t to store the data of each thread, and the array size is also NUM_THREADS. Then, calculate the log block size block_size to be processed by each thread, which is obtained by dividing the total log file size log_data_size by the number of threads NUM_THREADS and rounding up. Then, use a for loop to initialize the data of each thread and create the thread, which includes the steps: S301, in the loop, the data structure thread_data[i] of each thread is initialized, including the file pointer file, the starting offset start_offset, the ending offset end_offset and the thread ID thread_id; S302, the file pointer is opened through the fopen function, and the file specified in the command line parameter is opened in read-only binary mode ("rb"); if the file opening fails, an error message is output and the program exits; S303, calculating the start and end offsets of each thread to ensure that each thread processes a continuous block of the log file; S304. Use the pthread_create function to create a thread. The function executed by the thread is process_block. The parameter passed to the pthread_create function is the address of thread_data[i].

[0043] It should be noted that the SIMD (Single Instruction Multiple Data) instruction set of the present invention realizes vectorized processing of log data, and realizes fine-grained parallelization by using SIMD instructions inside each thread, which initializes a variable keyword_first_chars of type __m512i to store the mask of the first character of the keyword, uses the functions _mm512_setzero_si512() and _mm512_mask_storeu_epi8() in the Intel SIMD instruction set, creates a 512-bit vector of all zeros through _mm512_setzero_si512(), and stores the data in the 512-bit vector through _mm512_mask_storeu_epi8().

[0044] It should be noted that this method uses streaming processing to effectively reduce memory usage. It gradually reads and processes each block of the file through a loop. In this loop, the fread function is used to read data from the file into the buffer; BUFFER_SIZE bytes are read each time, and then the process_buffer function is called for processing; if the number of bytes read is less than BUFFER_SIZE, it means that the end of the file or the end of the block has been reached, and the loop is jumped out, thereby effectively avoiding reading the entire file at one time and realizing the gradual processing of each part of the file.

[0045] It should be noted that the present invention adopts a concurrent processing mechanism, wherein the process_block function is the entry point of the thread, responsible for processing a specific part of the file, and each thread will process its own log block, thereby realizing parallel processing.

[0046] In some embodiments, the thread processing logic of the present invention is: streaming reading and processing log blocks, in each thread, using fread to read the file content step by step, each time reading a data block of BUFFER_SIZE size, and then performing SIMD optimized keyword matching processing on each data block. The processing steps include: S410, initializing the SIMD vector to the first character of the keyword; S420, using SIMD instructions to quickly filter possible keyword starting positions; S430, performing detailed string matching on the starting position of the keyword to confirm whether the keyword is actually matched; S440, handle the situation where the last block is less than BLOCK_SIZE, and ensure that all data are processed correctly.

[0047] Finally, the matching results are output. If a match is found, relevant information is output, including the thread ID, the matching keyword and its position in the file.

[0048] Specifically, the function process_block accepts a void* type parameter arg, which is used to pass the data required for the thread to work; converts arg to a pointer data of thread_data_t* type to obtain the thread data structure; obtains the starting address block and block size block_size of the log block from the data structure; initializes a variable keyword_first_chars of type __m512i to store the mask of the first character of the keyword; traverses the keywords array, and if the keyword is not empty, uses the _mm512_mask_storeu_epi8 function to store the first character of the keyword in keyword_first_chars; tries to open a file named "large_log_file.log", and if it fails, outputs an error message and returns NULL; defines a character array buffer with a size of BUFFER_SIZE to store data read from the file; defines a variable bytes_read of type size_t to store the number of bytes read each time; calculates the file offset offset that the current thread should process; uses the fseek function to move the file pointer to the calculated offset position; enters a while Loop, use fread function to read data from the file into buffer, read BUFFER_SIZE bytes each time; call process_buffer function to process the read data, passing buffer, bytes_read, keyword_first_chars and data->thread_id as parameters. If the number of bytes read is less than BUFFER_SIZE, it means that the end of the file or the end of the block has been reached, and the loop is jumped out; close the file; the function returns NULL, indicating that the thread execution has ended.

[0049] Furthermore, the function process_buffer accepts the following parameters: const char *buffer, which is a pointer to the character buffer to be processed; size_t buffer_size, which is the size of the buffer; __m512i keyword_first_chars, which is a 512-bit vector that stores the information of the first character of the keyword for fast comparison; intthread_id, which is the ID of the current thread. Then, the offset of the last block, last_block_offset, is calculated, which is the remainder of the buffer size divided by the block size BLOCK_SIZE; the actual buffer size process_size to be processed is calculated, which is the buffer size minus the offset of the last block. Then, use a for loop to traverse the buffer, processing BLOCK_SIZE bytes each time. Specifically, use the _mm512_loadu_si512 function to load BLOCK_SIZE bytes into the __m512i type variable log_chunk; use the _mm512_cmpeq_epi8_mask function to compare log_chunk and keyword_first_chars, and store the result in __mmask64 result. Then, determine whether result is 0. If result is not 0, it means that the first character of the matching keyword has been found. Specifically, use a for loop to traverse BLOCK_SIZE again to find the specific matching position j, where if the jth bit of result is 1, it means that the first character of the keyword is found at position j; traverse the keywords array, for each non-empty keyword *kw, use the detailed_match function to check whether the buffer substring starting from position i + j matches the keyword *kw. If it matches, print a message including the thread ID, the found keyword, and the matching position.

[0050] It should be noted that the SIMD (Single Instruction Multiple Data) instruction set of the present invention realizes vectorized processing of log data, and uses SIMD instructions to realize fine-grained parallelization within each thread. It uses the functions mm512_loadu_si512() and _mm512_cmpeq_epi8_mask() in the Intel SIMD instruction set, loads 512-bit data into the vector through _mm512_loadu_si512(), compares two 512-bit vectors through _mm512_cmpeq_epi8_mask(), and generates a mask to represent the comparison result. The above operations are all performed within each thread, realizing fine-grained parallelization.

[0051] In some embodiments, the present invention executes 1 to N threads, which execute each thread block to be processed. Specifically, first, a for loop is used to initialize the thread and thread data, where the loop variable i starts from 0 until NUM_THREADS (the number of threads). Then, inside the loop, the data structure thread_data[i] of each thread is initialized, where the fopen function is used to open the file specified in the command line parameter in read-only binary mode ("rb"), where argv[1] is defined as the file path; if the file opening fails, the perror function is used to output an error message and exit the program with a failure status. Then, the start and end offsets of each thread are set, where start_offset is the starting position of the log block processed by the current thread, calculated as i * block_size; end_offset is the end position of the log block processed by the current thread, if the current thread is the last thread, it is processed to the end of the file, otherwise it is processed to the character before the start position of the next block. Then the ID of each thread is set, where thread_id is set to the index i of the current loop. Then create the thread, using the pthread_create function to create a new thread, which will execute the process_block function; pass &threads[i] as the thread identifier, NULL as the thread attribute, process_block as the thread function, and &thread_data[i] as the parameter of the thread function; if the thread creation fails, output an error message and exit the program.

[0052] In some embodiments, when cleaning and releasing resources of the present invention, it is necessary to clean up all threads and pointers, and close the log file. Specifically, first use a for loop to traverse all threads, and the variable i starts from 0 until NUM_THREADS (the number of threads). Then inside the loop, use the pthread_join function to wait for each thread to end, wherein the pthread_join(threads[i], NULL) function is used to wait for the thread numbered i to end. If the thread ends successfully, the function returns 0, and if an error occurs, it returns a non-zero value; if pthread_join returns a non-zero value, it means that an error occurred while waiting for the thread. At this time, the fprintf function is used to output the error information to the standard error output stderr, and strerror(errno) is used to obtain the error description; its output format is "Error joining thread%d: %s\n", wherein %d will be replaced by the thread number i, and %s will be replaced by the error description; the program exits with a failed status (exit(EXIT_FAILURE)). Then, after the loop ends, use fclose(file) to close the previously opened file; the program ends normally and returns 0 (return 0).

[0053] It should be appreciated that the method steps in the embodiments of the present invention can be implemented or implemented by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level process or object-oriented programming language to communicate with a computer system. However, if necessary, the program can be implemented in an assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, the program can be run on a programmed ASIC for this purpose.

[0054] In addition, the operations of the processes described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The processes described herein (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions that may be executed by one or more processors.

[0055] Further, the method can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, an RSM, a ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the process described herein. In addition, the machine-readable code, or portions thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention can also include the computer itself.

[0056] The computer program can be applied to input data to perform the functions described herein, thereby converting the input data to generate output data stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents physical and tangible objects, including specific visual depictions of physical and tangible objects produced on the display.

[0057] The above is only a preferred embodiment of the present invention. The present invention is not limited to the above implementation. As long as the technical effect of the present invention is achieved by the same means, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the scope of protection of the present invention. Within the scope of protection of the present invention, its technical scheme and / or implementation method may have various modifications and changes.

Claims

1. A method for detecting a log blacklist keyword, characterized in that: The method comprises the following steps: S100, initializing data and determining the log block size to be processed by each thread; S200, the main program starts, obtains the log file path and opens the file, calculates the file size, and performs log segmentation and thread startup in sequence; S300, creating threads and task allocation, calculating the task range of each thread and the file offset range that each thread needs to process, creating an independent file pointer for each thread and passing the corresponding task range; S400, streaming reading and processing of log blocks, in each thread, gradually reading the file content and then performing SIMD optimized keyword matching processing on each data block; S500, executing 1 to N log blocks through threads; S600, clean up and release memory.

2. The method according to claim 1, characterized in that The step S100 includes: S110, load the required library and header files; S120, defining constants and global variables, including setting BLOCK_SIZE for calculating the size of the log block to be processed by each thread and a predefined constant NUM_THREADS for indicating the number of threads to be created; and defining a function of init_keyword_first_chars for initializing a variable keyword_first_chars, wherein the variable keyword_first_chars is used to store a mask of the first character of the keyword; S130, defining a blacklist keyword list; S140, define auxiliary functions and structures; among them, detailed_match is used to check whether the complete keyword is matched; thread_data_t is used to store the processing result of each thread; init_keyword_first_chars is used to initialize the SIMD vector to make it the first character of the keyword.

3. The method according to claim 1, characterized in that The step S200 includes: S210, parsing and checking command line parameters to obtain a log file path; if the number of parameters is incorrect, outputting usage instructions and exiting the program in a failed state; S220, opening the log file, and attempting to open the file specified in the command line parameter in read-only binary mode; wherein, if the file opening fails, outputting an error message, and exiting the program in a failure state; S230, obtaining the file size; wherein, the file pointer is moved to the end of the file, the position of the current file pointer is obtained, and is stored in the log_data_size variable, and the file pointer is reset to the beginning of the file again for subsequent file reading operations.

4. The method according to claim 1, characterized in that: The step S300 includes: S310, calculating the task scope of each thread, and calculating the file offset range that each thread should process according to the file size and the number of threads; S320, create threads, create an independent file pointer for each thread, and pass the corresponding task scope.

5. The method according to claim 4, characterized in that In step S300, a for loop is used to initialize the data of each thread and create a thread, which includes the steps of: S301, in the loop, the data structure thread_data[i] of each thread is initialized; S302, opening the file pointer, and opening the file specified in the command line parameter in read-only binary mode; wherein, if the file opening fails, outputting an error message and exiting the program; S303, calculating the start and end offsets of each thread so that each thread processes a continuous block of the log file; S304. Use the pthread_create function to create a thread. The function executed by the thread is process_block. The parameter passed to the pthread_create function is the address of thread_data[i].

6. The method according to claim 2, characterized in that The step S400 includes: S410, initializing the SIMD vector to the first character of the keyword; S420, using SIMD instructions to quickly filter possible keyword starting positions; S430, performing detailed string matching on the starting position of the keyword to confirm whether the keyword is actually matched; S440, handle the situation where the last block is less than BLOCK_SIZE, and ensure that all data are processed correctly.

7. The method according to claim 2, characterized in that The step S500 includes: S510, using a for loop to initialize threads and thread data, wherein the loop variable i starts from 0 and goes up to NUM_THREADS; S520, inside the loop, the data structure thread_data[i] of each thread is initialized; S530, setting the start and end offsets of each thread; S540, setting the ID of each thread; wherein thread_id is set to the index i of the current loop; S550, then create a thread, wherein a pthread_create function is used to create a new thread, and the thread will execute the process_block function; &threads[i] is passed as the thread identifier, NULL as the thread attribute, process_block as the thread function, and &thread_data[i] as the parameter of the thread function.

8. The method according to claim 2, characterized in that: The step S600 includes: S610, use a for loop to traverse all threads, variable i starts from 0 until NUM_THREADS; S620, inside the loop, waiting for each thread to end; if the thread ends successfully, returning 0; if an error occurs, returning a non-zero value, indicating an error occurred while waiting for the thread, obtaining an error description, and exiting the program with a failed status; S630: After the loop ends, close the previously opened file.

9. A computer-readable storage medium, characterized in that: Program instructions are stored thereon, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.

10. A log blacklist keyword detection system, characterized in that: include: A computer device comprising a computer readable storage medium according to claim 9.

Citation Information

Patent Citations

  • Method and equipment for recognizing messages under mass flow

    CN112558948A

  • Instruction and Logic for Vector Permute

    US20170177357A1