A Detection Method and System for Log Blacklist Keywords

Through concurrent processing and streaming processing technology, and using SIMD instruction sets to perform parallel processing within each thread, the problem of inefficient processing of large-scale log files is solved, and efficient log analysis and extraction is achieved.

CN119961235BActive Publication Date: 2025-06-20POWERLEADER COMPUTER SYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510438274.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-06-20
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently process and analyze large-scale log files, resulting in high system resource utilization and low processing efficiency, which cannot meet the needs of fast and accurate analysis of log data.

Method used

Concurrency processing technology and streaming processing methods are adopted to process log files in parallel by multi-threading or multi-processing, and fine-grained parallel processing is performed using SIMD instruction set within each thread to realize vectorized processing of log data.

Benefits of technology

It significantly improves the processing efficiency of large-scale log files, reduces the occupation of system resources, and realizes efficient extraction and analysis of server logs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961235B_ABST
    Figure CN119961235B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for detecting log blacklist keywords. The method includes: initializing data to determine the size of the log block to be processed by each thread; starting the main program, obtaining the log file path and opening the file, calculating the file size, and sequentially performing log chunking and thread startup; creating threads and task allocation, calculating the task range for each thread and the file offset range to be processed by each thread, creating an independent file pointer for each thread and passing the corresponding task range; streaming reading and processing log chunks, in each thread, gradually reading the file content and then performing SIMD-optimized keyword matching processing on each data block; executing 1 to N log chunks through threads; cleaning up and releasing memory. The present invention provides an analysis method for server log files that enables efficient keyword search, streaming processing, and fine-grained parallelization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for detecting log blacklist keywords, and belongs to the field of communications. Background Art

[0002] As the core device of the network, the server undertakes key tasks such as data storage, processing, and transmission. During the daily operation of the server, a large amount of log information is generated. These log information record important information such as the running state of the server, user operations, and system events, and are of extremely important value for monitoring server performance, troubleshooting, ensuring network security, and optimizing system configuration.

[0003] However, with the continuous increase in the amount of data processed by the server and the improvement of business complexity, the scale of log files also shows an explosive growth trend. In practical applications, even the log files generated by small servers may reach several megabytes or even dozens of megabytes in size, while the log files generated by large servers or data centers may reach several gigabytes or even several terabytes in scale.

[0004] Currently, for small-scale log files, they can generally be opened and viewed through simple text editors or log viewing tools, and retrieved by manually entering keywords or blacklists to obtain the required information. However, this method is inadequate when dealing with large-scale log files. When attempting to open large log files, due to the overly large file content, the memory and processing power of the computer often cannot meet the requirements, resulting in slow computer operation or even freezing, greatly reducing work efficiency. In addition, the manual retrieval method is not only time-consuming and laborious, but also prone to missing important information and cannot meet the requirements for rapid and accurate analysis of log data. Summary of the Invention

[0005] The present invention provides a method and system for detecting log blacklist keywords, aiming to solve at least one of the technical problems existing in the prior art.

[0006] The technical solution of the present invention relates to a method for detecting log blacklist keywords. The method according to the present invention includes the following steps:

[0007] S100. Initialize data and determine the size of the log block to be processed by each thread;

[0008] S200. Start the main program, obtain the log file path and open the file, calculate the file size, and sequentially execute log chunking and thread startup;

[0009] S300. Create threads and task allocation, calculate the task range of each thread and the file offset range to be processed by each thread, create an independent file pointer for each thread and pass the corresponding task range;

[0010] S400 reads and processes log blocks in a streaming manner. In each thread, it gradually reads the file content and then performs SIMD-optimized keyword matching processing on each data block;

[0011] S500 executes 1 to N log blocks through threads;

[0012] S600 clears and releases memory.

[0013] Furthermore, the step S100 includes:

[0014] S110 loads the required libraries and header files;

[0015] S120 defines constants and global variables, including setting BLOCK_SIZE for calculating the size of log blocks to be processed by each thread and a predefined constant NUM_THREADS for representing the number of threads to be created; and defines a function init_keyword_first_chars for initializing the variable keyword_first_chars, which is used to store the mask of the first character of keywords.

[0016] S130 defines a list of blacklist keywords;

[0017] S140 defines auxiliary functions and structures; among them, detailed_match is used to check if a complete keyword is matched; thread_data_t is used to store the processing results of each thread; init_keyword_first_chars is used to initialize the SIMD vector to be the first character of keywords.

[0018] Furthermore, the step S200 includes:

[0019] S210 parses and checks command-line arguments to obtain the log file path; among them, if the number of arguments is incorrect, it outputs the usage method and exits the program with a failure status;

[0020] S220 opens the log file, attempts to open the file specified in the command-line arguments in read-only binary mode; among them, if the file opening fails, it outputs an error message and exits the program with a failure status;

[0021] S230 obtains the file size; among them, it moves the file pointer to the end of the file, obtains the current position of the file pointer, and stores it in the log_data_size variable, and then resets the file pointer to the beginning of the file for subsequent file reading operations.

[0022] Furthermore, the step S300 includes:

[0023] S310. Calculate the task range of each thread. According to the file size and the number of threads, calculate the file offset range that each thread should process;

[0024] S320. Create threads. Create an independent file pointer for each thread and pass the corresponding task range.

[0025] Furthermore, in the step S300, a for loop is used to initialize the data of each thread and create threads, which includes the steps of:

[0026] S301. In the loop, the data structure thread_data[i] of each thread is initialized;

[0027] S302. Open the file pointer. Open the file specified in the command line argument in read-only binary mode; wherein, if the file opening fails, output an error message and exit the program;

[0028] S303. Calculate the start and end offsets of each thread so that each thread processes a continuous block of the log file;

[0029] S304. Use the pthread_create function to create threads. The function executed by the threads is process_block, and the parameter passed to the pthread_create function is the address of thread_data[i].

[0030] Furthermore, the step S400 includes:

[0031] S410. Initialize the SIMD vector to the first character of the keyword;

[0032] S420. Use SIMD instructions to quickly filter the possible starting positions of keywords;

[0033] S430. Perform a detailed string match on the starting positions of the keywords to confirm whether the keywords are really matched;

[0034] S440. Process the case where the last block is less than BLOCK_SIZE to ensure that all data is correctly processed.

[0035] Furthermore, the step S500 includes:

[0036] S510. Use a for loop to initialize the threads and thread data, where the loop variable i starts from 0 until NUM_THREADS;

[0037] S520. Inside the loop, the data structure thread_data[i] of each thread is initialized;

[0038] S530. Set the start and end offsets of each thread;

[0039] S540. Set the ID of each thread; where thread_id is set to the index i of the current loop;

[0040] S550. Then create threads, where a new thread is created using the pthread_create function, and the thread will execute the process_block function; pass &threads[i] as the thread identifier, NULL as the thread attribute, process_block as the thread function, and &thread_data[i] as the parameter of the thread function.

[0041] Further, the step S600 includes:

[0042] S610. Use a for loop to iterate through all threads, with the variable i starting from 0 until NUM_THREADS;

[0043] S620. Inside the loop, wait for each thread to end; where if the thread ends successfully, return 0; if an error occurs, return a non-zero value indicating an error when waiting for the thread, obtain the error description, and exit the program with a failure status;

[0044] S630. After the loop ends, close the previously opened file.

[0045] The technical solution of the present invention also relates to a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the above-mentioned method is implemented.

[0046] The technical solution of the present invention also relates to a detection system for log blacklist keywords, and the system includes a computer device, and the computer device includes the above-mentioned computer-readable storage medium.

[0047] The beneficial effects of the present invention are as follows:

[0048] The detection method and system for log blacklist keywords of the present invention significantly improve the processing efficiency of log files, reduce the occupation of system resources, and achieve efficient extraction and analysis of server logs through efficient keyword search, concurrent processing, streaming processing, and fine-grained parallel processing.

[0049] By introducing the SIMD instruction set, the present invention realizes the vectorized processing of log data, enabling the keyword search process to be executed in parallel on multiple data blocks. The concurrent processing technology is adopted to perform parallel processing on the log file in the form of multi-threading or multi-processing. The log file is divided into multiple sub-tasks and assigned to multiple threads or processes for simultaneous processing. Considering that the log file may be extremely large, it is unrealistic to load the entire file into memory at once. Therefore, the present invention adopts a streaming processing method, reading the log file block by block, which reduces the memory occupancy. In addition to thread-level parallelization, the present invention also realizes fine-grained parallelization within each thread by using the SIMD instruction set, further tapping the potential of parallel processing and enabling each thread to more efficiently utilize CPU resources when processing log data. Brief Description of the Drawings

[0050] Figure 1 is the basic flowchart of the method according to the present invention. Detailed Embodiment

[0051] The concept, specific structure and technical effects of the present invention will be clearly and completely described below in conjunction with the embodiments and the drawings, so as to fully understand the purpose, solution and effects of the present invention.

[0052] It should be noted that unless otherwise specified, when a feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. The singular forms "a", "the" and "said" used herein are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this technology belongs. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0053] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language (such as "for example", "such as", etc.) provided herein is only intended to better illustrate the embodiments of the present invention, and unless otherwise required, will not impose a limitation on the scope of the present invention.

[0054] Refer to Figure 1, in some embodiments, the method for detecting log blacklist keywords according to the present invention at least includes the following steps:

[0055] S100. Initialize data and define configurations.

[0056] S200. Start the main program to open the log file to be read.

[0057] S300. After creating threads and allocating tasks, start executing the threads.

[0058] S400. Read and process log blocks in a streaming manner.

[0059] S500. Execute 1 to N process_log_blocks through the threads.

[0060] S600. After waiting for all threads to end, clean up and release the memory.

[0061] Aiming at the lack of analysis of large log files in the current technology, the inability to efficiently extract keywords and solve practical scenarios, the present invention proposes an efficient and low-resource-consuming log processing method. By efficiently searching for keywords, concurrent processing, streaming processing, and fine-grained parallel processing, the processing efficiency of large-scale log files is significantly improved, the occupation of system resources is reduced, and the efficient extraction and analysis of server logs are realized.

[0062] The present invention realizes the vectorized processing of log data by introducing the SIMD (Single Instruction Multiple Data) instruction set. In traditional log analysis methods, keyword search usually adopts the sequential scanning method byte by byte or character by character, which is inefficient in processing large-scale log files and has a low utilization rate of CPU resources. In contrast, the SIMD instruction set adopted by the present invention can process multiple data units simultaneously in one instruction cycle, thus significantly improving the data processing efficiency.

[0063] The present invention adopts concurrent processing technology to parallel process log files through multi-threading or multi-processing. Compared with the traditional single-threaded processing method, which is often limited by the processing power of a single CPU core, the present invention makes full use of the computing power of multi-core CPUs. Also, considering that log files may be very large, it is unrealistic to load the entire file into memory at once. The present invention adopts streaming processing technology to reduce the memory occupation by reading log files block by block.

[0064] Considering that log files can be extremely large, it is unrealistic to load the entire file into memory at once. The present invention adopts a streaming processing method. By reading the log file block by block, the memory occupancy is reduced. The core idea of streaming processing is to split the log file into multiple small blocks, process only one data block at a time, release the memory occupied by the data block after processing, and then read the next data block for continuous processing. The streaming processing method avoids the problem of insufficient memory caused by loading the entire file at once, enabling the present invention to efficiently process ultra-large-scale log files.

[0065] In addition to thread-level parallelization, the present invention also implements fine-grained parallelization within each thread using the SIMD instruction set. In traditional parallel processing, thread-level parallelization mainly relies on the number of cores of a multi-core CPU. However, the fine-grained parallelization of the present invention further improves the data processing efficiency by performing data parallel processing within each thread. For example, within a single thread, multiple data units can be processed simultaneously through the SIMD instruction set instead of the traditional one-by-one processing method.

[0066] In some embodiments, during the initialization and preparation stage of the present invention, necessary libraries and header files are first loaded, including C language libraries such as <immintrin.h>, <stdio.h>, <string.h>, <pthread.h>, etc. Then, constants and global variables are defined. Specifically, BLOCK_SIZE (aligned to 64 bytes), BUFFER_SIZE (e.g., 64KB), and NUM_THREADS (number of threads) are set. Then, a list of blacklist keywords keywords is defined, such as "error", "failure", "timeout", etc. Then, auxiliary functions and structures are defined. Specifically, detailed_match is used to check for a complete keyword match; thread_data_t is used to store the processing results of each thread; init_keyword_first_chars is used to initialize the SIMD vector to be the first character of the keyword.

[0067] Specifically, first, a string array named blacklist is defined, which contains three strings: "error", "failure", and "timeout". Then, a structure thread_data_t is defined to store thread-related data. Among them, the structure thread_data_t contains the following members: FILE *file, which is a pointer to a file and may be used for file operations; off_t start_offset, which is the starting offset of the file; off_t end_offset, which is the ending offset of the file; int thread_id, which is the identifier of the thread. Then, a function named detailed_match is defined to check string matching. Its parameters include: a pointer to a character block, an integer pos, and a pointer to a keyword. Further, the function uses the strncmp function to compare the substring of block starting from position pos with keyword. If they are the same, it returns 0; otherwise, it returns a non-zero value.

[0068] Then, a function named init_keyword_first_chars is defined to initialize the variable keyword_first_chars, which is used to store the mask of the first character of the keyword. It first initializes keyword_first_chars to all zeros using the _mm512_setzero_si512() function; then traverses the blacklist array. For each non-empty string, it checks if its length is greater than 0. If the length is greater than 0, it uses the _mm512_mask_storeu_epi8 function to store the first character of the string into keyword_first_chars. Further, using a mask operation, only the lower 6 bits of the character (& 0x3F) are stored and placed at the corresponding position in keyword_first_chars; finally, the initialized keyword_first_chars is returned.

[0069] It should be noted that the present invention adopts a concurrent processing mechanism. It defines a thread array and a thread data array, and a block_size for calculating the size of the log block to be processed by each thread. Among them, NUM_THREADS is a predefined constant representing the number of threads to be created. Thus, each thread will process a subset of the file, thereby achieving parallel processing.

[0070] In some embodiments, during the startup of the main program of the present invention, when the main function starts, it parses the command-line arguments to obtain the path of the log file, then attempts to open the file, calculates the file size, and then starts to execute log chunking and thread startup in sequence.

[0071] Specifically, first, it checks whether the number of command-line arguments is not equal to 2 (argc != 2). If so, it considers the number of arguments incorrect, and the program will output the usage method (Usage: %s <log_file>), where the symbol %s is replaced by the name of the program, and <log_file> is used to prompt the user to provide the path of a log file. The program exits with a failure status (exit(EXIT_FAILURE)). Then it opens the log file. Here, the fopen function is used to attempt to open the file specified in the command-line argument (argv[1]) in read-only binary mode ("rb"). If the file opening fails (file is NULL), the perror function is used to output the error message ("Failed to open file"), and the program also exits with a failure status. Then it obtains the file size. Here, the fseek function is used to move the file pointer to the end of the file (SEEK_END), the ftell function is used to obtain the position of the current file pointer, which is the file size, and it is stored in the log_data_size variable. The fseek function is used again to reset the file pointer to the beginning of the file (SEEK_SET) for subsequent file reading operations.

[0072] It should be noted that the present invention adopts a streaming processing method, which effectively reduces the memory occupation. It determines the size of the log chunk to be processed by each thread through the block_size variable, so that the file can be evenly distributed to all threads for processing.

[0073] It should be noted that the present invention adopts a concurrent processing mechanism. It assigns tasks and creates threads for each thread through a loop. In this loop, the file pointer, start and end offsets, and thread ID are initialized for each thread, and then the pthread_create function is used to create threads. Thus, each thread will execute the process_block function, which is responsible for processing the log chunk assigned to the thread.

[0074] In some embodiments, during the thread creation and task assignment of the present invention, first, it calculates the task range for each thread. According to the file size and the number of threads, it calculates the file offset range that each thread should process. Then, it creates threads. The pthread_create function is used to create an independent file pointer for each thread and pass the corresponding task range.

[0075] Specifically, first, an array threads of type pthread_t is declared to store the identifiers of the threads. The size of the array is NUM_THREADS, which is a predefined constant representing the number of threads to be created. Also, an array thread_data of type thread_data_t is declared to store the data for each thread, and the size of the array is also NUM_THREADS. Then, the size of the log block to be processed by each thread, block_size, is calculated by dividing the total log file size log_data_size by the number of threads NUM_THREADS and rounding up. Then, a for loop is used to initialize the data for each thread and create the threads, which includes the steps:

[0076] S301. In the loop, the data structure thread_data[i] for each thread is initialized, which includes the file pointer file, the start offset start_offset, the end offset end_offset, and the thread ID thread_id;

[0077] S302. The file pointer is opened using the fopen function to open the file specified in the command-line arguments in read-only binary mode ("rb"); if the file opening fails, an error message is output and the program exits;

[0078] S303. Calculate the start and end offsets for each thread to ensure that each thread processes a continuous block of the log file;

[0079] S304. Use the pthread_create function to create a thread. The function that the thread executes is process_block, and the argument passed to the pthread_create function is the address of thread_data[i].

[0080] It should be noted that the SIMD (Single Instruction Multiple Data) instruction set of the present invention realizes the vectorized processing of log data, and fine-grained parallelization is achieved within each thread using SIMD instructions. It initializes a variable keyword_first_chars of type __m512i to store the mask of the first character of the keyword, and uses the functions _mm512_setzero_si512() and _mm512_mask_storeu_epi8() in the Intel SIMD instruction set. A 512-bit vector of all zeros is created through _mm512_setzero_si512(), and data is stored into the 512-bit vector through _mm512_mask_storeu_epi8().

[0081] It should be noted that the present invention adopts a streaming processing method, which effectively reduces the memory occupation. It gradually reads and processes each block of the file through a loop. In this loop, the fread function is used to read data from the file into the buffer; each time BUFFER_SIZE bytes are read, and then the process_buffer function is called for processing; if the number of bytes read is less than BUFFER_SIZE, it means that the end of the file or the end of the block has been reached, and the loop is exited, thus effectively avoiding reading the entire file at once and realizing the step-by-step processing of each part of the file.

[0082] It should be noted that the present invention adopts a concurrent processing mechanism. Its process_block function is the entry point of the thread and is responsible for processing a specific part of the file. Each thread will process its own log block, thus realizing parallel processing.

[0083] In some embodiments, the thread processing logic of the present invention is as follows: streamingly read and process log blocks. In each thread, the fread function is used to gradually read the file content, and each time a data block of size BUFFER_SIZE is read, and then SIMD-optimized keyword matching processing is performed on each data block. Its processing steps include:

[0084] S410. Initialize the SIMD vector as the first character of the keyword;

[0085] S420. Use SIMD instructions to quickly filter the possible starting positions of keywords;

[0086] S430. Perform detailed string matching on the starting positions of keywords to confirm whether keywords are really matched;

[0087] S440. Process the case where the last block is less than BLOCK_SIZE to ensure that all data is correctly processed.

[0088] Among them, the final output is the matching result. If a matching item is found, relevant information is output, including the thread ID, the matching keyword, and its position in the file.

[0089] Specifically, the function process_block accepts a parameter arg of type void*, which is used to pass the data required for thread work; convert arg to a pointer data of type thread_data_t* to obtain the thread data structure; obtain the starting address block and block size block_size of the log block from the data structure; initialize a variable keyword_first_chars of type __m512i to store the mask of the first character of the keyword; traverse the keywords array, and if the keyword is not empty, use the _mm512_mask_storeu_epi8 function to store the first character of the keyword into keyword_first_chars; attempt to open the file named "large_log_file.log", and if it fails, output an error message and return NULL; define a character array buffer with a size of BUFFER_SIZE to store the data read from the file; define a variable bytes_read of type size_t to store the number of bytes read each time; calculate the file offset that the current thread should process; use the fseek function to move the file pointer to the calculated offset position; enter a while loop, use the fread function to read data from the file into the buffer, reading BUFFER_SIZE bytes each time; call the process_buffer function to process the read data, passing buffer, bytes_read, keyword_first_chars, and data->thread_id as parameters. If the number of bytes read is less than BUFFER_SIZE, it means the end of the file or the end of the block has been reached, and the loop is exited; close the file; the function returns NULL, indicating the end of thread execution.

[0090] Further, the function process_buffer accepts the following parameters: const char *buffer, which is a pointer to the character buffer to be processed; size_t buffer_size, which is the size of the buffer; __m512i keyword_first_chars, which is a 512-bit vector storing information about the first characters of keywords for fast comparison; int thread_id, which is the ID of the current thread. Then, calculate the offset of the last block last_block_offset, which is the remainder of the buffer size divided by the block size BLOCK_SIZE; calculate the actual buffer size to be processed process_size, which is the buffer size minus the offset of the last block. Then, use a for loop to iterate through the buffer, processing BLOCK_SIZE bytes each time. Specifically, use the _mm512_loadu_si512 function to load BLOCK_SIZE bytes into a variable log_chunk of type __m512i; use the _mm512_cmpeq_epi8_mask function to compare log_chunk and keyword_first_chars, and store the result in __mmask64 result. Then, check if result is 0. If result is not 0, it means the first character of a matching keyword has been found. Specifically, use another for loop to iterate through BLOCK_SIZE to find the specific matching position j. If the j-th bit of result is 1, it means the first character of the keyword has been found at position j. Iterate through the keywords array. For each non-empty keyword *kw, use the detailed_match function to check if the substring of the buffer starting from position i + j matches the keyword *kw. If there is a match, print a message including the thread ID, the found keyword, and the matching position.

[0091] It should be noted that the SIMD (Single Instruction Multiple Data) instruction set of the present invention realizes the vectorized processing of log data, and fine-grained parallelization is achieved within each thread by using SIMD instructions. It uses the functions mm512_loadu_si512() and _mm512_cmpeq_epi8_mask() in the Intel SIMD instruction set. 512-bit data is loaded into the vector through _mm512_loadu_si512(), and two 512-bit vectors are compared through _mm512_cmpeq_epi8_mask() to generate a mask representing the comparison result. The above operations are all carried out within each thread, realizing fine-grained parallelization.

[0092] In some embodiments, the present invention executes 1 to N threads, which execute each thread block to be processed. Specifically, first, a for loop is used to initialize the threads and thread data, where the loop variable i starts from 0 until NUM_THREADS (the number of threads). Then, inside the loop, the data structure thread_data[i] of each thread is initialized. Among them, the fopen function is used to open the file specified in the command-line argument in read-only binary mode ("rb"), where argv[1] is defined as the file path; if the file opening fails, the perror function is used to output an error message and the program exits with a failure status. Then, the start and end offsets of each thread are set. Among them, start_offset is the starting position of the log block processed by the current thread, and the calculation method is i * block_size; end_offset is the ending position of the log block processed by the current thread. If the current thread is the last thread, it processes to the end of the file, otherwise it processes to the character before the start of the next block. Then the ID of each thread is set, where thread_id is set to the index i of the current loop. Then threads are created. Among them, the pthread_create function is used to create a new thread, and the thread will execute the process_block function; &threads[i] is passed as the thread identifier, NULL as the thread attribute, process_block as the thread function, and &thread_data[i] as the parameter of the thread function; if the thread creation fails, an error message is output and the program exits.

[0093] In some embodiments, when cleaning up and releasing resources of the present invention, all threads and pointers need to be cleaned up, and the log file needs to be closed. Specifically, first, use a for loop to iterate through all threads, with the variable i starting from 0 until NUM_THREADS (the number of threads). Then, inside the loop, use the pthread_join function to wait for each thread to end. Among them, the pthread_join(threads[i], NULL) function is used to wait for the thread numbered i to end. If the thread ends successfully, this function returns 0; if an error occurs, it returns a non-zero value. If pthread_join returns a non-zero value, it means an error occurred while waiting for the thread. At this time, use the fprintf function to output the error message to the standard error output stderr, and use strerror(errno) to obtain the error description. Its output format is "Error joining thread %d: %s\n", where %d will be replaced by the thread number i, and %s will be replaced by the error description. The program exits with a failure status (exit(EXIT_FAILURE)). Then, after the loop ends, use fclose(file) to close the previously opened file. The program ends normally and returns 0 (return 0).

[0094] It should be recognized that the method steps in the embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or computer instructions stored in a non-transitory computer-readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if necessary, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can run on a dedicated integrated circuit programmed for this purpose.

[0095] In addition, the operations of the processes described herein can be performed in any suitable order, unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed commonly on one or more processors, by hardware, or a combination thereof. The computer program includes multiple instructions executable by one or more processors.

[0096] Further, the method may be implemented in any type of computing platform operatively connected to a suitable one, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention may be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RSM, ROM, etc., such that it is readable by a programmable computer and can be used to configure and operate the computer to perform the processes described herein when the storage medium or device is read by the computer. Additionally, the machine-readable code, or portions thereof, may be transmitted via wired or wireless networks. When such media includes instructions or programs that implement the above-described steps in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention may also include the computer itself.

[0097] A computer program can be applied to input data to perform the functions described herein, thereby transforming the input data to generate output data stored in non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the transformed data represents physical and tangible objects, including a specific visual depiction of the physical and tangible objects generated on the display.

[0098] As described above, these are only the preferred embodiments of the present invention, and the present invention is not limited to the above-described embodiments. As long as the same means are used to achieve the technical effects of the present invention, any modifications, equivalent replacements, improvements, etc., made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, its technical solutions and / or implementation manners may have various different modifications and variations.

Claims

1. A method for detecting a log blacklist keyword, characterized in that: The method comprises the following steps: S100, initializing data and determining the log block size to be processed by each thread; S200, the main program starts, obtains the log file path and opens the file, calculates the file size, and performs log segmentation and thread startup in sequence; S300, creating threads and task allocation, calculating the task range of each thread and the file offset range that each thread needs to process, creating an independent file pointer for each thread and passing the corresponding task range; S400, streaming reading and processing of log blocks, in each thread, gradually reading the file content and then performing SIMD optimized keyword matching processing on each data block; S500, executing 1 to N log blocks through threads; S600, clean up and release memory; Wherein, the step S100 includes: S110, load the required library and header files; S120, defining constants and global variables, including setting BLOCK_SIZE for calculating the size of the log block to be processed by each thread and a predefined constant NUM_THREADS for indicating the number of threads to be created; and defining a function init_keyword_first_chars for initializing a variable keyword_first_chars, wherein the variable keyword_first_chars is used to store a mask of the first character of the keyword; S130, defining a blacklist keyword list; S140, define auxiliary functions and structures; wherein detailed_match is used to check whether the complete keyword is matched; thread_data_t is used to store the processing result of each thread; init_keyword_first_chars is used to initialize the SIMD vector to make it the first character of the keyword; Wherein, the step S400 includes: S410, initializing the SIMD vector to the first character of the keyword; S420, using SIMD instructions to quickly filter possible keyword starting positions; S430, performing detailed string matching on the starting position of the keyword to confirm whether the keyword is actually matched; S440, handle the situation where the last block is less than BLOCK_SIZE, and ensure that all data are processed correctly.

2. The method according to claim 1, characterized in that: The step S200 includes: S210, parsing and checking command line parameters to obtain a log file path; if the number of parameters is incorrect, outputting usage instructions and exiting the program in a failed state; S220, opening the log file, and attempting to open the file specified in the command line parameter in read-only binary mode; wherein, if the file opening fails, outputting an error message, and exiting the program in a failure state; S230, obtaining the file size; wherein, the file pointer is moved to the end of the file, the position of the current file pointer is obtained, and is stored in the log_data_size variable, and the file pointer is reset to the beginning of the file again for subsequent file reading operations.

3. The method according to claim 1, characterized in that: The step S300 includes: S310, calculating the task scope of each thread, and calculating the file offset range that each thread should process according to the file size and the number of threads; S320, create threads, create an independent file pointer for each thread, and pass the corresponding task scope.

4. The method according to claim 3, characterized in that In step S300, a for loop is used to initialize the data of each thread and create a thread, which includes the steps of: S301, in the loop, the data structure thread_data[i] of each thread is initialized; S302, opening the file pointer, and opening the file specified in the command line parameter in read-only binary mode; wherein, if the file opening fails, outputting an error message and exiting the program; S303, calculating the start and end offsets of each thread so that each thread processes a continuous block of the log file; S304. Use the pthread_create function to create a thread. The function executed by the thread is process_block. The parameter passed to the pthread_create function is the address of thread_data[i].

5. The method according to claim 1, characterized in that The step S500 includes: S510, using a for loop to initialize threads and thread data, wherein the loop variable i starts from 0 and goes up to NUM_THREADS; S520, inside the loop, the data structure thread_data[i] of each thread is initialized; S530, setting the start and end offsets of each thread; S540, setting the ID of each thread; wherein thread_id is set to the index i of the current loop; S550, then create a thread, wherein a pthread_create function is used to create a new thread, and the thread will execute the process_block function; &threads[i] is passed as the thread identifier, NULL as the thread attribute, process_block as the thread function, and &thread_data[i] as the parameter of the thread function.

6. The method according to claim 1, characterized in that The step S600 includes: S610, use a for loop to traverse all threads, variable i starts from 0 until NUM_THREADS; S620, inside the loop, waiting for each thread to end; if the thread ends successfully, returning 0; if an error occurs, returning a non-zero value, indicating an error occurred while waiting for the thread, obtaining an error description, and exiting the program with a failed status; S630: After the loop ends, close the previously opened file.

7. A computer-readable storage medium, characterized in that: Program instructions are stored thereon, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

8. A log blacklist keyword detection system, characterized in that: include: A computer device comprising a computer readable storage medium according to claim 7.

Citation Information

Patent Citations

  • Method and equipment for recognizing messages under mass flow

    CN112558948A

  • Instruction and Logic for Vector Permute

    US20170177357A1