Log system performance parameter self-adaptive tuning method and system based on reinforcement learning driving

CN122673041APending Publication Date: 2026-09-01BEIJING YAN RONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610596525.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

当业务负载发生波动、日志类型变化或系统资源状态改变时,固定参数配置难以及时适配运行环境,容易出现日志写入阻塞、磁盘与内存资源浪费、延迟升高或日志丢失风险

Benefits of technology

[0011] Compared with existing technologies, the beneficial effects provided by this invention include: using the reinforcement learning-driven adaptive tuning method and system for log system performance parameters disclosed in this invention, the original log text stream generated by the target log system within a preset collection period is obtained and transformed into a sequence of log message units with time-sequence marking; semantic perception and parameter association parsing are performed on the log message unit sequence to obtain system performance state transition features and parameter configuration change response features; a reinforcement learning strategy generation network is used to perform parameter tuning decision deduction, generate a sequence of parameter adjustment actions, and form a tuning instruction set to be sent to the parameter control interface, thereby realizing dynamic adaptive adjustment of log system performance parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122673041A_ABST
    Figure CN122673041A_ABST
Patent Text Reader

Abstract

The application discloses a log system performance parameter self-adaptive tuning method and system based on reinforcement learning driving, comprising the following steps: firstly, obtaining original log text streams generated by a target log system in a preset collection period, and converting the original log text streams into log message unit sequences with time sequence marks; performing semantic perception and parameter association analysis on the log message unit sequences to obtain system performance state transition features and parameter configuration change response features; using a reinforcement learning strategy generation network to generate parameter tuning decision deduction, generate a parameter adjustment action sequence, and form a tuning instruction set to be sent to a parameter control interface, so that dynamic self-adaptive adjustment of log system performance parameters is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a method and system for adaptive tuning of performance parameters of a log system based on reinforcement learning. Background Technology

[0002] Log systems in cloud computing, distributed services, and large-scale business platforms are responsible for recording operational status, tracking anomalies, and performing performance analysis. As the volume of business requests, the frequency of log writes, and the scale of storage continue to grow, the performance parameters of the log system, such as buffer size, write batch size, refresh interval, compression strategy, and number of threads, will directly affect system throughput, latency, resource consumption, and stability.

[0003] Parameter tuning in existing logging systems typically relies on manual experience, fixed threshold rules, or offline performance test results. When business load fluctuates, log types change, or system resource states change, fixed parameter configurations struggle to adapt to the operating environment in a timely manner, easily leading to log write blocking, wasted disk and memory resources, increased latency, or the risk of log loss. While some automated tuning solutions can adjust based on monitoring metrics, they often lack joint analysis of the semantics of the original log text, the impact of parameter changes, and the relationship between system state transitions, making it difficult to accurately determine the impact of parameter adjustments on subsequent performance. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for adaptive tuning of performance parameters of a log system based on reinforcement learning.

[0005] In a first aspect, embodiments of the present invention provide a method for adaptive tuning of performance parameters of a log system based on reinforcement learning, the method comprising:

[0006] Acquire the raw log text stream generated by the target log system within a preset collection period, and convert the raw log text stream into a sequence of log message units with time sequence markers;

[0007] Semantic awareness and parameter association parsing are performed on the log message unit sequence to obtain the system performance state migration characteristics and parameter configuration change response characteristics of the log message unit sequence within the preset collection period.

[0008] The pre-built reinforcement learning strategy is invoked to generate a network to perform parameter tuning decision deduction on the system performance state transition characteristics and the parameter configuration change response characteristics, so as to obtain the parameter adjustment action sequence of the target log system;

[0009] Based on the parameter adjustment action sequence, a set of tuning instructions is generated, which includes the name of the target parameter to be adjusted, the direction identifier of the adjustment, and the adjustment step level. The set of tuning instructions is then sent to the parameter control interface of the target log system to trigger the dynamic parameter adjustment operation.

[0010] Secondly, embodiments of the present invention provide a log system performance parameter adaptive tuning system based on reinforcement learning, comprising at least one service node; the service node includes a storage unit and a computing unit; the storage unit is used to store program code; the computing unit is used to run the program code to execute the log system performance parameter adaptive tuning method based on reinforcement learning described in the first aspect.

[0011] Compared with existing technologies, the beneficial effects provided by this invention include: using the reinforcement learning-driven adaptive tuning method and system for log system performance parameters disclosed in this invention, the original log text stream generated by the target log system within a preset collection period is obtained and transformed into a sequence of log message units with time-sequence marking; semantic perception and parameter association parsing are performed on the log message unit sequence to obtain system performance state transition features and parameter configuration change response features; a reinforcement learning strategy generation network is used to perform parameter tuning decision deduction, generate a sequence of parameter adjustment actions, and form a tuning instruction set to be sent to the parameter control interface, thereby realizing dynamic adaptive adjustment of log system performance parameters. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A flowchart illustrating the steps of the reinforcement learning-driven adaptive tuning method for log system performance parameters provided in this embodiment of the invention.

[0014] Figure 2 This is a schematic diagram of the semantic association strength evolution trajectory and semantic coherence log segmentation provided in the embodiments of the present invention;

[0015] Figure 3 This is a schematic diagram illustrating the delay correlation between the parameter adjustment trigger time and the parameter adjustment response effect provided in an embodiment of the present invention.

[0016] Figure 4 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0018] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0019] In order to solve the technical problems mentioned in the background art Figure 1 This is a flowchart illustrating the adaptive tuning method for log system performance parameters based on reinforcement learning provided in this embodiment. The following is a detailed description of this adaptive tuning method for log system performance parameters based on reinforcement learning.

[0020] Step S201: Obtain the raw log text stream generated by the target log system within a preset collection period, and convert the raw log text stream into a log message unit sequence with time sequence markers;

[0021] Step S202: Perform semantic perception and parameter association parsing processing on the log message unit sequence to obtain the system performance state migration characteristics and parameter configuration change response characteristics corresponding to the log message unit sequence within the preset collection period;

[0022] Step S203: Invoke the pre-built reinforcement learning strategy generation network to perform parameter tuning decision deduction on the system performance state transition features and the parameter configuration change response features to obtain the parameter adjustment action sequence of the target log system;

[0023] Step S204: Generate a set of tuning instructions containing the name of the target parameter, the direction identifier, and the step size based on the parameter adjustment action sequence, and send the set of tuning instructions to the parameter control interface of the target log system to trigger the dynamic parameter adjustment operation.

[0024] In this embodiment of the invention, for example, a reinforcement learning-driven adaptive tuning method for log system performance parameters is provided, executed by a server. The server can be a tuning control server deployed in an operations and maintenance management platform, and the target log system can be a distributed log collection system, a log caching and forwarding system, or a centralized log writing system. During operation, the target log system involves multiple performance parameters such as log buffer size, batch sending count, disk flushing interval, compression level, number of log forwarding threads, and queue water level threshold. The server collects the log text stream output by the target log system, analyzes log semantics, changes in system performance status, and historical parameter adjustment response patterns, and invokes a reinforcement learning strategy to generate a network that generates parameter adjustment actions, thereby achieving adaptive tuning of the target log system's performance parameters.

[0025] Specifically, the server first acquires the raw log text stream generated by the target log system within a preset collection period, and converts the raw log text stream into a sequence of log message units with time-sequential markers. The preset collection period can be five minutes. The server deploys a log collection probe on the log output channel of the target log system, continuously capturing the raw log text stream within this period. For example, during peak business access periods, the target log system continuously outputs log content such as "queue usage 92%", "batch send delay 280ms", "flush task blocked", "compression cost rising", and "batch.size updated from 500 to 800". After receiving the above raw log text stream, the server performs log line segmentation, timestamp positioning, and time-sequential sorting to obtain a log text line queue.

[0026] After the log text line queue is formed, the server performs log line boundary detection based on the characteristic character sequence of the starting line of the log record. For log lines that begin with a timestamp, log level, and thread identifier, the server identifies them as the starting line of a log record; for exception stack lines, supplementary description lines, or consecutive text lines that do not contain an independent timestamp, the server merges them into the previous log record. Through the above processing, the server divides the log text line queue into multiple basic log fragment units. Subsequently, the server performs time interval calculation and content fragmentation processing on each basic log fragment unit, calculates the forward time interval between the current basic log fragment unit and the previous basic log fragment unit, and assigns it a sequential sequence number. For example, the server marks the basic log fragment unit corresponding to "10:15:01.125 queue usage 92%" as the 231st log message unit, marks the time interval between it and the previous log message unit as 22 milliseconds, and encapsulates the log text, timestamp, log level, service node identifier, forward time interval attribute value, and sequential sequence number together into a log message unit. After the server aggregates all log message units, it obtains a sequence of log message units with time-ordered markers. Then, the server performs semantic awareness and parameter association parsing on the log message unit sequence to obtain the system performance state transition characteristics and parameter configuration change response characteristics corresponding to the log message unit sequence within a preset collection period. The server inputs each log message unit in the log message unit sequence into a pre-trained log semantic awareness layer, which performs semantic vector encoding on individual log message units to generate a fixed-dimensional log semantic vector. In one optional implementation, the pre-trained log semantic awareness layer includes a text normalization layer, a field embedding layer, a context encoding layer, and a semantic vector output layer. The text normalization layer is used to normalize timestamps, numbers, percentages, delay units, thread identifiers, node identifiers, and exception stack keywords in the log message units. For example, it uniformly replaces timestamps of different formats with time markers, uniformly normalizes delay fields of different numerical forms into delay numerical markers, and normalizes queue levels in percentage form into queue occupancy numerical markers. The field embedding layer generates embedding vectors for the log body, log level, service node, forward time interval, and sequential number, respectively, and concatenates these embedding vectors to obtain the initial embedding representation of the log message unit. The context encoding layer can be implemented using a bidirectional recurrent neural network, a convolutional neural network, or a Transformer encoding network.

[0027] In this embodiment, the context encoding layer adopts a Transformer Encoder structure containing four encoding sub-layers. Each encoding sub-layer includes a multi-head self-attention sub-layer, a feedforward neural network sub-layer, residual connections, and layer normalization processing. The multi-head self-attention sub-layer is used to capture the semantic dependencies between different fields within a log message unit, and the feedforward neural network sub-layer is used to perform non-linear mapping on the attention output. The semantic vector output layer pools the hidden states output by the context encoding layer and outputs a fixed-dimensional log semantic vector through a fully connected mapping layer. For example, the dimension of the log semantic vector can be 128-dimensional, 256-dimensional, or 512-dimensional. The training samples of the log semantic perceptron consist of historical log message units and their corresponding semantic category labels. The semantic category labels include queue level anomaly, sending delay anomaly, disk flushing blockage, increased compression time, log forwarding timeout, parameter configuration change, recovery to stable state, and normal operating state. For log samples lacking manual annotation, weakly supervised labels can be generated based on preset keyword rules and historical fault work order records. During training, historical log message units are input into the log semantic perceptron, which outputs the semantic category prediction probability, using the cross-entropy loss function as the training objective. The cross-entropy loss function is: L_sem = -Σ_i y_i log(p_i), where y_i represents the true label of the i-th semantic category, and p_i represents the predicted probability of the i-th semantic category output by the log semantic perceptron. By minimizing the cross-entropy loss function, the log semantic perceptron learns the correspondence between different log texts and system performance semantics. After training, the semantic vector output layer of the log semantic perceptron is retained and used to convert log message units into fixed-dimensional log semantic vectors during online tuning. For example, for the log message "flush task blocked by high queue pressure," the server generates a log semantic vector representing disk flushing blockage, increased queue pressure, and increased write latency; for the log message "compression.levellowered to 1," the server generates a log semantic vector representing the parameter name, adjustment direction, and adjustment result. The server arranges all log semantic vectors in chronological order to form a log semantic vector sequence. Figure 2As shown, the server calculates the semantic association strength of the log semantic vector sequence based on a context-aware window and forms a semantic association strength evolution trajectory. When the semantic association strength evolution trajectory changes beyond a mutation threshold between adjacent time sequence windows, the corresponding position is determined as a semantic mutation breakpoint, and the log semantic vector sequence is segmented into multiple semantically coherent log segments based on the semantic mutation breakpoint. The server further performs context semantic association modeling on the log semantic vector sequence. The server slides a context-aware window of a preset span on the log semantic vector sequence and calculates the semantic association strength representation value between adjacent log semantic vectors within the window. When consecutive logs successively show "increased sending delay", "increased queue level", and "disk flushing blockage", the server calculates a high semantic association strength, indicating that the log segment reflects the same type of performance degradation process; when the log content suddenly changes from "successful batch sending" to "increased compression time and increased CPU load", the server calculates a significantly decreased semantic association strength and determines the corresponding position as a semantic mutation breakpoint. The server segments the log semantic vector sequence into multiple semantically coherent log segments according to the semantic mutation breakpoint.

[0028] In one optional implementation, let the log semantic vector sequence be V={v_1,v_2,...,v_n}, where v_i represents the log semantic vector corresponding to the i-th log message unit. The server slides a context-aware window with a preset span of w over the log semantic vector sequence, and obtains a local vector set W_t={v_t,v_{t+1},...,v_{t+w-1}} within the t-th window. For any two adjacent log semantic vectors v_i and v_{i+1} within the window W_t, the server calculates the semantic association strength representation value as follows: r_i=cos(v_i,v_{i+1})=(v_i·v_{i+1}) / (||v_i||·||v_{i+1}||). Where r_i represents the semantic association strength between two adjacent log semantic vectors, v_i·v_{i+1} represents the inner product of the two vectors, and ||v_i|| and ||v_{i+1}|| represent the magnitudes of the two vectors, respectively. For the t-th context-aware window, the server averages the semantic association strengths of all adjacent vectors within the window to obtain the window semantic association strength R_t: R_t=(1 / (w-1))Σ_{i=t} {t+w-2} r_iThe server arranges all windows according to their semantic association strength R_t in the window sliding order, forming a semantic association strength evolution trajectory R={R_1,R_2,...,R_m}. To extract semantic abrupt change breakpoints from the semantic association strength evolution trajectory, the server further calculates the difference between the semantic association strengths of adjacent windows: D_t=|R_t-R_{t-1}|. When D_t is greater than a preset abrupt change threshold θ_d, the server marks the position corresponding to the t-th window as a semantic abrupt change breakpoint. The preset abrupt change threshold θ_d can be determined based on the mean and standard deviation of the semantic association strength difference values ​​in the historical collection period, for example, θ_d=μ_D+λσ_D, where μ_D represents the mean of the historical difference values, σ_D represents the standard deviation of the historical difference values, and λ is the threshold adjustment coefficient, which can be set to 1.5 to 3. In this way, the server can determine semantic abrupt change breakpoints based on explicit vector similarity and difference threshold rules, avoiding reliance solely on manual rules for log segmentation. For each semantically coherent log segment, the server uses a preset system performance status dictionary to extract system performance status words. The system performance status dictionary can include performance status description phrases such as write latency, queue congestion, disk flushing blockage, compression timeout, forwarding timeout, throughput decrease, cache backlog, drop risk, and recovery stability. When a semantically coherent log segment contains log content such as "queueusage 92%", "batch send delay 280ms", and "flush task blocked", the server extracts performance status description phrases such as queue level increase, send latency increase, and disk flushing blockage from that segment. The server performs semantic fusion processing on the performance status description phrases extracted from the same segment to generate a system performance status representation vector for that semantically coherent log segment, and calculates the vector offset between the system performance status representation vectors of adjacent semantically coherent log segments. This vector offset is used as the system performance status transition feature. Thus, the server can represent the process of the target log system transitioning from a normal forwarding state to a queue congestion state, and then from a queue congestion state to a disk flushing blockage relief state. Simultaneously, the server performs parameter configuration change detection processing on each semantically coherent log segment. When the server detects log entries such as "batch.size increased to 800", "flush.interval.ms decreased from 1000 to 500", and "compression.level lowered to 1", it extracts the corresponding parameter name entity and parameter adjustment direction entity, and associates these entities with the forward time interval attribute tag of the semantically coherent log segment. Subsequently, based on the changes in log output after the parameter adjustment trigger time, the server constructs a latency-dependent mapping between the parameter adjustment trigger time and the parameter adjustment response effect.

[0029] In one alternative implementation, please refer to Figure 3 , Figure 3This is a schematic diagram illustrating the delay correlation mapping between the parameter adjustment trigger time point and the parameter adjustment response effect provided in this embodiment of the invention. The delay correlation mapping between the parameter adjustment trigger time point and the parameter adjustment response effect is constructed as follows. Suppose that the server detects a parameter adjustment action a at time point τ_a, where the parameter adjustment action a includes a parameter name entity p, an adjustment direction entity d, and an adjustment step size u. Starting from τ_a, the server reads subsequent semantically coherent log segments within a preset response observation window [τ_a, τ_a+T_r] and extracts the system performance state representation vector S_t for each subsequent semantically coherent log segment. The server sets the baseline state representation vector S_base before the parameter adjustment trigger to the mean of the system performance state representation vectors of a preset number of log segments before the trigger time point τ_a, and sets the response state representation vector S_resp after the trigger to the mean of the system performance state representation vectors of subsequent log segments within the response observation window. The server calculates the response effect vector E_a corresponding to the parameter adjustment action as follows: E_a = S_resp - S_base. To characterize the response delay, the server detects the time point τ_r where the first significant performance change occurs within the response observation window. When ||S_t - S_base|| is greater than the preset response change threshold θ_r, the server determines the corresponding time point as the first response time point τ_r and calculates the response delay Δτ: Δτ = τ_r - τ_a. The server combines the parameter name entity p, the adjustment direction entity d, the adjustment step size u, the parameter adjustment trigger time point τ_a, the response delay Δτ, and the response effect vector E_a into a single delay association map: Map_a = {p, d, u, τ_a, Δτ, E_a}. If the proportion of dimensions in the response effect vector E_a that are consistent with the preset performance optimization target direction is higher than the preset positive threshold, the server marks the delay association map as a positive response map; if the proportion of dimensions in the response effect vector E_a that are opposite to the preset performance optimization target direction is higher than the preset negative threshold, the server marks the delay association map as a negative response map. The server concatenates all delay association maps in the order of the parameter adjustment trigger time points τ_a to form the parameter configuration change response feature. For example, the server establishes a latency correlation mapping between "batch.size increased" and subsequent occurrences such as "send throughput increased" and "queue usage dropped from 92% to 67%", and similarly, it establishes a latency correlation mapping between "compression.level decreased" and subsequent occurrences such as "compression cost decreased" and "network payload increased slightly". The server concatenates all latency correlation mappings in chronological order to form a characteristic response to parameter configuration changes.After obtaining the system performance state transition characteristics and parameter configuration change response characteristics, the server invokes a pre-built reinforcement learning policy generation network to perform parameter tuning decision inference. The server inputs the system performance state transition characteristics into the state encoding module to generate a system state representation vector; it inputs the parameter configuration change response characteristics into the parameter response encoder to generate a parameter response context vector; and then fuses the two to obtain a joint state decision representation vector. This joint state decision representation vector reflects both the current performance state change trend of the target log system and the impact of historical parameter adjustment operations on the log output pattern. In a specific scenario, the server identifies problems such as a continuously rising queue level, increased sending latency, and increased compression time in the target log system based on the joint state decision representation vector. The policy inference module then generates action selection probability values ​​for multiple candidate parameter adjustment actions. Candidate actions include increasing the number of batch sends, decreasing the disk flushing interval, reducing the compression level, increasing the number of log forwarding threads, increasing the queue level threshold, and keeping the current parameters unchanged. The server selects the target parameter adjustment action based on the action selection probability values, forming a parameter adjustment action sequence. For example, when the probability values ​​for actions such as "increasing the number of log forwarding threads" and "reducing the compression level" are high, the server adds these actions to the parameter adjustment action sequence; when "increasing the compression level" may further increase the CPU load, the server reduces the probability of selecting this action.

[0030] In one optional implementation, the reinforcement learning policy generation network is implemented using an Actor-Critic structure, comprising a policy network (Actor) and a value network (Critic). The policy network (Actor) is used to adjust the action selection probability distribution of actions based on the current joint state decision representation vector by outputting candidate parameters. The value network (Critic) is used to estimate the state value corresponding to the current joint state decision representation vector. The policy network (Actor) includes an input layer, two fully connected hidden layers, and an action probability output layer, with each fully connected hidden layer followed by a nonlinear activation function. The value network (Critic) includes an input layer, two fully connected hidden layers, and a state value output layer. The policy network (Actor) and the value network (Critic) can share a front-end state encoding module or have independent parameters. The server uses system performance state transition characteristics, burst intensity distribution characteristics, parameter configuration change response characteristics, and optional parameter synergy topology characteristics to form the reinforcement learning state s_t. The state s_t includes at least: current queue level characteristics, sending delay characteristics, disk flushing blocking characteristics, compression time characteristics, log throughput change characteristics, log stream burst intensity characteristics, historical parameter adjustment response delay characteristics, and historical parameter adjustment response effect characteristics. The reinforcement learning action space A consists of adjustable parameters of the target log system, their adjustment directions, and step sizes. For example, action space A includes: increasing batch.size by one level, decreasing batch.size by one level, increasing flush.interval.ms by one level, decreasing flush.interval.ms by one level, increasing compression.level by one level, decreasing compression.level by one level, increasing send.thread.count by one level, decreasing send.thread.count by one level, increasing queue.high.watermark by one level, decreasing queue.high.watermark by one level, and actions that remain unchanged. Each candidate action corresponds to a parameter adjustment action type identifier. The policy network Actor outputs the action probability distribution π(a_t|s_t) based on the current state s_t. The server can select parameter adjustment actions using a probability sampling method or a maximum probability method. After the selected action a_t is executed by the target log system, the server collects the feedback log and calculates the immediate reward r_t based on the feedback log. The instant reward r_t is determined as follows: r_t = w_1·ΔDelay_t + w_2·ΔQueue_t + w_3·ΔDrop_t + w_4·ΔBlock_t - w_5·ΔCPU_t - w_6·ΔNet_t - w_7·Osc_t.Where ΔDelay_t represents the decrease in transmission delay, ΔQueue_t represents the decrease in queue level, ΔDrop_t represents the decrease in log drop risk, ΔBlock_t represents the decrease in disk flushing congestion, ΔCPU_t represents the increase in CPU usage, ΔNet_t represents the increase in network transmission load, Osc_t represents the parameter oscillation penalty term, and w_1 to w_7 are reward weight coefficients. If a certain indicator changes along a preset optimization direction, its corresponding term will make a positive contribution to the reward value; if a certain indicator changes along a preset deterioration direction, its corresponding term will make a negative contribution to the reward value. The value network Critic outputs the state value V(s_t), and the server calculates the time-series difference error δ_t based on the immediate reward r_t and the next state value V(s_{t+1}): δ_t = r_t + γV(s_{t+1}) - V(s_t). Where γ is a discount factor. The server updates the policy network Actor and the value network Critic based on the time-series difference error δ_t. The loss function of the policy network (Actor) can be expressed as: L_actor = -logπ(a_t|s_t)·δ_t. The loss function of the value network (Critic) can be expressed as: L_critic = (r_t + γV(s_{t+1}) - V(s_t))2. The server updates the network parameters of the reinforcement learning policy generation network by minimizing L_actor and L_critic, so that parameter adjustment actions that can generate positive rewards have a higher action selection probability under the same or similar system performance states, and parameter adjustment actions that generate negative rewards or cause parameter oscillations have a lower action selection probability. The server also generates an immediate reward signal based on the system log feedback observation results after executing the target parameter adjustment action. After the target log system executes "increase the number of log forwarding threads", the server continues to collect feedback log text streams and extracts system performance state phrases from the feedback log message unit sequence to generate performance indicator observation curves. If the feedback log shows a decrease in transmission latency, a drop in queue level, and a reduction in dropout risk, the server generates a positive immediate reward signal. If the feedback log shows an increase in CPU usage, an abnormal increase in network traffic, or aggravated queue fluctuations, the server generates a negative immediate reward signal. The server uses the immediate reward signals to update the reinforcement learning strategy generator network, increasing its probability of selecting effective parameter adjustment actions and decreasing its probability of selecting ineffective parameter adjustment actions in subsequent optimization decisions.

[0031] In one optional implementation, the reward signal generator extracts numerical performance indicators from the feedback log message unit sequence according to a preset indicator extraction rule. These numerical performance indicators include average transmission delay (Delay), queue level (Queue), disk flushing blocking duration (Block), log drop risk (Drop), CPU usage (CPU), network transmission load (Net), and log throughput (Throughput). For the baseline monitoring window before the optimization action is executed, the server calculates the baseline values ​​for each performance indicator; for the feedback monitoring window after the optimization action is executed, the server calculates the feedback values ​​for each performance indicator. The server calculates the changes in metrics as follows: ΔDelay = (Delay_base - Delay_feedback) / Delay_base; ΔQueue = (Queue_base - Queue_feedback) / Queue_base; ΔBlock = (Block_base - Block_feedback) / Block_base; ΔDrop = (Drop_base - Drop_feedback) / Drop_base; ΔCPU = (CPU_feedback - CPU_base) / CPU_base; ΔNet = (Net_feedback - Net_base) / Net_base; Δ Throughput = (Throughput_feedback - Throughput_base) / Throughput_base; where Delay_base, Queue_base, Block_base, Drop_base, CPU_base, Net_base, and Throughput_base represent the baseline metric values ​​before optimization, and Delay_feedback, Queue_feedback, Block_feedback, Drop_feedback, CPU_feedback, Net_feedback, and Throughput_feedback represent the feedback metric values ​​after optimization. The server uses reduced sending latency, reduced queue level, reduced disk flushing congestion, reduced log drop risk, and increased throughput as positive optimization goals, and increased CPU usage and abnormally high network transmission load as negative constraint goals.

[0032] The reward signal generator generates an instantaneous reward signal based on the changes in the aforementioned metrics: r = w_dΔDelay + w_qΔQueue + w_bΔBlock + w_lΔDrop + w_tΔThroughput - w_cΔCPU - w_nΔNet - w_oOsc. Here, w_d, w_q, w_b, w_l, w_t, w_c, w_n, and w_o represent the reward weights for each metric, and Osc represents the parameter oscillation penalty term. If two consecutive decision time steps adjust the same parameter in opposite directions, Osc takes a positive value; otherwise, Osc is 0. Through this reward signal generation method, the server can transform the performance changes in the feedback log into numerical training signals usable by the reinforcement learning policy generation network.

[0033] After obtaining the parameter adjustment action sequence, the server generates a set of tuning instructions containing the target parameter name, adjustment direction identifier, and adjustment step level. The server parses the parameter adjustment action type identifier corresponding to each decision time step in the parameter adjustment action sequence and maps it to the specific target parameter name, adjustment direction identifier, and adjustment step level according to the parameter name resolution table, adjustment direction resolution table, and step level resolution table. For example, the server maps action identifier A04 to the number of log forwarding threads, direction identifier UP to increase, and step level 1 to increase by one preset thread step; the server maps action identifier A03 to compression level, direction identifier DOWN to decrease, and step level 1 to decrease the compression level by one level. The server encapsulates the above content to obtain the tuning instruction sequence.

[0034] The server then simplifies and rearranges the tuning instruction sequence. For consecutive tuning instructions with the same target parameter name but opposite adjustment directions, the server cancels or merges them to avoid the target log system frequently executing invalid parameter tuning. For example, when two consecutive decision time steps generate "increase batch.size by one level" and "decrease batch.size by one level" respectively, the server simplifies it to keep batch.size unchanged. For parameters with dependencies, the server rearranges the order according to the preset parameter dependencies. For example, it first issues an instruction to increase the number of log forwarding threads, and then issues an instruction to increase the number of batches sent, so that the target log system first improves its log forwarding processing capacity and then expands the batch sending scale. After simplification and rearrangement, the server obtains the final set of tuning instructions.

[0035] Finally, the server sends the tuning instruction set to the parameter control interface of the target log system via the communication link. Upon receiving the tuning instruction set, the parameter control interface verifies the validity of the target parameter name, adjustment direction identifier, and adjustment step size, and updates the corresponding performance parameters in the runtime configuration management module of the target log system. The target log system then increases the number of log forwarding threads, adjusts the batch sending count, shortens the disk flushing interval, or reduces the compression level according to the tuning instructions. After the dynamic parameter adjustment operation is completed, the server continues to collect the feedback log text stream output by the target log system and repeatedly executes the processes of log message unit generation, semantic awareness and parameter association parsing, reinforcement learning decision inference, and tuning instruction issuance. Through this closed-loop processing, the server can continuously generate parameter tuning instructions adapted to the current operating state based on the real-time log performance of the target log system, historical parameter response patterns, and reinforcement learning feedback results, thereby improving the log system's operational stability and performance adaptability under scenarios such as high-concurrency writing, queue backlog, increased compression pressure, and forwarding latency.

[0036] In this embodiment of the invention, the step of acquiring the raw log text stream generated by the target log system within a preset collection period and converting the raw log text stream into a sequence of log message units with time sequence markings can be implemented through the following example.

[0037] A log collection probe is deployed on the log output channel of the target log system to capture the original log text stream flowing through the log output channel within the preset collection period.

[0038] The original log text stream is processed by log line segmentation to obtain a set of log text lines with independent line structures. The log line segmentation process is based on the newline boundary symbols in the original log text stream to divide the log text lines.

[0039] For each log text line in the log text line set, perform timestamp location processing, extract the timestamp string from the header character field of each log text line, and parse the timestamp string into a timestamp marker;

[0040] A time sequence index is constructed based on the timestamp marker, and the time sequence index is used to record the position number of each log text line in the chronological order.

[0041] The log text line set is sorted based on the time sequence index, and the log text lines are arranged in order from earliest to latest according to the timestamp markers to obtain a log text line queue with a chronological order.

[0042] Log line boundary detection processing is performed on the log text line queue. The characteristic character sequence of the starting line of the log record is identified in the log text line queue. Based on the characteristic character sequence, all log text lines between adjacent starting lines of log records are divided into the same basic log sharding unit, so that the log text line queue is transformed into a basic log sharding unit sequence.

[0043] For each basic log sharding unit in the basic log sharding unit sequence, time interval calculation processing is performed. The timestamp difference of each basic log sharding unit relative to its previous adjacent basic log sharding unit is calculated, and the timestamp difference is used as the forward time interval attribute value of the basic log sharding unit.

[0044] The log text content within each basic log shard unit is fragmented, and the entire log text within each basic log shard unit is cut into multiple log content short sentences according to semantic delimiters, so that each log content short sentence is an independent log message unit.

[0045] The forward time interval attribute value of each basic log shard unit is appended to all log message units cut out by that basic log shard unit, and each log message unit is assigned a sequential number in the sequence of the basic log shard units.

[0046] All log message units carrying forward time interval attribute values ​​and sequential sequence numbers are aggregated to form a log message unit sequence with time sequence marking, such that each log message unit in the log message unit sequence with time sequence marking carries a forward time interval attribute mark and a sequential sequence number.

[0047] In this embodiment of the invention, for example, the server acts as the execution entity, collecting and structuring the raw log text stream generated by the target log system within a preset collection period to form a sequence of log message units with time-sequence markings. The target log system is deployed in a business server cluster to receive operational logs generated by order services, payment services, and gateway services, and sends them to a centralized log platform through a log output channel. The server deploys a log collection probe on this log output channel, which continuously captures the raw log text stream flowing through the channel within a preset five-minute collection period. For example, the server collects continuous log content such as "2026-04-29 10:15:01.103 [INFO] batch senddelay 280ms", "2026-04-29 10:15:01.125 [WARN] queue usage 92%", and "2026-04-29 10:15:01.260 [ERROR] flush task blocked", and uses it as the raw log text stream to be processed.

[0048] After acquiring the raw log text stream, the server segments the log lines based on newline delimiters, dividing the continuous text into sets of log text lines with independent line structures. Then, the server performs timestamp localization on each log text line, extracting the timestamp string from the header segment of the log text line. For example, the server extracts "2026-04-29 10:15:01.125" from "2026-04-29 10:15:01.125 [WARN] queue usage 92%" and parses it into a uniformly formatted timestamp. For log arrival order discrepancies caused by network transmission delays between different service nodes, the server does not use the receiving order but instead constructs a time-sequence index based on the timestamps to record the position number of each log text line in chronological order.

[0049] The server sorts the log text lines based on a time-order index, arranging them from earliest to latest according to their timestamps, resulting in a sequential queue of log text lines. For example, a log line with a timestamp of 10:15:01.125 is placed before a log line with a timestamp of 10:15:01.260. Next, the server performs log line boundary detection on the queue, identifying characteristic character sequences of the starting lines of log records. Log text lines beginning with a timestamp, log level, or thread identifier are identified as starting lines; lines beginning with spaces, tabs, "at com.xxx", or "Caused by" are identified as continuations of the previous log record. Based on the positional relationship between adjacent starting lines, the server groups all log text lines into the same basic log shard unit, transforming the log text line queue into a sequence of basic log shard units.

[0050] Next, the server performs time interval calculation processing on each basic log shard unit in the basic log shard unit sequence. It calculates the timestamp difference between the current basic log shard unit and the previous adjacent basic log shard unit, and uses this timestamp difference as the forward time interval attribute value. For example, if the timestamp of the 230th basic log shard unit is 10:15:01.103 and the timestamp of the 231st basic log shard unit is 10:15:01.125, the server calculates 22 milliseconds and writes it to the 231st basic log shard unit. This attribute reflects the occurrence density between log events.

[0051] Subsequently, the server fragments the log text content within each basic log shard unit, dividing the entire log text into multiple short log sentences using semantic delimiters such as commas, semicolons, vertical bars, periods, and key-value pair separators. For example, the server splits "queue usage 92%; batch send delay 280ms; flush task blocked by high queue pressure" into "queue usage 92%", "batch send delay 280ms", and "flush task blocked by high queue pressure", treating each short sentence as an independent log message unit.

[0052] The server appends the forward time interval attribute value of each basic log shard unit to all log message units split from that shard, and assigns a sequential number to each log message unit. For example, the three log message units split from the 231st basic log shard unit all carry a forward time interval attribute of 22 milliseconds and are assigned the numbers 231-1, 231-2, and 231-3, respectively. Finally, the server aggregates all log message units carrying forward time interval attribute values ​​and sequential numbers to form a sequence of log message units with time-ordered tags. This sequence preserves the log generation order, event intervals, and fine-grained log content, providing a data foundation for subsequent semantic awareness, parameter correlation parsing, and reinforcement learning optimization decisions.

[0053] In this embodiment of the invention, the semantic awareness and parameter association parsing processing of the log message unit sequence to obtain the system performance state migration characteristics and parameter configuration change response characteristics of the log message unit sequence within the preset collection period can be implemented through the following example.

[0054] Obtain the sequence of log message units with time sequence markers, and input each log message unit in the log message unit sequence into a pre-trained log semantic perceptron, wherein the pre-trained log semantic perceptron is used to perform semantic vector encoding on a single log message unit;

[0055] The pre-trained log semantic perceptron is used to perform semantic vector encoding on each log message unit to generate a fixed-dimensional log semantic vector corresponding to each log message unit, and all log semantic vectors are combined into a log semantic vector sequence.

[0056] The log semantic vector sequence is subjected to context semantic association modeling processing. A context-aware window with a preset span is slid across the log semantic vector sequence, and the semantic association strength representation value between adjacent log semantic vectors within the context-aware window is calculated.

[0057] The semantic association strength representation values ​​calculated within each context-aware window are arranged in the window sliding order to form a semantic association strength evolution trajectory that runs through the preset collection period. The semantic association strength evolution trajectory is used to reflect the continuous or abrupt characteristics of the log semantic internal logic.

[0058] Extract the location information of semantic mutation breakpoints from the semantic association strength evolution trajectory, and divide the log semantic vector sequence into multiple semantically coherent log segments based on the location information of the semantic mutation breakpoints, with semantic consistency maintained within each semantically coherent log segment.

[0059] For each semantically coherent log segment, system performance status words are extracted. A preset system performance status dictionary is used to perform word-level matching on the log semantic vectors in the semantically coherent log segment to extract performance status description words that represent the system performance status.

[0060] All performance status description phrases extracted from the same semantically coherent log segment are semantically fused, and the system performance status representation vector of the semantically coherent log segment is generated by vector weighted averaging. The vector offset between the system performance status representation vectors of adjacent semantically coherent log segments is used as the system performance status transition feature.

[0061] For each semantically coherent log segment, parameter configuration change detection processing is performed. The process detects whether the semantically coherent log segment contains descriptive words of parameter adjustment actions. If it contains descriptive words of parameter adjustment actions, the parameter name entity and parameter adjustment direction entity appearing in the context window of the descriptive words of the parameter adjustment actions are extracted.

[0062] Associate the parameter name entity and the parameter adjustment direction entity with the forward time interval attribute tag of the semantically coherent log fragment to construct a delay association mapping between the parameter adjustment trigger time point and the parameter adjustment response effect;

[0063] The delay association mappings of all semantically coherent log fragments are concatenated in chronological order to form parameter configuration change response features. These features characterize the response transmission pattern between parameter adjustment operations and changes in system log output mode in the target log system.

[0064] In this embodiment of the invention, for example, after obtaining a sequence of log message units with time-ordered markers, the server performs semantic awareness and parameter association parsing processing on the sequence to obtain system performance state transition characteristics and parameter configuration change response characteristics within a preset collection period. Each log message unit in the sequence carries a forward time interval attribute marker and a sequential number. For example, log message unit number 231-1 is "queueusage 92%", log message unit number 231-2 is "batch send delay 280ms", and log message unit number 231-3 is "flush task blocked by high queue pressure". The server reads the above log message units sequentially according to the sequential number and inputs each log message unit into a pre-trained log semantic awareness unit.

[0065] The pre-trained log semantic perceptron is used to encode semantic vectors for individual log message units. The server uses this perceptron to encode key fields, performance descriptors, parameter names, action words, and numerical information within the log message units. For "queue usage 92%", the server generates a fixed-dimensional log semantic vector representing a high queue level; for "batch send delay 280ms", the server generates a log semantic vector representing an increased batch send delay; for "compression.level lowered to 1", the server generates a log semantic vector simultaneously representing the compression level parameter, the reduction direction, and the adjusted value. The server arranges the fixed-dimensional log semantic vectors corresponding to all log message units in their original chronological order, forming a log semantic vector sequence, thus transforming unstructured log text into a vector sequence capable of computational analysis.

[0066] Subsequently, the server performs contextual semantic association modeling on the log semantic vector sequence. The server slides a context-aware window with a preset span across the log semantic vector sequence; the preset span is a number of consecutive log message units. Each time the window is slid, the server calculates the semantic association strength between adjacent log semantic vectors within the window. In a specific scenario, when the window continuously displays "batch send delay 280ms," "queue usage 92%," and "flush task blocked by high queue pressure," the server recognizes that these logs all point to increased log writing and forwarding pressure, calculating a high semantic association strength. When a subsequent window suddenly changes from "batch send success" to "compression cost rising" and "CPU load high," the server recognizes that the log semantics have changed from sending link pressure to compression calculation pressure, calculating a significantly decreased semantic association strength.

[0067] The server arranges the semantic association strength values ​​calculated within each context-aware window according to the window sliding order, forming a semantic association strength evolution trajectory that spans a preset collection period. This trajectory reflects the continuity and abrupt changes in the inherent logic of log semantics. For example, in the first two minutes of the collection period, the trajectory remains in a high range, indicating that the log content continuously describes queue backlog and transmission delays; in the third minute, the trajectory drops significantly, indicating a topic shift in the log content, from queue congestion to increased compression time; in the fourth minute, the trajectory stabilizes again, indicating that the log system has entered a new stable operating phase. The server extracts the location information of semantic abrupt change breakpoints from this trajectory and divides the log semantic vector sequence into multiple semantically coherent log segments based on these breakpoints. Each semantically coherent log segment maintains semantic consistency; for example, the first segment corresponds to the queue congestion process, the second segment to the increased compression pressure process, and the third segment to the recovery process after parameter adjustment.

[0068] For each semantically coherent log segment, the server performs system performance status term extraction processing. The server uses a pre-defined system performance status dictionary to perform word-level matching on the log semantic vectors within the semantically coherent log segment. This dictionary includes performance status description phrases such as queue level increase, increased sending latency, disk flushing blockage, increased compression time, decreased throughput, limited network forwarding, cache backlog, and recovery to stability. In the queue congestion segment, the server extracts performance status description phrases such as queue level increase, increased sending latency, and disk flushing blockage from "queue usage 92%", "batch send delay 280ms", and "flush task blocked"; in the compression stress segment, the server extracts performance status description phrases such as increased compression time and increased CPU load from "compression costrising" and "CPU load high".

[0069] The server performs semantic fusion processing on all performance status description phrases extracted from the same semantically coherent log segment, and generates a system performance status representation vector for that semantically coherent log segment by means of vector weighted averaging.

[0070] In one optional implementation, for the k-th semantically coherent log segment P_k, the server extracts q performance status description phrases from the segment and obtains corresponding performance status phrase vectors s_{k,1},s_{k,2},...,s_{k,q}. The server calculates a fusion weight α_{k,j} for each performance status phrase, which is determined by the log level weight, the frequency weight, and the time density weight: α_{k,j}=Norm(β_level·g_{k,j}+β_freq·f_{k,j}+β_time·h_{k,j}). Where g_{k,j} represents the weight value of the log level corresponding to the j-th performance status description phrase, with error-level logs having a higher weight than warning-level logs, and warning-level logs having a higher weight than ordinary message-level logs; f_{k,j} represents the frequency of occurrence of the performance status description phrase in the semantically coherent log segment; h_{k,j} represents the time density weight of the log message unit in which the performance status description phrase is located, with a higher time density weight for shorter forward time intervals; β_level, β_freq, and β_time represent the weight adjustment coefficients for log level, frequency of occurrence, and time density, respectively; Norm represents normalization processing, making the sum of the fusion weights of all performance status description phrases within the same semantically coherent log segment equal to 1. The server calculates the system performance status representation vector S_k of the k-th semantically coherent log segment as follows: S_k=Σ_{j=1} {q} α_{k,j}s_{k,j} For two adjacent semantically coherent log segments P_k and P_{k+1}, the server calculates the vector offset between their system performance state representation vectors: M_k = S_{k+1} - S_k. M_k is then used as a component of the system performance state transition feature. To obtain the scalarized transition strength, the server can also calculate the magnitude of the vector offset: I_k = ||S_{k+1} - S_k||. Here, I_k represents the performance state transition strength between adjacent semantically coherent log segments. When I_k exceeds a preset transition strength threshold, the server determines that a significant performance state transition has occurred in the target log system at the corresponding time position.

[0071] The weights are determined based on log level, frequency of occurrence, and forward time interval attributes, with error-level logs and logs exhibiting performance anomalies in a short period receiving higher weights. The server further calculates the vector offset between the system performance state representation vectors of adjacent semantically coherent log segments and uses this vector offset as a system performance state transition feature. Through this feature, the server can record the change path of the target log system from a normal forwarding state to a queue congestion state, then from a queue congestion state to a compression pressure state, and finally to a stable recovery state.

[0072] Simultaneously, the server performs parameter configuration change detection processing on each semantically coherent log segment. The server checks whether the segment contains descriptive words for parameter adjustment actions, such as updated, increased, decreased, lowered, raised, changed, etc. When the server detects "batch.size increased from 500 to 800" in a segment, the server extracts the parameter name entity batch.size and the parameter adjustment direction entity increased from the context window of that descriptive word; when the server detects "flush.interval.ms decreased from 1000 to 500", the server extracts the parameter name entity flush.interval.ms and the parameter adjustment direction entity decreased; when the server detects "compression.level lowered to 1", the server extracts the parameter name entity compression.level and the parameter adjustment direction entity lowered.

[0073] The server associates the extracted parameter name entity and parameter adjustment direction entity with the forward time interval attribute tag of the semantically coherent log segment, and continues to read log segments after the parameter adjustment trigger time point to construct a latency correlation mapping between the parameter adjustment trigger time point and the parameter adjustment response effect. For example, if the server records that after batch.size increases, "send throughput increased" and "queue usage dropped from 92% to 67%" appear within the next 20 seconds, a positive latency correlation mapping is established between the increase in batch.size and the increase in throughput and the decrease in queue level; if the server records that after compression.level decreases, "compression cost decreased" and "network payload increased slightly" appear within the next 10 seconds, a response mapping is established between the decrease in compression.level and the decrease in compression time and the slight increase in network load.

[0074] Finally, the server concatenates the delay association mappings formed in all semantically coherent log fragments in chronological order to create a parameter configuration change response feature. This feature records the impact of different parameter adjustment actions triggered at different times on the system log output pattern, including the response occurrence time, the direction of performance change, and the duration of the impact. Thus, the server simultaneously obtains system performance state transition features and parameter configuration change response features, providing accurate state inputs and response basis for subsequent parameter tuning decision-making in the reinforcement learning policy generation network.

[0075] In this embodiment of the invention, the invocation of the pre-built reinforcement learning policy generation network performs parameter tuning decision deduction on the system performance state transition characteristics and the parameter configuration change response characteristics to obtain the parameter adjustment action sequence of the target log system, which can be implemented through the following example.

[0076] A reinforcement learning policy generation network is constructed, which includes a state encoding module, a policy reasoning module, and an action output module. The state encoding module receives the system performance state transition features as input, and the action output module outputs parameters to adjust the decision probability distribution of actions.

[0077] The system performance state transition features are input into the state encoding module, and the state encoding module maps the system performance state transition features into a fixed-dimensional system state representation vector. The system state representation vector is used to characterize the overall performance status of the target log system within the current preset collection period.

[0078] The parameter configuration change response features are input into the parameter response encoder, which is used to parse the structured information of the delay correlation mapping in the parameter configuration change response features and generate a parameter response context vector. The parameter response context vector is used to describe the causal effect pattern between historical parameter adjustment actions and the changes in log output patterns they cause.

[0079] The system state representation vector and the parameter response context vector are concatenated and fused to form a joint state decision representation vector, which simultaneously contains current system performance status information and historical parameter adjustment response experience knowledge.

[0080] The joint state decision representation vector is input into the policy reasoning module, which is composed of a multi-layer fully connected neural network. The policy reasoning module performs multi-layer nonlinear transformation processing on the joint state decision representation vector to generate decision score vectors for each candidate parameter adjustment action.

[0081] The decision scoring vector is input into the action output module. The action output module performs probability normalization on the decision scoring vector to generate action selection probability values ​​corresponding to each candidate parameter adjustment action. The candidate parameter adjustment actions include parameter increase action identifier, parameter decrease action identifier, and parameter remain unchanged action identifier.

[0082] Based on the action selection probability value, a strategy sampling operation is performed to select a target parameter adjustment action from all candidate parameter adjustment actions, and the selected target parameter adjustment action is added to the parameter adjustment action sequence.

[0083] The system log feedback observation results obtained after performing the target parameter adjustment action are input into the reward signal generator. The reward signal generator compares the system log feedback observation results with the preset performance optimization target direction and generates a numerical real-time reward signal.

[0084] The immediate reward signal is fed back to the reinforcement learning policy generation network, driving the policy inference module to update its network connection weights, thereby increasing the probability of the parameter adjustment action that improves the reward signal being selected in future decisions.

[0085] Repeat the sampling operation and network connection weight update process of the above strategy until all time steps that need to be decided within the preset collection period have been traversed, and output the final parameter adjustment action sequence, which includes the parameter adjustment action type identifier selected at each decision time step.

[0086] In this embodiment of the invention, for example, after obtaining system performance state transition characteristics and parameter configuration change response characteristics, the server invokes a pre-built reinforcement learning policy generation network to perform parameter tuning decision inference. The reinforcement learning policy generation network built by the server includes a state encoding module, a parameter response encoder, a policy inference module, and an action output module. The state encoding module is used to receive system performance state transition characteristics, and the action output module is used to output the decision probability distribution of parameter adjustment actions. Candidate parameter adjustment actions for the target log system include action identifiers such as increasing the number of batch sends, decreasing the disk flushing interval, reducing the compression level, increasing the number of log forwarding threads, increasing the queue water level threshold, and keeping the current parameters unchanged.

[0087] The server first inputs the system performance state transition characteristics into the state encoding module. The state encoding module maps these characteristics into a fixed-dimensional system state representation vector, used to characterize the overall performance status of the target log system within the current preset collection period. For example, if the server detects that the system performance state has transitioned from "normal forwarding" to "increased queue level," and then to "increased sending latency and disk flushing congestion," the state encoding module generates a system state representation vector indicating high queue pressure, write link congestion, and insufficient log forwarding capacity. This vector can centrally reflect the main performance bottlenecks currently existing in the log system.

[0088] Simultaneously, the server inputs the parameter configuration change response characteristics into the parameter response encoder. The encoder parses the latency-related mappings to generate a parameter response context vector. For example, the server might record in historical responses that "batch.size increased" resulted in "send throughput increased" and "queue usage dropped," "compression.level decreased" resulted in "compression cost decreased" and "network payload increased slightly," and "flush.interval.ms decreased" resulted in "flush delay reduced" but with increased disk write frequency. The encoder encodes these structured response relationships into a parameter response context vector, which describes the causal effect pattern between historical parameter adjustment actions and changes in log output patterns.

[0089] Next, the server concatenates and merges the system state representation vector with the parameter response context vector to form a joint state decision representation vector. This joint state decision representation vector simultaneously contains information about the current system performance status and historical parameter adjustment response experience. For example, if the current logs show that the queue level is consistently above 90% and the transmission latency reaches 280ms, historical responses show that increasing the number of log forwarding threads can lower the queue level, and reducing the compression level can reduce CPU compression time. The server provides both the current problem and historical effective tuning experience to the policy inference module through the joint state decision representation vector.

[0090] The server inputs the joint state decision representation vector into the policy inference module. The policy inference module, composed of a multi-layer fully connected neural network, performs multi-layer nonlinear transformations on the joint state decision representation vector to generate decision score vectors for each candidate parameter adjustment action. In a specific operational scenario, the policy inference module generates scores for candidate actions such as "increasing the number of log forwarding threads," "increasing batch size," "decreasing compression level," "decreasing flush interval ms," and "keeping it unchanged." Since the current state exhibits queue congestion and increased transmission latency, and historical records indicate that increasing the number of log forwarding threads and moderately increasing the batch size have positive effects, the server generates higher decision scores for these actions; conversely, increasing the compression level increases CPU load, so the server generates lower scores for these actions.

[0091] Subsequently, the server inputs the decision score vector into the action output module. The action output module performs probability normalization on the decision score vector, generating action selection probability values ​​corresponding to each candidate parameter adjustment action. The candidate parameter adjustment actions include parameter increase action identifiers, parameter decrease action identifiers, and parameter remain unchanged action identifiers. For example, the action output module outputs a probability of 0.38 for "send.thread.count increase", 0.26 for "batch.size increase", 0.21 for "compression.level decrease", 0.10 for "flush.interval.ms decrease", and 0.05 for "remain unchanged". The server performs a strategy sampling operation based on the action selection probability values, selects the target parameter adjustment action from the candidate parameter adjustment actions, and adds the selected target parameter adjustment action to the parameter adjustment action sequence.

[0092] After the server performs the selected target parameter adjustment action in the target log system, it continues to collect system log feedback observations and inputs these observations into the reward signal generator. The reward signal generator compares the performance changes reflected in the feedback logs with preset performance optimization target directions. The preset target directions include reducing send latency, reducing queue water level, reducing disk flushing congestion, maintaining stable CPU load, and avoiding log drop. After the server executes "send.thread.count increase", the feedback log shows "queue usage dropped from 92% to 68%" and "batch send delay reduced to 120ms", and the reward signal generator generates a positive immediate reward signal; after the server performs a certain action, the feedback log shows "CPU load high" and "network payload overflow risk", and the reward signal generator generates a negative immediate reward signal.

[0093] The server feeds back the immediate reward signal to the reinforcement learning policy generation network, driving the policy inference module to update the network connection weights. For parameter adjustment actions that improve the reward signal, the server increases the probability of that action being selected in similar system states in the future; for parameter adjustment actions that degrade performance, the server decreases their probability of being selected in the future. For example, in a state where the queue level is high and the transmission latency is high, "increasing the number of log forwarding threads" receives positive rewards multiple times. After updating the network weights, the server increases the probability of this action being selected in subsequent similar states.

[0094] The server repeatedly executes the strategy sampling operation and network connection weight update process according to the above procedure until all time steps requiring decision-making within the preset collection period have been traversed. Finally, the server outputs a parameter adjustment action sequence, which includes an identifier for the selected parameter adjustment action type at each decision time step, such as "Time Step 1: Increase send.thread.count", "Time Step 2: Decrease compression.level", "Time Step 3: Increase batch.size", and "Time Step 4: Remain unchanged". This parameter adjustment action sequence serves as the direct basis for subsequently generating a set of tuning instructions.

[0095] In this embodiment of the invention, the step of generating a set of tuning instructions containing the name of the target parameter, the direction identifier, and the step size based on the parameter adjustment action sequence, and sending the set of tuning instructions to the parameter control interface of the target log system to trigger the dynamic parameter adjustment operation can be implemented through the following example.

[0096] The parameter adjustment action sequence is parsed, and the parameter adjustment action type identifier corresponding to each parameter adjustment action record is extracted item by item in the order of the decision time step;

[0097] Each parameter adjustment action type identifier is entered into the parameter name resolution table for lookup. The parameter name resolution table records the mapping relationship between the parameter adjustment action type identifier and the name of the adjustable parameter in the target log system. Based on the mapping relationship, the name of the target parameter to which the parameter adjustment action points is determined.

[0098] Each parameter adjustment action type identifier is entered into the adjustment direction parsing table for lookup. The adjustment direction parsing table records the mapping relationship between the parameter adjustment action type identifier and the adjustment operation direction. Based on the mapping relationship, the adjustment direction identifier corresponding to the parameter adjustment action is determined. The adjustment direction identifier includes an upward adjustment direction identifier and a downward adjustment direction identifier.

[0099] Each parameter adjustment action type identifier is entered into the step level analysis table for lookup. The step level analysis table records the mapping relationship between the parameter adjustment action type identifier and the adjustment amplitude level. The adjustment step level corresponding to the parameter adjustment action is determined based on the mapping relationship.

[0100] The adjustment target parameter name, the adjustment direction identifier, and the adjustment step level are combined and encapsulated to generate a single parameter tuning instruction corresponding to the parameter adjustment action. The single parameter tuning instruction contains all the control information involved in the parameter adjustment.

[0101] The individual parameter tuning instructions generated by all decision time steps in the parameter adjustment action sequence are arranged and aggregated according to the order of the decision time steps to form a tuning instruction sequence.

[0102] The tuning instruction sequence is simplified and merged. Tuning instruction pairs that target the same adjustment target parameter name and have opposite adjustment direction identifiers at consecutive decision time steps are detected in the tuning instruction sequence. The detected tuning instruction pairs are removed from the tuning instruction sequence to eliminate back-and-forth oscillating adjustment operations, resulting in a simplified tuning instruction set.

[0103] The simplified tuning instruction set is subjected to dependency sorting. Based on the preset parameter dependency relationship between the names of the target parameters to be adjusted, the tuning instructions with dependencies are rearranged in order according to the principle of prioritizing the adjustment of the dependent parameters, thus obtaining the tuning instruction set.

[0104] Establish a communication link with the parameter control interface of the target log system, wherein the parameter control interface is the remote parameter adjustment entry exposed by the target log system.

[0105] The set of tuning instructions is sent sequentially to the parameter control interface through the communication link, so that the parameter control interface parses the target parameter name, adjustment direction identifier and adjustment step level in the tuning instructions one by one, and drives the target log system to perform the corresponding dynamic parameter adjustment operation.

[0106] In this embodiment of the invention, for example, after obtaining the parameter adjustment action sequence, the server converts the sequence into a set of tuning instructions executable by the target log system. The parameter adjustment action sequence is generated by the reinforcement learning policy network output, and parameter adjustment action type identifiers are recorded according to the decision time step. For example, the first time step is A04_UP_1, the second time step is A03_DOWN_1, the third time step is A02_DOWN_1, and the fourth time step is A01_UP_1. The server parses the sequence item by item according to the time step order, extracting the action type identifier from each parameter adjustment action record.

[0107] The server inputs the action type identifier for each parameter adjustment into the parameter name resolution table for lookup. The parameter name resolution table records the mapping relationship between the action type identifier and the adjustable parameter name of the target log system. For example, A01 corresponds to batch.size, A02 to flush.interval.ms, A03 to compression.level, A04 to send.thread.count, and A05 to queue.high.watermark. When the server parses A04_UP_1, it determines that the target parameter name to be adjusted is send.thread.count; when parsing A03_DOWN_1, it determines that the target parameter name to be adjusted is compression.level.

[0108] The server continues to look up the action type identifier in the adjustment direction resolution table. The adjustment direction resolution table records the mapping relationship between action type identifiers and adjustment operation directions, where UP indicates upward adjustment, DOWN indicates downward adjustment, and KEEP indicates no change. When the server parses A04_UP_1, it determines the adjustment direction identifier is upward adjustment, used to increase the number of log forwarding threads; when parsing A03_DOWN_1, it determines the adjustment direction identifier is downward adjustment, used to reduce the compression level. For KEEP type actions, the server marks them as maintaining the current parameter value and does not generate an actual change instruction.

[0109] The server further inputs the action type identifier into the step level resolution table for lookup. The step level resolution table records the mapping relationship between the action type identifier and the adjustment amplitude level. Step level 1 represents one basic adjustment unit, and step level 2 represents two basic adjustment units. For example, when the server resolves A04_UP_1, it determines that send.thread.count increases by one thread step; when resolving A01_UP_1, it determines that batch.size increases by one batch send level; and when resolving A02_DOWN_1, it determines that flush.interval.ms decreases by one interval level.

[0110] After determining the target parameter name, adjustment direction identifier, and adjustment step size, the server encapsulates these three elements into a single parameter tuning instruction. For example, A04_UP_1 is encapsulated as "Target parameter name: send.thread.count; Adjustment direction: UP; Adjustment step size: 1", and A03_DOWN_1 is encapsulated as "Target parameter name: compression.level; Adjustment direction: DOWN; Adjustment step size: 1". Each parameter tuning instruction contains complete control information required for one parameter adjustment. The server aggregates all single parameter tuning instructions generated at all decision time steps in chronological order to form a tuning instruction sequence.

[0111] Subsequently, the server streamlines and merges the tuning instruction sequence. The server checks if there are tuning instruction pairs in consecutive decision time steps that target the same parameter name but have opposite adjustment directions. For example, if time step 2 is "adjust batch.size up one level" and time step 3 is "adjust batch.size down one level," the server determines that this instruction pair will cause oscillating adjustments and removes it from the tuning instruction sequence, thus obtaining a streamlined set of tuning instructions.

[0112] The server continues to perform dependency sorting on the streamlined set of tuning instructions. Based on the preset parameter dependencies between the target parameter names, the server rearranges the instructions according to the principle of prioritizing the adjustment of dependent parameters. For example, since `send.thread.count` affects log forwarding capacity and `batch.size` affects the size of a single batch, the server first arranges the instructions to "increase send.thread.count" and then arranges the instructions to "increase batch.size," ensuring that the target log system first improves its forwarding capacity and then expands the batch sending size. For scenarios where `compression.level` is related to CPU load, the server first arranges the instructions to "decrease compression.level" and then arranges the queue water level threshold adjustment instructions to prioritize relieving compression computation pressure.

[0113] Finally, the server establishes a communication link with the target log system's parameter control interface. This interface serves as the remote parameter adjustment entry point exposed by the target log system. The server sends the set of tuning instructions to the parameter control interface in a rearranged order. The interface parses each tuning instruction, including the target parameter name, adjustment direction identifier, and adjustment step size, and drives the target log system to perform corresponding dynamic parameter adjustments, such as increasing the number of log forwarding threads, reducing the compression level, shortening the disk flushing interval, or increasing the batch sending number of records. After the adjustment is completed, the target log system continues to output feedback logs, which the server uses to enter the next round of closed-loop tuning.

[0114] In this embodiment of the invention, the method further includes:

[0115] After the original log text stream is converted into a sequence of log message units with time-ordered tags, a log stream burst pattern detection process is performed on the sequence of log message units with time-ordered tags to obtain burst intensity distribution characteristics. These burst intensity distribution characteristics are used to participate in the generation process of the system performance state transition characteristics.

[0116] The step of performing log stream burst pattern detection processing on the log message unit sequence with time sequence markers to obtain burst intensity distribution characteristics includes:

[0117] Based on the forward time interval attribute marker carried by each log message unit in the log message unit sequence with time sequence marker, a time interval change waveform within the preset collection period is constructed, wherein the horizontal axis of the time interval change waveform is the sequential number and the vertical axis of the time interval change waveform is the forward time interval attribute value.

[0118] The time interval change waveform is subjected to sliding window cutting processing. A cutting window with a preset window width slides along the direction of the sequential numbering, and a local time interval segment is cut out at each sliding window position.

[0119] Calculate the statistical characteristic value of the time interval density for each local time interval segment. The statistical characteristic value includes the maximum time interval, the minimum time interval, and the average time interval.

[0120] The deviation of the mean time interval of each local time interval segment is compared with the global mean time interval, and the deviation value of the mean of each local time interval segment is calculated. The global mean time interval is the arithmetic mean of all forward time interval attribute values ​​within the entire preset acquisition period.

[0121] Local time interval segments whose mean deviation exceeds a preset deviation threshold are marked as log stream burst candidate regions, and adjacent and continuous log stream burst candidate regions are merged into log stream burst segments. Each log stream burst segment has a segment start number and a segment end number.

[0122] Perform log level distribution statistics processing on the log message units within each log stream burst segment, and calculate the distribution vector of the proportion of the number of log message units at each level within the log stream burst segment to the total number of log message units within the log stream burst segment.

[0123] The proportion distribution vector is weighted and summed using a preset log level sensitivity weight vector to obtain the burst intensity metric of the burst segment of the log stream. Different log levels in the log level sensitivity weight vector correspond to different weight coefficients.

[0124] The burst intensity data points are formed by combining the segment start number and burst intensity metric of each log stream burst segment. The burst intensity data points of all log stream burst segments are arranged in order of segment start number to form burst intensity distribution characteristics.

[0125] The burst intensity distribution feature is concatenated into the system performance state transition feature, so that the system performance state transition feature simultaneously includes the vector offset information of the system performance state and the burst intensity distribution information of the log stream.

[0126] In this embodiment of the invention, for example, after completing the semantic awareness and parameter association parsing processing of the log message unit sequence, the server performs response trajectory association mining processing on the parameter configuration change response features to generate a parameter synergy impact graph, and uses this graph for subsequent parameter tuning decision deduction. The parameter configuration change response features contain multiple delay association mappings, such as response relationships like "increasing batch.size leads to a decrease in queue level", "increasing send.thread.count leads to a decrease in sending delay", and "decreasing compression.level leads to a decrease in compression time but an increase in network transmission volume".

[0127] The server first groups all latency-related mappings in the parameter configuration change response characteristics according to the parameter name entity. Latency-related mappings belonging to `batch.size` are grouped into the `batch.size` parameter response trajectory subset, latency-related mappings belonging to `send.thread.count` are grouped into the `send.thread.count` parameter response trajectory subset, and latency-related mappings belonging to `compression.level` are grouped into the `compression.level` parameter response trajectory subset. Subsequently, the server sorts the latency-related mappings in each parameter response trajectory subset according to the adjustment operation trigger time, forming a parameter adjustment response event sequence for the corresponding parameter name entity. For example, if `batch.size` is increased once at 10:15:20 and again at 10:18:40, the server forms a response event sequence for `batch.size` according to the chronological order.

[0128] The server continues to segment the response pattern of each parameter adjustment event sequence. By detecting time interval jumps between adjacent adjustment operation trigger points, the server divides the event sequence into multiple single response segments. Each single response segment corresponds to a complete response cycle from parameter adjustment to a stable log output pattern. For example, after adjusting batch.size from 500 to 800, log throughput increases, queue water level decreases, and stabilizes after twenty seconds; the server divides this process into a single response segment.

[0129] For each single response segment, the server extracts the initial response duration, peak change amplitude, and stable recovery duration to form a ternary response profile descriptor. The initial response duration represents the time when the log metric first changes after parameter adjustment; for example, after `send.thread.count` increases, the sending latency begins to decrease within five seconds. The peak change amplitude represents the maximum magnitude of the metric change; for example, the queue level drops from 92% to 65%. The stable recovery duration represents the time required for the log output to stabilize again; for example, the queue level remains below 70% after twenty-five seconds. The server aggregates all ternary response profile descriptors corresponding to the same parameter name entity, calculates the average initial response duration, average peak change amplitude, and average stable recovery duration, and obtains the response characteristic summary vector for that parameter name entity.

[0130] Next, the server compares the similarity of the response characteristic summary vectors of entities with different parameter names, calculating the cosine similarity value between any two parameters. When the cosine similarity between batch.size and send.thread.count is higher than a preset similarity threshold, the server marks them as a similar parameter pair in terms of response pattern, indicating that both have similar effects on throughput improvement and queue level reduction. The server also performs temporal cross-analysis on the latency correlation mapping of different parameters. When the response event sequence of batch.size and the response event sequence of send.thread.count have a temporal overlap region, and both exhibit increased throughput and decreased queue level within the overlap region, the server establishes a cooperative enhancement correlation edge between them; when the decrease in compression.level and the increase in batch.size exhibit an inverse restraining trend of increased network transmission pressure and decreased compression pressure within the same time period, the server establishes a cooperative suppression correlation edge between them.

[0131] Finally, the server constructs a parameter co-influence graph, using parameter name entities as nodes and similar parameter pairs, co-enhancing edges, and co-inhibiting edges as connections. Each node's attributes contain a summary vector of the corresponding parameter's response characteristics. When the reinforcement learning policy generation network performs parameter tuning decision-making, the server inputs this parameter co-influence graph into the graph attention perception module. The graph attention perception module extracts the topological features of the co-influence between parameters and integrates these features into the joint state decision representation vector. Through this processing, the server can prioritize parameter combinations with co-enhancing effects when selecting subsequent parameter adjustment actions and avoid generating mutually inhibiting tuning actions.

[0132] In this embodiment of the invention, the method further includes:

[0133] During the parameter tuning decision-making process of calling the pre-built reinforcement learning policy generation network, the parameter adjustment action sequence output by the reinforcement learning policy generation network is subjected to temporal consistency constraint processing to generate a smooth adjustment action sequence. The smooth adjustment action sequence replaces the parameter adjustment action sequence to generate the tuning instruction set.

[0134] The step of performing temporal consistency constraint processing on the parameter adjustment action sequence output by the reinforcement learning policy generation network to generate a smooth adjustment action sequence includes:

[0135] Obtain the parameter adjustment action sequence output by the reinforcement learning policy generation network at all decision time steps within the preset acquisition period, and expand the parameter adjustment action sequence into an action type identifier chain according to the decision time step order;

[0136] The action type identifier chain is subjected to high-frequency oscillation segment identification processing. A preset oscillation mode detection template is used to perform sliding matching on the action type identifier chain. The continuous decision time step subsequence that matches the oscillation mode detection template is marked as a high-frequency oscillation decision segment. The oscillation mode detection template is used to match the pattern of repeatedly switching parameters to adjust the action type between adjacent decision time steps.

[0137] The number of parameter adjustment action type identifiers in each high-frequency oscillation decision segment is statistically processed to count the frequency of occurrence of parameter upward adjustment action identifiers and parameter downward adjustment action identifiers in that high-frequency oscillation decision segment.

[0138] Compare the frequency of occurrence of the parameter upward adjustment action flag with the frequency of occurrence of the parameter downward adjustment action flag. If the frequency of occurrence of the parameter upward adjustment action flag is greater than the frequency of occurrence of the parameter downward adjustment action flag, then the high-frequency oscillation decision segment is classified as the upward adjustment dominant oscillation segment, and vice versa.

[0139] The direction locking process is performed on the dominant oscillation segment of the upward adjustment, and all parameter adjustment action type identifiers in the dominant oscillation segment of the upward adjustment are uniformly replaced with parameter upward adjustment action identifiers to eliminate the oscillation fluctuations in the dominant oscillation segment of the upward adjustment.

[0140] The direction locking process is performed on the downward adjustment dominant oscillation segment, and all parameter adjustment action type identifiers in the downward adjustment dominant oscillation segment are uniformly replaced with parameter downward adjustment action identifiers to eliminate oscillation fluctuations in the downward adjustment dominant oscillation segment;

[0141] The action type identifiers of non-high frequency oscillation decision segments are smoothed out. The action type identifiers of two adjacent non-high frequency oscillation decision segments are checked to see if they are consistent. If they are inconsistent, an action identifier with unchanged parameters is inserted at the adjacent boundary position as a transition buffer action.

[0142] All segments after direction locking and all segments after smooth transition are spliced ​​together in the original decision time step order to obtain the action type identifier smooth chain;

[0143] The smooth chain of the action type identifier is mapped back to the format of the parameter adjustment action sequence to obtain the smooth adjustment action sequence, so that the smooth adjustment action sequence no longer contains the parameter adjustment action type identifier that is switched frequently and repeatedly.

[0144] In this embodiment, after sending the set of tuning instructions to the parameter control interface of the target log system, the server continues to monitor the target log system through a log collection probe, and performs attribution analysis on the feedback log text stream after performing dynamic parameter adjustment operations, generating a parameter adjustment attribution evaluation report. This report is used to correct the reward signal of the reinforcement learning strategy-generated network, making subsequent parameter tuning decisions more consistent with the actual operating effect of the target log system.

[0145] Specifically, the server continuously captures the feedback log text stream within a monitoring time window following the sending of the tuning command set. This monitoring time window is three minutes after the tuning command is executed. After executing commands such as "send.thread.count increase by one level", "compression.level decrease by one level", and "batch.size increase by one level", the target log system continuously outputs feedback logs such as "queue usage dropped to 68%", "batch send delay reduced to 120ms", "compression cost decreased", and "network payload increased slightly". The server transforms the feedback log text stream into a sequence of feedback log message units according to the aforementioned log segmentation, timestamp positioning, boundary detection, and content fragmentation methods, ensuring that each feedback log message unit carries a time sequence marker.

[0146] Subsequently, the server performs system performance status term extraction processing on the feedback log message unit sequence. The server extracts performance status terms such as sending latency, queue level, compression time, disk flushing congestion, network load, and log drop risk from the feedback log and generates performance indicator observation curves in chronological order. For example, the server records that the queue level decreased from 92% to 68%, the sending latency decreased from 280 milliseconds to 120 milliseconds, the compression time decreased from 80 milliseconds to 40 milliseconds, while network transmission volume slightly increased. This performance indicator observation curve is used to describe the trajectory of system performance status changes over time within the monitoring time window.

[0147] The server compares the observed performance indicator curve with the baseline performance level before the optimization command was executed, calculates the performance change at each monitoring time point, and forms a performance change trajectory curve. The baseline performance level before optimization was: queue level 92%, transmission latency 280 milliseconds, and compression time 80 milliseconds. After optimization, at a certain time point, the queue level was 70% and the transmission latency was 130 milliseconds. The server calculated that the queue level decreased by 22 percentage points and the transmission latency decreased by 150 milliseconds, and recorded this change in the performance change trajectory curve.

[0148] Next, the server timestamps the time points on the performance change trajectory curve with the execution times of each tuning instruction in the tuning instruction set. The server identifies a significant decrease in sending latency within ten seconds of executing "send.thread.count increase by one level," and assigns the performance change segment within that time period to that instruction; similarly, it identifies a decrease in compression time but a slight increase in network load after executing "compression.level decrease by one level," and assigns the corresponding change segment to that instruction. Through this alignment process, the server establishes a correspondence between each tuning instruction and the performance change segment it triggers.

[0149] The server further analyzes the trend of performance change segments corresponding to each tuning instruction. If the performance change continues to evolve along the preset optimization direction, such as a continuous decrease in sending latency, a continuous drop in queue level, and a reduction in disk flushing congestion, the server marks the tuning instruction as a positive attribution instruction. If the performance change continues to evolve against the preset optimization direction, such as a continuous increase in network load, an abnormal increase in CPU usage, or an increase in log drop risk, the server marks the tuning instruction as a negative attribution instruction. The server quantifies the magnitude of change in performance change segments, calculates the cumulative integral value of performance change within the segment, and uses it as a contribution metric for the corresponding tuning instruction's tuning effect. For example, a high contribution metric for "increase send.thread.count by one level" indicates that it makes a significant contribution to reducing sending latency and queue level.

[0150] The server extracts the names and adjustment step sizes of the target parameters corresponding to all negative attribution instructions, generating a list of negative sensitivity parameters. For example, if the server records that "increasing batch.size by two levels" leads to increased memory usage in the current state, it adds batch.size and the two-level step size to the negative sensitivity parameter list. Simultaneously, the server extracts the names and adjustment step sizes of the target parameters corresponding to all positive attribution instructions, generating a list of positive sensitivity parameters. For example, it adds "increasing send.thread.count by one level" and "decreasing compression.level by one level" to the positive sensitivity parameter list. Finally, the server merges the positive sensitivity parameter list, the negative sensitivity parameter list, and the contribution metrics of each tuning instruction to generate a parameter tuning attribution evaluation report, which is then input into the reward signal generator. The reward signal generator uses this report to increase the reward weight of positive sensitivity parameters and decrease the selection weight of negative sensitivity parameters, guiding the reinforcement learning policy generation network to prioritize tuning targets with stable positive contributions in subsequent decisions.

[0151] In this embodiment, while the server performs semantic perception and parameter association parsing on the log message unit sequence, it also performs log cascading fault symptom mining to generate a cascading fault propagation path map. This map is used to assist the reinforcement learning strategy generation network in identifying key fault nodes in advance and to preventively adjust relevant parameters.

[0152] The server first performs anomaly log pattern injection detection on the log message unit sequence. The server calls a pre-built anomaly pattern feature library to perform feature matching scans on each log message unit. The anomaly pattern feature library stores anomaly log feature patterns associated with historical failure events, such as "queue overflow warning," "diskwrite timeout," "collector heartbeat lost," "send retry exceeded," and "flush blocked over threshold." After detecting entries like "queue usage 98%," "flushtask blocked over 5s," and "send retry exceeded limit" in the log message unit sequence, the server marks the successfully matching log message units as anomaly symptom log units.

[0153] Subsequently, the server constructs an anomaly event timeline based on the time sequence markers carried by the anomaly symptom log units, arranging all anomaly symptom log units from earliest to latest according to their occurrence time, forming an anomaly event sequence. For example, the server detects an abnormal queue level at 10:15:03, disk flushing blockage at 10:15:12, and log forwarding retry exceeding the limit at 10:15:25, thus forming an anomaly event sequence consisting of queue anomaly, disk flushing anomaly, and forwarding anomaly.

[0154] The server continues to perform fault topic clustering on the sequence of anomalous event symptoms. The server calculates the semantic vector distance between each anomalous log unit and aggregates log units with a semantic vector distance less than a preset aggregation threshold into the same fault topic cluster. For example, "queue usage 98%", "queue overflow warning", and "buffer nearly full" are aggregated into a queue backlog fault topic cluster; "flush task blocked" and "disk write timeout" are aggregated into a disk flushing blocking fault topic cluster; and "send retry exceeded" and "collector connection timeout" are aggregated into a log forwarding anomaly fault topic cluster. Each fault topic cluster corresponds to a potential system fault type.

[0155] For each faulty topic cluster, the server performs a time density distribution analysis. The server counts the frequency of abnormal symptoms within the faulty topic cluster according to preset time segments, forming a time density distribution curve. For example, the frequency of queue backlog faulty topic clusters increases rapidly between 10:15:00 and 10:15:10, the frequency of disk flushing blockage faulty topic clusters increases between 10:15:10 and 10:15:20, and the frequency of log forwarding anomaly faulty topic clusters increases after 10:15:20. This time density distribution curve reflects the occurrence rhythm of different faulty topics within the collection period.

[0156] The server performs time-series cross-correlation analysis on the time density distribution curves of any two different fault subject clusters, calculating the cross-correlation function values ​​at different time offsets. The server detects that the queue backlog curve rises first, followed by the disk flushing blockage curve, and that both reach their maximum cross-correlation function values ​​at positive time offsets. Therefore, the server determines that the queue backlog fault subject cluster is the predecessor fault subject, and the disk flushing blockage fault subject cluster is the successor fault subject. The server constructs a directed cascading propagation edge between them, pointing from the queue backlog to the disk flushing blockage. The server further detects that the log forwarding anomaly frequency increases after disk flushing blockage, so it constructs another directed cascading propagation edge between disk flushing blockage and log forwarding anomalies.

[0157] The server uses the time offset corresponding to the maxima of the cross-correlation function as the propagation delay attribute of the directed cascading propagation edges; for example, the propagation delay from queue backlog to disk flushing blockage is 12 seconds. It also uses the normalized value of the maxima of the cross-correlation function as the propagation strength attribute; for example, the propagation strength is 0.86. Subsequently, the server constructs a cascading fault propagation path graph, using each fault theme cluster as a cascading fault node and all directed cascading propagation edges as directed connections between nodes. Each cascading fault node contains a set of abnormal symptom log units and a potential system fault type identifier, such as "queue backlog," "disk flushing blockage," and "log forwarding anomaly."

[0158] Finally, the server analyzes the topological centrality metrics of each cascading failure node in the cascading failure propagation path graph, identifying critical failure nodes at the intersection of multiple propagation paths. For example, a node that is blocked from flushing simultaneously receives queue backlog propagation and continues to point to log forwarding anomalies is identified as a critical failure node by the server. The server maps the potential system failure type corresponding to this critical failure node to an associated set of adjustable parameters, such as mapping flush blocking to flush.interval.ms, batch.size, and disk.flush.thread.count, and mapping queue backlog to queue.high.watermark and send.thread.count, thereby generating a list of preventative tuning target parameters. During subsequent parameter tuning decision-making, the reinforcement learning policy generation network prioritizes generating preventative adjustment actions for the parameters in this list, thereby reducing the cascading propagation risk between queue backlog, flush blocking, and forwarding anomalies in advance.

[0159] In an embodiment of the invention, for example, during the parameter tuning decision-making process of the reinforcement learning policy generation network, the server introduces an adversarial perturbation state space exploration mechanism to generate a robust parameter adjustment action sequence. This mechanism enables the target log system to maintain stable parameter adaptive adjustment capabilities even when there are sudden fluctuations in log load, collection delays, or deviations in state observation.

[0160] Specifically, when the server performs state encoding processing on the system performance state transition characteristics, it constructs an adversarial perturbation generator. This adversarial perturbation generator is used to superimpose perturbation vectors onto the system performance state transition characteristics to generate perturbation state representations. For example, if the target log system's current logs show that the queue level has increased from 70% to 92% and the transmission latency has increased from 120 milliseconds to 280 milliseconds, the server has extracted system performance state transition characteristics representing queue congestion and increased transmission latency. To simulate scenarios such as a sudden increase in business requests, a momentary amplification of log output, or delayed arrival of some logs, the server inputs these system performance state transition characteristics into the adversarial perturbation generator.

[0161] The adversarial perturbation generator includes a perturbation intensity prediction subnetwork. The server uses this subnetwork to calculate the perturbation direction vector and magnitude based on the gradient information of the current system performance state transition characteristics, and generates an adversarial perturbation vector. For example, if the server identifies that queue level, transmission delay, and disk flushing congestion characteristics have a significant impact on action selection, the perturbation intensity prediction subnetwork generates more pronounced perturbations in these dimensions, resulting in a state where "queue level continues to rise, transmission delay fluctuates further, and the probability of disk flushing congestion increases." Subsequently, the server element-wise superimposes the adversarial perturbation vector with the system performance state transition characteristics to obtain a perturbation state representation. This perturbation state representation is used to simulate the system performance state observation errors that occur in the target logging system under scenarios of sudden fluctuations in log load.

[0162] The server inputs the perturbation state representation into the state encoding module of the reinforcement learning policy generation network, which generates a perturbation system state representation vector. Then, the server concatenates and fuses this vector with the parameter response context vector to form a joint perturbation state decision representation vector, which is then input into the policy inference module for forward inference. The policy inference module outputs a candidate set of adversarial parameter adjustment actions, containing the probability values ​​of each candidate action in the perturbation scenario. For example, in the original state, the probability of "increasing send.thread.count" is 0.38, and the probability of "decreasing compression.level" is 0.21; in the perturbation state, the server re-obtains the probabilities of each action to determine whether the network decision is stable.

[0163] The server further measures the difference in action probability distribution between the candidate set of adversarial parameter adjustment actions and the candidate set of original parameter adjustment actions generated under the condition of no superimposed perturbation, calculates the distribution distance metric between the two action probability distributions, and uses this distribution distance metric as an adversarial sensitivity index. When the action probability changes significantly before and after the perturbation, it indicates that the reinforcement learning policy generation network is more sensitive to log state fluctuations; when the same tuning direction is still preferentially selected before and after the perturbation, it indicates that the current policy has good robustness.

[0164] The server feeds back the adversarial sensitivity index to the perturbation intensity prediction subnetwork, driving it to adjust its generation strategy for perturbation direction vectors and perturbation magnitudes, continuously optimizing the perturbations in directions that better expose policy instability. Simultaneously, the reinforcement learning policy generation network is forced to learn to maintain consistent action decisions under perturbation conditions. For example, in a perturbation scenario of a sudden increase in queue level, the network consistently chooses to "increase the number of log forwarding threads" instead of frequently switching to "keep it unchanged" or "reduce the number of batches sent."

[0165] After reaching a preset number of adversarial training iterations, the server extracts a candidate set of parameter adjustment actions generated under multiple adversarial perturbation states. It then aggregates the mean of the action selection probability values ​​for each candidate action to obtain a robust action probability distribution. Based on this robust action probability distribution, the server performs policy sampling, selecting candidate actions with the smallest fluctuation in action selection probability and a high average probability across multiple perturbation scenarios as robust parameter adjustment actions. For example, "increases send.thread.count" consistently maintains a high probability under multiple perturbation scenarios, so the server identifies it as a robust parameter adjustment action.

[0166] Finally, the server sequentially combines the robust parameter adjustment actions selected at all decision time steps to form a robust parameter adjustment action sequence, and uses this sequence to replace the original parameter adjustment action sequence in the tuning instruction set generation process. The resulting tuning instruction set maintains a stable tuning direction even during sudden fluctuations in log load, reducing frequent reverse parameter tuning caused by state observation jitter, and improving the adaptive adjustment stability of the target log system under high-concurrency log writing scenarios.

[0167] This application further provides a lightweight log storage, retrieval, and alarm monitoring architecture. This architecture uses VictoriaLogs as the baseline component for log storage and retrieval, leveraging its keyword-based filtering and independent field storage mechanism to reduce storage, memory, and data read volumes during log querying. The system can combine VictoriaMetrics or Prometheus to collect performance metrics, alarm information, and error logs during server cluster operation, and push alarms via AlertManager. The architecture includes a server cluster, node-level alarm collection components, a VictoriaLogs single-node log service, and an internal cluster communication network. This solution enables high-throughput log access, low-resource-consumption querying, and multi-level alarm management under lightweight deployment conditions, improving the efficiency of operational fault location.

[0168] This invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned adaptive tuning method for log system performance parameters based on reinforcement learning. Figure 4 As shown, Figure 4This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. For illustrative purposes, the foregoing description is made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed.

Claims

1. A method for adaptive tuning of performance parameters of a log system based on reinforcement learning, characterized in that, The method includes: Acquire the raw log text stream generated by the target log system within a preset collection period, and convert the raw log text stream into a sequence of log message units with time sequence markers; Semantic awareness and parameter association parsing are performed on the log message unit sequence to obtain the system performance state migration characteristics and parameter configuration change response characteristics of the log message unit sequence within the preset collection period. The pre-built reinforcement learning strategy is invoked to generate a network to perform parameter tuning decision deduction on the system performance state transition characteristics and the parameter configuration change response characteristics, so as to obtain the parameter adjustment action sequence of the target log system; Based on the parameter adjustment action sequence, a set of tuning instructions is generated, which includes the name of the target parameter to be adjusted, the direction identifier of the adjustment, and the adjustment step level. The set of tuning instructions is then sent to the parameter control interface of the target log system to trigger the dynamic parameter adjustment operation.

2. The method according to claim 1, characterized in that, The process of acquiring the raw log text stream generated by the target log system within a preset collection period and converting the raw log text stream into a sequence of log message units with time-ordered markers includes: Deploy a log collection probe on the log output channel of the target log system to capture the raw log text stream within the preset collection period; The original log text stream is segmented into log lines, timestamps are located, and time order is sorted to obtain a log text line queue. Based on the characteristic character sequence of the starting line of the log record, the log text line queue is subjected to log bar boundary detection to obtain the basic log fragment unit sequence; For each basic log shard unit, time interval calculation and content fragmentation are performed to generate log message units carrying forward time interval attribute values ​​and sequentially arranged sequence numbers; The log message units are aggregated to form a sequence of log message units with time-order tags for subsequent semantic perception and parameter association parsing processing.

3. The method according to claim 2, characterized in that, The semantic awareness and parameter association parsing processing of the log message unit sequence yields the system performance state transition characteristics and parameter configuration change response characteristics corresponding to the log message unit sequence within the preset collection period, including: Obtain the sequence of log message units with time sequence markers, and input each log message unit in the log message unit sequence into a pre-trained log semantic perceptron, wherein the pre-trained log semantic perceptron is used to perform semantic vector encoding on a single log message unit; The pre-trained log semantic perceptron is used to perform semantic vector encoding on each log message unit to generate a fixed-dimensional log semantic vector corresponding to each log message unit, and all log semantic vectors are combined into a log semantic vector sequence. The log semantic vector sequence is subjected to context semantic association modeling processing. A context-aware window with a preset span is slid across the log semantic vector sequence, and the semantic association strength representation value between adjacent log semantic vectors within the context-aware window is calculated. The semantic association strength representation values ​​calculated within each context-aware window are arranged in the window sliding order to form a semantic association strength evolution trajectory that runs through the preset collection period. The semantic association strength evolution trajectory is used to reflect the continuous or abrupt characteristics of the log semantic internal logic. Extract the location information of semantic mutation breakpoints from the semantic association strength evolution trajectory, and divide the log semantic vector sequence into multiple semantically coherent log segments based on the location information of the semantic mutation breakpoints, with semantic consistency maintained within each semantically coherent log segment. For each semantically coherent log segment, system performance status words are extracted. A preset system performance status dictionary is used to perform word-level matching on the log semantic vectors in the semantically coherent log segment to extract performance status description words that represent the system performance status. All performance status description phrases extracted from the same semantically coherent log segment are semantically fused, and the system performance status representation vector of the semantically coherent log segment is generated by vector weighted averaging. The vector offset between the system performance status representation vectors of adjacent semantically coherent log segments is used as the system performance status transition feature. For each semantically coherent log segment, parameter configuration change detection processing is performed. The process detects whether the semantically coherent log segment contains descriptive words of parameter adjustment actions. If it contains descriptive words of parameter adjustment actions, the parameter name entity and parameter adjustment direction entity appearing in the context window of the descriptive words of the parameter adjustment actions are extracted. Associate the parameter name entity and the parameter adjustment direction entity with the forward time interval attribute tag of the semantically coherent log fragment to construct a delay association mapping between the parameter adjustment trigger time point and the parameter adjustment response effect; The delay association mappings of all semantically coherent log fragments are concatenated in chronological order to form parameter configuration change response features. These features characterize the response transmission pattern between parameter adjustment operations and changes in system log output mode in the target log system.

4. The method according to claim 1, characterized in that, The method of invoking a pre-built reinforcement learning policy generation network to perform parameter tuning decision deduction on the system performance state transition characteristics and the parameter configuration change response characteristics, thereby obtaining the parameter adjustment action sequence of the target log system, including: The system performance state transition features are input into the state encoding module to generate a system state representation vector. The parameter configuration change response features are input into the parameter response encoder to generate a parameter response context vector. By fusing the system state representation vector and the parameter response context vector, a joint state decision representation vector is obtained; and through the policy reasoning module and the action output module, action selection probability values ​​for candidate parameter adjustment actions are generated based on the joint state decision representation vector. Based on the probability value of the action selection, a target parameter adjustment action is selected, and an instant reward signal is generated by combining the observation results of the system log feedback after the target parameter adjustment action is executed; The reinforcement learning policy generation network is updated using the instant reward signal, and a sequence of parameter adjustment actions is output to generate a set of tuning instructions.

5. The method according to claim 1, characterized in that, The step of generating a set of tuning instructions based on the parameter adjustment action sequence, including the name of the target parameter, the adjustment direction identifier, and the adjustment step size, and sending the set of tuning instructions to the parameter control interface of the target log system to trigger the dynamic parameter adjustment operation, includes: Parse the parameter adjustment action type identifier corresponding to each decision time step in the parameter adjustment action sequence; Based on the parameter name resolution table, adjustment direction resolution table, and step size resolution table, the parameter adjustment action type identifier is mapped to the target parameter name, adjustment direction identifier, and adjustment step size, respectively. By encapsulating the adjustment target parameter name, the adjustment direction identifier, and the adjustment step size, a tuning instruction sequence is obtained; The continuous tuning instructions in the tuning instruction sequence that target the same parameter name and have opposite adjustment directions are simplified and rearranged according to a preset parameter dependency relationship to obtain a tuning instruction set. The set of tuning instructions is sent to the parameter control interface via a communication link to perform the corresponding dynamic parameter adjustment operation.

6. The method according to claim 4, characterized in that, The reinforcement learning policy generation network further includes a graph attention perception module, and the method further includes: After obtaining the parameter configuration change response characteristics, the delay association mapping is grouped according to the parameter name entity to obtain a subset of parameter response trajectories; The parameter response trajectory subset is sorted by time and divided into response segments. The initial reaction time, peak amplitude of change and stable recovery time are extracted to generate a response characteristic summary vector. Based on the similarity of response characteristic summary vectors of entities with different parameter names and the temporal cross-analysis of delayed association mapping, similar parameter pairs of response patterns, collaboratively enhanced association edges, and collaboratively suppressed association edges are determined. A parameter collaborative influence graph is constructed using parameter name entities as nodes and the similar parameter pairs of the response patterns, the collaborative enhancement association edges, and the collaborative inhibition association edges as connection relationships. The parameter synergy relationship graph is input into the graph attention perception module to extract the topological features of the synergy between parameters, and the topological features of the synergy between parameters are integrated into the joint state decision representation vector.

7. The method according to claim 1, characterized in that, The method further includes: After the set of tuning instructions is sent to the parameter control interface of the target log system, the feedback log text stream after the dynamic parameter adjustment operation is executed is continuously collected, and the feedback log text stream is converted into a sequence of feedback log message units. Extract system performance status phrases from the feedback log message unit sequence to generate performance indicator observation curves; The difference between the observed performance indicator curve and the baseline performance level before the optimization instruction is executed is compared to obtain the performance change trajectory curve. Align the performance change trajectory curve with the execution time of each optimization instruction using timestamps to determine the performance change segment corresponding to the optimization instruction. Based on the changing trends and magnitudes of the performance change segments, generate a list of positive sensitivity parameters, a list of negative sensitivity parameters, and a measurement of the contribution of optimization effect; The parameters are combined to form a parameter adjustment attribution evaluation report, which is then input into the reward signal generator to adjust the calculation weight of the instantaneous reward signal.

8. The method according to claim 1, characterized in that, The method further includes: During the semantic perception and parameter association parsing process of the log message unit sequence, the abnormal pattern feature library is used to perform feature matching scanning on the log message unit sequence to obtain abnormal symptom log units. The abnormal symptom log units are arranged according to time sequence to form an abnormal symptom event sequence; Based on semantic vector distance, the abnormal symptom event sequence is clustered into fault topic clusters to obtain fault topic clusters; Temporal density distribution analysis and time series cross-correlation analysis are performed on the fault theme clusters to determine the predecessor fault theme, successor fault theme, and directed cascade propagation edge; A cascaded fault propagation path graph is constructed using the fault theme cluster as cascaded fault nodes and the directed cascaded propagation edges as connection relationships. Based on the cascaded fault propagation path map, key fault nodes are identified, and the potential system fault type identifiers corresponding to the key fault nodes are mapped to the associated set of adjustable parameters to generate a list of preventive tuning target parameters for preventive adjustment of the reinforcement learning policy generation network.

9. The method according to claim 1, characterized in that, The method further includes: During the parameter tuning decision-making process of the reinforcement learning policy generation network, an adversarial perturbation state space exploration mechanism is introduced to generate a robust parameter adjustment action sequence. The robust parameter adjustment action sequence is used to enable the target log system to maintain the stability of parameter adaptive adjustment when facing sudden fluctuations in log load. The introduction of an adversarial perturbation state space exploration mechanism to generate robust parameter adjustment action sequences includes: In the process of state encoding processing of the system performance state transition features, an adversarial perturbation generator is constructed. The adversarial perturbation generator is used to superimpose adversarial perturbation vectors on the system performance state transition features to generate perturbation state representations. The system performance state transition features are input into the adversarial perturbation generator, which contains a perturbation intensity prediction sub-network. The perturbation intensity prediction sub-network calculates the perturbation direction vector and perturbation magnitude based on the gradient information of the current system performance state transition features, and generates an adversarial perturbation vector. The adversarial perturbation vector is superimposed on the system performance state transition feature element by element to obtain the perturbation state characterization, which is used to simulate the system performance state observation error that may occur in the target log system under the scenario of sudden fluctuation in log load. The perturbation state representation is input into the state encoding module of the reinforcement learning policy generation network, so that the state encoding module encodes the perturbation state representation to generate a perturbation system state representation vector. The disturbance system state representation vector and the parameter response context vector are concatenated and fused to form a disturbance joint state decision representation vector. The disturbance joint state decision representation vector is then input into the policy reasoning module for forward reasoning to obtain a candidate set of adversarial parameter adjustment actions. The candidate set of adversarial parameter adjustment actions includes the action selection probability value of each candidate parameter adjustment action under the adversarial disturbance scenario. The candidate set of adversarial parameter adjustment actions and the original candidate set of parameter adjustment actions generated under the condition of no superimposed adversarial perturbation vector are subjected to action probability distribution difference measurement processing, and the distribution distance metric between the two action probability distributions is calculated. The distribution distance metric is used as an adversarial sensitivity index. The adversarial sensitivity index is input into the perturbation intensity prediction subnetwork of the adversarial perturbation generator as a feedback signal, driving the perturbation intensity prediction subnetwork to adjust the generation strategy of the perturbation direction vector and the perturbation amplitude, so that the adversarial sensitivity index is optimized in the direction of improvement, thereby forcing the reinforcement learning strategy generation network to maintain the consistency of action decisions under adversarial perturbation conditions. After reaching the preset number of adversarial training iterations, multiple sets of adversarial parameter adjustment action candidate sets generated by the reinforcement learning policy generation network under adversarial perturbation are extracted. The action selection probability value of each candidate parameter adjustment action in the multiple sets of adversarial parameter adjustment action candidate sets is averaged and aggregated to obtain the robust action probability distribution. Based on the robust action probability distribution, a strategy sampling operation is performed, and the candidate parameter adjustment action with the smallest fluctuation amplitude of the action selection probability value under the adversarial perturbation scenario is selected as the robust parameter adjustment action. The robust parameter adjustment actions selected at all decision time steps are combined in sequence to form a robust parameter adjustment action sequence. The robust parameter adjustment action sequence is used to replace the original parameter adjustment action sequence, and the generation process of the tuning instruction set is input to ensure that the generated tuning instruction set maintains the stability of parameter tuning decisions when the target log system faces sudden fluctuations in log load.

10. A reinforcement learning-driven adaptive performance parameter tuning system for a log system, characterized in that, Includes at least one service node; The service node includes a storage unit and a computing unit; the storage unit is used to store program code. The computing unit is used to run the program code to execute the reinforcement learning-driven adaptive tuning method for log system performance parameters as described in any one of claims 1 to 9.