A software defect root cause automatic localization method combining operation log and machine learning
Patent Information
- Application Number
- CN202611000909.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]然而,嵌入式软件持续迭代更新、业务场景动态切换,新型隐性缺陷与连锁故障频发,给软件缺陷根因溯源带来极大技术挑战
1、本发明通过设计日志缺陷相关度融合算法与设备缺陷热度动态统计算法,实现了日志缺陷价值的自适应精准量化判定,解决了传统嵌入式日志缺陷判定规则僵化、固定阈值适配性差,难以匹配设备动态运行工况的问题。传统技术采用静态判定标准,无法适配软件迭代、工况波动与异常频次变化,易出现隐性缺陷漏判、高频异常误判的问题。本发明依托两组算法协同调控,从特征相似度、时序位置、模块故障属性、硬件约束四维维度自适应调整权重,精准量化每条日志的缺陷溯源价值,可有效捕捉隐性缺陷前置特征、过滤无效日志,从源头提升日志筛选精度与动态场景适配能力。
Smart Images

Figure CN122838151A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer software fault detection technology, and more specifically, to a method for automatically locating the root causes of software defects by combining runtime logs and machine learning. Background Technology
[0002] With the rapid iteration of embedded technology, various embedded terminal devices have been widely adopted in fields such as vehicle control, industrial automation, energy harvesting, and IoT edge sensing. These devices are mostly localized and operate independently, often in a state of unattended, uninterrupted operation. The built-in software of these devices undertakes core functions such as data acquisition, logic control, and status scheduling, and its operational stability directly determines the overall operational quality and safety of the entire system.
[0003] However, the continuous iteration and updates of embedded software, the dynamic switching of business scenarios, and the frequent occurrence of new hidden defects and cascading failures pose significant technical challenges to the root cause tracing of software defects. Traditional technologies, which use fixed log parsing templates and fixed threshold filtering rules, cannot adapt to the new log formats and anomaly types generated by software iterations, nor can they match the dynamically fluctuating operating states of devices and the lifespan degradation of Flash memory. Existing log processing models have inherent contradictions: writing all logs will exacerbate memory erase and write wear and shorten hardware lifespan, while fixed rule-based simplified filtering will lose a large number of logs related to hidden anomalies and new defects, resulting in incomplete tracing data. Incomplete log data can only rely on manual experience or simple rules to troubleshoot faults, making it difficult to uncover the correlation features of hidden faults and accurately locate lagging and cascading software defects. This not only significantly increases the difficulty of operation and maintenance troubleshooting and prevents localized intelligent offline location, but also leads to repeated failures, restricting the development of intelligent, accurate, and long-term troubleshooting technologies for embedded device software defects. In view of this, this invention proposes an automatic root cause localization method for software defects that combines operating logs and machine learning. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the prior art, adapt to practical needs, and provide an automatic root cause localization method for software defects that combines runtime logs and machine learning, aiming to solve the problems mentioned in the background art.
[0005] To address the aforementioned technical problems, this invention provides the following technical solution: an automatic root cause localization method for software defects combining runtime logs and machine learning, comprising the following steps: S100, Log Collection and Structured Parsing: Real-time collection of multi-dimensional operation logs of embedded software, splitting log fixed fields and dynamic fields, extracting core operation features and converting them into standardized log fields, and intelligently generating new templates based on log keywords to achieve dynamic iteration of the log template library; S200, Log Defect Value Quantification: Call the local historical defect sample library, rely on the log defect correlation fusion algorithm to match log and historical defect features, and combine the device defect heat dynamic statistical algorithm to adaptively quantify the defect tracing value of each log segment. S300, Adaptive Log Hierarchical Retention: Real-time monitoring of Flash memory operation status, quantification of hardware lifespan constraints through Flash remaining erase / write budget assessment algorithm, and classification of logs into three categories—critical, related, and ordinary—using a dual-constraint adaptive log hierarchical algorithm, and execution of differentiated hierarchical storage strategies. S400, Anomaly Sequence Construction and Regularization: After detecting a software anomaly, multiple types of retained logs are integrated to construct an anomaly event chain. A lightweight anomaly sequence regularization algorithm is used to clean and standardize the data, generating an anomaly event sequence that adapts to the model input. S500, Model Inference and Closed-Loop Iteration: Input the sequence into the machine learning model, select the optimal root cause through the root cause matching credibility evaluation algorithm, output the defect module, event type and root cause log fragment, and update the defect sample library at the same time to adaptively optimize the judgment and location rules of the whole process.
[0006] Preferably, the multi-dimensional operation log includes task operation log, interface call log, state transition log, exception prompt log and resource status log, which fully covers the software's normal operation, state transition and exception triggering throughout the entire life cycle of operation. Log structured parsing specifically includes: splitting the original log fields based on a preset log template library, removing valueless fixed constant fields, and retaining dynamically changing core fields; extracting nine core parameters: log time, log level, software module identifier, task identifier, event type, state transition information, error information, call return results, and resource status information, and uniformly converting them into standardized log fields; When the collected logs cannot match the existing templates, temporary templates are generated by extracting log keywords, calling objects, state change logic and error descriptions. After validating and removing invalid templates, the template library is updated to adapt to new log parsing scenarios after software iteration. Standardized log fields include module fields, event fields, status fields, error fields, call fields, and resource fields. These multi-dimensional fields work together to construct a complete software operation characteristic system.
[0007] Preferably, the formula for the log defect correlation fusion algorithm in step S200 is: ; in, The overall relevance of a single log entry's defect; The similarity coefficient for log features; Weights for log time sequence positions; This represents the module defect frequency coefficient. This is the penalty coefficient for Flash lifespan limitation; For log feature similarity dimension; For log time-series location dimension; For the dimension of frequent module defects; This is the dimension for Flash lifetime constraint penalties.
[0008] Preferably, the formula for the dynamic statistical algorithm of equipment defect heat in step S200 is: ; in, This refers to the real-time defect heat value of the equipment. The total number of all operation logs within a given statistical window period is calculated per unit. This refers to the duration of decay during abnormal times. It is a natural constant.
[0009] Preferably, the formula for the Flash remaining erase / write budget evaluation algorithm in step S300 is: ; in, The remaining valid erase / write budget for the memory; The maximum total erase / write capacity is factory-calibrated for the Flash memory; This refers to the aging and depreciation coefficient of the storage block. This counts the number of erase / write cycles the device has used.
[0010] Preferably, the formula for the dual-constraint log adaptive hierarchical algorithm in step S300 is: ; in, A threshold for dynamically determining high-value logs. The threshold for dynamically determining the value of logs. Threshold and The thresholds are adaptively adjusted bidirectionally based on the equipment defect heat level and the remaining Flash erase / write budget; The key logs are core log segments that are directly related to software exceptions, errors, and state transitions. The associated logs are auxiliary log segments that precede and follow the key logs and can reflect the evolution trend of software and hardware status. The ordinary log is a normal log segment that is repeatedly output and has no abnormal characteristics.
[0011] Preferred, differentiated tiered retention strategies specifically include: For the key logs that are crucial for defect localization, the entire original log content is completely preserved to ensure zero loss of core defect features. For the associated logs that assist in analyzing the abnormal evolution process, the full storage method is abandoned, and only standardized log features are retained. A four-dimensional context index of time series, module, call, and event is built to significantly reduce storage overhead and erase / write loss. For ordinary logs that have no value for defect analysis, the original data is not retained; only the statistical summary and count information of the scheduled operation are retained. Through a three-level differentiated retention mode, the dilemma of traditional log storage is solved, taking into account both the lifespan protection of Flash memory and the integrity of defect tracing logs.
[0012] Preferably, the formula for the lightweight outlier sequence normalization algorithm in step S400 is: ; in, Normalized and standardized outlier eigenvalues; These are the measured values of the original anomaly features; It represents the maximum value of the original feature within the abnormal event chain; It represents the minimum value of the original feature within the abnormal event chain; Weighting coefficients are used to filter features.
[0013] Preferably, the formula for the root cause matching reliability assessment algorithm in step S500 is: ; in, Root cause comprehensive credibility score; For software exception module matching degree; Matching degree of abnormal event type; The core root cause log feature matching degree; This is the adaptive weight adjustment coefficient.
[0014] Preferably, the closed-loop iteration mechanism specifically includes: The root cause localization results of software defects output by the machine learning model are compared and verified with the results of manual verification and confirmation and the results of equipment system recovery. The root cause software modules, root cause event types, root cause log fragments and corresponding standardized abnormal event sequences that have been verified are synchronously written into the local historical defect sample library to complete the iterative update of the sample library. Based on the updated sample data, the log defect correlation judgment rules, log hierarchical retention dynamic thresholds, and machine learning model inference parameters are adaptively optimized to continuously improve log processing accuracy and defect root cause localization capabilities, thereby achieving adaptive upgrades throughout the entire equipment lifecycle.
[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention achieves adaptive and accurate quantitative determination of log defect value by designing a log defect relevance fusion algorithm and a device defect heat dynamic statistical algorithm. This solves the problems of rigid traditional embedded log defect determination rules, poor adaptability of fixed thresholds, and difficulty in matching dynamic equipment operating conditions. Traditional technologies use static determination standards, which cannot adapt to software iterations, operating condition fluctuations, and changes in anomaly frequency, easily leading to missed detection of hidden defects and misjudgment of high-frequency anomalies. This invention relies on the coordinated control of two sets of algorithms to adaptively adjust weights from four dimensions: feature similarity, temporal position, module fault attributes, and hardware constraints. This accurately quantifies the defect tracing value of each log entry, effectively capturing pre-existing features of hidden defects and filtering invalid logs, thereby improving log screening accuracy and dynamic scenario adaptability from the source.
[0016] 2. This invention, through the design of an intelligent iterative update mechanism for log templates and an algorithm for assessing remaining Flash write / erase budget, achieves autonomous iterative log parsing and dynamic hardware lifespan protection based on accurate log selection. This further addresses the industry pain points of traditional templates being unable to adapt to software iterations and the severe, imperceptible hardware wear and tear. Traditional static log templates cannot adapt to new log formats generated by software updates, easily leading to parsing failures and the loss of defective logs. Simultaneously, they cannot detect Flash aging status, easily causing invalid writes and premature hardware degradation. This invention can automatically iterate and update the log template library, ensuring complete parsing of various logs. Simultaneously, it quantifies the remaining Flash write / erase budget and aging wear in real time, dynamically adjusting log write constraints. While ensuring effective log retention, it effectively delays hardware aging and reduces equipment maintenance costs.
[0017] 3. This invention achieves a balance between log traceability integrity and hardware lifespan protection by designing a dual-constraint adaptive log grading algorithm and a differentiated grading retention strategy. This effectively solves the core technical contradiction of traditional log storage: full writes damage hardware, and simplification filtering results in data loss. Traditional log storage models cannot simultaneously address hardware protection and defect traceability needs, exhibiting significant technical shortcomings. This invention integrates the dual constraints of defect traceability value and hardware lifespan. Through adaptive algorithmic classification, logs are divided into three categories: critical, related, and ordinary. Differentiated schemes—full retention, index retention, and summary retention—are matched to minimize Flash write / erase damage, fully preserve the evolutionary characteristics of the entire software anomaly process, and completely resolve the industry's storage dilemma, providing high-quality data support for accurate defect traceability.
[0018] 4. This invention achieves high-precision intelligent defect root cause localization in embedded low-computing-power scenarios by designing a lightweight anomaly sequence regularization algorithm, a root cause matching reliability assessment algorithm, and a closed-loop iterative optimization mechanism. This solves the problems of low efficiency, poor rule matching accuracy, and lack of autonomous optimization capabilities associated with traditional manual troubleshooting. Traditional defect tracing relies on human experience, making it difficult to identify cascading and hidden software faults, and the technical solutions are fixed and cannot be iteratively upgraded. This invention uses algorithms to perform lightweight regularization of anomaly data, adapting to embedded local low-computing-power offline inference scenarios. Simultaneously, it uses multi-dimensional weighted verification to select the optimal root cause result; relying on a closed-loop iterative mechanism, it continuously updates the sample library and optimizes algorithm and model parameters, achieving adaptive upgrades throughout the device's lifecycle and continuously improving the intelligence and accuracy of defect troubleshooting. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the overall workflow structure of the present invention.
[0020] Figure 2 This is a schematic diagram of the log structure parsing and template adaptive iteration process of the present invention.
[0021] Figure 3 This is a schematic diagram of the log adaptive hierarchical retention process structure under Flash hardware constraints according to the present invention.
[0022] Figure 4 This is a schematic diagram of the intelligent reasoning and closed-loop iterative optimization process structure of the present invention. Detailed Implementation
[0023] Example: Figures 1 to 4 As shown, the present invention relates to an automatic root cause localization method for software defects that combines runtime logs and machine learning, comprising the following steps: S100: Collects the operation logs generated during the operation of embedded software, performs structured parsing on the operation logs to obtain log features used to characterize the software's operating status. After the embedded device is powered on and running normally, it continuously monitors the software's full-dimensional operating status in real time and continuously collects various raw operation logs generated during the device's operation, covering the entire lifecycle data of normal device operation, critical anomalies, and fault triggering. S101. Real-time acquisition of various types of log data generated during the operation of embedded software, including task execution logs, interface call logs, state transition logs, exception message logs, and resource status logs. Specifically, the task execution log records the start, execution, suspension, and termination status of device control tasks; the interface call log records the call requests and response processes of internal and external interfaces of the device; the state transition log records the switching process of the device's working mode, operating level, and hardware / software status; the exception message log records explicit fault information such as program errors, task timeouts, and communication anomalies; and the resource status log records the fluctuation status of the device's memory, computing power, and storage resource usage. These multi-dimensional logs collaboratively cover both explicit and implicit characteristics of software defects. S102. Based on a pre-defined standardized log template library, the raw operation logs collected in real time are subjected to field identification and segmentation to distinguish between fixed constant fields and dynamically changing fields in the logs. Among them, fixed fields are identification information fixed in the device firmware and have no value for defect analysis; dynamic fields are status data, error data, and call data that change in real time with the software operation and are the core basis for subsequent defect judgment. Field segmentation achieves preliminary filtering of redundant information and reduces subsequent storage and computing overhead. S103. Accurately extract multi-dimensional core feature information from the split and purified log data, including log time, log level, software module identifier, task identifier, event type, state transition information, error information, call return results and resource status information, to build a comprehensive software operation status feature system and provide complete data support for defect correlation determination and anomaly evolution analysis. S104. Standardize and transform the extracted multidimensional and scattered information, and encapsulate it into a fixed format log field, specifically including module field, event field, status field, error field, call field and resource field, to eliminate the format differences of different types of logs and output logs of different software modules, and form standardized and structured log feature data to adapt to the input specifications of subsequent defect correlation calculation and machine learning models. S105. To adapt to dynamic scenarios such as embedded software iteration and upgrades, new business functions, and new exception types, an intelligent iterative update mechanism for log templates is set up. When the real-time collected operation logs cannot match any template in the existing log template library, the core keywords, calling objects, state change logic, and error descriptions in the logs are automatically extracted to intelligently generate temporary log templates. At the same time, the validity of the temporary templates is verified, invalid and redundant templates are eliminated, and compliant and valid temporary templates are merged and updated into the original log template library. This enables dynamic iterative upgrades of the template library, solving the technical defects of traditional fixed templates that cannot adapt to software iterations and cannot identify new exceptions, and ensuring the long-term effectiveness and comprehensiveness of log parsing.
[0024] S200. Based on the correspondence between log characteristics and historical software defect root causes, determine the defect correlation of each log segment. Adopt a dynamic adaptive defect correlation judgment mechanism. Combine the real-time defect occurrence frequency of the device, the historical defect distribution pattern, and the Flash storage life status to dynamically adjust the judgment criteria, so as to avoid minor defects being missed and high-frequency abnormal redundant storage problems caused by fixed thresholds. S201. Call the device’s locally fixed historical defect sample library. This sample library is continuously iterated and updated. It contains the device’s historical operation logs, historical abnormal event sequences, historical root cause software modules, historical root cause event types and historical root cause log fragments throughout the device’s entire life cycle. It fully records the pre-existing characteristics, triggering process and corresponding root cause information of various software defects of the device. S202. Perform a full-domain matching comparison between the standardized log features of the current single log segment and the log features corresponding to various root causes in the historical defect sample library, calculate the feature similarity, output accurate log feature matching results, and locate the basis of the association between the current log and historical defects. S203. Based on feature matching results, intelligently identify whether there are various software anomaly precursor features in the current log segment, including abnormal state transition, repeated call failure, abnormal fluctuation of resource status, continuous triggering of error messages, long-term task blocking, and frequent return of interface anomalies. This comprehensively captures the precursor signals of defect triggering. S204. The defect relevance of log segments is determined by a multi-dimensional fusion judgment logic. In order to achieve adaptive dynamic weight allocation, the relevance score is dynamically calculated by combining the real-time defect heat of the device and the lifespan constraint of the Flash hardware through the log defect relevance fusion algorithm. The defect analysis value of a single log is accurately quantified. The overall scale of the formula is unified as a dimensionless score to meet the compliance of mathematical operations. The formula for the log defect relevance fusion algorithm is: ; in, The overall relevance of a single log defect is dimensionless and is used to characterize the effective value of the current log segment in locating the root cause of software defects. , is the log feature similarity coefficient, which is dimensionless and used to characterize the degree of matching and fit between current log features and historical defect root cause log features; This is the log time-series position weight, which is dimensionless and used to characterize the positional priority of a log in an abnormal time-series chain. , is the module defect frequency coefficient, dimensionless, used to characterize the historical high-fault attribute of the software module to which the log belongs; , is the Flash lifetime limitation penalty coefficient, dimensionless, used to characterize the suppression weight of low-value logs under memory lifetime stress. For log feature similarity dimension; For log time-series location dimension; For the dimension of frequent module defects; This is the Flash lifespan constraint penalty dimension; addition is used to overlay the positive value weights of log feature matching, timing location, and module attributes; subtraction is used to deduct the negative interference weights caused by Flash lifespan stress; all parameters are dynamically adaptive and have no fixed values, and can be fine-tuned in real time according to device operating status, defect heat, and hardware lifespan status, resulting in the final output. The higher the score, the greater the reference value for locating the root cause of the defect in the current log. To accurately characterize the real-time abnormal operating status of equipment and dynamically adapt to the above-mentioned correlation judgment criteria, this invention simultaneously designs a dynamic statistical algorithm for equipment defect heat, quantifies the current overall frequency of abnormality of the equipment, provides a basis for threshold adaptive adjustment, and unifies the overall dimensions of the formula to dimensionless heat value. The formula for the dynamic statistical algorithm of equipment defect heat is: ; in, This is a dimensionless real-time defect heat value for the equipment, used to uniformly characterize the current overall frequency of anomalies in the equipment. The total number of abnormal logs within the statistical window period is counted per unit, in the form of logs. The total number of all operation logs within a given window period is counted per log entry. Abnormal time decay duration, in seconds, is used to characterize the time interval since the last abnormal device event; `<function>` is a natural constant, a fixed mathematical constant, used to realize the natural decay characteristic of defect heat over time; fractional operation is used to calculate the proportion of abnormal logs in the current window period; exponential decay operation is used to weaken the heat impact of long-standing abnormal events; through multi-dimensional calculations, the real-time abnormal heat of the device is dynamically output. The higher the value, the more frequently the device's software malfunctions. Based on the defect correlation output of the above algorithm With Defect Heat Dynamic fine-tuning log judgment criteria: When equipment experiences frequent anomalies and high defect intensity within a short period of time. When the temperature is high, the judgment criteria should be appropriately tightened, retaining only high-precision defect logs to avoid a large number of intermediate state logs consuming Flash erase / write lifespan; when the equipment operates stably for a long time and the defect heat is high... When the error rate is low, the judgment criteria can be appropriately relaxed to capture logs of minor anomalies, avoid missing hidden defects, and achieve dynamic adaptive optimization of defect relevance judgment.
[0025] S300 uses Flash memory erase / write lifetime, remaining erase / write budget, and storage occupancy status as core hard constraints, combined with the aforementioned defect correlation results of dynamic adaptive judgment, to implement differentiated and hierarchical log writing and retention strategies, thereby balancing Flash lifetime protection and defect log retention integrity from the root cause. S301. Real-time monitoring and collection of the hardware operation status of the dedicated log storage area of the Flash memory, including the storage space write occupancy ratio, the distribution of erase and write times of each storage block, the remaining available storage space, and the remaining erase and write budget, to fully understand the current lifespan consumption status and storage capacity of the Flash memory, and provide hardware constraints for the log writing strategy. S302. Based on the collected Flash memory hardware operating status data, extract the core constraint parameters that affect the log writing lifespan, remove invalid hardware status interference data, and retain only the effective parameters that are strongly correlated with log erasure and writing wear, storage capacity, and hardware aging, so as to provide an accurate data base for subsequent erasure and writing budget calculation and dynamic write constraint control. S303. By combining the core constraint parameters of Flash hardware with the real-time operating status of the device, a correlation mapping relationship between log writing and hardware lifespan loss is established. Based on the degree of hardware aging, remaining storage resources, and historical erase and write frequency, the basic constraint range of log writing is initially defined, providing pre-constraint conditions for the subsequent operation of the dual-constraint log adaptive hierarchical algorithm, and avoiding excessive hardware lifespan loss caused by indiscriminate log writing. To quantify the lifespan constraints of Flash hardware and achieve adaptive write strategy control, a Flash remaining erase / write budget evaluation algorithm is used to accurately calculate the effective erase / write resources that the memory can safely use at present, avoiding excessive erase / write losses. All variables in the formula are unified in the number of erase / write cycles, and the units are completely self-consistent. The formula for evaluating the remaining erase / write budget in Flash memory is: ; in, The remaining valid erase / write budget for the memory, in times, is used to characterize the number of valid erase / write operations that can be safely performed on a memory block at present. The maximum total erase / write capacity of the flash memory is factory-calibrated, in units of times, and is an inherent parameter of the device hardware. is the storage block aging and wear coefficient, which is dimensionless and used to characterize the degree of natural aging and wear caused by long-term erase and write iterations of hardware. The cumulative number of erase / write cycles consumed by the device is expressed in times, which is the total cumulative erase / write consumption counted in real time during device operation. Multiplication is used to deduct the inherent capacity loss caused by hardware aging. Subtraction is used to eliminate the erase / write resources consumed by the device. The calculation yields the effective erase / write budget that the device can currently use, providing hardware data support for the dynamic adjustment of the log writing strategy. Remaining valid erase / write budget based on algorithm output Dynamic matching of write constraint levels: when When there are sufficient numerical values and a large amount of free storage, write constraints can be appropriately relaxed to ensure the integrity of log information. when When the value is low, blocks are frequently erased and written, and storage resources are scarce, the write constraints are strictly tightened to minimize invalid erase and write operations and prioritize the protection of the memory's lifespan. To achieve accurate adaptive classification of logs, the algorithm integrates the dual constraints of defect value and hardware lifespan. It uses a dual-constraint log adaptive classification algorithm to automatically and accurately classify the three types of logs based on dynamic thresholds. There are no fixed standards, and all thresholds are dynamic dimensionless parameters. The formula for the dual-constraint log adaptive hierarchical algorithm is: ; in, The overall relevance of log defects is dimensionless and used to characterize the defect location value of a single log entry. A dimensionless threshold for dynamically determining high-value logs; used to define the minimum relevance standard for critical logs. The dynamic threshold for determining the value of logs is dimensionless and used to define the boundary standard for the relevance of related logs and ordinary logs; the segmentation condition judgment operation is used to automatically classify log levels based on the numerical range of log relevance. All are dynamic adaptive parameters, which can be adjusted according to the defect heat. Remaining erase / write budget The algorithm features two-way linkage adjustment without fixed manual settings; it can adaptively and accurately classify critical logs, related logs, and ordinary logs, providing a basis for determining differentiated log retention strategies. Based on the aforementioned dynamic algorithm judgment rules, all log fragments are adaptively and accurately divided into three categories: critical logs, related logs, and ordinary logs. The classification criteria are dynamically fine-tuned according to the device's operating status, defect popularity, and hardware lifespan, rather than being fixed: critical logs are core log fragments that directly correspond to software anomalies, system crashes, or task failures; related logs are auxiliary log fragments that are adjacent to critical logs in time sequence and can reflect the evolution trend of anomalies; and ordinary logs are regular log fragments that are from normal, repetitive device operation, have no abnormal characteristics, and have no defect correlation value. S304. Implement differentiated hierarchical write retention strategies for three types of logs to accurately adapt to Flash lifespan constraints: For critical logs indispensable for defect localization, retain all original log content to ensure zero loss of core defect features; for related logs that assist in analyzing abnormal processes, abandon the full write method, extract only core log features and establish time sequence, module, and call context indexes to significantly reduce the amount of data written and reduce Flash erase and write losses; for ordinary logs without defect value, do not write the original log content, only retain periodic statistical summaries and run count information to complete the running status record with the lowest hardware overhead. S305. Through the above-mentioned layered and differentiated retention methods, lightweight, high-value, and orderly log retention data is integrated, which not only completely avoids the problem of rapid Flash life decay caused by traditional full writing, but also avoids the problem of loss of key defect information caused by extreme deletion, providing accurate and effective data support for subsequent software anomaly root cause localization.
[0026] S400: The device monitors the running status of the software in real time. Once a software abnormal event such as abnormal restart, program freeze, communication interruption, control task failure, or abnormal resource exhaustion is detected, the root cause localization startup process is immediately triggered, and the complete abnormal evolution chain is reconstructed based on the hierarchically retained log data in Flash. S401. After a software anomaly is triggered, all log data persistently stored in the Flash memory is immediately read, including the original complete content of key logs, characteristics and context indexes of related logs, and statistical summaries of ordinary logs, to ensure that the retained data before and after the anomaly is retrieved in a complete time sequence without omissions or missing data. S402. Strictly follow the original time sequence of log generation to orderly splice and merge the three types of differentiated retained log data, and use context indexing to fill in the relationship between various logs to restore the complete running state before, during and at the moment of software exception. S403. Based on the software module call logic, state transition rules, error propagation paths, and resource state change trends contained in log data, sort out the correlation and causal logic of various log events, and construct an exception event chain that can completely and truthfully reflect the entire process of software exceptions from their inception and development to their final triggering. S404. Perform redundancy cleaning on the constructed abnormal event chain, remove duplicate log events, invalid state records, and unrelated normal operation data, retain the core abnormal event sequence that can support defect tracing, reduce the amount of data for subsequent model calculations, and adapt to embedded low computing power operation scenarios. To adapt to the low computing power and low memory limitations of embedded devices, redundant features are eliminated and core abnormal information is retained. A lightweight abnormal sequence regularization algorithm is used to standardize and lightweight process abnormal features, adapt to local machine learning inference, and the formula adopts normalized operation, with unified dimensions and no unit conflicts throughout the process. The formula for the lightweight outlier sequence normalization algorithm is: ; in, These are normalized and standardized outlier feature values, dimensionless, and are feature parameters in a unified format to adapt to the input of machine learning models. These are the measured values of the original abnormal characteristics, possessing the original physical dimensions, and are used to characterize the original abnormal parameters of equipment operation; The maximum value of all original features within the current exception event chain, and... Same physical dimensions; The minimum value of all original features within the current exception event chain, and... Same physical dimensions; The feature selection weight coefficients are dimensionless and used to differentiate the weight ratios of core abnormal features and redundant normal features; the difference operation is used to calculate the deviation between the original features and the feature extrema; the fractional normalization operation is used to eliminate the differences in the dimensions of the original features and achieve a unified standard for all features; the multiplication operation is used to complete the weighted strengthening of core features and the weakening and filtering of redundant features, and finally output standardized and lightweight model input features. The lightweight anomaly sequence normalization algorithm completes feature normalization and redundancy removal, unifies feature dimensions, compresses invalid feature dimensions, and converts it into anomaly event sequences that are compatible with the input format of machine learning models. This maximizes the preservation of core anomaly evolution features and is suitable for the limited computing resources of embedded devices.
[0027] S500. Input the abnormal event sequence into the machine learning model to obtain the candidate root cause matching results, and output the software defect root cause localization results based on the candidate root cause matching results. The software defect root cause localization results include the root cause software module, the root cause event type, and the root cause log fragment. S501: Input the standardized sequence of abnormal events into the offline-trained machine learning model. The model is adapted to the embedded low-computing-power and low-memory operating environment and can quickly complete feature parsing and inference operations. S502, the machine learning model performs comprehensive feature analysis on the abnormal event sequence, focusing on identifying the timing of events, the linkage and correlation of various software modules, the trend of changes in operating status, the attribution of error message types, and the causal relationship features of log context, and deeply mining the hidden root cause patterns behind the anomalies. S503. Based on the results of multi-dimensional feature analysis, the model accurately outputs candidate root cause software modules, candidate root cause event types, and corresponding key root cause log fragments, thus locking in the core scope of the defect occurrence. To address the issues of mixed results and low discrimination accuracy among multiple candidate root causes, a root cause matching credibility assessment algorithm is used to perform weighted verification and credibility scoring on multi-dimensional candidate root cause results, and to select the optimal root cause result. The entire process is dimensionless and has no fixed values. The formula for the root cause matching reliability assessment algorithm is as follows: ; in, The root cause comprehensive credibility score is dimensionless and is used to characterize the overall credibility of a single group of candidate root cause results. , is the software anomaly module matching degree, dimensionless, used to characterize the degree of matching between the features of candidate fault modules and historical fault modules; , is the anomaly event type matching degree, dimensionless, used to characterize the degree of type fit between the current anomaly event and the historical root cause event; The core root cause log feature matching degree is dimensionless and is used to characterize the degree of feature matching between retained key logs and historical fault root cause logs. The module matching adaptive weight adjustment coefficient is dimensionless and is used to adjust the scoring ratio of abnormal module matching in the software. This is an adaptive weighting coefficient for event types, dimensionless, used to adjust the scoring percentage of the matching degree of abnormal event types; The adaptive weight adjustment coefficient for log features is dimensionless and is used to adjust the scoring proportion of the matching degree of the core root cause log features; the multiplication operation is used to implement the weighted assignment of the matching degree of each dimension; the addition operation is used to superimpose the matching results of multiple dimensions and synthesize the final comprehensive credibility score. The higher the value, the more reliable the determination result of the corresponding candidate root cause; Comprehensive credibility score calculated based on algorithm Eliminate low-confidence invalid candidate results, retain high-confidence candidate root cause matching results, and generate an accurate and effective root cause candidate set; S504. Perform secondary verification and screening on each candidate root cause in the root cause candidate set. Combine the matching rules of historical defect samples, the propagation logic of abnormal events and the characteristics of module failures to eliminate invalid root causes with logical conflicts and mismatches, thereby further improving the accuracy and rationality of root cause location. S505, based on high-reliability matching results, finally outputs standardized software defect root cause localization results, clearly marking the root cause software module, root cause event type, and core root cause log fragment corresponding to the defect, providing accurate and practical basis for operation and maintenance personnel to troubleshoot and repair faults.
[0028] The embodiments disclosed in this invention are preferred embodiments, but are not limited thereto. Those skilled in the art can easily understand the spirit of this invention based on the above embodiments and make different extensions and variations, but as long as they do not depart from the spirit of this invention, they are all within the protection scope of this invention.
Claims
1. A method for automatically locating the root causes of software defects by combining runtime logs and machine learning, characterized in that, Includes the following steps: S100, Log Collection and Structured Parsing: Real-time collection of multi-dimensional operation logs of embedded software, splitting log fixed fields and dynamic fields, extracting core operation features and converting them into standardized log fields, and intelligently generating new templates based on log keywords to achieve dynamic iteration of the log template library; S200, Log Defect Value Quantification: Call the local historical defect sample library, rely on the log defect correlation fusion algorithm to match log and historical defect features, and combine the device defect heat dynamic statistical algorithm to adaptively quantify the defect tracing value of each log segment. S300, Adaptive Log Hierarchical Retention: Real-time monitoring of Flash memory operation status, quantification of hardware lifespan constraints through Flash remaining erase / write budget assessment algorithm, and classification of logs into three categories—critical, related, and ordinary—using a dual-constraint adaptive log hierarchical algorithm, and execution of differentiated hierarchical storage strategies. S400, Anomaly Sequence Construction and Regularization: After detecting a software anomaly, multiple types of retained logs are integrated to construct an anomaly event chain. A lightweight anomaly sequence regularization algorithm is used to clean and standardize the data, generating an anomaly event sequence that adapts to the model input. S500, Model Inference and Closed-Loop Iteration: Input the sequence into the machine learning model, select the optimal root cause through the root cause matching credibility evaluation algorithm, output the defect module, event type and root cause log fragment, and update the defect sample library at the same time to adaptively optimize the judgment and location rules of the whole process.
2. The method for automatic root cause localization of software defects combining runtime logs and machine learning according to claim 1, characterized in that, The multi-dimensional operation logs include task operation logs, interface call logs, state transition logs, exception prompt logs, and resource status logs, fully covering the software's normal operation, state transition, and exception triggering throughout the entire lifecycle of operation. Log structured parsing specifically includes: splitting the original log fields based on a preset log template library, removing valueless fixed constant fields, and retaining dynamically changing core fields; extracting nine core parameters: log time, log level, software module identifier, task identifier, event type, state transition information, error information, call return results, and resource status information, and uniformly converting them into standardized log fields; When the collected logs cannot match the existing templates, temporary templates are generated by extracting log keywords, calling objects, state change logic and error descriptions. After validating and removing invalid templates, the template library is updated to adapt to new log parsing scenarios after software iteration. Standardized log fields include module fields, event fields, status fields, error fields, call fields, and resource fields. These multi-dimensional fields work together to construct a complete software operation characteristic system.
3. The method for automatic root cause localization of software defects combining runtime logs and machine learning according to claim 1, characterized in that, The formula for the log defect correlation fusion algorithm in step S200 is: ; in, The overall relevance of a single log entry's defect; The similarity coefficient for log features; Weights for log time sequence positions; This represents the module defect frequency coefficient. This is the penalty coefficient for Flash lifespan limitation; For log feature similarity dimension; For log time-series location dimension; For the dimension of frequent module defects; This is the dimension for Flash lifetime constraint penalties.
4. The method for automatic root cause localization of software defects combining runtime logs and machine learning according to claim 1, characterized in that, The formula for the dynamic statistical algorithm of equipment defect heat in step S200 is as follows: ; in, This refers to the real-time defect heat value of the equipment. The total number of abnormal logs within a given statistical window period is calculated per unit. The total number of all operation logs within a given statistical window period is calculated per unit. This refers to the duration of decay during abnormal times. It is a natural constant.
5. The method for automatic root cause localization of software defects combining runtime logs and machine learning according to claim 1, characterized in that, The formula for evaluating the remaining erase / write budget of Flash in step S300 is as follows: ; in, The remaining valid erase / write budget for the memory; The maximum total erase / write capacity is factory-calibrated for the Flash memory; This refers to the storage block aging and loss coefficient. This counts the number of erase / write cycles the device has used.
6. The method for automatic root cause localization of software defects combining runtime logs and machine learning according to claim 3, characterized in that, The formula for the dual-constraint log adaptive hierarchical algorithm in step S300 is: ; in, A threshold for dynamically determining high-value logs. The threshold for dynamically determining the value of logs. Threshold and The thresholds are adaptively adjusted bidirectionally based on the equipment defect heat level and the remaining Flash erase / write budget; The key logs are core log segments that are directly related to software exceptions, errors, and state transitions. The associated logs are auxiliary log segments that precede and follow the key logs and can reflect the evolution trend of software and hardware status. The ordinary log is a normal log segment that is repeatedly output and has no abnormal characteristics.
7. The method for automatic root cause localization of software defects combining runtime logs and machine learning according to claim 1, characterized in that, Differentiated tiered storage strategies specifically include: For the key logs that are crucial for defect localization, the entire original log content is completely preserved to ensure zero loss of core defect features. For the associated logs that assist in analyzing the abnormal evolution process, the full storage method is abandoned, and only standardized log features are retained. A four-dimensional context index of time series, module, call, and event is built to significantly reduce storage overhead and erase / write loss. For ordinary logs that have no value for defect analysis, the original data is not retained; only the statistical summary and count information of the scheduled operation are retained. By adopting a three-tiered differentiated retention model, the dilemma of traditional log storage is solved, taking into account both the lifespan protection of flash memory and the integrity of defect traceability logs.
8. The method for automatic root cause localization of software defects combining runtime logs and machine learning according to claim 1, characterized in that, The formula for the lightweight anomaly sequence normalization algorithm in step S400 is: ; in, Normalized and standardized outlier eigenvalues; These are the measured values of the original anomaly features; It represents the maximum value of the original feature within the abnormal event chain; It represents the minimum value of the original feature within the abnormal event chain; Weighting coefficients are used to filter features.
9. The method for automatic root cause localization of software defects combining runtime logs and machine learning according to claim 1, characterized in that, The formula for the root cause matching reliability evaluation algorithm in step S500 is as follows: ; in, Root cause comprehensive credibility score; For software exception module matching degree; Matching degree for abnormal event types; For core root cause log feature matching degree; This is the adaptive weight adjustment coefficient.
10. The method for automatic root cause localization of software defects combining runtime logs and machine learning according to claim 1, characterized in that, The closed-loop iterative mechanism specifically includes: The root cause localization results of software defects output by the machine learning model are compared and verified with the results of manual verification and confirmation and the results of equipment system recovery. The root cause software modules, root cause event types, root cause log fragments and corresponding standardized abnormal event sequences that have been verified are synchronously written into the local historical defect sample library to complete the iterative update of the sample library. Based on the updated sample data, the log defect correlation judgment rules, log hierarchical retention dynamic thresholds, and machine learning model inference parameters are adaptively optimized to continuously improve log processing accuracy and defect root cause localization capabilities, thereby achieving adaptive upgrades throughout the entire equipment lifecycle.