Log processing method, model training method, related device and storage medium
By training the log location model and fault repair model, the exception log is identified based on the probability distribution of the log token and the repair suggestions are generated, which solves the problem that log analysis tools in the existing technology cannot flexibly locate equipment failures, and achieves efficient fault location and repair.
Patent Information
- Application Number
- CN202410090384.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-22
- Publication Date
- 2025-07-22
AI Technical Summary
Existing log analysis tools are limited to mining predefined log events, and cannot perceive potential problems in the log, resulting in the inability to flexibly locate device failures or abnormal events.
By training the log positioning model, using long-stable logs to generate log tokens, learning the probability distribution of log tokens based on the Transformer architecture, predicting exception logs, and combining the fault repair model to generate repair suggestions to realize self-closed loop repair of non-code faults.
It improves the flexibility and accuracy of log location, can accurately identify abnormal logs, and perform fault analysis and repair locally on the client, reduces dependence on back-end support and improves fault repair efficiency.
Smart Images

Figure CN120353670A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a log processing method, a model training method, related devices, and a storage medium. Background Art
[0002] Log data is a widely available data resource used to record the runtime system states and key events in various software systems. Developers usually utilize log data to obtain system states, detect anomalies, and locate root causes. Currently, developers often use log localization tools to detect the logs generated in devices. Log localization tools need to define a finite number of fault events in advance based on experience, and then classify and mine problems based on the preset events. The design of such log analysis tools aims to automatically detect patterns, events, or anomalies in logs, thereby supporting system management, monitoring, and fault diagnosis and repair.
[0003] The core of a log analysis tool is a set of predefined rules formulated by system administrators or developers. The log analysis tool first parses the raw logs generated by the system, including operations such as splitting the logs into fields, extracting key information, and identifying keywords, for subsequent rule matching. The parsed log data is matched with the predefined rules. When the log data meets the conditions of a certain rule, the engine triggers corresponding actions, including operations such as logging the log to a specific file, generating an alert, triggering a script, or notifying the administrator. Once the log analysis tool matches a specific event in the log, it can take predefined response measures.
[0004] However, log analysis tools are limited to mining predefined log events and detecting possible device failures or abnormal events based on templates, and are unable to perceive potential problems in logs beyond the predefined rules. Summary of the Invention
[0005] This application discloses a log processing method, a model training method, related devices, and a storage medium for locating abnormal logs and repairing faults.
[0006] In the first aspect of the present application, a log processing method is provided. The execution subject of this method can be a network device, or a component or device applied to the network device (such as a processor, a chip, or a chip system, etc.), or a logic module or software that can implement all or part of the functions of the network device. In this method, the network device pre-trains a log positioning model according to the long-term stable logs generated when the device is operating normally without faults. When an abnormal operation occurs in the network device and the user triggers the log positioning function, the network device converts the original log into log tokens, and the generation order of the log tokens is the same as that of the original log. The log tokens obtained by the network device at least include a first log token and a second log token, where the generation order of the second log token is after that of the first log token. The network device predicts a third log token that appears after the first log token according to the log positioning model. If the predicted third log token is different from the actually generated second log token, the network device determines that the target log corresponding to the second log token is an abnormal log.
[0007] In the embodiment of the present application, since the log positioning model is trained based on long-term stable logs, the third log token predicted by the log positioning model is the log token that should be generated when the device is operating normally without faults. If the actually generated second log token is different from the third log token, the abnormal log can be accurately located, improving the flexibility of log positioning.
[0008] In some optional embodiments, the second log token is the next log token of the first log token, and the third log token is the next log token of the first log token predicted by the network device based on the log positioning model.
[0009] In the embodiment of the present application, when both the second log token and the third log token are the next log token of the first log token, the comparison accuracy can be improved, thereby more accurately locating the abnormal log.
[0010] In some optional embodiments, the third log token includes N log tokens, and these N log tokens are the log tokens predicted by the network device based on the log positioning model and having an appearance probability higher than a preset value after the first log token, where N is a positive integer. Specifically, if the second log token is different from each of these N log tokens, the network device determines that the original log corresponding to the second log token is an abnormal log.
[0011] In some alternative embodiments, the network device also generates a first prompt according to the second log token, and this prompt is used as the input of the fault repair model to indicate the fault repair model to classify the faults of the target log. The network device receives the first feedback information from the fault repair model, and the first feedback information includes the fault summary and the fault type corresponding to the target log. The fault type is divided into code type and non-code type.
[0012] In the embodiments of the present application, the network device can also input the original text of the corresponding abnormal log into the fault repair model to obtain the corresponding fault summary and fault type, so that the user can directly perform fault analysis locally, improving the repair efficiency.
[0013] In some alternative embodiments, when the fault type of the target log is non-code type, the network device also generates a second prompt, and this second prompt is used as the input of the fault repair model to indicate the fault repair model to generate the repair suggestion corresponding to the target log. The network device receives the second feedback information from the fault repair model, and the second feedback information includes this repair suggestion.
[0014] In the embodiments of the present application, when the fault type of the abnormal log is non-code type, the network device can adjust or configure parameter modification according to the repair suggestion to complete the self-closed loop of the network device repair, improving the fault repair efficiency.
[0015] A model training method is proposed in the second aspect of the present application. In this method, the trainer inputs the long-term stable log generated by the device in the fault-free normal operation state into the model training device, and the model training device denoises the long-term stable log to obtain a structured log template. The log template is converted into multiple long-term stable log tokens, and a log positioning model is trained according to the multiple long-term stable log tokens.
[0016] In some alternative embodiments, at least the first long-term stable log token, the second long-term stable log token, and the third long-term stable log token are included in the multiple long-term stable log tokens. When the preceding token is the first long-term stable log token, the probability that the second long-term stable log token appears after the first long-term stable log token is the first probability, and the probability that the third long-term stable log token appears after the first long-term stable log token is the second probability. The model training device trains a log positioning model according to the first probability and the second probability, so that the log positioning model can predict the probability distribution of the log tokens that appear after the first long-term stable log token.
[0017] In the embodiments of the present application, the model training device locates the probability distribution of the next long - stable log token that appears after a long - stable log token appears through training logs, so that the log location model can locate abnormal logs according to the actually appeared log tokens and the predicted log tokens, without relying on pre - set rules, and improves the accuracy of abnormal log location.
[0018] In the third aspect of the present application, a model training method is proposed. In this method, the model training device obtains abnormal logs and corresponding repair suggestions, converts the abnormal logs into abnormal log tokens, and converts the repair suggestions into repair suggestion tokens. The model training device uses the abnormal log tokens and repair suggestion tokens as inputs to train a fault repair model.
[0019] In some alternative embodiments, the repair suggestion tokens include a first repair suggestion token and a second repair suggestion token. When the preceding token is an abnormal log token, the probability that the first repair suggestion token appears after the abnormal log token is the third probability, and the probability that the second repair suggestion token appears after the abnormal log token is the fourth probability. The model training device trains a fault repair model according to the third probability and the fourth probability, so that the fault repair model can predict the probability distribution of the repair suggestion tokens that appear after the abnormal log tokens.
[0020] In the embodiments of the present application, the model training device predicts the probability of the repair suggestion token corresponding to an abnormal log token after the abnormal log token appears by training the fault repair model, so that the fault repair model can obtain the corresponding repair suggestion token according to the input abnormal log token. The generated repair suggestions are not based on template predetermination but can vary according to different logs, improving the flexibility of abnormal log repair.
[0021] In the fourth aspect of the present application, a network device is proposed, including:
[0022] A conversion module, configured to convert the log original text into log tokens according to the log location model. The generation order of the log tokens is the same as the generation order of the log original text. The log location model is trained based on long - stable logs. The long - stable logs include multiple logs generated when the device operates normally without faults. The log tokens include a first log token and a second log token, and the second log token is after the first log token;
[0023] A prediction module, configured to predict a third log token according to the log location model. The third log token is after the first log token;
[0024] A determination module, configured to determine that the log original text corresponding to the second log token is the target log if the second log token is different from the third log token, and the target log belongs to the abnormal log.
[0025] Based on the fourth aspect, in some optional embodiments, the determination module is specifically configured to:
[0026] If the second log token is different from each of the N log tokens, determine that the log original text corresponding to the second log token is the abnormal log.
[0027] Based on the fourth aspect, in some optional embodiments, the network device further includes:
[0028] A generation module, configured to generate a first prompt according to the second log token, where the first prompt is used to instruct a fault repair model to perform fault classification on the target log, and the fault repair model is trained based on the abnormal log and the corresponding repair suggestions;
[0029] A receiving module, configured to receive first feedback information generated by the fault repair model, where the first feedback information includes a fault summary and a fault type corresponding to the second log, and the fault type includes code type and non-code type.
[0030] Based on the fourth aspect, when the fault type of the target log is non-code type,
[0031] The generation module is further configured to generate a second prompt, where the second prompt is used to instruct the fault repair model to generate a repair suggestion corresponding to the target log;
[0032] The receiving module is further configured to receive second feedback information generated by the fault repair model, where the second feedback information includes the repair suggestion.
[0033] A fifth aspect of the present application provides a model training device, including:
[0034] A denoising module, configured to denoise the long-term stable log to obtain a structured log template, where the long-term stable log includes multiple logs generated when the device operates normally without faults;
[0035] A conversion module, configured to convert the log template into multiple long-term stable log tokens;
[0036] A training module, configured to train a log positioning model according to multiple long-term stable log tokens.
[0037] Based on the fifth aspect, in some alternative embodiments, the multiple long-term stable log tokens include a first long-term stable log token, a second long-term stable log token, and a third long-term stable log token. The training module is specifically configured to:
[0038] When the previous token is the first long-term stable log token, obtain a first probability that the generation order of the second long-term stable log token appears after the first long-term stable log token and a second probability that the generation order of the third long-term stable log token appears after the first long-term stable log token;
[0039] Obtain a log positioning model based on the first probability and the second probability. The log positioning model is used to determine the probability of the long-term stable log token appearing after the first long-term stable log token.
[0040] A sixth aspect of the present application proposes a model training device, including:
[0041] An acquisition module, configured to acquire abnormal logs and repair suggestions, where the repair suggestions correspond to the abnormal logs;
[0042] A conversion module, configured to convert the abnormal logs into abnormal log tokens and convert the repair suggestions into repair suggestion tokens;
[0043] A training module, configured to train the correspondence between the abnormal log tokens and the repair suggestion tokens to obtain a fault repair model.
[0044] Based on the sixth aspect, in some alternative embodiments, the repair suggestion tokens include a first repair suggestion token and a second repair suggestion token. The training module is specifically configured to:
[0045] When the previous token is an abnormal log token, obtain a third probability that the generation order of the first repair suggestion token appears after the abnormal log token and a fourth probability that the generation order of the second repair suggestion token appears after the abnormal log token;
[0046] Obtain a fault repair model based on the third probability and the fourth probability. The fault repair model is used to determine the probability distribution of multiple repair suggestion tokens corresponding to the abnormal log tokens.
[0047] A seventh aspect of the present application proposes a model training device, which includes a processor, a memory, an input / output device, and a bus; computer instructions are stored in the memory; when the processor executes the computer instructions in the memory, the computer instructions are stored in the memory; when the processor executes the computer instructions in the memory, it is used to implement any implementation manner of the first aspect.
[0048] The eighth aspect of the present application proposes a model training device, which includes: a processor, a memory, an input / output device, and a bus; computer instructions are stored in the memory; when the processor executes the computer instructions in the memory, computer instructions are stored in the memory; when the processor executes the computer instructions in the memory, it is used to implement any implementation manner of the second aspect.
[0049] The ninth aspect of the present application proposes a model training device, which includes: a processor, a memory, an input / output device, and a bus; computer instructions are stored in the memory; when the processor executes the computer instructions in the memory, computer instructions are stored in the memory; when the processor executes the computer instructions in the memory, it is used to implement any implementation manner of the third aspect.
[0050] The tenth aspect of the present application proposes a chip system, which includes a processor and an input / output port. The processor is used to implement the processing functions involved in the method described in any of the first, second, or third aspects above, and the input / output port is used to implement the transceiver functions involved in the method described in any of the first, second, or third aspects above.
[0051] In a possible design, the chip system further includes a memory, which is used to store program instructions and data for implementing the functions involved in the method described in any of the first, second, or third aspects above.
[0052] The chip system can be composed of chips or can include chips and other discrete devices.
[0053] The eleventh aspect of the present application proposes a computer-readable storage medium, including instructions, which when run on a computer, cause the computer to execute the method described in the foregoing first aspect, or the method described in the foregoing second aspect, or the method described in the foregoing third aspect.
[0054] The twelfth aspect of the present application proposes a computer program product containing instructions, which when run on a computer, cause the computer to execute the method described in the foregoing first aspect, or the method described in the foregoing second aspect, or the method described in the foregoing third aspect.
[0055] The beneficial effects from the fourth aspect to the twelfth aspect can be understood by referring to the beneficial effects of the first, second, and third aspects and their corresponding implementation manners, and will not be elaborated here specifically. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a schematic diagram of an embodiment of the system architecture in an embodiment of the present application;
[0057] Figure 2 It is a schematic diagram of an application scenario for the log processing method applicable in the embodiments of the present application;
[0058] Figure 3 It is a schematic diagram of an embodiment of the model training method in the embodiments of the present application;
[0059] Figure 4 It is a schematic diagram of an embodiment of the probability distribution of the log positioning model prediction in the embodiments of the present application;
[0060] Figure 5 It is a schematic diagram of another embodiment of the model training method in the embodiments of the present application;
[0061] Figure 6 It is a schematic diagram of an embodiment of the probability distribution of the fault repair model prediction in the embodiments of the present application;
[0062] Figure 7 It is a schematic diagram of an embodiment of the log processing method in the embodiments of the present application;
[0063] Figure 8 It is a schematic diagram of an embodiment of the model training device in the embodiments of the present application;
[0064] Figure 9 It is a schematic diagram of another embodiment of the model training device in the embodiments of the present application;
[0065] Figure 10 It is a schematic diagram of an embodiment of the network device in the embodiments of the present application;
[0066] Figure 11 It is a schematic diagram of another embodiment of the network device in the embodiments of the present application;
[0067] Figure 12 It is a schematic diagram of another embodiment of the model training device in the embodiments of the present application;
[0068] Figure 13 It is a schematic diagram of another embodiment of the model training device in the embodiments of the present application;
[0069] Figure 14 It is a schematic diagram of an embodiment of the chip in the embodiments of the present application. Detailed implementation manners
[0070] The embodiments of the present application provide a log processing method, a model training method, related devices and a storage medium, which are applied to the field of communication technologies, and can enable a network device to locate abnormal logs based on long-term stable logs rather than based on predefined rules, so as to locate abnormal logs more flexibly.
[0071] The embodiments of the present application will be described below in conjunction with the accompanying drawings. As is known to those of ordinary skill in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0072] The terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, and this is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0073] In the present application, "for indicating" may include for direct indication and for indirect indication. When describing that a certain indication information is used to indicate A, it may include that the indication information directly indicates A or indirectly indicates A, and does not mean that A must be carried in the indication information.
[0074] In addition, the specific indication methods may also be various existing indication methods, such as, but not limited to, the above-mentioned indication methods and their various combinations, etc. The specific details of various indication methods can refer to the prior art and will not be elaborated herein. As described above, for example, when it is necessary to indicate multiple information of the same type, there may be a situation where the indication methods of different information are different. In the specific implementation process, the required indication method can be selected according to specific needs. The embodiments of the present application do not limit the selected indication method. In this way, the indication methods involved in the embodiments of the present application should be understood to cover various methods that can enable the party to be indicated to obtain the information to be indicated.
[0075] In the embodiments of the present application, descriptions such as "when...", "in the case of...", "if" and "if" all refer to that the device will perform corresponding processing under a certain objective situation, which does not limit the time, and does not require the device to have a judgment action when implementing, nor does it mean that there are other limitations.
[0076] The embodiments of the present application are related to the related applications of large language models (LLMs). To better understand the solutions of the embodiments of the present application, the related terms and concepts of LLMs that may be involved in the embodiments of the present application will be introduced below.
[0077] (1) LLM;
[0078] LLM is a deep learning model designed to understand and generate natural language text. These models are trained on large amounts of text data and learn how to generate reasonable sentences or paragraphs by analyzing the structure and patterns of language. At the core of LLM are deep neural networks, especially the self-attention mechanism and the transformer architecture, which enable LLM to capture the dependencies between words in the input sequence and generate contextually relevant outputs.
[0079] (2) token;
[0080] A token is usually used to represent a discrete unit in language. In LLM, tokens play an important role as the basic units for the model to process and understand natural language. In LLM, tokens can cover multiple levels, including words, punctuation marks, numbers, or other language elements. These tokens form the basis of the input text and are the cornerstone for the model to perform semantic analysis and generate text. By splitting the text into a series of tokens, LLM can better capture and understand the grammar, semantics, and context information of the language.
[0081] When processing natural language, LLM first splits the input text into a series of tokens. These tokens can be words, phrases, or symbols, and can also depend on the specific task and the settings of the model. Tokens also have semantic meanings, and each token represents a concept, entity, or expression in language. By analyzing the semantics and context information of tokens, the model can understand the internal logic and structure of the language, thereby generating more accurate and contextually appropriate text content. In addition, there are some technical details involved in tokens in LLM. For example, in the preprocessing stage, the input text needs to be tokenized, splitting the continuous text into discrete tokens. In addition, the tokens need to be encoded, converting them into numerical vectors that the model can process.
[0082] (3) Encoding and segmentation;
[0083] Encoding and segmentation in LLM are two key technical steps. Encoding is to convert the input text into numerical vectors that the model can process, usually using word embedding techniques to convert each token into a fixed-length vector. This vector can capture the semantic information of the token and help the model understand the internal logic and structure of the language. Segmentation refers to the process of splitting the input text into a series of tokens. Depending on the specific task and the settings of the model, different tokenization methods can be selected, such as segmentation based on space, punctuation, and abbreviation rules. The purpose of segmentation is to convert the continuous text into a discrete sequence that the model can process for subsequent processing and calculation.
[0084] (4) Transformer architecture;
[0085] The Transformer architecture is a deep learning model originally proposed to solve sequence-to-sequence (seq2seq) learning problems. It uses self-attention mechanisms and multi-head attention mechanisms, which can capture the dependencies between words in the input sequence, thereby better understanding the context information of the language. This feature makes the Transformer architecture an ideal choice for LLMs. For example, the GPT series models are built based on the Transformer architecture. By stacking multiple Transformer layers, LLMs can learn more complex language patterns and context information. Each layer of the Transformer encodes the input sequence and passes it to the next layer for deeper learning and generation. This hierarchical structure enables large language models to generate coherent and meaningful text content.
[0086] (5) Prompt
[0087] A prompt is the initial input used to guide the model to generate text in an LLM. In natural language processing tasks, a prompt is usually a short text snippet or question used to stimulate the model to generate an output related to the topic. By providing a clear starting point, the prompt can help the model better understand the intention of the task and guide it to generate text content that meets the requirements. At the same time, the prompt can also provide context information to help the model understand the internal logic and structure of the language, thereby generating more coherent and meaningful text.
[0088] The following introduces the system architecture provided by the embodiments of this application.
[0089] As Figure 1 shown, the model training device 101 and the model training device 102 are used to train the log localization model and the fault repair model, and send the trained models to the client 103. When a running fault occurs in the client 103, the abnormal logs in the device logs are located according to the log localization model and the fault repair model, and the abnormal logs are classified for faults. If the abnormal log corresponds to a non-code class fault, the fault repair model generates a repair suggestion; if the abnormal log corresponds to a code class fault, it is recommended that the user send the abnormal log back to the backend 104 for repair by development and operation and maintenance personnel.
[0090] In the embodiments of the present application, both the log localization model and the fault repair model are generative models based on the Transformer architecture. It should be noted that in practical applications, the training data of the log localization model and the fault repair model can come from the client 103 or be provided by the trainers of the backend 104, and specific details are not limited here.
[0091] Figure 1 The client in [description] can be a network device, or a component or device applied to a network device (such as a processor, a chip, or a chip system, etc.), or a logical module or software that can implement all or part of the functions of a network device.
[0092] Figure 2 Figure [number] shows an application scenario applicable to the present application. As Figure 2 shown, the trainer trains the log localization model with stable logs and trains the fault repair model with abnormal logs and corresponding repair suggestions. The software tool for running the log localization model and the fault repair model is called the log localization and repair engine (hereinafter referred to as the engine for short). The engine reads the log information of the device at the client. When an abnormality / fault occurs during the operation / operation of the client, the engine is triggered. The engine detects the abnormal logs in the device logs and locates the root cause of the abnormal logs. Through the fault repair model, the engine performs root cause reasoning on the abnormal logs and generates repair guidance. The client obtains the root cause summary and repair guidance generated by the engine, and the operating user can operate according to the repair guidance or upload the abnormal logs back to the backend.
[0093] Please refer to Figure 3 , a model training method in the embodiments of the present application includes:
[0094] 301. Obtain stable log tokens;
[0095] The model trainer uses the stable logs generated by the device in the fault-free normal operation state as the input of the model training engine. Since some of the logs in the stable logs include garbled or invalid information, the model training device needs to denoise the stable logs to obtain high-quality log fragments, and obtain a structured log template based on these log fragments. The model training device encodes and splits the log template to obtain stable log tokens. Specifically, the log template is a string of characters. The model training device maps the string to a hash value and splits the hash value to construct semantically meaningful words.
[0096] In practical applications, there can be various ways of encoding and splitting by the model training device, and specific details are not limited here.
[0097] 302. Train to obtain a log localization model;
[0098] The model training device learns the probability distribution of long-term stable log tokens based on the transformer architecture according to multiple long-term stable log tokens when the device is operating normally. Specifically, as Figure 4 shown, Log A to Log J all represent log tokens corresponding to different logs. When Log A is the preceding token, the model training device learns the probability distribution of the log that comes after Log A. For example, the probability that Log B appears after Log A is 35%, the probability that Log C appears after Log A is 38%, and the probability that Log D appears after Log A is 17%. After obtaining the probability distribution after Log A, change the preceding token and continue to train the log positioning model. For example, when Log A and Log B are the preceding tokens, the probability that Log E appears after Log A and Log B is 62%, and the probability that Log F appears after Log A and Log B is 32%; when Log A and Log C are the preceding tokens, the probability that Log G appears after Log A and Log C is 53%, and the probability that Log H appears after Log A and Log C is 31%; when Log A and Log D are the preceding tokens, the probability that Log I appears after Log A and Log D is 56%, and the probability that Log J appears after Log A and Log D is 34%. And so on. After multiple rounds of iterative learning, the log positioning model is obtained. In the embodiment of the present application, each long-term stable log token corresponds to a complete log. This log positioning model can predict the probability of the next log token that conforms to the normal log under the condition that all preceding tokens are known. It should be understood that Figure 4 the number and probability of log tokens in
[0099] are only examples. In actual applications, the log positioning model will learn the probability distribution of all log tokens that appear after Log A, and log tokens with an appearance probability less than the preset value will not be determined as normal long-term stable log tokens by the log positioning model.
[0100] In this embodiment, the client does not need to locate the abnormal log according to predefined rules, but trains a log positioning model through the transformer architecture to obtain the context relationship between multiple logs, and predicts the probability distribution between log tokens to determine the abnormal log. Therefore, it is possible to locate unstructured log messages or abnormal logs with new abnormal patterns, improving the accuracy of abnormal log positioning.
[0101] Please refer to Figure 5, another model training method in the embodiment of the present application includes:
[0102] 501. Obtain abnormal log tokens and repair suggestion tokens;
[0103] The model trainer uses the abnormal log and the corresponding repair suggestion as the input of the model training device. The model training device encodes and segments the abnormal log and the repair suggestion to obtain abnormal log tokens and repair suggestion tokens. Specifically, the abnormal log tokens are similar to the aforementioned long-term stable log tokens, and the repair suggestion tokens are the tokens obtained by the model training device encoding and segmenting the description text corresponding to the abnormal log.
[0104] 502. Train to obtain a fault repair model;
[0105] The model trainer performs enhanced training on the LLM, inputs the abnormal log tokens and repair suggestion tokens into the LLM, and trains the probability distribution of the repair suggestion tokens corresponding to each abnormal log token based on the transformer architecture. As Figure 6 shown, the LLM uses the abnormal log tokens as the preceding tokens and learns the probability distribution of the repair suggestion tokens that appear after the abnormal log tokens. For example, the probability of repair suggestion A appearing after the abnormal log is 46%, and the probability of repair suggestion B appearing is 45%. After multiple rounds of iterative updating of the model parameters, the model training device obtains a fault repair model. Figure 6 In the above, repair suggestion A and repair suggestion B are only examples. In actual applications, the fault repair model will obtain the probabilities of all repair suggestions corresponding to the abnormal log, and only use the repair suggestion tokens with probabilities higher than the preset value as the repair suggestions corresponding to the abnormal log.
[0106] It should be understood that during inference, the input of the fault repair model is the same as that of the LLM, which is the prompt text constructed by the client and used to describe the abnormal log. The fault repair model obtains the corresponding feedback information according to this prompt, where the feedback information includes the fault type and repair suggestions.
[0107] In this embodiment, the fault repair model is trained through the transformer architecture, so that the generated repair suggestions are not based on template predetermination, but can vary according to different abnormal logs, improving the flexibility of fault repair.
[0108] Please refer to Figure 7 , a log processing method in the embodiment of the present application includes:
[0109] 701. Obtain log tokens;
[0110] When the device has an anomaly, the network device obtains the logs on the device and converts the device logs into log tokens according to a log template. Each log token corresponds to one log.
[0111] 702. Locate the anomaly log according to the log location model;
[0112] The network device inputs the log tokens into the log location model. The log location model starts from the first log token and predicts the log token that should appear next under normal circumstances. For example, the first log token is the pre-token, the second log token is the actually appeared log token, and the third log token is the log token predicted by the log location model. If the second log token is different from the third log token, the log corresponding to the second log token is the anomaly log. The log location model decodes the second log token to obtain the original text of the anomaly log.
[0113] 703. Obtain the fault type according to the fault repair model;
[0114] The network log constructs a prompt based on the original text of the anomaly log for input into the fault repair model. Specifically, the prompt constructed based on the anomaly log can be:
[0115] "Prompt: The device / system has a fault, and the anomaly log segment corresponding to this fault is {X}. Please give a description of the cause and manifestation of this fault and classify the fault (code type or non-code type)."
[0116] Among them, {X} represents the original text of the anomaly log obtained by the log location model, and the fault types include code type problems and non-code type problems. The large fault repair model gives two parts of answers according to the Prompt: the fault and anomaly summary and the fault classification.
[0117] It should be understood that the text content of the prompt in this step is only an example, and there can be various construction methods for the prompt constructed based on the anomaly log, which are not specifically limited here.
[0118] 704. Determine the repair suggestions according to the fault type.
[0119] If the fault type is code type, the feedback information returned to the user includes the fault summary and repair suggestions. Among them, the repair suggestions can include, for example: uploading the logs, contacting the operation and maintenance personnel, etc.
[0120] If the fault type is non-code type, continue to construct the prompt:
[0121] "Prompt: The device / system has a fault, and the corresponding abnormal log fragment is {X}. The cause of the fault is a non-code problem on the client side. Please give repair suggestions for operations on the client side."
[0122] The fault repair model generates repair suggestions based on this prompt, such as adjusting the hardware device switch, adjusting parameter configurations, etc.
[0123] In this embodiment, through the log positioning model and the fault repair model, users do not need to perform a series of complex operations such as contacting customer service and back-end R & D for problem positioning. It can be done locally on the client side. At the same time, users can complete the repair loop on the client side through the repair guidance, improving the efficiency of fault repair.
[0124] The following describes the model training device and network device in the embodiments of the present application with reference to the accompanying drawings.
[0125] Please refer to Figure 8 , an embodiment of the model training device in the embodiments of the present application includes:
[0126] A denoising module 801, configured to denoise the long-term stable logs to obtain a structured log template. The long-term stable logs include multiple logs generated when the device is operating normally without faults;
[0127] A conversion module 802, configured to convert the log template into multiple long-term stable log tokens;
[0128] A training module 803, configured to train a log positioning model based on multiple long-term stable log tokens.
[0129] Optionally, the multiple long-term stable log tokens include a first long-term stable log token, a second long-term stable log token, and a third long-term stable log token. The training module 803 is specifically configured to:
[0130] When the previous token is the first long-term stable log token, obtain a first probability that the generation order of the second long-term stable log token appears after the first long-term stable log token and a second probability that the generation order of the third long-term stable log token appears after the first long-term stable log token;
[0131] Obtain a log positioning model based on the first probability and the second probability. The log positioning model is used to determine the probability that a long-term stable log token appears after the first long-term stable log token.
[0132] Please refer to Figure 9 , another embodiment of the model training device in the embodiments of the present application includes:
[0133] An acquisition module 901, configured to acquire an exception log and a repair suggestion, where the repair suggestion corresponds to the exception log;
[0134] A conversion module 902, configured to convert the exception log into an exception log token and convert the repair suggestion into a repair suggestion token;
[0135] A training module 903, configured to train the correspondence between the exception log token and the repair suggestion token to obtain a fault repair model.
[0136] Optionally, the repair suggestion token includes a first repair suggestion token and a second repair suggestion token. Specifically, the training module 903 is configured to:
[0137] When the preceding token is the exception log token, obtain a third probability that the generation order of the first repair suggestion token appears after the exception log token and a fourth probability that the generation order of the second repair suggestion token appears after the exception log token;
[0138] Obtain a fault repair model according to the third probability and the fourth probability. The fault repair model is used to determine the probability distribution of multiple repair suggestion tokens corresponding to the exception log token.
[0139] Please refer to Figure 10 , an embodiment of the network device in the embodiment of the present application includes:
[0140] A conversion module 1001, configured to convert the log original text into a log token according to a log positioning model. The generation order of the log token is the same as the generation order of the log original text. The log positioning model is trained based on stable logs. The stable logs include multiple logs generated when the device is operating normally without faults. The log token includes a first log token and a second log token, and the second log token is after the first log token;
[0141] A prediction module 1002, configured to predict a third log token according to the log positioning model. The third log token is after the first log token;
[0142] A determination module 1003, configured to determine that the log original text corresponding to the second log token is a target log if the second log token is different from the third log token, and the target log belongs to an exception log.
[0143] Optionally, the determination module 1003 is specifically configured to:
[0144] If the second log token is different from each of the N log tokens, determine that the log original text corresponding to the second log token is an abnormal log.
[0145] Optionally, the network device further includes:
[0146] A generation module 1004, configured to generate a first prompt according to the second log token, where the first prompt is used to instruct a fault repair model to perform fault classification on a target log, and the fault repair model is trained based on the abnormal log and the repair suggestion corresponding to the abnormal log;
[0147] A receiving module 1005, configured to receive first feedback information generated by the fault repair model, where the first feedback information includes a fault summary and a fault type corresponding to the second log, and the fault type includes code type and non-code type.
[0148] Based on the fourth aspect, when the fault type of the target log is non-code type,
[0149] The generation module 1004 is further configured to generate a second prompt, where the second prompt is used to instruct the fault repair model to generate a repair suggestion corresponding to the target log;
[0150] The receiving module 1005 is further configured to receive second feedback information generated by the fault repair model, where the second feedback information includes a repair suggestion.
[0151] Next, a network device provided by an embodiment of the present application will be introduced. Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of the network device provided by the embodiment of the present application.
[0152] The network device specifically includes:
[0153] A processor 1101, a memory 1102, an input / output unit 1103, and a bus 1104;
[0154] The processor 1101 is connected to the memory 1102, the input / output unit 1103, and the bus 1104;
[0155] The memory 1102 stores a program;
[0156] The processor 1101 executes the program in the memory 1102, so that the network device executes the method in the foregoing embodiment.
[0157] Next, a model training device provided by an embodiment of the present application will be introduced. Please refer to Figure 12 , Figure 12 which is a schematic structural diagram of the model training device provided by the embodiment of the present application.
[0158] The model training device specifically includes:
[0159] A processor 1201, a memory 1202, an input / output unit 1203, and a bus 1204;
[0160] The processor 1201 is connected to the memory 1202, the input / output unit 1203, and the bus 1204;
[0161] A program is stored in the memory 1202;
[0162] The processor 1201 executes the program in the memory 1202, so that the model training device executes the method in the foregoing embodiment.
[0163] Next, a model training device provided in an embodiment of the present application will be introduced. Please refer to Figure 13 , Figure 13 which is a schematic structural diagram of a model training device provided in an embodiment of the present application.
[0164] The model training device specifically includes:
[0165] A processor 1301, a memory 1302, an input / output unit 1303, and a bus 1304;
[0166] The processor 1301 is connected to the memory 1302, the input / output unit 1303, and the bus 1304;
[0167] A program is stored in the memory 1302;
[0168] The processor 1301 executes the program in the memory 1302, so that the model training device executes the method in the foregoing embodiment.
[0169] An embodiment of the present application also provides a computer-readable storage medium. A program is stored in the computer-readable storage medium. When the program runs on a computer, the computer executes the steps in the method described in the foregoing embodiment.
[0170] An embodiment of the present application also provides a model determination device, which can also be referred to as a digital processing chip or a chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is used to execute the method steps executed by the model training device shown in the foregoing embodiment.
[0171] The embodiments of the present application also provide a digital processing chip. Circuits for implementing the functions of the aforementioned processor 1201, processor 1301, or both processor 1201 and processor 1301, and one or more interfaces are integrated in the digital processing chip. When a memory is integrated in the digital processing chip, the digital processing chip can complete the method steps of any one or more of the foregoing embodiments. When a memory is not integrated in the digital processing chip, it can be connected to an external memory through a communication interface. The digital processing chip implements the actions performed by the model training device in the foregoing embodiments according to the program code stored in the external memory.
[0172] The embodiments of the present application also provide a computer program product, which, when running on a computer, causes the computer to execute the method steps described in the foregoing embodiments.
[0173] The model training device or model determination device provided by the embodiments of the present application may be a chip, which includes a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, or the like. The processing unit can execute the computer execution instructions stored in the storage unit to cause the chip in the server to execute the method described in the foregoing embodiments. Optionally, the storage unit is a storage unit inside the chip, such as a register, a cache, etc. The storage unit may also be a storage unit outside the chip in the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0174] Specifically, the foregoing processing unit or processor may be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0175] Exemplarily, please refer to Figure 14, Figure 14 This is a schematic structural diagram of a chip provided by an embodiment of the present application. The chip can be embodied as a neural network processor NPU 1400. The NPU 1400 is mounted on a main CPU (Host CPU) as a coprocessor, and tasks are assigned by the Host CPU. The core part of the NPU is the arithmetic circuit 1403. The arithmetic circuit 1403 is controlled by a controller 1404 to extract matrix data from a memory and perform multiplication operations.
[0176] In some implementations, the arithmetic circuit 1403 includes multiple processing units (process engine, PE) inside. In some implementations, the arithmetic circuit 1403 is a two-dimensional systolic array. The arithmetic circuit 1403 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1403 is a general matrix processor.
[0177] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 1402 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 1401 and performs a matrix operation with matrix B. The partial results or final results of the obtained matrix are stored in an accumulator 1408.
[0178] The unified memory 1406 is used to store input data and output data. The weight data is directly transported through a direct memory access controller (DMAC) 1405, and the DMAC transports it to the weight memory 1402. The input data is also transported to the unified memory 1406 through the DMAC.
[0179] A bus interface unit (BIU) 1410 is used for the interaction between the AXI bus and the DMAC and the instruction fetch buffer (IFB) 1409.
[0180] The bus interface unit 1410 (bus interface unit, BIU) is used for the instruction fetch buffer 1409 to obtain instructions from an external memory, and is also used for the storage unit access controller 1405 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0181] The DMAC is mainly used to transport the input data in the external memory DDR to the unified memory 1406, or transport the weight data to the weight memory 1402, or transport the input data to the input memory 1401.
[0182] The vector calculation unit 1407 includes multiple operation processing units, which, if necessary, further process the output of the operation circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as batch normalization, pixel-level summation, upsampling of the feature plane, etc.
[0183] In some implementations, the vector calculation unit 1407 can store the processed output vector into the unified memory 1406. For example, the vector calculation unit 1407 can apply a linear function and / or a non-linear function to the output of the operation circuit 1403, such as performing linear interpolation on the feature plane extracted by the convolutional layer, or for example, a vector of accumulated values to generate activation values. In some implementations, the vector calculation unit 1407 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1403, such as for use in subsequent layers in a neural network.
[0184] The instruction fetch buffer 1409 connected to the controller 1404 is used to store the instructions used by the controller 1404;
[0185] The unified memory 1406, the input memory 1401, the weight memory 1402, and the instruction fetch memory 1409 are all On-Chip memories. The external memory is private to this NPU hardware architecture.
[0186] Wherein, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the program of the above Figure 3 、 Figure 5 or Figure 8 method.
[0187] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. One can select some or all of the modules according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0188] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0189] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0190] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0191] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0192] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.
[0193] The terms "first", "second", "third", "fourth", etc. (if any) in the description, claims and the above drawings of the present application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that shown or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0194] Finally, it should be noted that the above is only the specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should be covered within the protection scope of the present application.
Claims
1. A log processing method, characterized in that, Including: Converting the original log into log tokens according to a log positioning model. The generation order of the log tokens is the same as that of the original log. The log positioning model is trained based on stable logs, and the stable logs include multiple logs generated when the device is operating normally without faults. The log tokens include a first log token and a second log token, and the second log token comes after the first log token. Predicting a third log token according to the log positioning model, and the third log token comes after the first log token. If the second log token is different from the third log token, determining the original log corresponding to the second log token as the target log, and the target log belongs to an abnormal log.
2. The log processing method according to claim 1, wherein, The second log token is the next log token after the first log token, and the third log token is the next log token after the first log token predicted by the log positioning model.
3. The log processing method according to claim 2, wherein The third log token includes N log tokens, and the probabilities of the N log tokens appearing after the first log token are all higher than a preset value, where N is a positive integer. The step of, if the second log token is different from the third log token, determining the original log corresponding to the second log token as an abnormal log, includes: If the second log token is different from each of the N log tokens, determining the original log corresponding to the second log token as an abnormal log.
4. The log processing method according to any one of claims 1 to 3, characterized in that, The method further includes: Generating a first prompt according to the second log token, where the first prompt is used to instruct a fault repair model to classify the target log, and the fault repair model is trained based on the abnormal log and the corresponding repair suggestions. Receiving a first feedback message generated by the fault repair model, where the first feedback message includes a fault summary and a fault type corresponding to the second log, and the fault type includes code type and non-code type.
5. The abnormal log processing method according to claim 4, wherein, When the fault type of the second log is non-code type, the method further includes: Generating a second prompt, where the second prompt is used to instruct the fault repair model to generate a repair suggestion corresponding to the target log. Receiving a second feedback message generated by the fault repair model, where the second feedback message includes the repair suggestion.
6. A model training method, characterized in that, The method includes: Denosing the stable logs to obtain a structured log template, where the stable logs include multiple logs generated when the device is operating normally without faults. Converting the log template into multiple stable log tokens. Training a log positioning model according to the multiple stable log tokens.
7. The model training method according to claim 6, wherein The multiple long - stable log tokens include a first long - stable log token, a second long - stable log token, and a third long - stable log token. The log location model trained according to the multiple long - stable log tokens includes: When the preceding token is the first long - stable log token, obtain a first probability that the generation order of the second long - stable log token appears after the first long - stable log token and a second probability that the generation order of the third long - stable log token appears after the first long - stable log token; Obtain the log location model according to the first probability and the second probability. The log location model is used to determine the probability of the long - stable log token appearing after the first long - stable log token.
8. A model training method, characterized in that, The method includes: Obtain an abnormal log and a repair suggestion, where the repair suggestion corresponds to the abnormal log; Convert the abnormal log into an abnormal log token and convert the repair suggestion into a repair suggestion token; Train the corresponding relationship between the abnormal log token and the repair suggestion token to obtain a fault repair model.
9. The model training method according to claim 8, wherein The repair suggestion token includes a first repair suggestion token and a second repair suggestion token. The training of the corresponding relationship between the abnormal log token and the repair suggestion token to obtain a fault repair model includes: When the preceding token is the abnormal log token, obtain a third probability that the generation order of the first repair suggestion token appears after the abnormal log token and a fourth probability that the generation order of the second repair suggestion token appears after the abnormal log token; Obtain the fault repair model according to the third probability and the fourth probability. The fault repair model is used to determine the probability distribution of multiple repair suggestion tokens corresponding to the abnormal log token.
10. A network device, characterized in that, Includes: A conversion module for converting a log original text into log tokens according to a log location model. The generation order of the log tokens is the same as that of the log original text. The log location model is trained based on long - stable logs. The long - stable logs include multiple logs generated when the device operates normally without faults. The log tokens include a first log token and a second log token, and the second log token is after the first log token; A prediction module for predicting a third log token according to the log location model. The third log token is after the first log token; A determination module for determining that the log original text corresponding to the second log token is a target log if the second log token is different from the third log token. The target log belongs to an abnormal log.
11. The network device according to claim 10, characterized in that, The determination module is specifically used for: If the second log token is different from each of the N log tokens, determine that the log original text corresponding to the second log token is an abnormal log.
12. The network device according to claim 10 or 11, characterized in that, The network device further includes: A generation module, configured to generate a first prompt according to the second log token, where the first prompt is used to instruct a fault repair model to perform fault classification on the target log, and the fault repair model is trained based on the abnormal log and the repair suggestion corresponding to the abnormal log; A receiving module, configured to receive first feedback information generated by the fault repair model, where the first feedback information includes a fault summary and a fault type corresponding to the second log, and the fault type includes code type and non-code type.
13. The network device according to any one of claims 10 to 12, characterized in that, When the fault type of the second log is non-code type, The generation module is further configured to generate a second prompt, where the second prompt is used to instruct the fault repair model to generate a repair suggestion corresponding to the target log; The receiving module is further configured to receive second feedback information generated by the fault repair model, where the second feedback information includes the repair suggestion.
14. A model training device, characterized in that, Includes: A denoising module, configured to denoise a long-term stable log to obtain a structured log template, where the long-term stable log includes a plurality of logs generated when the device operates normally without faults; A conversion module, configured to convert the log template into a plurality of long-term stable log tokens; A training module, configured to train a log positioning model according to the plurality of long-term stable log tokens.
15. The model training device according to claim 14, wherein The plurality of long-term stable log tokens include a first long-term stable log token, a second long-term stable log token, and a third long-term stable log token. Specifically, the training module is configured to: When the preceding token is the first long-term stable log token, obtain a first probability that the generation order of the second long-term stable log token appears after the first long-term stable log token and a second probability that the generation order of the third long-term stable log token appears after the first long-term stable log token; Obtain the log positioning model according to the first probability and the second probability, where the log positioning model is used to determine the probability that a long-term stable log token appears after the first long-term stable log token.
16. A model training device, characterized in that, Includes: An acquisition module, configured to acquire an abnormal log and a repair suggestion, where the repair suggestion corresponds to the abnormal log; A conversion module, configured to convert the abnormal log into an abnormal log token and convert the repair suggestion into a repair suggestion token; A training module, configured to train the correspondence between the abnormal log token and the repair suggestion token to obtain a fault repair model.
17. The model training device according to claim 16, wherein The repair suggestion token includes a first repair suggestion token and a second repair suggestion token. Specifically, the training module is configured to: When the prefix token is the exception log token, obtain a third probability that the generation order of the first repair suggestion token appears after the exception log token and a fourth probability that the generation order of the second repair suggestion token appears after the exception log token; Obtain the fault repair model according to the third probability and the fourth probability, where the fault repair model is used to determine the probability distribution of a plurality of the repair suggestion tokens corresponding to the exception log token.
18. A network device, characterized in that, Comprising: A processor and a memory, the processor being coupled to the memory; The memory is used to store programs; The processor is configured to execute the program in the memory, so that the network device executes the method according to any one of claims 1 to 5.
19. A model training device, characterized in that, Comprising: A processor and a memory, the processor being coupled to the memory; The memory is used to store programs; The processor is configured to execute the program in the memory, so that the model training device executes the method according to any one of claims 6 to 7.
20. A model training device, characterized in that, Comprising: A processor and a memory, the processor being coupled to the memory; The memory is used to store programs; The processor is configured to execute the program in the memory, so that the model training device executes the method according to any one of claims 8 to 9.
21. A computer-readable storage medium, comprising instructions that, when run on a computer, cause the computer to execute the method according to any one of claims 1 to 5, or the method according to any one of claims 6 to 7, or the method according to any one of claims 8 to 9.
22. A computer program product containing instructions that, when run on a computer, cause the computer to execute the method according to any one of claims 1 to 5, or the method according to any one of claims 6 to 7, or the method according to any one of claims 8 to 9.