Method and device for generating log and computing equipment
By fine-tuning the large model with instructions, log statements that meet user needs are generated, solving the problems of low efficiency and accuracy in log generation in existing technologies, and achieving efficient and accurate log generation.
Patent Information
- Application Number
- CN202410605233.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2025-11-18
AI Technical Summary
Existing automated logging tools cannot adjust the output according to user needs, resulting in low log generation efficiency and accuracy.
By employing artificial intelligence models, especially large models, and fine-tuning them through instructions, log statements that meet user needs are generated, including recommendations for location, level, and information.
It improves the efficiency and accuracy of log generation, saves labor costs, and meets users' personalized needs.
Smart Images

Figure CN120973615A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology, and more specifically, to a method, apparatus, and computing device for generating logs. Background Technology
[0002] Logs, also known as log data or log records, are mechanisms that record the code execution process in detail. For example, they record events or messages that occur during the operation of a system or application. High-quality log data helps developers quickly locate the location and cause of faults, and thus quickly fix them.
[0003] A logging statement, also known as logging code, is a line of code in the source code used to record program execution information. By writing and inserting logging statements into the source code, corresponding log data can be generated during program runtime. This log data is then collected, stored, and analyzed to support various system management and development tasks. Writing logging statements is an important task in ensuring software reliability; developers write logging statements in the source code to record system status, and operations and maintenance personnel use the logs generated by these statements to monitor and repair system faults. Currently, most unreported software faults are mainly caused by the lack of logging statements in critical locations, or by excessive logging statements that obscure valuable data. Therefore, designing necessary and correct logging statements during development is crucial for the subsequent operation and maintenance of the software system.
[0004] Related technical solutions propose automated logging tools to assist developers in automatically generating logging statements. In one solution, the developer inputs the source code for which logging statements are to be generated (the source code currently contains no logging statements), and the automated logging tool outputs the entire source code with the added logging statements. In another solution, the automated logging tool directly outputs the complete logging statement and its location within the source code. Both of these solutions have relatively limited application scenarios; they focus only on the complete logging statement and ignore user needs, failing to adjust the output according to user requirements. This results in low efficiency and accuracy in logging statement generation.
[0005] Therefore, improving the efficiency and accuracy of log generation has become a pressing technical problem that needs to be solved. Summary of the Invention
[0006] This application provides a method, apparatus, and computing device for generating logs, which can improve the efficiency and accuracy of log generation.
[0007] In a first aspect, a method for generating logs is provided, the method comprising: obtaining a first source code and a first instruction; obtaining one or more of the following based on the first source code and the first instruction using a first model: a recommended log statement for the first source code and the position information of the log statement in the first source code; and generating a corresponding log for the first source code based on the log statement.
[0008] The first instruction described above is used to instruct one or more of the following tasks: recommending a logging statement for the first source code, and recommending the location information of the logging statement in the first source code.
[0009] The above log statements include at least one of the following: log level, log message.
[0010] The input information of the first model includes the first source code and the first instruction, and the output information of the first model includes the log statement recommended for the first source code, and / or the location information of the log statement in the first source code.
[0011] The above logs can also be called log records or log data, etc.
[0012] The above technical solution can generate corresponding log statements according to user needs, which greatly saves labor costs and improves the efficiency and accuracy of log generation.
[0013] In conjunction with the first aspect, in some implementations of the first aspect, the first source code includes a first logging statement that lacks a logging level, a first instruction that indicates that the logging level is recommended for the first logging statement, and the output information of the first model includes the recommended logging level for the first logging statement.
[0014] In conjunction with the first aspect, in some implementations of the first aspect, the first source code includes a first logging statement that lacks logging information, the first instruction being used to indicate that the logging information is recommended for the first logging statement, and the output information of the first model including the logging information recommended for the first logging statement.
[0015] In conjunction with the first aspect, in some implementations of the first aspect, the first instruction is used to instruct a first log statement to be recommended at a target location of the first source code, and the output information of the first model includes the first log statement recommended for the first source code at the target location.
[0016] In conjunction with the first aspect, in some implementations of the first aspect, the first instruction is used to indicate the location information of a recommended first log statement for the first source code, and the output information of the first model includes the location information of the recommended first log statement.
[0017] In conjunction with the first aspect, in some implementations of the first aspect, the first instruction is used to indicate a recommended first logging statement and the location information of the first logging statement in the first source code, and the output information of the first model includes the recommended first logging statement and the location information of the first logging statement in the first source code.
[0018] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: fine-tuning the second model with instructions to obtain the first model, wherein the second model is a pre-trained model.
[0019] In conjunction with the first aspect, in some implementations of the first aspect, the second model is a large model.
[0020] Secondly, an apparatus for generating logs is provided, the apparatus comprising: an acquisition module, a recommendation module, and a generation module, wherein the acquisition module is used to acquire a first source code and a first instruction; the recommendation module is used to obtain one or more of the following based on the first source code and the first instruction using a first model: a log statement recommended for the first source code, and the position information of the log statement in the first source code; and the generation module is used to generate a corresponding log for the first source code based on the log statement.
[0021] The first instruction described above is used to instruct one or more of the following tasks: recommending a logging statement for the first source code, and recommending the location information of the logging statement in the first source code.
[0022] The above log statements include at least one of the following: log level, log message.
[0023] The input information of the first model includes the first source code and the first instruction, and the output information of the first model includes the log statement recommended for the first source code, and / or the location information of the log statement in the first source code.
[0024] In conjunction with the second aspect, in some implementations of the second aspect, the first source code includes a first logging statement that lacks a logging level, a first instruction that indicates a recommended logging level for the first logging statement, and the output information of the first model includes the recommended logging level for the first logging statement.
[0025] In conjunction with the second aspect, in some implementations of the second aspect, the first source code includes a first logging statement that lacks logging information, the first instruction being used to indicate that the logging information is recommended for the first logging statement, and the output information of the first model including the logging information recommended for the first logging statement.
[0026] In conjunction with the second aspect, in some implementations of the second aspect, the first instruction is used to instruct the recommendation of a first log statement at a target location of the first source code, and the output information of the first model includes the first log statement recommended for the first source code at the target location.
[0027] In conjunction with the second aspect, in some implementations of the second aspect, the first instruction is used to indicate the location information of a first log statement recommended for the first source code, and the output information of the first model includes the location information of the recommended first log statement.
[0028] In conjunction with the second aspect, in some implementations of the second aspect, the first instruction is used to indicate the recommendation of a first logging statement and the location information of the first logging statement in the first source code, and the output information of the first model includes the recommended first logging statement and the location information of the first logging statement in the first source code.
[0029] In conjunction with the second aspect, in some implementations of the second aspect, the device further includes: a training module for fine-tuning the second model to obtain the first model, wherein the second model is a pre-trained model.
[0030] In conjunction with the second aspect, in some implementations of the second aspect, the second model is a large model.
[0031] It should be understood that the beneficial effects of the second aspect and its various implementations can be found in the first aspect and its various implementations, and will not be elaborated here.
[0032] Thirdly, a computing device is provided, including a processor and a memory, and optionally, an input / output interface. The processor controls the input / output interface to send and receive information, the memory stores a computer program, and the processor retrieves and runs the computer program from the memory, causing the program to execute the method of the first aspect or any possible implementation thereof.
[0033] Optionally, the processor can be a general-purpose processor, which can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, integrated circuit, etc.; when implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. This memory can be integrated into the processor or located outside the processor and exist independently.
[0034] Fourthly, a computing device cluster is provided, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, such that the computing device cluster performs the method of the first aspect or any possible implementation thereof.
[0035] Fifthly, a chip is provided that acquires and executes instructions to implement the methods described in the first aspect and any implementation thereof.
[0036] Optionally, as one implementation, the chip includes a processor and a data interface, through which the processor reads instructions stored in the memory and executes the methods in the first aspect and any implementation thereof.
[0037] Optionally, as one implementation, the chip may further include a memory storing instructions, and the processor is used to execute the instructions stored in the memory. When the instructions are executed, the processor is used to perform the method in the first aspect and any implementation thereof.
[0038] In a sixth aspect, a computer program product containing instructions is provided, which, when executed by a computing device, cause the computing device to perform the methods described in the first aspect and any implementation thereof.
[0039] In a seventh aspect, a computer program product containing instructions is provided, which, when run by a cluster of computing devices, cause the cluster of computing devices to perform the methods described in the first aspect and any implementation thereof.
[0040] Eighthly, a computer-readable storage medium is provided, including computer program instructions that, when executed by a computing device, perform the method as described in the first aspect and any implementation thereof.
[0041] As examples, these computer-readable storage devices include, but are not limited to, one or more of the following: read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), flash memory, electrically EPROM (EEPROM), and hard drive.
[0042] Alternatively, as one implementation method, the aforementioned storage medium can specifically be a non-volatile storage medium.
[0043] A ninth aspect provides a computer-readable storage medium including computer program instructions that, when executed by a cluster of computing devices, perform the method as described in the first aspect and any implementation thereof.
[0044] As examples, these computer-readable storage devices include, but are not limited to, one or more of the following: read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), flash memory, electrically EPROM (EEPROM), and hard drive.
[0045] Alternatively, as one implementation method, the aforementioned storage medium can specifically be a non-volatile storage medium. Attached Figure Description
[0046] Figure 1 This is a schematic block diagram of a cloud scenario applicable to embodiments of this application.
[0047] Figure 2 This is a schematic flowchart illustrating a model training method provided in an embodiment of this application.
[0048] Figure 3 This is a schematic flowchart illustrating a method for generating logs provided in an embodiment of this application.
[0049] Figure 4 This is a schematic block diagram of a log generation device 400 provided in an embodiment of this application.
[0050] Figure 5 This is a schematic diagram of the architecture of a computing device 1500 provided in an embodiment of this application.
[0051] Figure 6 This is a schematic diagram of the architecture of a computing device cluster provided in an embodiment of this application.
[0052] Figure 7 This is a schematic diagram showing the connection between computing devices 1500A and 1500B via a network, as provided in the embodiments of this application. Detailed Implementation
[0053] The technical solutions in this application will now be described with reference to the accompanying drawings.
[0054] This application will present various aspects, embodiments, or features relating to systems comprising multiple devices, components, modules, etc. It should be understood and appreciated that individual systems may include additional devices, components, modules, etc., and / or may not include all devices, components, modules, etc. discussed in conjunction with the accompanying drawings. Furthermore, combinations of these approaches are also possible.
[0055] Furthermore, in the embodiments of this application, the words "exemplary," "for example," etc., are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the term "exemplary" is intended to present the concept in a concrete manner.
[0056] In the embodiments of this application, "corresponding" and "corresponding" can sometimes be used interchangeably. It should be noted that when the distinction is not emphasized, their intended meanings are consistent.
[0057] The business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0058] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0059] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0060] For ease of description, the concepts involved in the embodiments of this application will be explained below.
[0061] 1. Log
[0062] Logs, also known as log data or log records, are mechanisms that record the code execution process in detail. For example, they record events or messages that occur during the operation of a system or application. High-quality log data helps developers quickly locate the location and cause of faults, and thus quickly fix them.
[0063] The following sections describe logs from the perspectives of their purpose, type, and content.
[0064] The purposes of logs may include, but are not limited to:
[0065] 1) Monitoring: Tracking the status and operation of a system or application;
[0066] 2) Debugging: Helps developers identify and resolve problems;
[0067] 3) Auditing: Record user activity for security audits or compliance checks;
[0068] 4) Prediction: Predict future system behavior or performance issues by analyzing log data.
[0069] Log types can include, but are not limited to:
[0070] 1) System Log: Records events from the operating system or underlying hardware;
[0071] 2) Application Logs: Record events and states within the application;
[0072] 3) Security Log: Records security-related events, such as login attempts and permission changes;
[0073] 4) Access Log: Records user access to resources.
[0074] The log content may include, but is not limited to:
[0075] 1) Timestamp: Records the time when an event occurred;
[0076] 2) Event Description: Describe the specific details of the event;
[0077] 3) Source information: Indicates the source of the event (e.g., which user, which application, which server, etc.);
[0078] 4) Log levels: These typically include DEBUG, INFO, WARNING, ERROR, and FATAL.
[0079] 2. Log statement
[0080] A logging statement, also known as logging code, is a line of code in the source code used to record information about program execution. By writing and inserting logging statements into the source code, corresponding log data can be generated during program runtime. This log data is then collected, stored, and analyzed to support various system management and development tasks.
[0081] For example, a logging statement typically consists of two parts: the log level and the log message. The log level defines the severity and priority of the event being logged. Common log levels include DEBUG (debugging information, usually used only during development), INFO (general information, used to record normal program runtime information), WARNING (warning message, indicating a potential problem that may not cause program interruption), ERROR (error message, indicating an error has occurred during program runtime), and FATAL (fatal error, indicating the program can no longer continue running). The log message typically records information such as the occurrence of a system event, abnormal system states, and key variable values. This information helps developers understand the program's operation and promptly identify and resolve problems.
[0082] 3. Logging
[0083] Log recording, also known as log statement generation, aims to write log statements in the source code so that when the system runs to a code segment, it can output corresponding logs by executing the log statements, helping development and maintenance personnel understand the system's operating status.
[0084] Writing logging statements is a crucial task in ensuring software reliability. Developers write logging statements in the source code to record system status, and operations and maintenance personnel use the logs generated by these statements to monitor and repair system faults. Currently, most unreported software faults are caused by the lack of logging statements in critical locations, or by excessive logging statements that obscure valuable data. Therefore, designing necessary and correct logging statements during development is crucial for the subsequent operation and maintenance of the software system.
[0085] For developers, manually writing logging statements requires deciding whether to log a section of source code, where to log it, and what information to log. Therefore, manually writing logging statements is error-prone and time-consuming. Related technical solutions have proposed automated logging statement tools to assist developers in automatically generating logging statements. The following section introduces relevant solutions for automated logging statement generation.
[0086] In one related technical solution, developers input source code that needs to generate log statements (the source code itself does not contain log statements), and an automated log statement tool outputs the entire source code with the log statements added. First, this requires extensive pre-training; second, besides adding log statements, the output source code may also be arbitrarily modified.
[0087] Another related technical solution involves an automated logging tool that directly outputs the complete logging statement and its location in the source code. This solution has a limited application scenario; it focuses solely on the complete logging statement while neglecting user needs and failing to adjust the output according to user requirements, resulting in unsatisfactory practicality. For example, if a user wants to generate a logging statement at a specific location in the source code, this solution cannot generate the statement at that location, leading to low efficiency and accuracy in logging generation.
[0088] In view of this, embodiments of this application provide a method for generating logs, which can generate corresponding log statements according to user needs, greatly saving manpower costs and improving the efficiency and accuracy of log generation.
[0089] In one possible implementation, the method provided in this application embodiment can be applied to a cloud service scenario, where the method is executed by a cloud management platform within the cloud service scenario. For ease of description, the following will first refer to... Figure 1 It provides a detailed description of cloud service scenarios.
[0090] Figure 1 This is a schematic block diagram illustrating a cloud scenario applicable to embodiments of this application. For example... Figure 1As shown, the cloud scenario may include: cloud management platform 110, Internet 120, and client 130.
[0091] like Figure 1 As shown, the cloud management platform 110 is used to manage the infrastructure that provides multiple cloud services. The infrastructure includes multiple cloud data centers, each cloud data center includes multiple servers, and each server includes cloud service resources to provide corresponding cloud services to tenants.
[0092] The cloud management platform 110 can be located in a cloud data center and provides access interfaces (such as user interfaces or application program interfaces, APIs). Tenants can use client 130 to remotely access the cloud management platform 110, register a cloud account and password, and log in. After successful authentication of the cloud account and password, the tenant can further select and purchase virtual machines of specific specifications (processor, memory, disk) on the cloud management platform 110. After successful purchase, the cloud management platform 110 provides the remote login account and password for the purchased virtual machine, allowing client 130 to remotely log in and install and run the tenant's applications. Therefore, tenants can create, manage, log in to, and operate virtual machines in the cloud data center through the cloud management platform 110. Virtual machines can also be referred to as Elastic Compute Service (ECS) or Elastic Instances (different cloud service providers may use different names).
[0093] It should be understood that cloud service tenants can be individuals, businesses, schools, hospitals, government agencies, etc.
[0094] The cloud management platform 110 includes, but is not limited to, a user console, compute management services, network management services, storage management services, authentication services, and image management services. The user console provides an interface or API for interaction with tenants. The compute management services manage servers running virtual machines and containers, as well as bare metal servers. The network management services manage network services (such as gateways and firewalls). The storage management services manage storage services (such as data bucket services). The authentication services manage tenant account passwords. The image management services manage virtual machine images. Tenants can log in to the cloud management platform 110 via client 130 and the internet 120 to manage their rented cloud services.
[0095] In this embodiment, the aforementioned log statements can be automatically generated for the user using an artificial intelligence (AI) model. It should be understood that an AI model is a general term for mathematical algorithms built upon the principles of artificial intelligence. These algorithms can automatically summarize and learn potential patterns or features from data, thereby achieving a level of thinking similar to that of humans. Depending on the specific methods and / or technologies used to implement artificial intelligence, an AI model can also be specifically referred to as a machine learning model, a deep learning model, or a reinforcement learning model.
[0096] Machine learning is a method for achieving artificial intelligence. Its goal is to design and analyze algorithms (i.e., models) that allow computers to automatically "learn." These designed algorithms are called machine learning models. Machine learning models are algorithms that automatically analyze data to obtain patterns and use these patterns to predict unknown data. Machine learning models are diverse. Based on whether model training relies on the labels corresponding to the training data, machine learning models can be divided into: 1. Supervised learning models; 2. Unsupervised learning models.
[0097] 1. Supervised learning models: These are models obtained by determining the parameters of an initial AI model based on data from a given training dataset and the labels corresponding to each data point. The process of determining the parameters of the initial AI model using the data and their labels in the training dataset is also called supervised learning (or supervised training). The labels on the data in the training dataset are usually manually labeled to identify the correct answer for a specific task. Typical supervised learning models include: Support Vector Machines, Neural Network Models, Logistic Regression Models, Decision Trees, Naive Bayes Models, and Gaussian Discriminant Models. Supervised learning models are commonly used for classification or regression.
[0098] 2. Unsupervised learning models: These are models obtained by determining the parameters of an initial AI model using unlabeled data from a given training dataset. The process of determining the parameters of the initial AI model using unlabeled training data is also called unsupervised learning (or unsupervised training). Through unsupervised learning, the model can discover meaningful information and correlations in the data, thereby making predictions. There are many types of unsupervised learning models, some of the more commonly used ones being: clustering models, principal component analysis (PCA), anomaly detection models, autoencoders, and generative adversarial networks (GANs).
[0099] Deep learning is a new technological field that emerged during machine learning research. In machine learning methods, almost all features need to be determined by industry experts and then encoded. Currently, the typical structure of deep learning models is a deep neural network. A neural network is a mathematical or computational model that mimics the structure and function of a biological neural network (the central nervous system of animals, especially the brain). Neural networks consist of a large number of interconnected neurons performing computations. A neural network can include multiple layers with different functions, each containing parameters and computational rules. Depending on the computational formula or function, different layers in a neural network have different names; for example, the layer performing convolutional calculations is called a convolutional layer, which is often used for feature extraction from input signals (e.g., images). A neural network can also be composed of multiple sub-neural networks. Different neural network structures can be applied to different scenarios (e.g., classification, recognition) or provide different results when used in the same scenario. The specific differences in neural network structures include one or more of the following: different numbers of network layers, different order of network layers, and different weights, parameters, or computational formulas in each network layer. There are already many different neural networks with high accuracy used for applications such as recognition or classification. Some neural networks can be trained on specific datasets and used alone to complete a task or combined with other neural networks (or other functional modules) to complete a task.
[0100] In other words, deep learning models are actually machine learning models with complex neural network structures. Based on whether deep learning models need to rely on the labels of the training data during training, they can also be divided into supervised learning models and unsupervised learning models, which will not be elaborated upon here. Classic deep learning models include convolutional neural networks (CNNs), recurrent neural networks (RNNs), and recursive neural networks (RNNs).
[0101] Reinforcement learning is a special field within machine learning. It involves an agent continuously learning optimal policies, making sequential decisions, and maximizing rewards through the interaction between the agent and the environment. Reinforcement learning teaches "what to do (i.e., how to map the current situation into actions) to maximize numerical rewards." The agent is not told what actions to take; instead, it must discover, through trial and error, which actions yield the greatest rewards.
[0102] Before any AI model can be used to solve a specific technical problem, it needs to be trained. AI model training refers to using a specified initial model to compute on training data, and then adjusting the parameters of the initial model based on the computation results, so that the model gradually learns certain patterns and acquires specific functions. Once trained and possessing stable functionality, the AI model can be used for inference. AI model inference is the process of using the trained AI model to compute on input data and obtain predicted inference results.
[0103] The most common approach is supervised training of AI models. For example, most deep learning models are trained using supervised training methods. During the training phase, a training set for the deep learning model needs to be constructed based on the objective. The training set includes multiple training data points, each labeled. The label of a training data point represents the correct answer to a specific question, and the label can indicate the objective of training the deep learning model using the training data. For example, to train a deep learning model that can identify different animals, the training set could include images of multiple different animals (i.e., training data). Each image could have a label identifying the type of animal it contains, such as cat or dog. In this example, the type of animal corresponding to each image is the label of that training data point.
[0104] Large models are neural network models with a large number of parameters, pre-trained on massive corpora, capable of understanding and generating natural language text. Specifically, large models are typically based on neural network techniques, learning the syntax, semantics, and contextual information of a language through training on large amounts of text data. During training, the model continuously optimizes its parameters to improve its ability to understand and generate text. Due to their powerful ability to understand natural language, large models have been widely applied in many fields to solve natural language understanding and generation problems. Large models have broad applications in artificial intelligence, such as natural language processing, machine translation, and dialogue systems.
[0105] Because the large model has a large number of parameters, in order to accelerate training and reduce training difficulty, this embodiment of the application can perform instruction fine-tuning on the large model, so that the large model after instruction fine-tuning can execute different log statement generation tasks according to the user's instructions. The following is combined with Figure 2 The method for fine-tuning large models using instructions is described in detail.
[0106] Figure 2 This is a schematic flowchart illustrating a model training method provided in an embodiment of this application. Figure 2 As shown, the method may include steps 210-220, which will be described in detail below.
[0107] Step 210: Construct the training dataset.
[0108] In this embodiment of the application, an open-source code dataset can be used to construct datasets corresponding to different log statement generation tasks. These datasets are used for subsequent fine-tuning of the large model, enabling the large model to execute different log statement generation tasks according to user instructions.
[0109] As an example, the following lists the datasets corresponding to several different log statement generation tasks.
[0110] 1. The source code for the logging statement is missing;
[0111] 2. The source code for the logging statement is available, but the logging statement lacks logging levels;
[0112] 3. The source code for the logging statement is available, but the logging statement lacks logging information;
[0113] 4. The source code for a logging statement is missing at a user-specified location;
[0114] Specifically, the phrase "missing source code for logging statements" corresponds to a scenario or task where the user needs to be recommended the location of the logging statement within the code, or a scenario or task where the user needs to be recommended a complete logging statement and its location within the source code. The phrase "source code for logging statements is available, but the logging statement lacks a logging level" corresponds to a scenario or task where the user needs to be recommended the logging level within the logging statement. The phrase "source code for logging statements is available, but the logging statement lacks logging information" corresponds to a scenario or task where the user needs to be recommended the logging information within the logging statement. The phrase "source code for logging statements is missing at a user-specified location" corresponds to a scenario or task where the user needs to be recommended a complete logging statement at that user-specified location.
[0115] It should be understood that the complete logging statement above includes both the log level and the log message. For a detailed description of log levels and log messages, please refer to the explanation above; it will not be repeated here.
[0116] Optionally, in some embodiments, the aforementioned code dataset may also be cleaned and reformatted to ensure data quality and prevent potential data leakage from large models.
[0117] For example, in this embodiment of the application, the Java method dataset provided by LANCE et al. can be used as the above-mentioned code dataset. This Java method dataset was collected from a total of 1,465 high-quality GitHub open repositories.
[0118] Step 220: Fine-tune the large model according to the training dataset and user instructions.
[0119] In this embodiment, to enable the large model to generate results in accordance with user instructions, it is necessary to fine-tune the large model according to the instructions. Instruction fine-tuning is a supervised fine-tuning method that matches the pre-trained knowledge of the large model with the user's intent (which is reflected in the input user instructions), enabling the large model to generate the corresponding results according to the user instructions.
[0120] As an example, fine-tuning a large model requires a large amount of data in the form of instructions. Specifically, a standard instruction-based data set includes the instructions themselves—the functions or goals the large model needs to achieve. The input to the large model is the information it needs to complete the instructions, and the output is the expected correct result. During instruction fine-tuning, the instructions and input information are input into the large model in text form using a given template format. The parameters of the large model are then optimized by analyzing the difference between its output and the expected true result. Due to the sheer size of the large model's parameters, fine-tuning is typically done on a smaller number of parameters rather than all of the large model's parameters.
[0121] For example, when generating log statements using a large model, the input to the large model includes instructions containing user requirements and the corresponding training dataset. The input to the large model also includes the corresponding log statements output based on the user's instructions and / or the positions of the log statements.
[0122] For example, in this embodiment, CodeLlaMA-7B can be selected as the pre-trained large model, which is built on the basis of the LlaMA2 model. Since the LlaMA2 model performs well on natural language data, the CodeLlaMA model goes a step further by incorporating code corpus training, with a particular focus on code completion. This unique training method enables the CodeLlaMA model to effectively handle the mixture of programming language (e.g., source code) and natural language (e.g., user-inputted instructions) tags in log statements, making it more suitable for log statement generation tasks.
[0123] In this embodiment of the application, in order to meet different user needs, the large model can be fine-tuned by instructions, so that the large model can achieve different training objectives related to log statements according to different instructions input by the user.
[0124] The following are some different training objectives or training tasks.
[0125] 1. The training task is to recommend the location of log statements to users.
[0126] In the training task described above, the input to the large model includes: "source code lacking log statements" and an instruction, which could be, for example, "recommend the location of log statements in the input source code for the user." The output of the large model includes: the location information of the log statements in the source code (e.g., line number). <x>).
[0127] 2. The training task is to recommend complete log statements to users and the location of those complete log statements in the source code.
[0128] In the training task described above, the input to the large model includes: "source code missing log statement" and instructions, such as "recommend a complete log statement to the user in the input source code and the location of the complete log statement in the source code". The output of the large model includes: a complete log statement and the location information of the complete log statement in the source code.
[0129] 3. The training task is to recommend log levels to users.
[0130] In the training task described above, the input to the large model includes: "the source code of a log statement, but the log statement lacks a log level" and an instruction, such as "recommend a log level for the log statement in the source code". The output of the large model includes: the log level of the log statement.
[0131] 4. The training task is to recommend log level information to users.
[0132] In the training task described above, the input to the large model includes: "the source code of a log statement, but the log statement is missing log information" and an instruction, such as "recommend log information for the log statement in the source code". The output of the large model includes: the log information of the log statement.
[0133] 5. The training task is to recommend complete log statements to the user at a user-specified location in the source code.
[0134] In the training task described above, the input to the large model includes: "source code, in which a user-specified location is missing a log statement" and an instruction, such as "recommend a complete log statement at a user-specified location in the source code". The output of the large model includes: the complete log statement.
[0135] In the above technical solution, the large model learns from the different training tasks, enabling it to not only execute each training task according to specific instructions, but also to master the semantic information of each log statement component, thereby improving the generation effect of complete log statements.
[0136] It should be understood that the above-mentioned log statement components include, but are not limited to: the location of the log statement, the log level, and the log information.
[0137] In this embodiment of the application, by means of Figure 2 After the method shown completes the training process of the model, the model can be used for inference. The following section combines... Figure 3 This paper describes in detail one implementation process for reasoning about the model.
[0138] It should be understood that, for ease of description, the following can be... Figure 2 The model obtained through training is called the first model.
[0139] Figure 3 This is a schematic flowchart of a method for generating logs provided in an embodiment of this application. The method may include steps 310-330, which will be described in detail below.
[0140] Step 310: Obtain the first source code and the first instruction.
[0141] In this embodiment of the application, the first source code and the first instruction input by the user can be obtained. The first source code and the first instruction will be explained below.
[0142] The aforementioned first source code may be source code containing log statements (or log statement code), or it may be source code not containing log statements (or log statement code). This application embodiment does not specifically limit this.
[0143] The aforementioned first instruction can be used to instruct one or more of the following tasks: recommending logging statements for the first source code, and recommending the location information of the logging statements in the first source code. The logging statements include at least one of the following: log level, and log information.
[0144] That is, the first instruction may be used to instruct the task to be performed, including but not limited to any one or more of the following:
[0145] Task 1: Recommend the location information of the first log statement.
[0146] In Task 1 above, the first source code input by the user is the source code that lacks the first log statement.
[0147] Task 2: Recommend the first logging statement and its location in the first source code.
[0148] In Task 2 above, the first source code input by the user is the source code that lacks the first log statement.
[0149] Task 3: Recommended log level.
[0150] In Task 3 above, the first source code input by the user is the source code containing the first log statement, which lacks a log level.
[0151] Task 4: Recommend log information.
[0152] In Task 4 above, the first source code input by the user is the source code containing the first log statement, which lacks log information.
[0153] Task 5: Recommend the first log statement at a specified location in the first source code.
[0154] In Task 5 above, the first source code input by the user is the source code that is missing the first log statement at a specified location.
[0155] Step 320: Based on the first source code and the first instruction, use the first model to obtain one or more of the following: recommend log statements for the first source code, and the location information of the recommended log statements in the first source code.
[0156] In this embodiment of the application, one or more of the following can be obtained using the first model based on the first source code and the first instruction: recommended log statements for the first source code, and the location information of the recommended log statements in the first source code.
[0157] One example is that, based on the first source code and the first instructions, a first model can be used to obtain a recommended logging statement for the first source code. Another example is that, based on the first source code and the first instructions, the first model can be used to obtain the location information of the logging statement within the first source code. Yet another example is that, based on the first source code and the first instructions, the first model can also be used to obtain a recommended logging statement for the first source code and the location information of that logging statement within the first source code.
[0158] For example, the first source code and the first instruction can be used as input information for the first model. After receiving the input information, the first model can perform reasoning based on the input information to obtain output information. The output information includes log statements recommended for the first source code and / or the location information of the log statements in the first source code.
[0159] It should be understood that the first model can output corresponding information based on the input first instruction and the first source code. The following examples illustrate the output information of the first model using different instructions of the first instruction.
[0160] Example 1: Suppose the user inputs first source code that lacks a first log statement. The first instruction is used to instruct the first model to recommend the location information of the first log statement. The output information of the first model may include the location information of the first log statement recommended for the first source code.
[0161] It should be understood that the location information of the first log statement mentioned above refers to the specific location of the first log statement in the first source code.
[0162] For example, the following lists several possible locations where the first log statement might be placed in the first source code.
[0163] 1. Place the first log statement in the main entry point of the first source code (such as the main function or the application's startup class) to record the startup time, parameters, configuration information, etc. of the first source code.
[0164] 2. Place the first log statement at the point where the first source code is closed or exited, and record the closing time, execution status, etc. of the first source code.
[0165] 3. Place the first log statement at the beginning and / or end of the function or method in the first source code to record the start and end times of the function's execution, as well as the input parameters and return value.
[0166] 4. Place the first log statement before and after the key business logic code segment in the first source code to record the execution flow of the business logic, the values of key variables, the results of business operations, etc.
[0167] 5. Place the first logging statement in the catch block of the try-catch statement in the first source code to record the captured exception type, exception information, stack trace, etc., to help developers locate the problem.
[0168] 6. Place the first log statement at the beginning and end of the loop in the first source code to record the number of loop iterations, the value of the loop variable, etc.
[0169] 7. Place the first log statement before and after the database query, update, and delete operations in the first source code to record the execution status, execution time, and return results of the SQL statements.
[0170] 8. Place the first log statement in the first source code when sending and receiving network requests, recording the request URL, parameters, response status code, response content, etc.
[0171] 9. Place the first log statement at the performance bottleneck of the first source code to record the execution time and resource consumption of key operations for performance analysis and optimization.
[0172] 10. Place the first log statement in the first source code at the locations of user input, output, and interface operations to record user operation behavior, interface status, etc., for user behavior analysis and problem troubleshooting.
[0173] 11. Place the first log statement at the start, commit, and rollback points of the transaction in the first source code to record the execution status of the transaction, the database operations involved, etc., which helps with transaction management and problem diagnosis.
[0174] Example 2: Suppose the user inputs a first source code that lacks a first logging statement. The first instruction instructs the first model to recommend a first logging statement for the first source code, along with the location information of that first logging statement within the first source code. The output information of the first model may include the recommended first logging statement for the first source code and the location information of that first logging statement within the first source code.
[0175] For example, the output information of the first model is shown below:
[0176] <line4>log.warn("No instance found,fallback port:{}",fallbackPort)
[0177] As mentioned above <line4>The "" indicates the location of the log statement in the source code; "log.warn("No instance found,fallback port:{}",fallbackPort)" is the log statement, where "warn" indicates the log level and "Noinstance found,fallback port:{}",fallbackPort" indicates the log message.
[0178] Example 3: Suppose the user inputs first source code containing a first logging statement that lacks a logging level. A first instruction is used to instruct a first model to recommend a logging level for the first logging statement. The output information of the first model may include the recommended logging level for the first logging statement.
[0179] Example 4: Suppose the user inputs first source code containing a first log statement that lacks log information. The first instruction is used to instruct the first model to recommend the log information from the first log statement. The output information of the first model may include the recommended log information for the first log statement.
[0180] Example 5: Suppose the first source code input by the user is source code where a first log statement is missing at a user-specified location. The first instruction instructs the first model to recommend a first log statement at the user-specified location in the first source code. The output information of the first model may include the recommended first log statement at the user-specified location in the first source code.
[0181] Step 330: Generate corresponding logs for the first source code based on the log statements.
[0182] In this embodiment of the application, a corresponding log can be generated for the first source code based on the log statements recommended for the first source code, and / or the location information of the log statements in the first source code.
[0183] In one possible implementation, the log statement is placed at a corresponding location (e.g., a specific line) in the first source code, based on its position information. During the execution of the first source code, when the log statement code in a specific line of the first source code is executed, the corresponding log (also called log data or log record) is generated for that source code. For example, by executing the log statement code, a corresponding log entry is generated, and this generated log entry is added to a log file to form a complete log record.
[0184] The above technical solution supports both semi-automatic and fully automatic log generation, and can complete log recording in various scenarios according to user intent, which greatly saves labor costs and improves code quality and operation and maintenance capabilities.
[0185] The above text combined Figures 1 to 3 The method provided in the embodiments of this application is described in detail below. Figures 4-7 The embodiments of the apparatus of this application are described in detail below. It should be understood that the descriptions of the method embodiments correspond to the descriptions of the apparatus embodiments; therefore, any parts not described in detail can be referred to the foregoing method embodiments.
[0186] Figure 4 This is a schematic block diagram of a log generation device 400 provided in an embodiment of this application. The device 400 can be implemented by software, hardware, or a combination of both. The device 400 provided in this embodiment can implement the embodiments of this application. Figure 2 or Figure 3 The method flow shown includes the following: the device 400 comprises an acquisition module 410, a recommendation module 420, and a generation module 430. The acquisition module 410 is used to acquire a first source code and a first instruction. The recommendation module 420 is used to obtain one or more of the following based on the first source code and the first instruction using a first model: a recommended log statement for the first source code and the location information of the log statement in the first source code. The generation module 430 is used to generate a corresponding log for the first source code based on the log statement.
[0187] The first instruction described above is used to instruct one or more of the following tasks: recommending a logging statement for the first source code, and recommending the location information of the logging statement in the first source code.
[0188] The above log statements include at least one of the following: log level, log message.
[0189] The input information of the first model includes the first source code and the first instruction, and the output information of the first model includes the log statement recommended for the first source code, and / or the location information of the log statement in the first source code.
[0190] Optionally, the first source code includes a first logging statement that lacks a logging level, the first instruction being used to indicate that the logging level is recommended for the first logging statement, and the output information of the first model including the recommended logging level for the first logging statement.
[0191] Optionally, the first source code includes a first logging statement that lacks logging information, the first instruction is used to indicate that the logging information is recommended for the first logging statement, and the output information of the first model includes the logging information recommended for the first logging statement.
[0192] Optionally, the first instruction is used to instruct a first log statement to be recommended at the target location of the first source code, and the output information of the first model includes the first log statement recommended for the first source code at the target location.
[0193] Optionally, the first instruction is used to indicate the location information of a first log statement recommended for the first source code, and the output information of the first model includes the location information of the recommended first log statement.
[0194] Optionally, the first instruction is used to indicate a recommended first logging statement and the location information of the first logging statement in the first source code, and the output information of the first model includes the recommended first logging statement and the location information of the first logging statement in the first source code.
[0195] Optionally, the device 400 further includes a training module for fine-tuning the second model to obtain the first model, wherein the second model is a pre-trained model.
[0196] Optionally, the second model is a large model.
[0197] The device 400 here can be embodied in the form of a functional module. The term "module" here can be implemented in software and / or hardware, without specific limitations.
[0198] For example, a "module" can be a software program, a hardware circuit, or a combination of both that implements the above functions. For instance, the implementation of module 410 will be described below using module 410 as an example. Similarly, the implementation of other modules, such as recommendation module 420 and generation module 430, can refer to the implementation of module 410.
[0199] As an example of a software functional unit, the acquisition module 410 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the acquisition module 410 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0200] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0201] As an example of a hardware functional unit, the acquisition module 410 may include at least one computing device, such as a server. Alternatively, the acquisition module 410 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0202] The multiple computing devices included in the acquisition module 410 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the acquisition module 410 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the acquisition module 410 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0203] Therefore, the modules of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0204] It should be noted that the above embodiments of the device, when executing the above methods, are only illustrative examples of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. For example, the acquisition module 410 can be used to execute any step in the above methods, the recommendation module 420 can be used to execute any step in the above methods, and the generation module 430 can be used to execute any step in the above methods. The steps implemented by the acquisition module 410, the recommendation module 420, and the generation module 430 can be specified as needed. By implementing different steps in the above methods through the acquisition module 410, the recommendation module 420, and the generation module 430, all the functions of the above device can be realized.
[0205] Furthermore, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments above, which will not be repeated here.
[0206] The method provided in this application can be executed by a computing device, which can also be referred to as a computer system. It includes a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as processing units, memory, and memory control units; the functions and structure of this hardware will be described in detail later. The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software. Optionally, the computer system can be a handheld device such as a smartphone, or a terminal device such as a personal computer; this application does not particularly limit this, as long as the method provided in this application can be used. The executing entity of the method provided in this application can be a computing device, or a functional module within the computing device capable of calling and executing programs.
[0207] The following is combined with Figure 5 This application provides a detailed description of a computing device provided in an embodiment.
[0208] Figure 5 This is a schematic diagram of the architecture of a computing device 1500 provided in an embodiment of this application. The computing device 1500 can be a server, a computer, or other device with computing capabilities. Figure 5 The computing device 1500 shown includes at least one processor 1510 and a memory 1520.
[0209] It should be understood that this application does not limit the number of processors and memories in the computing device 1500.
[0210] The processor 1510 executes instructions in the memory 1520, causing the computing device 1500 to implement the method provided in this application. Alternatively, the processor 1510 executes instructions in the memory 1520, causing the computing device 1500 to implement the various functional modules provided in this application, thereby implementing the method provided in this application.
[0211] Optionally, the computing device 1500 also includes a communication interface 1530. The communication interface 1530 uses a transceiver module, such as, but not limited to, a network interface card or a transceiver, to enable communication between the computing device 1500 and other devices or communication networks.
[0212] Optionally, the computing device 1500 further includes a system bus 1540, wherein the processor 1510, memory 1520, and communication interface 1530 are respectively connected to the system bus 1540. The processor 1510 can access the memory 1520 through the system bus 1540; for example, the processor 1510 can perform data read / write or code execution in the memory 1520 through the system bus 1540. The system bus 1540 is a peripheral component interconnect express (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The system bus 1540 is divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0213] In one possible implementation, the processor 1510 primarily functions to interpret the instructions (or code) of a computer program and process data within the computer software. The instructions of the computer program and the data within the computer software can be stored in memory 1520 or cache 1516.
[0214] Optionally, processor 1510 may be an integrated circuit chip with signal processing capabilities. By way of example and not limitation, processor 1510 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor, etc. For example, processor 1510 may be a central processing unit (CPU).
[0215] Optionally, each processor 1510 includes at least one processing unit 1512 and a memory control unit 1514.
[0216] Optionally, the processing unit 1512, also known as the core, is the most important component of the processor. The processing unit 1512 is manufactured from single-crystal silicon using a specific production process. All calculations, command reception, command storage, and data processing are performed by the core. Each processing unit independently executes program instructions, utilizing parallel computing capabilities to accelerate program execution. Various processing units have fixed logical structures; for example, a processing unit includes logical units such as a Level 1 cache, a Level 2 cache, an execution unit, an instruction-level unit, and a bus interface.
[0217] In one implementation example, the memory control unit 1514 controls the data interaction between the memory 1520 and the processing unit 1512. Specifically, the memory control unit 1514 receives memory access requests from the processing unit 1512 and controls access to memory based on the memory access requests. By way of example and not limitation, the memory control unit is a device such as a memory management unit (MMU).
[0218] In one implementation example, each memory control unit 1514 addresses the memory 1520 via the system bus. An arbitrator is configured in the system bus. Figure 5 (Not shown in the image), the arbitrator is responsible for handling and coordinating competing accesses of multiple processing units 1512.
[0219] In one implementation example, the processing unit 1512 and the memory control unit 1514 are connected via internal chip connection lines, such as address lines, thereby enabling communication between the processing unit 1512 and the memory control unit 1514.
[0220] Optionally, each processor 1510 also includes a cache 1516, which is a buffer for data exchange (called a cache). When the processing unit 1512 needs to read data, it first looks for the required data in the cache. If the data is found, it is executed directly; otherwise, it looks for the data in memory. Since the cache operates much faster than memory, its purpose is to help the processing unit 1512 run faster.
[0221] The memory 1520 provides runtime space for processes in the computing device 1500. For example, the memory 1520 stores the computer program (specifically, the program code) used to generate the process. After the computer program is run by the processor to generate a process, the processor allocates corresponding storage space for the process in the memory 1520. Furthermore, the aforementioned storage space further includes text segments, initialized data segments, bit initialized data segments, stack segments, heap segments, etc. The memory 1520 stores data generated during the process's execution, such as intermediate data or process data, in the aforementioned process-specific storage space.
[0222] Optionally, the memory, also known as RAM, is used to temporarily store the data processed by the processor 1510, as well as data exchanged with external storage devices such as hard disks. As long as the computer is running, the processor 1510 will load the data that needs to be processed into RAM for processing, and after the processing is completed, the processing unit 1512 will send the result out.
[0223] By way of example and not limitation, memory 1520 is volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory is read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory is random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DRRAM). It should be noted that the memory 1520 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0224] The above-described structure of the computing device 1500 is merely illustrative and is not intended to limit the application. The computing device 1500 in this application includes various hardware components found in existing computer systems. For example, the computing device 1500 may also include other memories besides the memory 1520, such as disk storage. Those skilled in the art should understand that the computing device 1500 may also include other devices necessary for normal operation. Furthermore, depending on specific needs, those skilled in the art should understand that the computing device 1500 may also include hardware devices for implementing other additional functions. Moreover, those skilled in the art should understand that the computing device 1500 may only include the devices necessary for implementing the embodiments of this application, and may not necessarily include... Figure 5 All the devices shown.
[0225] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server. In some embodiments, the computing device may also be a desktop computer, a laptop computer, or a smartphone, or other terminal device.
[0226] like Figure 6 As shown, the computing device cluster includes at least one computing device 1500. The memory 1520 of one or more computing devices 1500 in the computing device cluster may store the same instructions for performing the methods described above.
[0227] In some possible implementations, the memory 1520 of one or more computing devices 1500 in the computing device cluster may also each store a portion of the instructions for executing the above-described methods. In other words, a combination of one or more computing devices 1500 can jointly execute the instructions of the above-described methods.
[0228] It should be noted that the memory 1520 in different computing devices 1500 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the aforementioned device. That is, the instructions stored in the memory 1520 of different computing devices 1500 can implement the functions of one or more modules within the aforementioned device.
[0229] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 7 One possible implementation is shown. For example... Figure 7 As shown, the two computing devices 1500A and 1500B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device.
[0230] It should be understood that Figure 7 The functions of computing device 1500A shown can also be performed by multiple computing devices 1500. Similarly, the functions of computing device 1500B can also be performed by multiple computing devices 1500.
[0231] In this embodiment, a computer program product containing instructions is also provided. The computer program product may be a software or program product containing instructions capable of running on a computing device or stored on any usable medium. When run on a computing device, it causes the computing device to perform the methods provided above, or causes the computing device to perform the functions of the apparatus provided above.
[0232] In this embodiment, a computer-readable storage medium is also provided. This computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that, when executed on a computing device, cause the computing device to perform the method described above.
[0233] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0234] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0235] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0236] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0237] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0238] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0239] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0240] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims. < / x>
Claims
1. A method of generating a log, characterized by, The method comprises: obtaining a first source code and a first instruction, the first instruction being used to indicate one or more tasks of recommending a log statement for the first source code, recommending position information of the log statement in the first source code; wherein the log statement comprises at least one of the following: a log level, log information; obtaining one or more of the following: the log statement recommended for the first source code, the position information of the log statement in the first source code, according to the first source code and the first instruction, by using a first model; wherein the input information of the first model comprises the first source code and the first instruction, and the output information of the first model comprises the log statement recommended for the first source code, and / or the position information of the log statement in the first source code; generating a corresponding log for the first source code according to the log statement.
2. The method of claim 1, wherein, The first source code comprises a first log statement, the first log statement lacks a log level, the first instruction is used to indicate that the log level is recommended for the first log statement, and the output information of the first model comprises the log level recommended for the first log statement.
3. The method of claim 1, wherein, The first source code comprises a first log statement, the first log statement lacks log information, the first instruction is used to indicate that the log information is recommended for the first log statement, and the output information of the first model comprises the log information recommended for the first log statement.
4. The method of claim 1, wherein, The first instruction is used to indicate that a first log statement is recommended at a target position of the first source code, and the output information of the first model comprises the first log statement recommended for the first source code at the target position.
5. The method of claim 1, wherein, The first instruction is used to indicate that position information of a first log statement is recommended for the first source code, and the output information of the first model comprises the position information of the recommended first log statement.
6. The method of claim 1, wherein, The first instruction is used to indicate that a first log statement and position information of the first log statement in the first source code are recommended for the first source code, and the output information of the first model comprises the recommended first log statement and the position information of the first log statement in the first source code.
7. The method according to any one of claims 1 to 6, characterized in that, The method further comprises: instruction fine-tuning a second model to obtain the first model, wherein the second model is a pre-trained model.
8. The method of claim 7, wherein, The second model is a large model.
9. An apparatus for generating a log, the apparatus comprising: The device comprises: an obtaining module configured to obtain a first source code and a first instruction, the first instruction being used to indicate one or more tasks of recommending a log statement for the first source code, recommending position information of the log statement in the first source code; wherein the log statement comprises at least one of the following: a log level, log information; The recommendation module is configured to obtain one or more of the following based on the first source code and the first instruction by using a first model: a recommended log statement for the first source code, and position information of the log statement in the first source code; wherein input information of the first model comprises the first source code and the first instruction, and output information of the first model comprises the recommended log statement for the first source code and / or the position information of the log statement in the first source code. The generation module is configured to generate a corresponding log for the first source code based on the log statement.
10. The apparatus of claim 9, wherein, The first source code comprises a first log statement, the first log statement lacks a log level, the first instruction is used to instruct to recommend the log level for the first log statement, and the output information of the first model comprises the recommended log level for the first log statement.
11. The apparatus of claim 9, wherein, The first source code comprises a first log statement, the first log statement lacks log information, the first instruction is used to instruct to recommend the log information for the first log statement, and the output information of the first model comprises the recommended log information for the first log statement.
12. The apparatus of claim 9, wherein, The first instruction is used to instruct to recommend a first log statement at a target position of the first source code, and the output information of the first model comprises the recommended first log statement at the target position of the first source code.
13. The apparatus of claim 9, wherein, The first instruction is used to instruct to recommend position information of a first log statement for the first source code, and the output information of the first model comprises the recommended position information of the first log statement.
14. The apparatus of claim 9, wherein, The first instruction is used to instruct to recommend a first log statement and position information of the first log statement in the first source code for the first source code, and the output information of the first model comprises the recommended first log statement and the position information of the first log statement in the first source code.
15. The apparatus of any one of claims 9 to 14, wherein, The apparatus further comprises: The training module is configured to perform instruction fine-tuning on a second model to obtain the first model, wherein the second model is a pre-trained model.
16. The apparatus of claim 15, wherein, The second model is a large model.
17. A cluster of computing devices, characterized in that, The at least one computing device comprises a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the computing device cluster to perform the method of any one of claims 1 to 8.
18. A computer program product comprising instructions, characterized in that, The instructions, when executed by the computing device cluster, cause the computing device cluster to perform the method of any one of claims 1 to 8.
19. A computer-readable storage medium, characterized in that, The computer program instructions, when executed by the computing device cluster, cause the computing device cluster to perform the method of any one of claims 1 to 8.