Program code processing method, device and equipment
By constructing operation topology diagrams and functional block description information, and using the big model to generate analysis logs, the communication cost problem of adding program code analysis logs in resource-intensive environments is solved, and efficient and easy-to-maintain analysis log generation is achieved.
Patent Information
- Application Number
- CN202510218680.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-26
AI Technical Summary
In a platform-based deployment environment with tight resources, the huge communication costs between algorithm-side technicians and model-deployment platform technicians are unable to understand the algorithm, especially when adding analysis logs to program code.
By obtaining the target program code, building operation topology diagram and function block description information, generating prompt information and inputting a large model to generate analysis logs, improving the success rate and easy maintenance of the analysis logs.
It greatly improves the success rate and ease of maintenance of adding analytical logs, reduces the performance overhead and cost of analytical logs, and reduces the communication cost and the understanding cost of program code.
Smart Images

Figure CN120029634A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of computer technology, and in particular to a method, device and equipment for processing program codes. Background Art
[0002] In the platform deployment work, algorithm-side technicians deploy inference program codes on the platform. However, in the context of tight resources, there is often a common situation of insufficient resources. In addition to blindly adding resources, the option is often to analyze the performance of the deployed program code and look for optimization space.
[0003] In the model deployment platform, each model is deployed by the technicians on the algorithm side. In order to protect the user's program code and other privacy data, the technicians on the model deployment platform need not invade the user's program code. When the resources of the technicians on the algorithm side need to be optimized, if the algorithm cannot be understood, then the optimization cannot be performed. Therefore, there is a huge communication cost. The purpose of performance analysis of a program code is to find the performance bottleneck of the program code. Customized optimization is required for the performance bottleneck, and this process relies on adding analysis logs to the program code. This is also the biggest communication cost between the technicians on the algorithm side and the technicians on the model deployment platform. To this end, it is necessary to provide a better technical solution for setting analysis logs for program code, so as to improve the success rate and maintainability of adding analysis logs and reduce the performance overhead and cost of analysis logs. Summary of the invention
[0004] The purpose of the embodiments of this specification is to provide a better technical solution for setting analysis logs for program code, thereby improving the success rate and maintainability of adding analysis logs and reducing the performance overhead and cost of analysis logs.
[0005] In order to implement the above technical solution, the embodiments of this specification are implemented as follows: A method for processing program code provided by an embodiment of the present specification includes: obtaining a target program code to be processed, wherein the target program code includes a program code of a target model, wherein the target model includes one or more different model operations; determining to construct an operation topology relationship diagram corresponding to the target program code based on the model operations contained in the target program code, and determining the function blocks contained in the target program code and the description information of each function block based on the operations contained in the target program code and the operation topology relationship diagram corresponding to the target program code; constructing first prompt information based on the function blocks contained in the target program code and the description information of each function block, and the target program code provided with a code line identifier, and obtaining, from a code database, a first program code whose similarity with the target program code is greater than a preset threshold based on the target program code; generating second prompt information based on the first prompt information, the first program code and the target program code, and inputting the second prompt information into a large model to obtain an analysis log of the target program code.
[0006] The embodiment of the present specification provides a program code processing device, the device comprising: a code acquisition module, which acquires a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations; a code processing module, which determines to construct an operation topology relationship diagram corresponding to the target program code based on the model operations contained in the target program code, and determines the function blocks contained in the target program code and the description information of each function block based on the operations contained in the target program code and the operation topology relationship diagram corresponding to the target program code; a retrieval enhancement module, which constructs first prompt information based on the function blocks contained in the target program code and the description information of each function block, and the target program code with a code line identifier, and obtains a first program code whose similarity with the target program code is greater than a preset threshold from a code database based on the target program code; a log generation module, which generates second prompt information based on the first prompt information, the first program code and the target program code, and inputs the second prompt information into a large model to obtain an analysis log of the target program code.
[0007] An embodiment of the present specification provides a program code processing device, the program code processing device comprising: a processor; and a memory arranged to store computer executable instructions, wherein the executable instructions, when executed, cause the processor to: obtain a target program code to be processed, wherein the target program code includes a program code of a target model, wherein the target model includes one or more different model operations; determine to construct an operation topology relationship diagram corresponding to the target program code based on the model operations included in the target program code, and determine the function blocks included in the target program code and the description information of each function block based on the operations included in the target program code and the operation topology relationship diagram corresponding to the target program code; construct first prompt information based on the function blocks included in the target program code and the description information of each function block, and the target program code provided with a code line identifier, and obtain, from a code database, a first program code whose similarity with the target program code is greater than a preset threshold based on the target program code; generate second prompt information based on the first prompt information, the first program code, and the target program code, and input the second prompt information into a large model to obtain an analysis log of the target program code.
[0008] The embodiments of the present specification also provide a storage medium, which is used to store computer-executable instructions. When the executable instructions are executed by a processor, the following process is implemented: obtaining a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations; based on the model operations included in the target program code, determining to construct an operation topology relationship diagram corresponding to the target program code, and based on the operations included in the target program code and the operation topology relationship diagram corresponding to the target program code, determining the function blocks included in the target program code and the description information of each function block; based on the function blocks included in the target program code and the description information of each function block, and the target program code with a code line identifier, constructing first prompt information, and based on the target program code, obtaining a first program code whose similarity with the target program code is greater than a preset threshold from a code database; generating second prompt information based on the first prompt information, the first program code and the target program code, and inputting the second prompt information into the large model to obtain an analysis log of the target program code.
[0009] The embodiment of the present specification also provides a computer program product, including a computer program, which implements the following process when executed by a processor: obtaining a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations; based on the model operations included in the target program code, determining to construct an operation topology relationship diagram corresponding to the target program code, and based on the operations included in the target program code and the operation topology relationship diagram corresponding to the target program code, determining the function blocks included in the target program code and the description information of each of the function blocks; based on the function blocks included in the target program code and the description information of each of the function blocks, and the target program code with a code line identifier, constructing first prompt information, and based on the target program code, obtaining a first program code whose similarity with the target program code is greater than a preset threshold from a code database; generating second prompt information based on the first prompt information, the first program code and the target program code, and inputting the second prompt information into the large model to obtain an analysis log of the target program code. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings required for use in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative labor. Figure 1 This is a schematic diagram of the analysis and optimization of a program code in this specification; Figure 2 This is another schematic diagram of analyzing and optimizing the program code of this specification; Figure 3 This is another schematic diagram of analyzing and optimizing the program code of this specification; Figure 4 This is an embodiment of a method for processing program code of this specification; Figure 5 A schematic diagram of a program code processing process of this specification; Figure 6 Another embodiment of a method for processing program codes of this specification; Figure 7 This is another embodiment of a method for processing program codes of this specification; Figure 8 This is another embodiment of a method for processing program codes of this specification; Fig. 9 A schematic diagram of another processing process of program code in this specification; Fig.10 An embodiment of a processing device for program code of the present specification; Fig.11 This is an embodiment of a program code processing device of the present specification. DETAILED DESCRIPTION
[0011] The embodiments of this specification provide a method, apparatus and device for processing program codes.
[0012] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0013] The embodiments of this specification provide a mechanism for setting the analysis log of program code based on a large model. In the platform deployment work, the technicians on the algorithm side deploy the reasoning program code on the platform with a high degree of freedom. In the context of tight resources, there is often a common situation where resources are insufficient, such as Figure 1 As shown in Figure 1, in addition to simply adding resources, an alternative approach is to analyze the performance of the deployed program code and find room for optimization. In the model deployment platform, each model is deployed by the technicians on the algorithm side, and the technicians on the model deployment platform do not intrude into the user's program code. Figure 2 As shown in Figure 1, when the resources of the technical staff on the algorithm side need to be optimized, if the algorithm is not understood, then the optimization cannot be performed. The technical staff on the model deployment platform do not understand the user's program code, so there is a huge communication cost. Figure 3 As shown in the figure, the purpose of performance analysis of a program code is to find the performance bottleneck of the program code. Customized optimization is required to address the performance bottleneck. This process relies on adding analysis logs to the program code, which is also the biggest communication cost between the technicians on the algorithm side and the technicians on the model deployment platform.
[0014] At present, analysis logs can be added to program code according to expert experience and regular algorithms. Although adding analysis logs according to expert experience and regular algorithms is indeed an effective strategy, it does have some obvious shortcomings and potential problems. For example, the above method will have inaccurate and incomplete recognition, poor maintainability, large performance overhead, and context information loss. In addition, analysis logs can also be added for each line in the target program code, but the above method will have log confusion. After adding the analysis log, the analysis cost is also very high. In addition, for the loop code statement therein, if the analysis log is added, it will cause phenomena such as log explosion. For this reason, it is necessary to provide a better technical solution for setting analysis logs for program code, by abstracting the functional blocks of program code from fine to coarse, and combining each line of program code with the functional blocks from coarse to fine, greatly improving the success rate of adding analysis logs, and, based on the framework of RAG, constantly enriching the correctly verified expert experience and historical cases, and realizing the positive cycle of the success rate of analysis logs. For specific processing, please refer to the specific content in the following embodiment.
[0015] like Figure 4 As shown, an embodiment of this specification provides a method for processing program code, and the execution subject of the method may be a terminal device or a server, etc., wherein the terminal device may be a mobile terminal device such as a mobile phone, a tablet computer, or a computer device such as a laptop or a desktop computer, or an IoT device (specifically such as a smart watch, a car-mounted device, etc.), etc., wherein the server may be an independent server, or a server cluster composed of multiple servers, etc., and the server may be a background server in the financial field or the online shopping field, etc., or a background server of an application, etc. In this embodiment, the execution subject is taken as an example for detailed description. For the case where the execution subject is a terminal device, please refer to the following server case processing, which will not be repeated here. The method may specifically include the following steps: In step S402, the target program code to be processed is obtained, where the target program code includes the program code of the target model, and the target model includes one or more different model operations.
[0016] Among them, the target program code can be any program code, and the target program code can be a program code written in a specified programming language, wherein the programming language can include C language, C++, Java language, Python programming language, etc., which can be specifically set according to actual conditions. In this embodiment, the target program code can include the program code of the target model. In addition, it can also include program codes other than the program code of the target model, for example, program codes for collecting and obtaining certain specified data, and program codes for processing the above-mentioned specified data to obtain input data to be input into the target model. For another example, it can also include program codes for processing the output data of the target model for subsequent specified application fields, etc., which can be specifically set according to actual conditions. The target model can be any model, for example, the target model can be a model for risk prevention and control, a model for facial recognition, a model for natural language processing, etc. The target model can be constructed by a variety of different algorithms or networks, for example, the target model can be constructed by a specified neural network (such as a convolutional neural network, a recurrent neural network, etc.), a decision tree algorithm, a support vector machine algorithm, a BERT algorithm, etc., which can be specifically set according to actual conditions. Model operations can include multiple types. For example, if the target model is built based on a convolutional neural network, the model operations can include convolution, activation function, pooling, etc. The target models built with different algorithms or networks may have different corresponding model operations, which can be set according to actual conditions.
[0017] In implementation, the target program code to be processed can be obtained in a variety of different ways. For example, a user writes a target program code including the program code of a certain model (i.e., the target model). In order to set the analysis log of the target program code, the analysis log setting page for the program code can be obtained through the terminal device. The page may include a program code input box, a confirmation button and a cancel button, and a result output box, etc. The user can enter the target program code in the program code input box in the page. After the input is completed, the confirmation button can be clicked. At this time, the terminal device can obtain the target program code entered by the user in the program code input box, and can generate a code log setting request based on the obtained target program code, and can send the code log setting request to the server. The server can receive the code log setting request and can extract the target program code to be processed from the code log setting request. For another example, a technician can write a target program code including the program code of a certain model (i.e., the target model), and then directly input the target program code into the server, and the server can obtain the target program code to be processed. The specific setting can be based on the actual situation, and the embodiments of this specification are not limited to this.
[0018] In step S404, based on the model operations contained in the target program code, an operation topology relationship diagram corresponding to the target program code is determined, and based on the operations contained in the target program code and the operation topology relationship diagram corresponding to the target program code, the function blocks contained in the target program code and the description information of each function block are determined.
[0019] The operation topology diagram can be used to show the dependency relationship between different model operations. For example, model operation 1-model operation 2-model operation 3 can indicate that model operation 2 depends on model operation 1, and model operation 3 depends on model operation 2, etc. The specific settings can be made according to the actual situation. The operations included in the target program code can include model operations, and in addition, other operations included in the program code other than the program code of the target model can also be included.
[0020] In implementation, in order to reduce the communication cost and the cost of understanding the program code, and to understand the context information of the program code, the logic of the program code of the target model contained in the target program code can be analyzed, and the key paths and important events therein can be identified. To this end, the target program code can be parsed, and based on the parsing results, the model operations contained in the target program code can be determined. The logic of the program code of the target model can be analyzed to determine the dependency between different model operations in the target model, and based on the dependency between different model operations in the target model, an operation topology diagram corresponding to the target program code can be constructed.
[0021] Considering that the model operation is a more specific or detailed structure in the target model, it is difficult for people who do not understand the structure of the target model to know the more specific or detailed structure. In order to enable the large model to understand the context information of the program code, the above structure can be abstracted. Specifically, the target program code can be parsed, and based on the parsing results, the operations contained in the target program code can be determined. The operations contained in the target program code can be classified to determine the operations that can be classified into the same category. For example, model operations such as feature extraction and convolution operations can be classified and abstracted to obtain abstracted functions (such as image information extraction). For another example, operations such as normalization, data enhancement, and image conversion can be abstracted into a function, which can be image data preprocessing. The corresponding function block can be determined based on the above-mentioned functions, and can be set according to actual conditions. Through the above-mentioned classification and abstract processing, the functional blocks contained in the target program code can be determined based on the operations contained in the target program code and the operation topology relationship diagram corresponding to the target program code, such as the functional block for image information extraction, the functional block for image information preprocessing, etc. In addition, the function and role of each functional block can be described to obtain the description information of each functional block, and then the functional blocks contained in the target program code and the description information of each functional block can be obtained.
[0022] In step S406, first prompt information is constructed based on the function blocks contained in the target program code and the description information of each function block, and the target program code with code line identifiers, and based on the target program code, a first program code whose similarity with the target program code is greater than a preset threshold is obtained from a code database.
[0023] Among them, the prompt information can be Prompt, and the first prompt information can be the prompt information Prompt for the large model. The preset threshold can be set according to the actual situation, such as 80% or 90%. The code database can be a database that stores a large amount of program code. In actual applications, the code database can only store program code. In addition, in addition to storing program code, the Embedding vector (or characterization vector) or identifier corresponding to the program code can also be stored. In this way, the corresponding program code can be indexed by the Embedding vector (or characterization vector) or identifier, which can be set according to the actual situation.
[0024] In implementation, Figure 5As shown in FIG. 1 , the RAG (Retrieval-Augmented Generation) model can be used to complete the processing of setting the analysis log for the target program code. The RAG model can combine the language model with the information retrieval mechanism. When it is necessary to generate text or answer questions, the RAG model will first retrieve relevant information from a large database, and then use the retrieved information to guide the generation of text or answers, thereby improving the quality and accuracy of the prediction. The processing process of the RAG model can specifically include data extraction-vectorization (such as embedding processing, etc.)-create indexes-retrieval-automatic sorting (Rerank)-large model induction generation. Specifically, for data extraction processing, it can include cleaning and processing the target program code (such as removing redundant content, format conversion, etc.), information extraction (such as file name, time, image, etc.). For vectorization processing, the target program code can be vectorized to obtain the Embedding vector corresponding to the target program code. In the process of vectorization processing, the target program code can also be divided into blocks to obtain multiple block codes of the target program code, and then each block code is vectorized. Through the above-mentioned vectorization processing, the corresponding relationship between the embedding vector and the block code can be constructed, thereby obtaining the corresponding index information. For the retrieval processing, a huge database (i.e., code database) can be pre-set, and the retrieval processing can be performed from the code database based on the target program code. The retrieval method can include multiple methods, such as similarity-based retrieval, keyword-based retrieval, SQL-based retrieval, etc. For example, based on the target program code, a first program code whose similarity with the target program code is greater than a preset threshold can be retrieved from the code database through a retrieval method based on similarity (the similarity can be determined by the Euclidean distance similarity algorithm, the Manhattan distance similarity algorithm, the cosine distance similarity algorithm, or the kNN algorithm to determine whether two program codes are similar, etc.), or, based on the target program code, a retrieval processing can be performed from the code database through a keyword-based retrieval method. Specifically, keywords can be extracted from the target program code, and the retrieved keywords can be used to perform retrieval processing from the code database to obtain a first program code that is similar or close to the target program code. The specific method can be set according to actual conditions.
[0025] Usually, the search results obtained through the search process are not ideal because the search dimensions are not necessarily optimal. The results of a search may not be so ideal in terms of relevance. At this time, some strategies are needed to re-rank the search results to make them more suitable for the application scenario. You can pre-set a judge to review the relevance to trigger re-ranking. Automatic reranking can make the search results more suitable for the application scenario.
[0026] In addition, in addition to the retrieval process, the RAG model also includes an enhancement process. Specifically, a prompt information template can be pre-set, and the function blocks contained in the target program code and the description information of each function block, as well as the target program code, can be used together to embed the above prompt information template. Specifically, a code line identifier (such as code line number, etc.) can be set for each line in the target program code. Then, the function blocks contained in the target program code and the description information of each function block, as well as the target program code with the code line identifier, can be embedded in the above prompt information template to generate the first prompt information. In addition to the retrieval process and the enhancement process, the RAG model also includes an important process, namely, the generation process (that is, the large model inductive generation). For details, please refer to the process of step S408 below.
[0027] In step S408, second prompt information is generated based on the first prompt information, the first program code and the target program code, and the second prompt information is input into the large model to obtain an analysis log of the target program code.
[0028] Among them, the big model can include multiple types, such as a generative big model or a discriminative big model, etc. The big model in this embodiment can be a generative big model, specifically a big language model, such as ChatGPT3.5, ChatGPT3, ChatGLM, etc., which can be set according to actual conditions. The second prompt information can be a prompt information Prompt that can be directly used for the big model.
[0029] In implementation, Figure 5 As shown, new prompt information can be regenerated based on the first prompt information, the first program code and the target program code to obtain the second prompt information. Since the first prompt information and the first program code contain rich information related to the target program code, the second prompt information generated based on the first prompt information, the first program code and the target program code can make the large model more accurately identify the key paths and important events in the target program code, and then identify important log points, reducing the risk of missing key logs. The second prompt information can be input into the large model, and the large model can accurately understand the context information of the target program code through the second prompt information, analyze its logic, accurately identify the key paths and important events, and finally generate an analysis log of the target program code.
[0030] The embodiment of the present specification provides a method for processing program code, by obtaining a target program code to be processed, the target program code includes a program code of a target model, and the target model includes one or more different model operations, then, based on the model operations included in the target program code, it is determined to construct an operation topology relationship diagram corresponding to the target program code, and based on the operations included in the target program code and the operation topology relationship diagram corresponding to the target program code, it is determined that the function blocks included in the target program code and the description information of each function block, then, based on the function blocks included in the target program code and the description information of each function block, and the target program code with code line identifiers, first prompt information can be constructed, and based on the target program code, a code database is obtained that has a similarity greater than a predetermined value with the target program code. Set the first program code with a threshold value. Finally, the second prompt information can be generated based on the first prompt information, the first program code and the target program code, and the second prompt information can be input into the large model to obtain the analysis log of the target program code. In this way, by abstracting the functional blocks of the program code from fine to coarse, and combining each line of the program code with the functional block from coarse to fine, the success rate of adding analysis logs is greatly improved. Moreover, based on the RAG framework, the correctly verified expert experience and historical cases are continuously enriched to achieve a positive cycle of the success rate of analysis logs. In addition, the large model can understand the contextual information of the program code, analyze its logic, and accurately identify critical paths and important events, so that the model can identify important log points in various situations and reduce the risk of missing key logs, rather than just relying on surface patterns.
[0031] In practical applications, the specific processing methods for determining the construction of the operation topology diagram corresponding to the target program code based on the model operation contained in the target program code in the above step S404 can be various. An optional processing method is provided below, such as Figure 6 As shown, the processing may specifically include the following steps S40402 to S40406.
[0032] In step S40402, the application programming interface API provided by the target model is obtained.
[0033] In implementation, taking the target model as a neural network model as an example, the neural network model may include an input layer, a hidden layer, and an output layer, wherein the input layer is used to receive input data, the hidden layer includes many intermediate layers, such as convolutional layers, activation layers, etc., and the output layer can be used to output processing results (i.e., output data). For each model operation of the neural network model (such as convolution, activation function, pooling, etc.), a topological relationship needs to be established. Specifically, the API provided by the target model can be determined through the API provided by the deep learning framework (such as TensorFlow, PyTorch, etc.) (such as tf.function, torch.nn.Module, etc.).
[0034] In step S40404, based on the API provided by the target model, the model operation included in the target program code is determined.
[0035] In implementation, the model operations contained in the target program code can be extracted according to the specific information of the API provided by the target model.
[0036] In step S40406, based on the model operations included in the target program code, it is determined to construct an operation topology relationship diagram corresponding to the target program code.
[0037] The specific processing method of step S40406 can be found in the above-mentioned related content and will not be repeated here.
[0038] In actual applications, key-value pairs for the operation types of model operations may also be constructed, and code line identifiers may be set for target program codes. For details, see the processing of the following steps A2 and A4.
[0039] In step A2, the target program code is parsed to obtain the code line identifier corresponding to each model operation.
[0040] The code line identifier may include multiple types. For example, the code line identifier may be a line number or code of a line in the program code, and may be specifically set according to actual conditions.
[0041] In implementation, the target program code may be parsed in a variety of ways. For example, the target program code may be parsed using an AST (Abstract Syntax Tree) tool to obtain a code line identifier corresponding to each model operation.
[0042] In step A4, based on the operation type of each model operation and the code line identifier corresponding to each model operation, a key-value pair for the operation type of the model operation is constructed, and based on the code line identifier corresponding to each model operation, the target program code with the code line identifier is determined.
[0043] In implementation, an association between a specific model operation and a code line identifier can be established, so that a dictionary or other data structure can be established. Based on this, the operation type to which each model operation belongs can be used as a key, and the code line identifier corresponding to the model operation can be used as a value to construct a key-value pair for the operation type of the model operation, so that the corresponding code line can be quickly found. In addition, based on the code line identifier corresponding to each model operation, a corresponding code line identifier can be set for each line of program code in the target program code, so that the target program code with the code line identifier set can be obtained.
[0044] In practical applications, the specific processing methods for determining the function blocks contained in the target program code and the description information of each function block based on the operations contained in the target program code and the operation topology diagram corresponding to the target program code in the above step S404 can be various. The following is an optional processing method, such as Figure 7 As shown, the processing may specifically include the following steps S40408 and S40410.
[0045] In step S40408, based on the operation type and attribute information of the operations contained in the target program code, the model operations in the operation topology diagram corresponding to the target program code are clustered to determine the information of each first clustering category and the information of the model operations belonging to different first clustering categories.
[0046] In implementation, the model operation can be attributed to a higher granularity relative to the model operation, such as network granularity, etc. Specifically, the model operations in the operation topology diagram corresponding to the target program code can be clustered according to the operation type and attribute information (such as the name and / or characteristics of the operation, etc.) to which the operation contained in the target program code belongs, so as to classify operations of different operation types into some larger categories. For example, the model operations in all network layers related to convolution can be clustered into convolution operations, and the first clustering category at this time is convolution operations. For another example, the model operations in the network layer containing the Transformer module can be clustered into transformation operations, and the first clustering category at this time is transformation operations. For another example, the model operations in the network layer containing maximum pooling and mean pooling can be clustered into pooling operations, and the first clustering category at this time is pooling operations, etc. In the above manner, the model operations in the operation topology diagram corresponding to the target program code can be clustered, so as to determine the information of each first clustering category and the information of model operations belonging to different first clustering categories.
[0047] In step S40410, based on the information of each first cluster category and the information of the model operations belonging to different first cluster categories, the function blocks included in the target program code and the description information of each function block are determined.
[0048] In implementation, the first clustering category can be directly set as the corresponding function block. For example, if the first clustering category is a convolution operation, the corresponding function block can be set as a function block of a convolution operation, etc. In the above manner, the function blocks included in the target program code can be determined based on the information of each first clustering category. In addition, the information of each first clustering category and the information of model operations belonging to different first clustering categories can be analyzed to determine the function and role of each function block, and the description information of each function block can be generated based on the information of the function and role of each function block obtained by the analysis.
[0049] It should be noted that other methods may be used to implement the processing of the function blocks contained in the target program code and the description information of each function block based on the information of each first clustering category and the information of model operations belonging to different first clustering categories. For example, a specified algorithm may be set in advance, and the information of each first clustering category and the information of model operations belonging to different first clustering categories may be processed by the specified algorithm. The function blocks contained in the target program code and the description information of each function block may be determined based on the processing results obtained. The specific settings may be based on actual conditions.
[0050] In practical applications, the specific processing methods of the above step S40410 can be varied. An optional processing method is provided below, such as Figure 8 As shown, the processing may specifically include the following steps S404102 and S404104.
[0051] In step S404102, based on the information of each first cluster category and the information of model operations belonging to different first cluster categories, the determined categories are clustered to obtain information of each second cluster category and information of first cluster categories belonging to different second cluster categories.
[0052] In implementation, in actual applications, the first clustering category obtained by the above-mentioned clustering may not be able to reach a coarser-grained category. Therefore, further clustering processing can be performed on the basis of the first clustering. Specifically, the determined categories can be clustered according to the information of each first clustering category and the information of the model operations belonging to different first clustering categories, so as to further abstract the different determined categories into larger categories. For example, normalization, data enhancement, image conversion, etc. can be clustered as image data preprocessing, and the second clustering category at this time is image data preprocessing. For another example, feature extraction, convolutional layer, Transformer layer, etc. can be clustered as image information extraction, and the second clustering category at this time is image information extraction, etc. In the above manner, the determined categories can be clustered to determine the information of each second clustering category and the information of the first clustering categories belonging to different second clustering categories.
[0053] In step S404104, based on the information of each second cluster category and the information of the first cluster category belonging to a different second cluster category, the function blocks included in the target program code and the description information of each function block are determined.
[0054] In implementation, the second clustering category can be directly set as the corresponding function block. For example, if the second clustering category is image data preprocessing, the corresponding function block can be set as the image data preprocessing function block, etc. In the above manner, the function blocks included in the target program code can be determined based on the information of each second clustering category. In addition, the information of each second clustering category and the information of the first clustering categories belonging to different second clustering categories can be analyzed to determine the function and role of each function block, and the description information of each function block can be generated based on the information of the function and role of each function block obtained by the analysis.
[0055] It should be noted that other methods can be used to implement the processing of determining the function blocks contained in the target program code and the description information of each function block based on the information of each second cluster category and the information of the first cluster category belonging to different second cluster categories. For example, a specified algorithm can be pre-set, and the information of each second cluster category and the information of the first cluster category belonging to different second cluster categories can be processed by the specified algorithm, and the function blocks contained in the target program code and the description information of each function block can be determined by the obtained processing results, which can be set specifically according to actual conditions. In addition, clustering processing can be further performed to obtain more abstract cluster categories, and then obtain the corresponding function blocks and the description information of each function block, which can be set specifically according to actual conditions, and the embodiments of this specification do not limit this.
[0056] In practical applications, in order to further improve the performance of recording analysis logs, the large model can be preheated (i.e., warmed up), such as Fig. 9 As shown, please refer to the processing of the following steps B2 and B4 for details.
[0057] In step B2, a first number of second program codes are randomly collected.
[0058] The first number may be set according to actual conditions, for example, the first number may be 500 or 1000, or the first number may be 4% or 10% of the total number of program codes included in the code database.
[0059] In implementation, a first number of program codes may be randomly collected from a code database, and the collected program codes may be used as the second program codes.
[0060] In step B4, the large model is warmed up based on the second program code to obtain the warmed-up large model.
[0061] In implementation, warming up the large model can not only further improve the recording performance of the analysis log, but also improve the performance of the large model on specific tasks, and enhance the general analogy ability of the large model (that is, finding the common features of the limited context information according to each instruction in the prompt information). Specifically, for each warm-up training, the kNN algorithm, the Euclidean distance similarity algorithm, etc. can be used to obtain the program code in the code database whose similarity with the program code in the warm-up training is greater than a preset threshold, and obtain the label information (that is, the analysis log) of the above program code (that is, the obtained program code whose similarity with the program code in the warm-up training is greater than the preset threshold), and the above program code, label information and specified instructions can be packaged together to generate prompt information, and the prompt information can be input into the large model to adjust the model parameters of the large model, so as to warm up the large model, and finally, the large model after warm-up can be obtained.
[0062] Based on the processing of the above steps B2 and B4, the processing of the above step S408 may include: generating second prompt information based on the first prompt information, the first program code and the target program code, and inputting the second prompt information into the large model after warm-up to obtain an analysis log of the target program code.
[0063] The above specific processing process can be found in the above related content, which will not be repeated here.
[0064] In practical applications, if a log is printed in a program loop code statement whose traversal times exceeds a preset threshold, multiple logs (up to hundreds of logs) will be generated, and data objects with a data volume greater than a preset data volume threshold may be printed. Therefore, the analysis log generated above can also be verified. For details, please refer to the processing of steps C2 and C4 below.
[0065] In step C2, the analysis log of the target program code is input into the large model to determine whether there is a data object in the analysis log that prints logs in program loop code statements whose traversal times exceed a preset threshold or whose print data volume exceeds a preset data volume threshold, and obtain a corresponding judgment result.
[0066] The data object may include multiple types, such as images, videos, etc., which may be set according to actual conditions. The program loop code statement may include multiple types, such as for loop code statements, etc.
[0067] In implementation, the analysis log of the target program code can be input into the large model, and the large model can be used to determine whether there is a log printed in the analysis log in a program loop code statement whose traversal times exceeds a preset threshold. The above determination can prevent the analysis log from having too many log data generated due to printing logs in program loop code statements, and can determine whether to print a data object with a data volume greater than a preset data volume threshold, thereby preventing the printing of data objects with a data volume greater than the preset data volume threshold. Data related to shape shape or size size can be printed, but the entire array array should not be input into the analysis log. The corresponding judgment result can be obtained through the above processing.
[0068] In step C4, the analysis log is adjusted based on the judgment result to obtain an adjusted analysis log.
[0069] In implementation, if the judgment result indicates that the analysis log contains a program loop code statement whose traversal times exceed a preset number of times threshold, then the multiple log data generated in the analysis log due to the printing of logs in the program loop code statement can be deleted or simplified. If the judgment result indicates that the analysis log contains a data object whose print data volume is greater than a preset data volume threshold, then the relevant log data in the analysis log can be deleted or simplified, thereby adjusting the analysis log to obtain an adjusted analysis log. If the judgment result indicates that the analysis log does not contain a program loop code statement whose traversal times exceed a preset number of times threshold, and there is no data object whose print data volume is greater than the preset data volume threshold, then there is no need to adjust the analysis log, and the analysis log can be used as the final analysis log of the target program code.
[0070] In practical applications, in addition to providing the adjusted analysis log to the user or technician, the adjusted analysis log may also be added to the above code database.
[0071] In practical applications, the analysis log of the target program code includes a processing level and a processing purpose. The processing level includes one or more of Fatal, Error, Warning, Info, Debug, and Try-excepe. The processing purpose includes one or more of time consumption, determining the size of the tensor Tensor, and determining whether the hardware position of the tensor Tensor has changed.
[0072] The embodiment of the present specification provides a method for processing program code, by obtaining a target program code to be processed, the target program code includes a program code of a target model, and the target model includes one or more different model operations, then, based on the model operations included in the target program code, it is determined to construct an operation topology relationship diagram corresponding to the target program code, and based on the operations included in the target program code and the operation topology relationship diagram corresponding to the target program code, it is determined that the function blocks included in the target program code and the description information of each function block, then, based on the function blocks included in the target program code and the description information of each function block, and the target program code with code line identifiers, first prompt information can be constructed, and based on the target program code, a code database is obtained that has a similarity greater than a predetermined value with the target program code. Set the first program code with a threshold value. Finally, the second prompt information can be generated based on the first prompt information, the first program code and the target program code, and the second prompt information can be input into the large model to obtain the analysis log of the target program code. In this way, by abstracting the functional blocks of the program code from fine to coarse, and combining each line of the program code with the functional block from coarse to fine, the success rate of adding analysis logs is greatly improved. Moreover, based on the RAG framework, the correctly verified expert experience and historical cases are continuously enriched to achieve a positive cycle of the success rate of analysis logs. In addition, the large model can understand the contextual information of the program code, analyze its logic, and accurately identify critical paths and important events, so that the model can identify important log points in various situations and reduce the risk of missing key logs, rather than just relying on surface patterns.
[0073] In addition, a verification process for analysis logs is added to avoid the addition of unexpected logs. Moreover, the log types are subdivided to achieve more effective log addition, so that logs can be added in places where time fluctuations, large amount of calculations, unknown input and output data volumes, and critical impacts on model stability may occur, so as to obtain better analysis logs.
[0074] The above is a method for processing program code provided in the embodiments of this specification. Based on the same idea, the embodiments of this specification also provide a device for processing program code, such as Fig.10 shown.
[0075] The program code processing device includes: a code acquisition module 1001, a code processing module 1002, a search enhancement module 1003 and a log generation module 1004, wherein: A code acquisition module 1001 acquires a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations; The code processing module 1002 determines to construct an operation topology relationship diagram corresponding to the target program code based on the model operation included in the target program code, and determines the function blocks included in the target program code and the description information of each function block based on the operation included in the target program code and the operation topology relationship diagram corresponding to the target program code; The search enhancement module 1003 constructs first prompt information based on the function blocks contained in the target program code and the description information of each of the function blocks, and the target program code with code line identifiers, and obtains a first program code whose similarity with the target program code is greater than a preset threshold from a code database based on the target program code; The log generation module 1004 generates second prompt information based on the first prompt information, the first program code and the target program code, and inputs the second prompt information into the large model to obtain an analysis log of the target program code.
[0076] In the embodiment of this specification, the code processing module 1002 includes: An interface acquisition unit, which acquires an application programming interface API provided by the target model; A model operation determination unit, which determines the model operation included in the target program code based on the API provided by the target model; A topology unit determines to construct an operation topology relationship diagram corresponding to the target program code based on the model operation included in the target program code.
[0077] In the embodiment of this specification, the device further includes: A parsing module, which parses the target program code to obtain a code line identifier corresponding to each model operation; The data processing module constructs a key-value pair for the operation type of the model operation based on the operation type of each model operation and the code line identifier corresponding to each model operation, and determines the target program code with the code line identifier based on the code line identifier corresponding to each model operation.
[0078] In the embodiment of this specification, the code processing module 1002 includes: A first clustering unit, based on the operation type and attribute information of the operation contained in the target program code, clusters the model operations in the operation topology diagram corresponding to the target program code, and determines information of each first cluster category and information of model operations belonging to different first cluster categories; The function determination unit determines the function blocks contained in the target program code and the description information of each of the function blocks based on the information of each first cluster category and the information of the model operations belonging to different first cluster categories.
[0079] In an embodiment of the present specification, the function determination unit clusters the determined categories based on the information of each first clustering category and the information of model operations belonging to different first clustering categories to obtain information of each second clustering category and information of first clustering categories belonging to different second clustering categories; based on the information of each second clustering category and the information of first clustering categories belonging to different second clustering categories, determines the function blocks contained in the target program code and the description information of each of the function blocks.
[0080] In the embodiment of this specification, the device further includes: A sampling module randomly collects a first number of second program codes from a code database; A warm-up module, performing a warm-up process on the large model based on the second program code to obtain a warmed-up large model; The log generation module 1004 generates second prompt information based on the first prompt information, the first program code and the target program code, and inputs the second prompt information into the warm-up large model to obtain the analysis log of the target program code.
[0081] In the embodiment of this specification, the device further includes: The log processing module inputs the analysis log of the target program code into the large model to determine whether there is a data object in the analysis log that prints logs in a program loop code statement whose traversal times exceed a preset number threshold or whose print data volume exceeds a preset data volume threshold, and obtains a corresponding determination result; The log adjustment module adjusts the analysis log based on the judgment result to obtain an adjusted analysis log.
[0082] In the embodiment of this specification, the analysis log of the target program code includes a processing level and a processing purpose, the processing level includes one or more of Fatal, Error, Warning, Info, Debug, and Try-excepe, and the processing purpose includes one or more of time consumption, determining the size of the tensor Tensor, and determining whether the hardware position of the tensor Tensor has changed.
[0083] The embodiment of the present specification provides a program code processing device, which obtains a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations. Then, based on the model operations included in the target program code, an operation topology relationship diagram corresponding to the target program code can be determined, and based on the operations included in the target program code and the operation topology relationship diagram corresponding to the target program code, a function block included in the target program code and description information of each function block can be determined. Then, based on the function block included in the target program code and the description information of each function block, and the target program code with a code line identifier, first prompt information can be constructed, and based on the target program code, a code database can be used to obtain a code whose similarity to the target program code is greater than a predetermined value. Set the first program code with a threshold value. Finally, the second prompt information can be generated based on the first prompt information, the first program code and the target program code, and the second prompt information can be input into the large model to obtain the analysis log of the target program code. In this way, by abstracting the functional blocks of the program code from fine to coarse, and combining each line of the program code with the functional block from coarse to fine, the success rate of adding analysis logs is greatly improved. Moreover, based on the RAG framework, the correctly verified expert experience and historical cases are continuously enriched to achieve a positive cycle of the success rate of analysis logs. In addition, the large model can understand the contextual information of the program code, analyze its logic, and accurately identify critical paths and important events, so that the model can identify important log points in various situations and reduce the risk of missing key logs, rather than just relying on surface patterns.
[0084] In addition, a verification process for analysis logs is added to avoid the addition of unexpected logs. Moreover, the log types are subdivided to achieve more effective log addition, so that logs can be added in places where time fluctuations, large amount of calculations, unknown input and output data volumes, and critical impacts on model stability may occur, so as to obtain better analysis logs.
[0085] The above is a program code processing device provided in the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a program code processing device, such as Fig.11 shown.
[0086] The processing device of the program code may provide a terminal device or a server, etc. for the above-mentioned embodiment.
[0087] The processing device of the program code may have relatively large differences due to different configurations or performances, and may include one or more processors 1101 and memory 1102, and one or more storage applications or data may be stored in the memory 1102. Among them, the memory 1102 may be a short-term storage or a permanent storage. The application stored in the memory 1102 may include one or more modules (not shown in the figure), and each module may include a series of computer executable instructions in the processing device of the program code. Furthermore, the processor 1101 may be configured to communicate with the memory 1102, and execute a series of computer executable instructions in the memory 1102 on the processing device of the program code. The processing device of the program code may also include one or more power supplies 1103, one or more wired or wireless network interfaces 1104, one or more input and output interfaces 1105, and one or more keyboards 1106.
[0088] Specifically in this embodiment, the program code processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer executable instructions in the program code processing device, and the one or more programs are configured to be executed by one or more processors, including the following computer executable instructions: Acquire a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations; Based on the model operations included in the target program code, determine to construct an operation topology relationship graph corresponding to the target program code, and based on the operations included in the target program code and the operation topology relationship graph corresponding to the target program code, determine the function blocks included in the target program code and description information of each of the function blocks; Based on the function blocks contained in the target program code and the description information of each of the function blocks, and the target program code with code line identifiers, construct first prompt information, and based on the target program code, obtain from a code database a first program code whose similarity with the target program code is greater than a preset threshold; Second prompt information is generated based on the first prompt information, the first program code and the target program code, and the second prompt information is input into the large model to obtain an analysis log of the target program code.
[0089] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the program code processing device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0090] The embodiment of the present specification provides a program code processing device, which obtains a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations. Then, based on the model operations included in the target program code, an operation topology relationship diagram corresponding to the target program code can be determined, and based on the operations included in the target program code and the operation topology relationship diagram corresponding to the target program code, a function block included in the target program code and description information of each function block can be determined. Then, based on the function block included in the target program code and the description information of each function block, and the target program code with a code line identifier, first prompt information can be constructed, and based on the target program code, a code database can be used to obtain a code whose similarity to the target program code is greater than a predetermined value. Set the first program code with a threshold value. Finally, the second prompt information can be generated based on the first prompt information, the first program code and the target program code, and the second prompt information can be input into the large model to obtain the analysis log of the target program code. In this way, by abstracting the functional blocks of the program code from fine to coarse, and combining each line of the program code with the functional block from coarse to fine, the success rate of adding analysis logs is greatly improved. Moreover, based on the RAG framework, the correctly verified expert experience and historical cases are continuously enriched to achieve a positive cycle of the success rate of analysis logs. In addition, the large model can understand the contextual information of the program code, analyze its logic, and accurately identify critical paths and important events, so that the model can identify important log points in various situations and reduce the risk of missing key logs, rather than just relying on surface patterns.
[0091] Furthermore, based on the above Figures 4 to 9 In one embodiment, the present specification further provides a storage medium for storing computer executable instruction information. In a specific embodiment, the storage medium may be a USB flash drive, an optical disk, a hard disk, etc. When the computer executable instruction information stored in the storage medium is executed by the processor, the following process can be implemented: Acquire a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations; Based on the model operations included in the target program code, determine to construct an operation topology relationship graph corresponding to the target program code, and based on the operations included in the target program code and the operation topology relationship graph corresponding to the target program code, determine the function blocks included in the target program code and description information of each of the function blocks; Based on the function blocks contained in the target program code and the description information of each of the function blocks, and the target program code with code line identifiers, construct first prompt information, and based on the target program code, obtain from a code database a first program code whose similarity with the target program code is greater than a preset threshold; Second prompt information is generated based on the first prompt information, the first program code and the target program code, and the second prompt information is input into the large model to obtain an analysis log of the target program code.
[0092] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0093] The embodiment of the present specification provides a storage medium, which obtains a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations. Then, based on the model operations included in the target program code, an operation topology relationship diagram corresponding to the target program code can be determined, and based on the operations included in the target program code and the operation topology relationship diagram corresponding to the target program code, the function blocks included in the target program code and the description information of each function block can be determined. Then, based on the function blocks included in the target program code and the description information of each function block, and the target program code with code line identifiers, first prompt information can be constructed, and based on the target program code, a code database can be used to obtain a code whose similarity with the target program code is greater than a preset threshold. The first program code can eventually generate the second prompt information based on the first prompt information, the first program code and the target program code, and input the second prompt information into the large model to obtain the analysis log of the target program code. In this way, by abstracting the functional blocks of the program code from fine to coarse, and combining each line of the program code with the functional block from coarse to fine, the success rate of adding analysis logs is greatly improved. Moreover, based on the RAG framework, the correctly verified expert experience and historical cases are continuously enriched to achieve a positive cycle of the success rate of analysis logs. In addition, the large model can understand the contextual information of the program code, analyze its logic, and accurately identify critical paths and important events, so that the model can identify important log points in various situations and reduce the risk of missing key logs, rather than just relying on surface patterns.
[0094] Furthermore, based on the above Figures 4 to 9 In one or more embodiments of the present specification, a computer program product is provided, including a computer program. When the computer program in the computer program product is executed by a processor, the following process can be implemented: Acquire a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations; Based on the model operations included in the target program code, determine to construct an operation topology relationship graph corresponding to the target program code, and based on the operations included in the target program code and the operation topology relationship graph corresponding to the target program code, determine the function blocks included in the target program code and description information of each of the function blocks; Based on the function blocks contained in the target program code and the description information of each of the function blocks, and the target program code with code line identifiers, construct first prompt information, and based on the target program code, obtain from a code database a first program code whose similarity with the target program code is greater than a preset threshold; Second prompt information is generated based on the first prompt information, the first program code and the target program code, and the second prompt information is input into the large model to obtain an analysis log of the target program code.
[0095] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned computer program product embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0096] The embodiment of the present specification provides a computer program product, which obtains a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations. Then, based on the model operations included in the target program code, an operation topology relationship diagram corresponding to the target program code can be determined, and based on the operations included in the target program code and the operation topology relationship diagram corresponding to the target program code, the function blocks included in the target program code and the description information of each function block can be determined. Then, based on the function blocks included in the target program code and the description information of each function block, and the target program code with code line identifiers, first prompt information can be constructed, and based on the target program code, a code database can be used to obtain a code database having a similarity with the target program code greater than a preset value. The first program code of the threshold value can eventually generate the second prompt information based on the first prompt information, the first program code and the target program code, and input the second prompt information into the large model to obtain the analysis log of the target program code. In this way, by abstracting the functional blocks of the program code from fine to coarse, and combining each line of the program code with the functional block from coarse to fine, the success rate of adding analysis logs is greatly improved. Moreover, based on the RAG framework, the correctly verified expert experience and historical cases are continuously enriched to achieve a positive cycle of the success rate of analysis logs. In addition, the large model can understand the contextual information of the program code, analyze its logic, and accurately identify critical paths and important events, so that the model can identify important log points in various situations and reduce the risk of missing key logs, rather than just relying on surface patterns.
[0097] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0098] In the 1990s, it was very clear whether the improvement of a technology was hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0099] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.
[0100] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0101] For the convenience of description, the above devices are described in terms of functions and are divided into various units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0102] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, one or more embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0103] The embodiments of this specification are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable fraud case serial and parallel device to produce a machine, so that the instructions executed by the processor of the computer or other programmable fraud case serial and parallel device generate instructions for implementing the processes in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0104] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable fraud case serial and parallel device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0105] These computer program instructions may also be loaded onto a computer or other programmable device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0106] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0107] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0108] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0109] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0110] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, one or more embodiments of this specification may be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0111] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0112] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0113] The above description is only an embodiment of this specification and is not intended to limit this document. For those skilled in the art, this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included in the scope of the claims of this specification.
Claims
1. A method for processing a program code, the method comprising: Acquire a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations; Based on the model operations included in the target program code, determine to construct an operation topology relationship graph corresponding to the target program code, and based on the operations included in the target program code and the operation topology relationship graph corresponding to the target program code, determine the function blocks included in the target program code and description information of each of the function blocks; Based on the function blocks contained in the target program code and the description information of each of the function blocks, and the target program code with code line identifiers, construct first prompt information, and based on the target program code, obtain from a code database a first program code whose similarity with the target program code is greater than a preset threshold; Second prompt information is generated based on the first prompt information, the first program code and the target program code, and the second prompt information is input into the large model to obtain an analysis log of the target program code.
2. The method according to claim 1, wherein determining to construct an operation topology diagram corresponding to the target program code based on the model operation contained in the target program code comprises: Obtaining an application programming interface API provided by the target model; Determining the model operation included in the target program code based on the API provided by the target model; Based on the model operation included in the target program code, it is determined to construct an operation topology relationship diagram corresponding to the target program code.
3. The method according to claim 1, further comprising: Parsing the target program code to obtain a code line identifier corresponding to each model operation; Based on the operation type of each model operation and the code line identifier corresponding to each model operation, a key-value pair for the operation type of the model operation is constructed, and based on the code line identifier corresponding to each model operation, a target program code with a code line identifier is determined.
4. The method according to claim 1, wherein determining the function blocks contained in the target program code and the description information of each function block based on the operations contained in the target program code and the operation topology relationship diagram corresponding to the target program code comprises: Based on the operation type and attribute information of the operation contained in the target program code, clustering the model operations in the operation topology relationship diagram corresponding to the target program code, and determining information of each first cluster category and information of model operations belonging to different first cluster categories; Based on the information of each first cluster category and the information of the model operations belonging to different first cluster categories, the function blocks included in the target program code and the description information of each of the function blocks are determined.
5. The method according to claim 4, wherein the determining the function blocks contained in the target program code and the description information of each function block based on the information of each first cluster category and the information of the model operations belonging to different first cluster categories comprises: Based on the information of each first cluster category and the information of the model operations belonging to different first cluster categories, clustering the determined categories to obtain information of each second cluster category and information of the first cluster categories belonging to different second cluster categories; Based on the information of each second cluster category and the information of the first cluster category belonging to a different second cluster category, the function blocks contained in the target program code and the description information of each of the function blocks are determined.
6. The method according to claim 5, further comprising: randomly collecting a first number of second program codes from a code database; Performing a warm-up process on the large model based on the second program code to obtain a warmed-up large model; The step of generating second prompt information based on the first prompt information, the first program code, and the target program code, and inputting the second prompt information into a large model to obtain an analysis log of the target program code includes: Second prompt information is generated based on the first prompt information, the first program code and the target program code, and the second prompt information is input into the warm-up large model to obtain an analysis log of the target program code.
7. The method according to claim 6, further comprising: Inputting the analysis log of the target program code into the large model to determine whether there is a data object in the analysis log that prints logs in a program loop code statement whose traversal times exceed a preset number threshold or whose print data volume exceeds a preset data volume threshold, and obtaining a corresponding determination result; The analysis log is adjusted based on the judgment result to obtain an adjusted analysis log.
8. According to the method of claim 7, the analysis log of the target program code includes a processing level and a processing purpose, the processing level includes one or more of Fatal, Error, Warn, Info, Debug and Try-excepe, and the processing purpose includes one or more of time consumption, determining the size of the tensor Tensor and determining whether the hardware position of the tensor Tensor has changed.
9. A program code processing device, the device comprising: A code acquisition module, which acquires a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations; A code processing module determines to construct an operation topology relationship graph corresponding to the target program code based on the model operation included in the target program code, and determines the function blocks included in the target program code and description information of each of the function blocks based on the operation included in the target program code and the operation topology relationship graph corresponding to the target program code; The search enhancement module constructs first prompt information based on the function blocks contained in the target program code and the description information of each of the function blocks, and the target program code with code line identifiers, and obtains a first program code whose similarity with the target program code is greater than a preset threshold from a code database based on the target program code; The log generation module generates second prompt information based on the first prompt information, the first program code and the target program code, and inputs the second prompt information into the large model to obtain an analysis log of the target program code.
10. A program code processing device, the program code processing device comprising: processor; as well as a memory arranged to store computer executable instructions which, when executed, cause the processor to: Acquire a target program code to be processed, wherein the target program code includes a program code of a target model, and the target model includes one or more different model operations; Based on the model operations included in the target program code, determine to construct an operation topology relationship graph corresponding to the target program code, and based on the operations included in the target program code and the operation topology relationship graph corresponding to the target program code, determine the function blocks included in the target program code and description information of each of the function blocks; Based on the function blocks contained in the target program code and the description information of each of the function blocks, and the target program code with code line identifiers, construct first prompt information, and based on the target program code, obtain from a code database a first program code whose similarity with the target program code is greater than a preset threshold; Second prompt information is generated based on the first prompt information, the first program code and the target program code, and the second prompt information is input into the large model to obtain an analysis log of the target program code.
Citation Information
Patent Citations
Log code generation method and device, computer system and readable storage medium
CN111221521A
Data processing method, device and equipment
CN115758471A
Code processing method and device, electronic equipment and storage medium
CN116991412A
User instruction execution method and device based on large language model
CN118446224A
Hierarchical code abstract generation method based on large language model thinking chain
CN118550579A
Cited By
Method and device for constructing API topological relation graph, equipment and medium
CN121350516A