Abnormal data generation method and abnormal code generation method
By using the target generation model to insert exception content corresponding to exception description information into the target data fragment, the problem of low efficiency in generating exception data and exception codes in the prior art is solved, and more efficient and automated exception content generation is achieved.
Patent Information
- Application Number
- CN202510310552.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-17
AI Technical Summary
In the prior art, generating exception data and exception codes requires a lot of manpower, and the generation efficiency is low, resulting in limited data size and diversity.
By obtaining the target data fragment and exception description information, the target generation model is used to insert the exception content corresponding to the exception description information into the target data fragment to generate the exception data fragment or exception code fragment.
It realizes more accurate and automated insertion of exception content, improves the generation scale and efficiency of exception data and exception code, and can adapt to data warehouses of different sizes and complexities.
Smart Images

Figure CN119807019B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and in particular to a method for generating abnormal data and a method for generating abnormal codes. Background Art
[0002] With the rapid development of computer technology and artificial intelligence technology, large models are being used more and more widely and are now able to perform various complex language tasks, including code generation, writing, translation, dialogue, question-answering, etc. When debugging large models, a large number of high-quality debugging data sets are required, which often need to include correct data, abnormal data, etc.
[0003] In the existing technology, abnormal data is often manually written or inserted into the correct data based on rules to obtain abnormal data. Manually writing abnormal data requires a lot of manpower and also introduces uncertain factors, resulting in limited scale and diversity of abnormal data generation and extremely low generation efficiency. Therefore, an efficient and accurate abnormal data generation solution is urgently needed. Summary of the invention
[0004] In view of this, an embodiment of this specification provides an abnormal data generation method. One or more embodiments of this specification also relate to an abnormal code generation method, an information processing method based on a target generation model, an abnormal data generation device, an abnormal code generation device, a computing device, an electronic device, a computer-readable storage medium and a computer program product to solve the technical defects existing in the prior art.
[0005] According to a first aspect of an embodiment of this specification, a method for generating abnormal data is provided, comprising:
[0006] Acquire task information of an abnormal data generation task, wherein the task information includes a target data segment and abnormal description information, the target data segment is an original data segment in a target data unit into which abnormal content is to be inserted, and the target data segment is determined based on test information recorded by testing the target data unit;
[0007] Based on the target data segment and the exception description information, at least one exception data segment is generated by using a target generation model corresponding to the exception data generation task, wherein the exception data segment is a data segment obtained by inserting the exception content corresponding to the exception description information into the target data segment;
[0008] Based on at least one abnormal data segment corresponding to each target data segment of the target data unit, an abnormality generation result of the target data unit is obtained.
[0009] According to a second aspect of an embodiment of this specification, there is provided a method for generating an exception code, comprising:
[0010] Acquire task information of an abnormal code generation task, wherein the task information includes a target code segment and abnormal description information, the target code segment is an original code segment to be inserted into the target test unit into which abnormal content is to be inserted, and the target code segment is determined based on call information recorded by code testing of the target test unit;
[0011] Based on the target code snippet and the exception description information, at least one exception code snippet is generated through a target generation model corresponding to the exception code generation task, wherein the exception code snippet is a code snippet obtained by inserting the exception content corresponding to the exception description information into the target code snippet;
[0012] Based on at least one abnormal code segment corresponding to each target code segment of the target test unit, an abnormality generation result of the target test unit is obtained.
[0013] According to a third aspect of an embodiment of this specification, there is provided an information processing method based on a target generation model, which is applied to a task platform and includes:
[0014] Receiving a model request sent by a terminal device, wherein the model request includes a scene identifier of a target scene, scene input data of the target scene, and at least one of a model specification parameter;
[0015] Based on the model request, a corresponding target generation model is determined from at least one generation model, wherein the target generation model is used to execute the generation process of the abnormal data fragment in the above-mentioned abnormal data generation method.
[0016] According to a fourth aspect of the embodiments of this specification, there is provided an abnormal data generating device, including:
[0017] A first acquisition module is configured to acquire task information of an abnormal data generation task, wherein the task information includes a target data segment and abnormal description information, the target data segment is an original data segment in a target data unit into which abnormal content is to be inserted, and the target data segment is determined based on test information recorded by testing the target data unit;
[0018] A first generating module is configured to generate at least one abnormal data segment based on the target data segment and the abnormal description information by using a target generation model corresponding to the abnormal data generation task, wherein the abnormal data segment is a data segment obtained by inserting abnormal content corresponding to the abnormal description information into the target data segment;
[0019] The first obtaining module is configured to obtain an abnormality generation result of the target data unit based on at least one abnormal data segment corresponding to each target data segment of the target data unit.
[0020] According to a fifth aspect of the embodiments of this specification, there is provided an exception code generating device, comprising:
[0021] A second acquisition module is configured to acquire task information of an abnormal code generation task, wherein the task information includes a target code snippet and abnormal description information, the target code snippet is an original code snippet in a target test unit to be inserted with abnormal content, and the target code snippet is determined based on call information recorded by code testing of the target test unit;
[0022] A second generation module is configured to generate at least one abnormal code snippet based on the target code snippet and the abnormal description information through a target generation model corresponding to the abnormal code generation task, wherein the abnormal code snippet is a code snippet obtained by inserting abnormal content corresponding to the abnormal description information into the target code snippet;
[0023] The second obtaining module is configured to obtain the exception generation result of the target test unit based on at least one exception code fragment corresponding to each target code fragment of the target test unit.
[0024] According to a sixth aspect of an embodiment of this specification, a computing device is provided, including:
[0025] Memory and processor;
[0026] Among them, the memory is used to store computer programs / instructions, and the processor is used to execute computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the above-mentioned abnormal data generation method or abnormal code generation method or information processing method based on the target generation model are implemented.
[0027] According to a seventh aspect of the embodiments of this specification, an electronic device is provided, including:
[0028] A memory and a processor, wherein the memory and the processor are connected via a bus;
[0029] Among them, the memory is used to store computer programs / instructions, and the processor is used to execute computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the above-mentioned abnormal data generation method or abnormal code generation method or information processing method based on the target generation model are implemented.
[0030] According to an eighth aspect of an embodiment of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned exception data generation method or exception code generation method or information processing method based on a target generation model.
[0031] According to the ninth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned exception data generation method or exception code generation method or information processing method based on a target generation model.
[0032] An embodiment of the present specification provides an abnormal data generation method, which obtains task information of an abnormal data generation task, wherein the task information includes a target data segment and abnormal description information, the target data segment is an original data segment in a target data unit into which abnormal content is to be inserted, and the target data segment is determined based on test information recorded by testing the target data unit; based on the target data segment and the abnormal description information, at least one abnormal data segment is generated through a target generation model corresponding to the abnormal data generation task, wherein the abnormal data segment is a data segment obtained by inserting the abnormal content corresponding to the abnormal description information into the target data segment; based on at least one abnormal data segment corresponding to each target data segment of the target data unit, an abnormal generation result of the target data unit is obtained.
[0033] An embodiment of the present specification realizes that, based on the test information recorded by testing the target data unit, the target data segment to be inserted with the abnormal content is determined from the target data unit, and the abnormal content corresponding to the abnormal description information can be inserted into the target data segment by using the target generation model in combination with the target data segment and the abnormal description information to obtain the abnormal data segment, thereby obtaining the abnormal generation result of the target data unit. In this way, based on the abnormal description information, the target generation model can generate more diverse abnormal data segments, improve the data quality, and can effectively and quickly locate the target data segment in the target data unit where the abnormal content needs to be inserted based on the test information recorded by testing the target data unit, thereby achieving more accurate and automated abnormal content insertion, providing a full-process automation solution from testing to abnormal content insertion, greatly improving the scale and efficiency of abnormal generation, and can adapt to data warehouses of different sizes and complexities, with good scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is an application architecture diagram of an abnormal data generation task provided by an embodiment of this specification;
[0035] Figure 2is a flow chart of a method for generating abnormal data provided by an embodiment of this specification;
[0036] Figure 3 is a schematic diagram of a target data segment provided by an embodiment of this specification;
[0037] Figure 4 is a schematic diagram of a target data segment and a corresponding abnormal data segment provided by an embodiment of this specification;
[0038] Figure 5 is a flow chart of an exception code generation method provided by an embodiment of this specification;
[0039] Figure 6 It is a schematic diagram of a processing process of an exception code generation method provided by an embodiment of this specification;
[0040] Figure 7 is a flowchart of an information processing method based on a target generation model provided by an embodiment of this specification;
[0041] Figure 8 It is a structural schematic diagram of an abnormal data generating device provided by an embodiment of this specification;
[0042] Fig. 9 It is a structural schematic diagram of an abnormal code generating device provided by an embodiment of this specification;
[0043] Fig.10 is a structural block diagram of a computing device provided by one embodiment of this specification;
[0044] Fig.11 It is a structural block diagram of an electronic device provided by an embodiment of this specification. DETAILED DESCRIPTION
[0045] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.
[0046] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0047] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0048] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0049] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, which usually contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than 10 trillion model parameters. A large model can also be called a foundation model / foundation model. The large model is pre-trained with large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as a large-scale language model (LLM, Large Language Model), a multi-modal pre-training model, etc.
[0050] When the big model is used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. The big model can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image description (IC, Image Caption), image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the big model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0051] First, the terms involved in one or more embodiments of this specification are explained.
[0052] AST (Abstract Syntax Tree): It is an intermediate representation used in compilers and interpreters to represent the structure of source code. It is a tree-like data structure that facilitates code analysis and conversion, where each node represents a grammatical structure in the source code, such as an expression, statement, or declaration. For example, a function call, variable declaration, conditional statement, etc. can all be a node. The root node is the starting point of the abstract syntax tree, usually representing the entire program or file. Leaf nodes are nodes without child nodes, usually specific values or identifiers. Branch nodes contain other nodes as child nodes, usually representing composite structures, such as conditional statements, loops, etc.
[0053] Pytest framework: is a powerful and easy-to-use testing framework for the Python programming language, used to write and execute tests. It is widely used in unit testing, integration testing, and end-to-end testing. It has concise syntax, a rich plug-in ecosystem, and a powerful feature set, which can significantly improve the writing efficiency and maintainability of test code.
[0054] Tracing plug-in: It can be a tool of the programming language. It can set a tracing function when the code is running to monitor and trace the execution process of the code, including function calls and execution line numbers. Through this tracing plug-in, you can track events such as function calls, line executions, exceptions, etc.
[0055] Token-level detection: In natural language processing (NLP) and related fields, it often refers to some form of analysis or processing of each element (token) in the data. This analysis can include syntax checking, named entity recognition (NER), part-of-speech tagging (POS tagging), sentiment analysis, etc. In the code generation scenario, token-level detection can refer to operations that are refined to the basic element level of the code (such as operators and variable names) during code synthesis or analysis.
[0056] Bug: refers to an error or defect in the code that may cause the code to not run properly or produce incorrect results.
[0057] Code Repository: It is a system for storing and managing source code. The code repository can contain code files of multiple projects or modules and support collaborative development by multiple people. Repository-level operations include cloning repositories, submitting changes, and pulling new code.
[0058] Unit Test Suite: Usually refers to the unit test suite, which is a collection of unit tests used to automatically run. Unit testing is a testing method in the software development process that aims to verify that each individual part (usually a function or method) works as expected. Unit test suites can help developers quickly and automatically check the correctness of the code, especially when the code base is constantly changing and expanding.
[0059] It should be noted that, taking the error code generation scenario as an example, LLM has been widely used in code-related tasks such as code generation and translation. However, due to the lack of high-quality training and evaluation of code debugging datasets, the debugging capabilities of LLM at the repository level are still not fully explored.
[0060] In one implementation, the error code can be manually written or inserted based on rules, but the scale and diversity of the constructed code set are limited. In another implementation, a tool or framework for debugging and benchmarking can collect function-level source data from an online platform that provides programming challenges related to algorithms and data structures, and use a large model to implant bugs into code snippets. For implanting multiple bugs, multiple single bugs can be merged. Subsequently, invalid code can be automatically filtered based on rules and manually checked. In another implementation, a bug data warehouse and benchmarking tool designed for Python projects can collect known defects from multiple open source Python projects and provide corresponding repair patches. It can provide warehouse-level Python code data, but its code data is manually constructed, and the solution is difficult to expand to any code warehouse on a large scale and continuously.
[0061] In the above implementation, the code debugging dataset is mainly function-level code. Compared with function-level code, warehouse-level code generally has more code files and more single test suites. Function-level code cannot be directly applied to the construction of warehouse-level code, and it is difficult to meet the needs of training and testing LLM on warehouse-level code debugging tasks. In addition, the code in the code warehouse (such as GitHub) is usually relatively complete and cannot reflect the fine-grained bugs encountered during the development and debugging process. It relies on manual and rule restrictions, and it is difficult to generate real and diverse error code data on a large scale. When merging a single bug, it is more likely that the bug positions will overlap, resulting in the actual number of bugs not meeting the requirements. Manual comparison and screening will take a lot of time.
[0062] The embodiments of this specification provide an exception data generation scheme, which can automatically synthesize error codes based on the code warehouse, obtain a warehouse-level code debugging data set, dynamically capture the call chain of the code warehouse during testing, use AST for precise code block parsing and LLM for automatic error implantation, especially in the case of synthesizing multiple bugs, token-level position detection can be used to ensure the accuracy of the synthesis, and LLM can be used to generate more diverse error codes, and these errors are closer to real bugs, avoiding the problems of frequent manual intervention and low efficiency, and at the same time filling the gap in error data of synthetic warehouse-level code, providing strong data support for the debugging capabilities of the code processing platform.
[0063] To solve the above technical problems, in this specification, a method for generating abnormal data is provided. One or more embodiments of this specification also relate to a method for generating abnormal code, an information processing method based on a target generation model, an abnormal data generating device, an abnormal code generating device, a computing device, an electronic device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.
[0064] Considering the huge number of model parameters of the large model and the limited computing resources of the mobile terminal, the abnormal data generation method provided in the embodiment of the present application can be applied to Figure 1 The application architecture diagram shown in the figure is not limited to this. Figure 1 In the application architecture shown, the large model is deployed in the server 10, and the server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. The client device 20 here may include but is not limited to: a smart phone, a tablet computer, a laptop computer, a PDA, a personal computer, a smart home device, a vehicle-mounted device, etc. The client device 20 can interact with the user through a graphical user interface to implement the call of the large model, thereby implementing the method provided in the embodiment of this specification.
[0065] In the embodiment of the present specification, the system composed of the client device and the server can perform the following steps: the client device 20 executes sending an abnormal data generation task to the server 10. The server 10 executes to obtain task information of the abnormal data generation task, wherein the task information includes a target data segment and abnormal description information, the target data segment is a data segment to be inserted into the target data unit into which abnormal content is to be inserted, and the target data segment is determined based on the test information recorded by testing the target data unit; based on the target data segment and the abnormal description information, at least one abnormal data segment is generated through the target generation model corresponding to the abnormal data generation task, wherein the abnormal data segment is a data segment obtained by inserting abnormal content corresponding to the abnormal description information into the target data segment; based on at least one abnormal data segment corresponding to each target data segment of the target data unit, the abnormal generation result of the target data unit is obtained.
[0066] In an optional embodiment of the present specification, the server is further used to return an exception generation result to the client; the client is further used to receive the exception generation result sent by the server.
[0067] By applying the scheme of the embodiments of this specification, more diverse abnormal data fragments can be generated based on the abnormal description information using the target generation model, thereby improving data quality. In addition, the test information recorded by the test of the target data unit can be used to effectively and quickly locate the target data fragment in the target data unit where the abnormal content needs to be inserted, thereby achieving more accurate and automated insertion of abnormal content. A full-process automation solution from testing to insertion of abnormal content is provided, which greatly improves the scale and efficiency of abnormal generation, can adapt to data warehouses of different sizes and complexities, and has good scalability.
[0068] It should be noted that, when the operating resources of the client device can meet the deployment and operating conditions of the large model, the embodiments of the present application can be carried out in the client device.
[0069] See also Figure 2 , Figure 2 A flowchart of a method for generating abnormal data according to an embodiment of the present specification is shown, which specifically includes the following steps 202-206.
[0070] Step 202: Obtain task information of the abnormal data generation task, wherein the task information includes a target data segment and abnormal description information. The target data segment is a data segment in the target data unit to be inserted with abnormal content, and the target data segment is determined based on test information recorded by testing the target data unit.
[0071] The embodiments of this specification are applied to applications, websites or mini-programs that have the function of processing abnormal data generation tasks. On the application, website or mini-program, the abnormal data generation task processing function is implemented. For example, a website with a target generation model deployed can implement the processing functions of abnormal data generation tasks such as abnormal code generation and abnormal text generation. For example, a third-party application can implement the corresponding abnormal data generation task processing function by calling the deployed target generation model through the application programming interface (Application Programming Interface, referred to as API).
[0072] Specifically, the exception data generation task is a task to be processed that generates exception data, and the exception data can be exception text, exception code, etc. For example, in the exception code generation scenario, the exception data generation task is an exception code generation task, which is used to generate exception code corresponding to the correct code; in the exception text generation scenario, the exception data generation task is an exception text generation task, which is used to generate exception text corresponding to the correct text.
[0073] In actual implementation, the task information of the exception data generation task includes the target data segment and exception description information. The target data segment is the original data segment in the target data unit of the data warehouse where the exception content is to be inserted, that is, the target data segment is the correct data segment where the target data unit is located in the data warehouse. The exception data segment corresponding to the target data segment can be generated subsequently.
[0074] It should be noted that the data warehouse may include multiple files to be tested, and each file to be tested may include multiple data units. In the scenario of abnormal code generation, the data unit refers to a single test suite, which is mainly used to verify whether the smallest testable unit (usually a function or method) in the code works as expected; in the scenario of abnormal text generation, the data unit may refer to the smallest test unit of each file in the data warehouse to verify whether it works as expected. In specific implementation, the data units in each test file can be tested in parallel, and the original data fragments corresponding to each data unit in the data warehouse are determined as the corresponding target data fragments. The target data unit is any one of the data units, and the target data unit can correspond to at least one target data fragment, which is convenient for the subsequent generation of corresponding abnormal data fragments for each target data fragment based on the configured abnormal description information.
[0075] Among them, the exception description information refers to the detailed parameters of the exception content to be inserted. The exception description information can be pre-specified, such as the exception description information can include the exception type, exception content and / or exception position, etc. The detailed parameters of the exception content to be inserted can be specified through the exception description information to generate specified exception data. By configuring different exception description information, exception data of different exception types, different exception contents and / or different exception positions can be generated.
[0076] It should be noted that for abnormal code generation scenarios, the original correct code fragment in the code warehouse can be determined through the call information recorded during the code testing process as the target data fragment; for abnormal text generation scenarios, the target data unit can be tested and analyzed to determine the fragment in the data warehouse that includes the specified content as the target data fragment.
[0077] In an optional implementation of this embodiment, the task information further includes a data description; obtaining the task information of the abnormal data generation task includes:
[0078] Acquire test information recorded by testing the target data unit, and locate the corresponding target data segment in the data warehouse based on the test information;
[0079] Analyze the target data segment to determine the data description of the target data segment;
[0080] Get the set exception description information.
[0081] It should be noted that the task information of the abnormal data generation task may also include a data description, which may be a summary description of the target data segment. If the target data segment is a code segment, the data description may be a code description, such as a code comment; if the target data segment is a text segment, the data description may be a text summary.
[0082] In an embodiment of the present specification, test information recorded in the test of the target data unit can be obtained, and based on the test information, the corresponding original data segment can be located in the data warehouse as the target data segment, and then the target data segment can be analyzed to determine the data description of the target data segment, and the target data segment, data description and exception description information can be used as task information of the exception data generation task, so as to facilitate the subsequent insertion of corresponding exception content in the target data segment based on the target data segment, data description and exception description information, and obtain the corresponding exception data segment.
[0083] In actual implementation, taking the scenario of generating exception codes as an example, the target data segment is the target code segment, and the code comments in the target code segment are extracted as the data description. Specifically, select an appropriate programming language and the corresponding libraries or tools. For example, use the re module or the tokenize module of the Python programming language to process the target code segment, identify the comment formats in the corresponding programming language. Common comment types include single-line comments (such as / / or #) and multi-line comments (such as / *... * / ); then, read the target code segment and use regular expressions or parsers to extract comments and identify the comment part as the corresponding data description.
[0084] Taking the scenario of generating exception texts as an example, the target data segment is the target text segment. Semantic analysis is performed on the target text segment, and the text summary of the target text segment is extracted as the corresponding data description. Specifically, the target text segment can be preprocessed first, such as removing irrelevant characters (such as punctuation marks and special symbols), converting to lowercase to ensure consistency, tokenizing (splitting the text into words or phrases), removing stop words (such as common but meaningless words like "of", "is", "in", etc.), and performing stemming or lemmatization (normalizing different forms of words to their basic forms); then, by analyzing the content of each sentence, identify the key sentences and paragraphs that contain important information. This can be done by calculating the keyword frequency in the sentence, evaluating the position of the sentence in the overall text (for example, the sentences at the beginning and end are usually more important), and using natural language processing techniques to judge the importance of the sentence, etc.; after that, semantic analysis techniques can be used to identify the main themes and concepts in the target text segment. The hidden theme distribution in the target text segment can be discovered through topic modeling methods, or named entity recognition techniques can be used to identify entities such as person names, place names, and organization names, or the semantic relationship between words can also be captured through word vector models to generate the text summary of the target text segment.
[0085] For example, Figure 3 is a schematic diagram of a target data segment provided by an embodiment of this specification. As Figure 3 shown, the target data segment is the target code segment. Analyzing the target code segment as Figure 3 shown, the corresponding code comment can be obtained as the data description. The code comment is: "This is a simple function for calculating the sum of two numbers. The function accepts two parameters: a and b, and returns their sum. Check if the input is of numeric type (integer or floating point). If the input is not of numeric type, an exception is thrown. Calculate the sum of the two numbers. Return the calculation result."
[0086] In the embodiments of the present specification, the target data segment can be analyzed and the corresponding data description can be extracted as the task information of the abnormal data generation task. Subsequently, the target data segment, data description and abnormal description information can be combined to facilitate the target generation model to more accurately understand the target data segment, insert the corresponding abnormal content into the target data segment, obtain the corresponding abnormal data segment, and improve the generation accuracy of the abnormal data segment.
[0087] In an optional implementation of this embodiment, the target data segment is a target code segment, and the test information is call information of a code test process; acquiring the test information recorded by testing the target data unit, and locating the corresponding target data segment in a data warehouse based on the test information, includes:
[0088] Obtain call information recorded by executing a target data unit in a test framework, where the target data unit is any test unit in a code repository;
[0089] Based on the call information and the grammatical structure of the target data unit, a target code snippet corresponding to the execution statement in the target data unit in the code repository is determined.
[0090] It should be noted that, taking the abnormal code generation scenario as an example, the target data fragment is the target code fragment, and the test information is the call information of the code test process. In actual implementation, the call information recorded by executing the target data unit in the test framework can be obtained, and then the target code fragment corresponding to the execution statement in the target data unit in the code repository can be determined by combining the call information and the grammatical structure of the target data unit. Among them, the call information is a test log or test tracking information, which may include various runtime data and metadata generated during the test process. For example, the call information may include the actual execution path and line number of each execution statement in the target data unit. Of course, in actual implementation, other information may also be included, such as the test case name, execution time, call order, input parameters and output results, environment information, log output, and assertion information.
[0091] In one implementation, other code execution tools can be used to obtain the call information recorded by each data unit in the execution code repository; in another implementation, each data unit in the code repository can be executed in a test framework and the corresponding call information can be recorded.
[0092] In actual implementation, after obtaining the call information, the call information and the grammatical structure of the target data unit can be parsed to determine the target code snippet corresponding to the execution statement in the target data unit in the code warehouse. Specifically, the test files in the entire code warehouse can be read and parsed, and the abstract syntax tree of each data unit in each test file can be generated. The syntax structure tree of the target data unit can be traversed to identify and extract the function definition, class definition, global variable declaration, and import statement, etc., locate the execution statement, and associate these execution statements with the corresponding code snippets in the code warehouse (such as the line code that sets the value before and after the execution statement), and locate the specific code snippet to which each execution statement belongs in the code warehouse (for example, the statement in a function or class method) as the target code snippet.
[0093] In the embodiments of the present specification, by obtaining the call information recorded when executing the target data unit in the test framework, the code information actually passed through when the target data unit is executed can be obtained. By utilizing the grammatical structure of the target data unit, the code in the target data unit can be parsed more accurately in a fine-grained dimension, and the appropriate target code fragment can be accurately located according to the actual execution path and line number for subsequent abnormal content implantation. The execution statement can be effectively located in a larger-scale and more complex code repository, and the target code fragment into which the abnormal content can be implanted can be accurately obtained, avoiding the validity problem caused by the random position of the exception, solving the problem of implanting abnormal content in the warehouse-level code, and filling the gap in the synthesis of warehouse-level abnormal code.
[0094] In an optional implementation of this embodiment, obtaining the call information recorded by executing the target data unit in the test framework includes:
[0095] Registering a tracing plug-in in the test framework, wherein the tracing plug-in is used to record call information during the code execution process;
[0096] Each test unit of the code repository is executed in the test framework, and the call information of the target data unit that passes the test is recorded through the tracking plug-in.
[0097] Specifically, a tracing plug-in can be registered in the test framework, and the tracing plug-in can capture the triggering event, line number, variable value and other information of the execution code, so as to record the actual execution path and line number and other call information of each execution statement, and execute each test file in the code repository in parallel in the test framework, and each test file includes multiple data units. The data unit that fails to execute indicates that there is an exception in its code itself, and the correct code snippet cannot be obtained. There is no need to record the call information of the data unit that fails to execute. Only the call information of the target data unit that passes the execution can be recorded. The call information can indicate the actual execution path and line number of each execution statement in the target data unit, so as to determine the target code snippet corresponding to the execution statement in the target data unit in the code repository based on the call information and the grammatical structure of the target data unit.
[0098] As an example, a tracing plug-in can be registered in the Pytest test framework. The tracing plug-in can capture information such as trigger events, line numbers, and variable values, thereby obtaining the actual execution path and line number of each execution statement. In the Pytest test framework, each data unit (single test suite) of each code file is run in parallel, and the call information of the target data unit that has passed the execution is recorded. Then, the syntax structure tree (AST) of the target data unit is used to parse the code files in the code repository at the granularity of function definition, class definition, global variable declaration, and import statement, and the code block where the execution statement is located is matched as the target code snippet.
[0099] In the embodiments of the present specification, by registering a tracing plug-in in the test framework, the call information of the code repository execution can be dynamically captured to obtain the code information actually passed when the target data unit is executed, thereby utilizing the grammatical structure of the target data unit to more accurately parse the code in the target data unit in a fine-grained dimension and accurately locate the appropriate target code fragment for subsequent abnormal content implantation.
[0100] In another implementation method, taking the abnormal text generation scenario as an example, the target data segment can be a target text segment, and the test information is the specified content of the target text segment. Specifically, each data unit of each test file in the data warehouse can be analyzed in parallel, and the specified content in the data unit can be identified (the specified content can be pre-configured, such as nouns, verbs, place names, names of people, etc.), and then the text segment where the specified content is located in the data unit is located as the target text segment, and the target text segment where the abnormal content can be implanted is accurately located.
[0101] Step 204: Based on the target data segment and the exception description information, at least one exception data segment is generated through the target generation model corresponding to the exception data generation task, wherein the exception data segment is a data segment obtained by inserting the exception content corresponding to the exception description information into the target data segment.
[0102] Specifically, the target generation model corresponding to the abnormal data generation task is a trained large model, which is used to perform the abnormal data generation task. It can insert abnormal content corresponding to the abnormal description information at any position of the input target data segment, and output the abnormal data segment corresponding to the target data segment. In the abnormal code generation scenario, the abnormal content can be a bug inserted into the code, causing the execution result of the code to report an error; in the abnormal text generation scenario, the abnormal content can be an error content (such as a typo, an incorrect character, an incorrect punctuation mark), etc., inserted into the text, causing the semantic analysis of the text or other natural language processing results to report an error.
[0103] It should be noted that the target data segment is the correct data segment stored in the data warehouse, and the abnormal data segment generated for the target data segment is the erroneous data segment with abnormal content inserted. In one implementation, the target data segment and the abnormal description information are input into the target generation model, and the target generation model can be used to insert the abnormal content into the target data segment to obtain the corresponding abnormal data segment; in another implementation, the target data segment, the data description of the target data segment and the abnormal description information can be input into the target generation model, and the target generation model can be used to analyze the target data segment based on the data description, so as to insert the abnormal content corresponding to the abnormal description information at any position of the target data segment to obtain the corresponding abnormal data segment.
[0104] In the embodiments of the present specification, the target generation model is used to insert the abnormal content corresponding to the abnormal description information into the target data segment to obtain the corresponding abnormal data segment. The abnormal parameters of the abnormal data segment to be generated can be specified, and the trained target generation model can be used to generate abnormal data segments with different abnormal parameters, thereby improving the diversity of the abnormal data segments.
[0105] In an optional implementation of this embodiment, based on the target data segment and the exception description information, at least one exception data segment corresponding to the target data segment is generated by using a target generation model corresponding to the exception data generation task, including:
[0106] Generate first input data based on the target data segment and the abnormal description information, input the first input data into the target generation model at least once, and obtain at least one abnormal data segment corresponding to the target data segment; and / or,
[0107] Acquire at least one type of abnormal description information, generate at least one second input data based on each type of abnormal description information and the target data segment, input the at least one second input data into the target generation model, and obtain at least one corresponding abnormal data segment;
[0108] The abnormal insertion position of at least one abnormal data segment is the same or different, and the abnormal content inserted into at least one abnormal data segment is the same or different.
[0109] In an optional implementation, the target data segment and the exception description information can be used as the first input data, and the first input data can be input into the target generation model at least once, respectively, to obtain at least one exception data segment corresponding to the output target data segment, and the exception insertion position of at least one exception data segment can be the same or different. The number of times the first input data is input into the target generation model is the number of exception data segments obtained, which can be configured based on the number of exceptions that need to be generated for the target data unit in the exception data generation task.
[0110] For example, taking the exception description information as exception type A as an example, the target data segment and exception type A are input into the target generation model to obtain the corresponding exception data segment 1; then, the target data segment and exception type A are input into the target generation model again to obtain the corresponding exception data segment 2; thereafter, the target data segment and exception type A are input into the target generation model again to obtain the corresponding exception data segment 3. Three exception data segments corresponding to the target data segment can be obtained, and the exception insertion positions of the three exception data segments are the same or different, but the exception contents inserted into the three exception data segments are all exception type A, but the specific inserted exception contents can be the same or different, so that multiple exception data segments of the same exception type of the target data segment can be obtained.
[0111] In another optional implementation, at least one type of exception description information can be obtained, and at least one second input data is generated based on each exception description information and the target data segment, and the at least one second input data is input into the target generation model to obtain at least one corresponding exception data segment, and the exception insertion position of at least one exception data segment can be the same or different. The type of exception description information is the number of exception data segments to be generated, and the number can be configured based on the number of exceptions that need to be generated for the target data unit according to the exception data generation task.
[0112] For example, taking the exception description information as the exception type, assume that three exception types are configured. Input the target data segment and exception type A into the target generation model to obtain the corresponding exception data segment 1; then, input the target data segment and exception type B into the target generation model to obtain the corresponding exception data segment 2; after that, input the target data segment and exception type C into the target generation model again to obtain the corresponding exception data segment 3. Three exception data segments corresponding to the target data segment can be obtained, and the exception insertion positions of the three exception data segments are the same or different, and the exception types of the exception contents inserted into the three exception data segments are different, that is, the inserted exception contents are different, so that multiple exception data segments of the target data segment under different exception types can be obtained.
[0113] As an example, Figure 4 is a schematic diagram of a target data segment and a corresponding abnormal data segment provided by an embodiment of this specification, such as Figure 4 As shown in the figure, the target data segment is the target code segment (i.e., the correct code segment), which is used to calculate the squares of all numbers in a list, and the abnormal data segment is the abnormal code segment (i.e., the error code). Assuming that the target code segment and the abnormal description information (logical error) are input into the target generation model, the target generation model can output the following: Figure 4 The abnormal code snippet with logical error shown in the figure does not square number when accumulating, but directly adds number to square_sum. In addition, assuming that the target code snippet and the abnormal description information (operator error) are input into the target generation model, the target generation model can output the following: Figure 4 The operator error exception code snippet shown in the figure multiplies square_sum by number instead of adding the square of number in the loop body. Figure 4 As shown, two types of abnormal code snippets corresponding to the target code snippet can be obtained.
[0114] In the embodiments of the present specification, the abnormal parameters of the abnormal data segment to be generated can be flexibly configured to generate multiple abnormal data segments of the target data segment under the same abnormal parameters or different abnormal parameters, which is convenient for the subsequent synthesis of any number of abnormal contents under the same or different abnormal parameters and has good scalability.
[0115] In an optional implementation of this embodiment, after generating at least one abnormal data segment corresponding to the target data segment by using a target generation model corresponding to the abnormal data generation task based on the target data segment and the abnormal description information, the method further includes:
[0116] identifying abnormal content in at least one abnormal data segment;
[0117] Filter out the abnormal invalid data segments in at least one abnormal data segment based on the set filtering rules and abnormal content.
[0118] Specifically, the abnormal content refers to the error content inserted in the abnormal data segment, and the abnormal invalid data segment refers to the abnormal data segment in which the inserted abnormal content is invalid. That is, although abnormal content is inserted in a certain abnormal data segment, the execution result of this abnormal data segment does not report an error, and this abnormal data segment is an invalid abnormal data segment and needs to be filtered out.
[0119] It should be noted that after generating at least one abnormal data segment corresponding to the target data segment through the target generation model, the abnormal content in at least one abnormal data segment can also be identified. Based on the set filtering rules, this abnormal content is analyzed, and the abnormal invalid data segments in at least one abnormal data segment are filtered out. Among them, the set filtering rules can be pre-configured based on the scenario of the target data segment. For example, in the abnormal code generation scenario, the set filtering rules can be configured as abnormal code segments with obvious abnormal prompts inserted; in the abnormal text generation scenario, the set filtering rules can be configured as abnormal text segments where the natural semantics before and after the insertion of abnormal content do not change.
[0120] Exemplarily, taking the abnormal code generation scenario as an example, assume that after inserting the position of the abnormal content (bug) in an obtained abnormal code segment, a comment "#Here is bug data, please skip" is also inserted accordingly. That is, although abnormal content is inserted in this abnormal code segment, a corresponding prompt comment is also inserted, so that the execution result of this abnormal code segment does not report an error, and this abnormal code segment is an abnormal invalid data segment. Taking the abnormal text generation scenario as an example, assume that in an obtained abnormal text segment, abnormal content (assume inserting "le" at the end of "I'm going to have dinner" to get "I'm going to have dinner le") is inserted, and the natural semantics before and after the insertion of abnormal content do not change, so that the natural semantic analysis result of this abnormal text segment does not go wrong, and this abnormal text segment is an abnormal invalid data segment.
[0121] In the embodiments of this specification, after obtaining at least one abnormal data segment corresponding to the target data segment by using the target generation model, the obvious abnormal invalid data segments in each abnormal data segment can also be filtered out based on the configured set filtering rules, ensuring that the obtained abnormal data segments all insert real and effective abnormal content, which can be used as the error data corresponding to the target data segment, and improving the accuracy and reliability of the abnormal data segment.
[0122] In an optional implementation of this embodiment, the abnormal data segment is an abnormal code segment; based on the target data segment and the abnormal description information, after generating at least one abnormal data segment corresponding to the target data segment by using the target generation model corresponding to the abnormal data generation task, the method further includes:
[0123] Execute each abnormal code fragment corresponding to the target data unit to obtain the execution result of each abnormal code fragment in the target data unit;
[0124] According to the execution results of each abnormal code fragment, the abnormal invalid code fragment in the target data unit is filtered out.
[0125] It should be noted that if it is an abnormal code generation scenario, the obtained abnormal data fragment is an abnormal code fragment. At this time, the abnormal code fragments corresponding to the target data unit can also be executed to obtain the execution results of each abnormal code fragment in the target data unit. According to the execution results of each abnormal code fragment, the abnormal invalid code fragments in the target data unit are filtered out.
[0126] In actual implementation, the target data unit may correspond to at least one target code fragment, and each target code fragment corresponds to at least one abnormal code fragment. Specifically, for each target code fragment included in the target data unit, any abnormal code fragment corresponding to the target code fragment can be randomly selected as the abnormal code unit corresponding to the target data unit. The abnormal code unit is executed, and the execution results of each abnormal code fragment can be obtained. If the execution result of the abnormal code fragment reports an error, it means that the corresponding abnormal code fragment is an abnormal valid code fragment. If the execution result of the abnormal code fragment does not report an error, it means that the corresponding abnormal code fragment is an abnormal invalid code fragment. Afterwards, the other abnormal code fragments corresponding to the target code fragment are executed and tested to determine whether the inserted abnormal content is valid.
[0127] In the embodiments of the present specification, the abnormal code snippets corresponding to each test unit of each test file in the code repository can be executed, and the abnormal code snippets that generate errors are retained, ensuring that the obtained abnormal code snippets can generate errors during the execution process, further improving the accuracy and reliability of the abnormal data snippets.
[0128] In addition, by executing each abnormal code snippet corresponding to the target data unit in the code repository, the execution result of the target data unit can indicate the execution status of the corresponding abnormal code snippet, and obtain the corresponding error information, which greatly reduces the process of manual intervention and manual screening, and avoids the need to run each data unit of each test file in the code repository to detect an abnormal content once, shortens the generation time, and improves the efficiency of synthesizing abnormal codes in complex repositories.
[0129] It should be noted that the actual error parameters, that is, the actual exception parameters, can be recorded according to the execution results of each abnormal code snippet, which can be used to construct a subsequent debugging data set and provide richer data support for subsequent model training and debugging.
[0130] Step 206: Obtain an abnormality generation result of the target data unit based on at least one abnormal data segment corresponding to each target data segment of the target data unit.
[0131] Specifically, the target data unit is the original correct data unit, and the abnormal generation result of the target data unit refers to the corresponding erroneous data unit.
[0132] It should be noted that any target data segment of the target data unit can generate at least one corresponding abnormal data segment, so for any target data segment, the corresponding at least one abnormal data segment can be merged to obtain the abnormal generation result corresponding to the target data segment; then, the abnormal generation results of each target data segment of the target data unit are merged to obtain the abnormal generation result of the target data unit, and then, the abnormal generation results of each data unit in each test file in the data warehouse are merged to obtain the complete abnormal generation result of the data warehouse. A method for generating warehouse-level abnormal data is provided to make up for the problem of missing warehouse-level data in debugging tasks.
[0133] In an optional implementation of this embodiment, obtaining an abnormality generation result of the target data unit based on at least one abnormal data segment corresponding to each target data segment of the target data unit includes:
[0134] For any target data segment of the target data unit, when there are at least two corresponding abnormal data segments, based on the data difference between the target data segment and the at least two abnormal data segments, determine the abnormal insertion positions of the at least two abnormal data segments; based on the abnormal insertion positions of the at least two abnormal data segments, merge the at least two abnormal data segments to obtain the abnormal merged segment corresponding to the target data segment;
[0135] The abnormal fusion segment corresponding to each target data segment of the target data unit is used as the abnormal generation result of the target data unit.
[0136] It should be noted that if at least two abnormal data fragments are generated corresponding to any target data fragment in the target data unit, the data difference between the target data fragment and each abnormal data fragment can be compared respectively to determine the abnormal insertion position of each abnormal data fragment, and the abnormal data fragments with different abnormal insertion positions can be fused to obtain the abnormal fused fragment corresponding to the target data fragment. For the abnormal data fragments with the same abnormal insertion position, it means that the abnormal content is repeated and can be discarded.
[0137] In one implementation, the target data segment and the corresponding abnormal data segment can be compared to determine the data row where the abnormal content is inserted. For the abnormal data segment whose abnormal insertion position is the same data row, further element-by-element detection is performed to determine the abnormal element position, thereby realizing the fusion of different abnormal data segments. In another implementation, the target data segment and the abnormal data segment can be directly compared character by character or element by element to determine the specific element position where the abnormal content is inserted, thereby realizing the fusion of different abnormal data segments.
[0138] In an embodiment of the present specification, at least two abnormal data fragments can be fused based on the abnormal insertion positions of each abnormal data fragment to obtain an abnormal fused fragment corresponding to the target data fragment, and then the abnormal fused fragment corresponding to each target data fragment of the target data unit is used as the abnormal generation result of the target data unit. In this way, different abnormal data fragments are fused based on the abnormal insertion position, which improves the diversity and flexibility of the abnormal generation result of the target data unit, and the precise abnormal insertion and synthesis strategy reduces invalid data caused by reasons such as overlapping abnormal positions, thereby improving data quality.
[0139] In an optional implementation of this embodiment, based on the abnormal insertion positions of the at least two abnormal data segments, the at least two abnormal data segments are merged to obtain the abnormal merged segment corresponding to the target data segment, including:
[0140] Determine a first abnormal data segment whose abnormal insertion position is a different data row, and a second abnormal data segment whose abnormal insertion position is a same data row, among at least two abnormal data segments;
[0141] Abnormal contents located in different data rows in the first abnormal data segment are fused, and the second abnormal data segment is fused based on the abnormal element position in the second abnormal data segment to obtain an abnormal fused segment corresponding to the target data segment.
[0142] In actual implementation, the target data segment and any abnormal data segment are compared line by line to identify the changed data row, that is, the data row where the abnormal content is inserted, that is, the identified abnormal insertion position. Specifically, the difference of each row can be calculated (for example, using string comparison or hash value comparison) to determine which rows are modified, added or deleted; then, the row number of the changed row in the original data segment is recorded. Specifically, for each difference, the starting and ending row numbers in the original data segment can be recorded. If the data segment is structured, the column changes can be further analyzed, and the identified row number is output as the specific data row where the abnormal content is inserted.
[0143] As an example, taking the abnormal code generation scenario as an example, the target data segment is the target code segment, and the abnormal data segment is the abnormal code segment. The version control system (such as Git) can automatically complete the code difference between the target code segment and the abnormal code segment before and after the insertion of the abnormal content, analyze the output code difference, identify the newly added, deleted or modified code lines, and for each difference record, extract its corresponding line number information. Specifically, the difference record usually contains the file name, line number range and specific change content. By parsing this information, it can be determined which data rows have changed due to the insertion of the abnormal content, and record the final data row as the insertion position of the abnormal content.
[0144] It should be noted that it can be determined that the first abnormal data segment with abnormal insertion position is a different data row, that is, the first abnormal data segment is an abnormal data segment with abnormal content inserted in different data rows, and the abnormal content located in different data rows in the first abnormal data segment can be directly merged.
[0145] In addition, it is also possible to determine that the second abnormal data segment whose abnormal insertion position is the same data row is a second abnormal data segment. In one implementation method, the second abnormal data segment whose abnormal insertion position is the same data row can be directly discarded; in another implementation method, the second abnormal data segment whose abnormal insertion position is the same data row can be further detected for the second abnormal element position in the second abnormal data segment, that is, the element position where the abnormal content is located is determined through token-level detection to achieve fusion.
[0146] In the embodiments of the present specification, according to the different data segments before and after the insertion of the abnormal content, the data row where the abnormal insertion position is located is obtained, and the abnormal data segments located in different data rows can be directly merged to obtain abnormal data segments with multiple abnormal contents inserted. By automatically capturing test information in a large-scale data warehouse, accurately locating the target data segment, automatically inserting the abnormal content and synthesizing the abnormal data segments with multiple abnormal contents, accurate and reliable abnormal data are provided, and the precise abnormal insertion and synthesis strategy reduces invalid data caused by reasons such as overlapping abnormal positions, thereby improving data quality.
[0147] In an optional implementation of this embodiment, fusing the second abnormal data fragment based on the position of the abnormal element in the second abnormal data fragment includes:
[0148] Performing element-by-element detection on each second abnormal data segment to determine the position of the abnormal element in each second abnormal data segment;
[0149] The second abnormal data segments with the same abnormal element positions are discarded, and the abnormal contents in the second abnormal data segments with different abnormal element positions are merged.
[0150] In actual implementation, for the second abnormal data segment whose abnormal insertion position is the same data row, element-by-element detection can be performed to determine the abnormal element position in each second abnormal data segment, and the abnormal content of different abnormal element positions can be merged. Specifically, the data row where the abnormal content is inserted in the abnormal data segment and the normal data segment is identified, and the data row contains the inserted abnormal content. The elements in the data row in the target data segment and the abnormal data segment are compared element by element to detect whether each element is consistent. The consistency of the elements can be identified by directly comparing the element values at the corresponding positions, or calculating the differences between the elements (such as numerical differences or string distances), and the inconsistent elements and their specific element positions in the row (such as column indexes) are recorded. If the data is structured (such as tabular data), the element position of the abnormal content can be further confirmed in combination with contextual information (such as the relationship between adjacent elements).
[0151] For example, assuming that the target data segment corresponds to 5 abnormal data segments, abnormal data segment 1 inserts abnormal content in the 4th row, abnormal data segment 2 inserts abnormal content in the 7th row, abnormal data segment 3 inserts abnormal content in the 9th row, and abnormal data segment 4 and abnormal data segment 5 insert abnormal content in the 11th row. At this time, the abnormal content inserted in abnormal data segment 1, abnormal data segment 2 and abnormal data segment 3 can be directly fused, and the abnormal element positions in abnormal data segment 4 and abnormal data segment 5 can be further identified. Assuming that the abnormal element position in abnormal data segment 4 is element 10-element 15, and the abnormal element position in abnormal data segment 5 is element 21-element 28, the abnormal content inserted at different element positions in abnormal data segment 4 and abnormal data segment 5 can be fused, and the obtained abnormal fusion segment is the 4th row, 7th row, and 9th row of the target data segment respectively inserted with corresponding abnormal content. And the elements 10-element 15 and elements 21-element 28 of the 11th row are respectively inserted with corresponding abnormal content.
[0152] In the embodiments of the present specification, for abnormal contents located in the same data row, in-depth detection can be performed at the element level, thereby ensuring the accuracy when synthesizing multiple abnormal contents. Abnormal data segments with any number of abnormal contents can be fused, avoiding overlapping problems, and ensuring the accuracy of abnormal generation results of multiple abnormal contents.
[0153] In an optional implementation of this embodiment, after obtaining the abnormality generation result of the target data unit based on at least one abnormal data segment corresponding to each target data segment of the target data unit, the method further includes:
[0154] A debugging data set is constructed based on the target data segments and abnormal data segments corresponding to each target data unit, the abnormal type of the abnormal data segment, and the test information recorded during the test process, wherein the abnormal type is determined based on the execution result of the abnormal data segment, or based on the abnormal description information generated by the abnormal data segment; the debugging data set is used to train the model to be trained in the task platform.
[0155] It should be noted that after executing the abnormal data segment, the corresponding execution result can be recorded, and the actual exception type can be recorded based on the execution result to construct a debugging data set; or, the exception type of the abnormal data segment can be determined based on the exception description information of the abnormal data segment to construct a debugging data set.
[0156] In actual implementation, based on the target data segments and abnormal data segments corresponding to each target data unit in the data warehouse, the abnormal type of the abnormal data segment, and the test information recorded during the test process, a debugging data set can be constructed to train the model to be trained in the task platform. The model to be trained can be a model to perform the corresponding generation task in the task platform, or it can be a target generation model that is currently performing the abnormal data generation task. Based on the debugging data set, the target generation model is further debugged or optimized for training.
[0157] It should be noted that more diverse and accurate abnormal data can be automatically generated based on the data warehouse, and warehouse-level debugging data set construction can be realized, avoiding the problems of frequent manual intervention and low efficiency, and providing strong data support for the debugging capabilities of the models to be trained in any task platform.
[0158] An embodiment of the present specification provides a method for generating abnormal data. Based on the abnormal description information, a target generation model is used to generate more diverse abnormal data fragments, thereby improving data quality. The method can also dynamically capture the test information recorded by the target data unit during the test process, effectively and quickly locate the target data fragment in the target data unit where the abnormal content needs to be inserted, thereby achieving more accurate and automated insertion of abnormal content. A full-process automation solution from testing to insertion of abnormal content is provided, which greatly improves the scale and efficiency of abnormal generation, can adapt to data warehouses of different sizes and complexities, and has good scalability.
[0159] The following combination Figure 5 , taking the application of the abnormal data generation method provided in this specification in the code generation scenario as an example, the abnormal data generation method is further explained. Figure 5 A flowchart of an exception code generation method provided by an embodiment of the present specification is shown, which specifically includes the following steps.
[0160] Step 502: Obtain task information of the exception code generation task, wherein the task information includes a target code snippet and exception description information. The target code snippet is an original code snippet in the target test unit where the exception content is to be inserted. The target code snippet is determined based on the call information recorded during the code test of the target test unit.
[0161] Step 504: Based on the target code snippet and the exception description information, at least one exception code snippet is generated through the target generation model corresponding to the exception code generation task, wherein the exception code snippet is a code snippet obtained by inserting the exception content corresponding to the exception description information into the target code snippet.
[0162] Step 506: Obtain an exception generation result of the target test unit based on at least one exception code segment corresponding to each target code segment of the target test unit.
[0163] It should be noted that the implementation of step 502 to step 506 is the same as the implementation of step 202 to step 206 described above, and the embodiments of this specification do not impose any limitation on this.
[0164] By applying the scheme of the embodiments of this specification, more diverse exception code snippets can be generated based on the exception description information using the target generation model, thereby improving the quality of the exception code. In addition, the target code snippet in the target test unit where the exception content needs to be inserted can be effectively and quickly located based on the call information recorded by the test of the target test unit, thereby achieving more accurate and automated insertion of exception content. A full-process automation solution from testing to insertion of exception content is provided, which greatly improves the scale and efficiency of exception generation, can adapt to code repositories of different sizes and complexities, and has good scalability.
[0165] See also Figure 6 , Figure 6A schematic diagram of a processing process of an abnormal code generation method provided by an embodiment of the present specification is shown. Taking the abnormal code generation scenario as an example, the code repository includes source code files (source code), and the source code files may include configuration files (setup files) and test files (test files). An abnormal code generation task (issue, indicating a corresponding problem or task) can be created based on the code repository. In response to the abnormal code generation task, each test file (test file) can be run through multiple processes in a test framework (Pytest framework). A tracing plug-in is registered in the test framework. Call information of the code execution process can be recorded in the test framework. Specifically, test file 1 (test file 1)-test file k (test file k) can be executed. Each test file includes multiple single test suites (that is, data units), such as single-sided suite 1 (case 1)-single test suite n (-case n). "×" indicates that the corresponding single-sided suite has not been executed, and "√" indicates that the corresponding single-sided suite has been executed. The call information of the single-sided suites that have been executed is recorded, such as single test suite n in test file 1, single test suite 1 and single test suite n in test file 2. The call information may include events (Event), function names or function names (function name), line numbers (line number), file names (file name), variables or values (local value) defined and used by functions, methods, code blocks, etc.
[0166] Then, the corresponding exception code snippet is generated using the big model. Based on the call information, the target code snippet corresponding to the execution statement of any single test suite in the code repository can be determined. The input prompt of the big model (target generation model) is constructed based on the target code snippet, the corresponding code description and the exception type. For example, the type of code description can include a function description (generated by the big model); a description of changes between two commit instructions (commit) (which can be generated by the big model), requiring the exception code to cover the changed code; the repository submission description (also known as the repository pr (Pull Request) description, a collaborative mechanism for submitting code changes and requesting that these changes be merged into the main branch or other target branches of the project. When creating a Pull Request, a text description needs to be provided to describe in detail the content, purpose and related background information of this submission), requiring the exception code snippet to cover the changed code.
[0167] The input prompt is input into the corresponding big model, and the big model can output the abnormal code snippet corresponding to the abnormal type. The exception type can be a function-level compilation syntax error, a function-level runtime error, a warehouse-level runtime error, a function-level logic error, a warehouse-level logic error, a mixed error, etc. Execute the abnormal code snippet corresponding to each single-side suite to obtain the corresponding execution result, and filter out the abnormal invalid code snippet based on the execution result to obtain the valid abnormal code snippet. Based on the target code snippet and the abnormal code snippet corresponding to each single test suite, the execution result of the abnormal code snippet (that is, the error type) and the call information recorded during the test process, a debugging data set is constructed, such as Figure 6 As shown, one data in the debugging data set includes <{correct code description, unit suite test, error code, call information, execution result}, correct code>.
[0168] By applying the solution of the embodiments of this specification, a large model can be used to generate exception code snippets with diverse exception types, thereby improving the quality of exception code, and being able to dynamically capture the call information recorded by the single-sided suite during the code testing process, effectively and quickly locate the target code snippet in the single test suite where the exception content needs to be inserted, thereby achieving more accurate and automated exception content insertion, and providing a full-process automation solution from testing to exception content insertion, which greatly improves the scale and efficiency of exception generation, can adapt to code repositories of different sizes and complexities, and has good scalability.
[0169] See also Figure 7 , Figure 7 A flowchart of an information processing method based on a target generation model provided according to an embodiment of the present specification is shown, which is applied to a task platform and specifically includes the following steps 702-704.
[0170] Step 702: Receive a model request sent by a terminal device, wherein the model request includes a scene identifier of a target scene, scene input data of the target scene, and at least one of a model specification parameter.
[0171] Step 704: Based on the model request, determine a corresponding target generation model from at least one generation model, wherein the target generation model is used to execute the generation process of the abnormal data fragment in the above-mentioned abnormal data generation method.
[0172] It should be noted that based on the model request, the corresponding target generation model is determined from at least one generation model. One optional method is: based on the model request, the corresponding target generation model is searched from at least one generation model included in the model library; another optional method is: based on the model request, the target generation model is trained; and another optional method is: based on the model request, the target generation model is constructed, which is not limited here.
[0173] For example, based on the scene identification of the target scene, you can first search for at least one pre-trained generation model from the model library, then filter out a generation model of corresponding size from at least one generation model based on the model specification parameters, and then train the generation model of corresponding size based on the scene input data of the target scene to obtain a target generation model suitable for user needs.
[0174] In an optional implementation of this embodiment, the model request includes a scene identifier of the target scene; based on the model request, determining a corresponding target generation model from at least one generation model includes:
[0175] Based on the scene identifier of the target scene, a target generation model adapted to the target scene is searched from a model library, wherein the model library stores at least one generation model adapted to different generation task scenes.
[0176] It should be noted that the model library is a database for storing and managing various pre-trained deep learning models. Multiple generation models adapted to different abnormal data generation scenarios cover different application scenarios and needs. The model library allows users to select appropriate models according to their needs, or directly use the model to process abnormal data generation tasks through API calls.
[0177] The multiple generation models adapted to different abnormal data generation scenarios are multiple models specially designed for different abnormal data generation scenarios stored in the model library, and each model is optimized for a specific application environment. For example, based on the scenario identifier "abnormal code generation" of the target scenario, the target generation model adapted to the abnormal code generation scenario can be searched from the model library.
[0178] In the embodiments of this specification, based on the scenario requirements, the target generation model suitable for the scenario is accurately found through the scenario identification, so that the exception generation result is more accurate and fits the scenario, thereby improving the user experience and the processing quality of the exception data generation task.
[0179] As an example, the task platform can provide target generation models for a variety of scenarios. For example, in an exception code generation scenario, it can provide a corresponding target generation model based on a model request sent by a smart terminal to implement corresponding exception code generation task processing.
[0180] In an optional implementation of this embodiment, the model request includes scene input data of the target scene; based on the model request, determining a corresponding target generation model from at least one generation model includes:
[0181] Determining an initial generative model adapted to the target scenario from at least one generative model;
[0182] Based on the scene input data of the target scene, the initial generation model is trained to obtain the target generation model.
[0183] In actual implementation, the model request may include scene input data of the target scene, and the target generation model is a generation model suitable for the target scene.
[0184] Exemplarily, the general generation model is a basic generation model that is trained to be adaptable to different abnormal data generation scenarios, but is not optimized for any specific scenario. For example, based on the scenario input data of the abnormal code generation scenario, the general generation model is trained to obtain a target generation model that is adapted to the abnormal code generation scenario.
[0185] In the embodiments of this specification, based on scenario requirements, the general generation model is further trained through scenario input data to obtain a target generation model adapted to the scenario, so that the exception generation results are more accurate and fit the scenario, thereby improving the user experience and the processing quality of the exception data generation task.
[0186] In an optional implementation of this embodiment, the model request includes model specification parameters; based on the model request, determining a corresponding target generation model from at least one generation model includes:
[0187] Based on the model specification parameters, a corresponding target generation model is searched from a model library, wherein the model library stores a plurality of generation models with different model specification parameters.
[0188] The model specification parameter may be a model size, such as searching for a target generation model of a corresponding size from a model library based on the model size: 32 GB.
[0189] In the embodiments of this specification, based on the model specification requirements, the corresponding target generation model is accurately found through the model specification parameters, which ensures the efficient and stable operation of the target generation model and improves the user experience.
[0190] In an optional implementation of this embodiment, after determining a corresponding target generation model from at least one generation model based on the model request, the method further includes:
[0191] Deploy the target generation model, and build a task processing interface based on the target generation model so that the terminal device schedules the target generation model to perform the corresponding abnormal data generation task.
[0192] It should be noted that the task processing interface is an interactive programming interface for the intelligent terminal scheduling target generation model, which is usually provided in the form of an API. Through the task processing interface, users can initiate abnormal data generation tasks and call the corresponding target generation model.
[0193] In actual implementation, an optional way to deploy the target generation model is to deploy the target generation model on the distributed system of the task platform. For example, the target generation model is deployed on the distributed system of the task platform, and based on the target generation model, a task processing interface is constructed and provided to the intelligent terminal, so that the intelligent terminal schedules the target generation model to execute the corresponding abnormal data generation task.
[0194] In the embodiments of this specification, efficient terminal calling is achieved, the processing of abnormal data generation tasks is optimized, and the processing quality and response speed of abnormal data generation tasks are improved.
[0195] The information processing method based on the target generation model provided in the embodiment of this specification is applied to obtain the target generation model according to user needs, realize personalized model service, provide users with an efficient, flexible and easy-to-use model service method, and improve user experience.
[0196] Corresponding to the above method embodiment, this specification also provides an abnormal data generating device embodiment, Figure 8 FIG. 1 is a schematic diagram showing the structure of an abnormal data generating device provided by an embodiment of the present specification. Figure 8 As shown, the device comprises:
[0197] A first acquisition module 802 is configured to acquire task information of an abnormal data generation task, wherein the task information includes a target data segment and abnormal description information, wherein the target data segment is an original data segment in a target data unit into which abnormal content is to be inserted, and the target data segment is determined based on test information recorded by testing the target data unit;
[0198] The first generating module 804 is configured to generate at least one abnormal data segment based on the target data segment and the abnormal description information by using the target generation model corresponding to the abnormal data generation task, wherein the abnormal data segment is a data segment obtained by inserting the abnormal content corresponding to the abnormal description information into the target data segment;
[0199] The first obtaining module 806 is configured to obtain an abnormality generation result of the target data unit based on at least one abnormal data segment corresponding to each target data segment of the target data unit.
[0200] Optionally, the task information also includes a data description; the first acquisition module 802 is further configured to:
[0201] Acquire test information recorded by testing the target data unit, and locate the corresponding target data segment in the data warehouse based on the test information;
[0202] Analyze the target data segment to determine the data description of the target data segment;
[0203] Get the set exception description information;
[0204] Accordingly, the first generating module 804 is further configured to:
[0205] Based on the target data segment, the data description and the exception description information, at least one exception data segment is generated through a target generation model corresponding to the exception data generation task.
[0206] Optionally, the target data segment is a target code segment, and the test information is call information of a code test process; the first acquisition module 802 is further configured to:
[0207] Obtain call information recorded by executing a target data unit in a test framework, where the target data unit is any test unit in a code repository;
[0208] Based on the call information and the grammatical structure of the target data unit, a target code snippet corresponding to the execution statement in the target data unit in the code repository is determined.
[0209] Optionally, the first acquisition module 802 is further configured to:
[0210] Registering a tracing plug-in in the test framework, wherein the tracing plug-in is used to record call information during the code execution process;
[0211] Each test unit of the code repository is executed in the test framework, and the call information of the target data unit that passes the test is recorded through the tracking plug-in.
[0212] Optionally, the first generating module 804 is further configured to:
[0213] Generate first input data based on the target data segment and the abnormal description information, input the first input data into the target generation model at least once, and obtain at least one abnormal data segment corresponding to the target data segment; and / or,
[0214] Acquire at least one type of abnormal description information, generate at least one second input data based on each type of abnormal description information and the target data segment, input the at least one second input data into the target generation model, and obtain at least one corresponding abnormal data segment;
[0215] The abnormal insertion position of at least one abnormal data segment is the same or different, and the abnormal content inserted into the at least one abnormal data segment is the same or different.
[0216] Optionally, the first obtaining module 806 is further configured to:
[0217] For any target data segment of the target data unit, when there are at least two corresponding abnormal data segments, based on the data difference between the target data segment and the at least two abnormal data segments, determine the abnormal insertion positions of the at least two abnormal data segments; based on the abnormal insertion positions of the at least two abnormal data segments, merge the at least two abnormal data segments to obtain the abnormal merged segment corresponding to the target data segment;
[0218] The abnormal fusion segment corresponding to each target data segment of the target data unit is used as the abnormal generation result of the target data unit.
[0219] Optionally, the first obtaining module 806 is further configured to:
[0220] Determine a first abnormal data segment whose abnormal insertion position is a different data row, and a second abnormal data segment whose abnormal insertion position is a same data row, among at least two abnormal data segments;
[0221] Abnormal contents located in different data rows in the first abnormal data segment are fused, and the second abnormal data segment is fused based on the abnormal element position in the second abnormal data segment to obtain an abnormal fused segment corresponding to the target data segment.
[0222] Optionally, the first obtaining module 806 is further configured to:
[0223] Performing element-by-element detection on each second abnormal data segment to determine the position of the abnormal element in each second abnormal data segment;
[0224] The second abnormal data segments with the same abnormal element positions are discarded, and the abnormal contents in the second abnormal data segments with different abnormal element positions are merged.
[0225] Optionally, the device further includes a first filtering module configured to:
[0226] identifying abnormal content in at least one abnormal data segment;
[0227] Based on the set filtering rules and abnormal content, an abnormal invalid data segment in at least one abnormal data segment is filtered out.
[0228] Optionally, the abnormal data segment is an abnormal code segment; the device further includes a second filtering module configured to:
[0229] Execute each abnormal code fragment corresponding to the target data unit to obtain the execution result of each abnormal code fragment in the target data unit;
[0230] According to the execution results of each abnormal code fragment, the abnormal invalid code fragment in the target data unit is filtered out.
[0231] Optionally, the device further comprises a construction module configured to:
[0232] A debugging data set is constructed based on the target data segments and abnormal data segments corresponding to each target data unit, the abnormal type of the abnormal data segment, and the test information recorded during the test process, wherein the abnormal type is determined based on the execution result of the abnormal data segment, or based on the abnormal description information generated by the abnormal data segment; the debugging data set is used to train the model to be trained in the task platform.
[0233] An embodiment of the present specification provides an exception data generation device, which can generate more diverse exception data fragments based on the exception description information using a target generation model, thereby improving data quality, and can effectively and quickly locate the target data fragment in the target data unit where the exception content needs to be inserted based on the test information recorded by the test of the target data unit, thereby achieving more accurate and automated insertion of exception content, and providing a full-process automation solution from testing to insertion of exception content, which greatly improves the scale and efficiency of exception generation, can adapt to data warehouses of different sizes and complexities, and has good scalability.
[0234] The above is a schematic scheme of an abnormal data generating device of this embodiment. It should be noted that the technical scheme of the abnormal data generating device and the technical scheme of the abnormal data generating method described above belong to the same concept, and the details not described in detail in the technical scheme of the abnormal data generating device can be referred to the description of the technical scheme of the abnormal data generating method described above.
[0235] Corresponding to the above method embodiment, this specification also provides an exception code generation device embodiment, Fig. 9 FIG. 1 shows a schematic diagram of the structure of an abnormal code generating device provided by an embodiment of the present specification. Fig. 9 As shown, the device comprises:
[0236] A second acquisition module 902 is configured to acquire task information of an abnormal code generation task, wherein the task information includes a target code segment and abnormal description information, wherein the target code segment is an original code segment to be inserted into the target test unit into which the abnormal content is to be inserted, and the target code segment is determined based on call information recorded by code testing of the target test unit;
[0237] The second generation module 904 is configured to generate at least one abnormal code snippet based on the target code snippet and the abnormal description information by using the target generation model corresponding to the abnormal code generation task, wherein the abnormal code snippet is a code snippet obtained by inserting abnormal content corresponding to the abnormal description information into the target code snippet;
[0238] The second obtaining module 906 is configured to obtain the exception generation result of the target test unit based on at least one exception code snippet corresponding to each target code snippet of the target test unit.
[0239] An embodiment of the present specification provides an exception code generation device, which can generate more diverse exception code snippets based on exception description information and utilize a target generation model, thereby improving code quality. It can also effectively and quickly locate the target code snippet in the target test unit where exception content needs to be inserted based on the call information recorded by the code test of the target test unit, thereby achieving more accurate and automated exception content insertion, and providing a full-process automation solution from testing to exception content insertion, which greatly improves the scale and generation efficiency of exception codes, can adapt to code repositories of different sizes and complexities, and has good scalability.
[0240] The above is a schematic scheme of an abnormal code generating device of this embodiment. It should be noted that the technical scheme of the abnormal code generating device and the technical scheme of the abnormal data generating method described above belong to the same concept, and the details not described in detail in the technical scheme of the abnormal code generating device can be referred to the description of the technical scheme of the abnormal data generating method described above.
[0241] Fig.10 A structural block diagram of a computing device provided by an embodiment of the present specification is shown.
[0242] The computing device 1000 includes:
[0243] Memory 1010 and processor 1020;
[0244] The memory 1010 is used to store computer programs / instructions, and the processor 1020 is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor 1020, the steps of the above-mentioned abnormal data generation method or abnormal code generation method or information processing method based on the target generation model are implemented.
[0245] In one or more embodiments of the present specification, the computing device 1000 may be understood as an integrated intelligent terminal, including but not limited to a server, a desktop computer, a PC (Personal Computer), an all-in-one model machine, a mobile phone, a tablet computer or other portable intelligent terminal, etc., and the computing device may be pre-installed with the models in the above embodiments of the present application.
[0246] Specifically, the computing device 1000 can preset multiple types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, data processing, multimodal task processing, etc., so as to provide a variety of model selections. In different product forms, the computing device 1000 can support one or more model usage methods, including but not limited to model training, model calling, model fine-tuning, model deployment, model reasoning and application, etc. In some product forms, the computing device 1000 also supports model management, including but not limited to multi-type model management (supporting the management of multiple types of models such as discriminants and generative models), model version control (supporting the control of different model versions), model evaluation (based on model evaluation tools, evaluating the performance and effect of models), etc. In other product forms, the computing device 1000 can also create applications based on models, provide API (Application Programming Interface, application programming interface) calling capabilities, and can call models to the created applications through API interfaces, while providing application management tools to achieve management and monitoring of applications.
[0247] Furthermore, the computing device 1000 may also include data management (supporting the creation and management of model tuning data sets), a training center (providing rich training resources to help users learn artificial intelligence technology), and basic management and control capabilities (providing enterprise-level basic management and control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive, integrated artificial intelligence development, training, deployment and application device is provided.
[0248] Fig.11 A structural block diagram of an electronic device provided by an embodiment of the present specification is shown.
[0249] The memory 1110 and the processor 1120 are connected via a bus 1130;
[0250] The memory 1110 is used to store computer programs / instructions, and the processor 1120 is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor 1120, the steps of the above-mentioned abnormal data generation method or abnormal code generation method or information processing method based on the target generation model are implemented.
[0251] Specifically, the components of the electronic device 1100 include but are not limited to a memory 1110 and a processor 1120. The processor 1120 and the memory 1110 may be connected via a bus 1130.
[0252] The electronic device 1100 may also include an access device 1140, which enables the electronic device 1100 to communicate with a database 1150 storing data via one or more networks 1160. Examples of these networks 1160 include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1140 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).
[0253] In one embodiment of the present specification, the above components of the electronic device 1100 and Fig.11 Other components not shown in the figure may also be connected to each other, for example, via bus 1130. It should be understood that Fig.11 The electronic device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0254] The electronic device 1100 may be any type of stationary or mobile electronic device, including a mobile computer or mobile electronic device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable electronic device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary electronic device such as a desktop computer or a personal computer (PC). The electronic device 1100 may also be a mobile or stationary server.
[0255] The above is a schematic scheme of an electronic device of the present embodiment. It should be noted that the technical scheme of the electronic device and the technical scheme of the above-mentioned abnormal data generation method or abnormal code generation method or information processing method based on the target generation model belong to the same concept, and the details not described in detail in the technical scheme of the electronic device can be referred to the description of the technical scheme of the above-mentioned abnormal data generation method or abnormal code generation method or information processing method based on the target generation model.
[0256] An embodiment of the present specification also provides a computer-readable storage medium storing a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned abnormal data generation method or abnormal code generation method or information processing method based on a target generation model.
[0257] The above is a schematic scheme of a computer-readable storage medium of the present embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned abnormal data generation method or abnormal code generation method or information processing method based on the target generation model belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the above-mentioned abnormal data generation method or abnormal code generation method or information processing method based on the target generation model.
[0258] An embodiment of the present specification also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned abnormal data generation method or abnormal code generation method or information processing method based on a target generation model.
[0259] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the above-mentioned abnormal data generation method or abnormal code generation method or information processing method based on the target generation model belong to the same concept, and the details not described in detail in the technical scheme of the computer program product can be referred to the description of the technical scheme of the above-mentioned abnormal data generation method or abnormal code generation method or information processing method based on the target generation model.
[0260] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0261] Computer programs / instructions include computer program data, which may be in the form of source data, object data, executable files, or some intermediate forms. Computer-readable media may include: any entity or device capable of carrying computer program data, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROM), random access memories (RAM), electric carrier signals, telecommunication signals, and software distribution media. It should be noted that the content contained in computer-readable media may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0262] It should be noted that, for the convenience of description, the aforementioned method embodiments are all described as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and module segments involved are not necessarily required by the embodiments of this specification.
[0263] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0264] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A method for generating abnormal data, comprising: Acquire test information recorded by testing the target data unit, and locate the corresponding target data segment in the data warehouse based on the test information; Analyze the target data segment to determine the data description of the target data segment; obtain the set exception description information, and use the target data segment, the data description and the exception description information as task information of the exception data generation task, wherein the target data segment is the original data segment in the target data unit into which the exception content is to be inserted; Based on the target data segment, the data description and the exception description information, at least one exception data segment is generated by using a target generation model corresponding to the exception data generation task, wherein the exception data segment is a data segment obtained by inserting the exception content corresponding to the exception description information into the target data segment; Based on at least one abnormal data segment corresponding to each target data segment of the target data unit, an abnormality generation result of the target data unit is obtained.
2. The method according to claim 1, wherein the target data segment is a target code segment, and the test information is call information of a code test process; The obtaining of test information recorded by testing the target data unit and locating the corresponding target data segment in the data warehouse based on the test information includes: Acquire call information recorded by executing the target data unit in the test framework, wherein the target data unit is any test unit in the code repository; Based on the call information and the grammatical structure of the target data unit, a target code snippet corresponding to the execution statement in the target data unit in the code repository is determined.
3. The method according to claim 2, wherein obtaining the call information recorded by executing the target data unit in the test framework comprises: Registering a tracing plug-in in the test framework, wherein the tracing plug-in is used to record call information during the code execution process; Each test unit of the code repository is executed in the test framework, and the call information of the target data unit that passes the test is recorded through the tracking plug-in.
4. The method according to claim 1, wherein the step of generating at least one abnormal data segment corresponding to the target data segment based on the target data segment and the abnormal description information by using a target generation model corresponding to the abnormal data generation task comprises: Generate first input data based on the target data segment and the abnormal description information, input the first input data into the target generation model at least once, and obtain at least one abnormal data segment corresponding to the target data segment; and / or, Acquire at least one type of abnormal description information, generate at least one second input data based on each abnormal description information and the target data segment, input the at least one second input data into the target generation model, and obtain at least one corresponding abnormal data segment; The abnormal insertion position of the at least one abnormal data segment is the same or different, and the abnormal content inserted into the at least one abnormal data segment is the same or different.
5. The method according to any one of claims 1 to 4, wherein obtaining the abnormality generation result of the target data unit based on at least one abnormal data segment corresponding to each target data segment of the target data unit comprises: For any target data segment of the target data unit, when there are at least two corresponding abnormal data segments, determining abnormal insertion positions of the at least two abnormal data segments based on the data difference between the target data segment and the at least two abnormal data segments; based on the abnormal insertion positions of the at least two abnormal data segments, merging the at least two abnormal data segments to obtain an abnormal fused segment corresponding to the target data segment; The abnormal fusion segment corresponding to each target data segment of the target data unit is used as the abnormal generation result of the target data unit.
6. The method according to claim 5, wherein the step of fusing the at least two abnormal data segments based on the abnormal insertion positions of the at least two abnormal data segments to obtain the abnormal fused segment corresponding to the target data segment comprises: Determine a first abnormal data segment whose abnormal insertion position is a different data row, and a second abnormal data segment whose abnormal insertion position is a same data row, among the at least two abnormal data segments; Abnormal contents in different data rows in the first abnormal data segment are fused, and the second abnormal data segment is fused based on the abnormal element position in the second abnormal data segment to obtain an abnormal fused segment corresponding to the target data segment.
7. The method according to claim 6, wherein fusing the second abnormal data fragment based on the position of the abnormal element in the second abnormal data fragment comprises: Performing element-by-element detection on each second abnormal data segment to determine the position of the abnormal element in each second abnormal data segment; The second abnormal data segments with the same abnormal element positions are discarded, and the abnormal contents in the second abnormal data segments with different abnormal element positions are merged.
8. The method according to any one of claims 1 to 4, after generating at least one abnormal data segment corresponding to the target data segment based on the target data segment and the abnormal description information by using the target generation model corresponding to the abnormal data generation task, further comprising: identifying abnormal content in the at least one abnormal data segment; Based on the set filtering rules and the abnormal content, the abnormal invalid data segment in the at least one abnormal data segment is filtered out.
9. The method according to any one of claims 1 to 4, wherein the abnormal data segment is an abnormal code segment; after generating at least one abnormal data segment corresponding to the target data segment through a target generation model corresponding to the abnormal data generation task based on the target data segment and the abnormal description information, the method further comprises: Execute each abnormal code fragment corresponding to the target data unit to obtain the execution result of each abnormal code fragment in the target data unit; According to the execution results of the abnormal code fragments, the abnormal invalid code fragments in the target data unit are filtered out.
10. The method according to any one of claims 1 to 4, after obtaining the abnormality generation result of the target data unit based on at least one abnormal data segment corresponding to each target data segment of the target data unit, further comprising: A debugging data set is constructed based on the target data segments and abnormal data segments corresponding to each target data unit, the abnormal type of the abnormal data segment, and the test information recorded during the test process, wherein the abnormal type is determined based on the execution result of the abnormal data segment, or based on the abnormal description information generated by the abnormal data segment; the debugging data set is used to train the model to be trained in the task platform.
11. A method for generating an exception code, comprising: Acquire call information recorded during code testing of a target test unit, and locate a corresponding target code snippet in a code repository based on the call information; Analyzing the target code fragment to determine a code description of the target code fragment; Acquire the set exception description information, and use the target code snippet, the code description and the exception description information as task information of the exception code generation task, wherein the target code snippet is the original code snippet in the target test unit into which the exception content is to be inserted; Based on the target code snippet, all code snippets and the exception description information, at least one exception code snippet is generated through a target generation model corresponding to the exception code generation task, wherein the exception code snippet is a code snippet obtained by inserting the exception content corresponding to the exception description information into the target code snippet; Based on at least one abnormal code segment corresponding to each target code segment of the target test unit, an abnormality generation result of the target test unit is obtained.
12. An information processing method based on a target generation model, applied to a task platform, comprising: Receiving a model request sent by a terminal device, wherein the model request includes a scene identifier of a target scene, scene input data of the target scene, and at least one of a model specification parameter; Based on the model request, a corresponding target generation model is determined from at least one generation model, wherein the target generation model is used to execute a generation process of abnormal data fragments in the abnormal data generation method according to any one of claims 1 to 10.
13. The method according to claim 12, wherein the model request includes a scene identifier of a target scene; and determining a corresponding target generative model from at least one generative model based on the model request comprises: Based on the scene identifier of the target scene, searching a target generation model adapted to the target scene from a model library, wherein the model library stores at least one generation model adapted to different generation task scenes; The model request includes scene input data of a target scene; and determining a corresponding target generation model from at least one generation model based on the model request includes: Determining, from at least one generative model, an initial generative model adapted to the target scenario; Based on the scene input data of the target scene, the initial generation model is trained to obtain a target generation model; The model request includes model specification parameters; and determining a corresponding target generation model from at least one generation model based on the model request includes: Based on the model specification parameters, a corresponding target generation model is searched from a model library, wherein the model library stores a plurality of generation models with different model specification parameters.
14. The method according to claim 12, after determining a corresponding target generative model from at least one generative model based on the model request, further comprising: The target generation model is deployed, and based on the target generation model, a task processing interface is constructed so that the terminal device schedules the target generation model to execute the corresponding abnormal data generation task.
15. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 14 are implemented.
16. An electronic device, comprising: A memory and a processor, wherein the memory and the processor are connected via a bus; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 14 are implemented.
17. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 14.
18. A computer program product, comprising a computer program / instruction, which implements the steps of the method according to any one of claims 1 to 14 when executed by a processor.
Citation Information
Patent Citations
Sample processing method and device, computing equipment and computer readable storage medium
CN117911803A