Software defect positioning and distributing method and device based on large model

Through the software defect positioning and allocation method based on large models, using large models to perform defect inference and deal with human reasoning, the inefficiency and inaccuracy caused by artificial dependence in traditional methods are solved, and efficient and accurate defect positioning and allocation are achieved.

CN120029804APending Publication Date: 2025-05-23GUANGZHOU FAISCO INFORMATON TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510047079.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Traditional software defect positioning and allocation methods rely on manual experience, have low accuracy, low efficiency, and are prone to omissions or misjudgment, resulting in unfair distribution.

Method used

The software defect positioning and allocation method based on the big model is adopted, and the initial code submission records are obtained, preprocessed and vectorized storage is performed. The big model is used to inference the project where the defect is located, similar code retrieval, correlation reasoning and processing human reasoning to achieve automated defect positioning and allocation.

Benefits of technology

It improves the accuracy and efficiency of software defect positioning, realizes automated defect positioning and distribution, reduces manual intervention, and improves the fairness of distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029804A_ABST
    Figure CN120029804A_ABST
Patent Text Reader

Abstract

The invention discloses a software defect positioning and distributing method and device based on a large model. The method comprises the following steps: acquiring an initial code submission record; preprocessing to obtain a preprocessed code submission record; the preprocessing code submission record is stored in a vectorization database; generating a defect information list of the target project by utilizing a preset software interface; selecting a piece of target defect information; performing reasoning processing on the item where the defect is located by using the large model to obtain a group reasoning result; performing defect similar code submission record retrieval processing on the vectorized database to obtain an initial code submission record list; performing relevance submission record reasoning processing by utilizing the large model to obtain a target code submission record list; and performing defect reference information and processor reasoning processing by using the large model to obtain a defect reference information and processor reasoning result. According to the method, software defect positioning and distribution are realized, and the accuracy and the efficiency are improved. The method can be widely applied to the technical field of computer software.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer software technology, and in particular to a method and device for locating and allocating software defects based on a large model. Background Art

[0002] Defect management is an important part of software development. Software defects are problems, errors, or hidden functional defects in computer software or programs that damage the normal operation ability. The existence of software defects will cause the software product to fail to meet the needs of users to some extent. Traditional defect location and allocation methods mainly rely on manual experience, which is prone to omissions or misjudgments, low accuracy, cumbersome manual processing, and may lead to unfair allocation and low efficiency.

[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the invention

[0004] The embodiments of the present invention provide a method and device for locating and allocating software defects based on a large model, which effectively improves accuracy and efficiency.

[0005] On the one hand, an embodiment of the present invention provides a software defect location and allocation method based on a large model, comprising the following steps:

[0006] Get the initial code submission record;

[0007] Preprocessing the initial code submission record to obtain a preprocessed code submission record;

[0008] Using a vectorized storage method to store the preprocessed code submission record in a vectorized database;

[0009] Generate a defect information list of a target project using a preset software interface, wherein the defect information list includes a plurality of items of defect information to be processed, and the defect information to be processed includes a defect title, defect details, and a defect module;

[0010] Selecting one item of defect information from the defect information list as target defect information;

[0011] According to the target defect information, the large model is used to perform reasoning processing on the item where the defect is located to obtain a group reasoning result;

[0012] According to the group reasoning result, the vectorized database is searched for defect-similar code submission records to obtain an initial code submission record list;

[0013] According to the initial code submission record list, the large model is used to perform associative submission record reasoning processing to obtain a target code submission record list;

[0014] According to the target code submission record list, the large model is used to perform defect reference information and handler reasoning processing to obtain defect reference information and handler reasoning results, which are used for software defect location and allocation.

[0015] In some embodiments, preprocessing the initial code submission record to obtain a preprocessed code submission record includes:

[0016] Filtering the initial code submission record according to the invalid character information to obtain a first record, wherein the invalid character information includes space character information or line feed character information;

[0017] According to the mapping table, mapping processing is performed on the first record to obtain a second record, wherein the mapping processing is used to map the code submitter code to the company personnel name;

[0018] The second record is format converted to obtain the preprocessing code submission record.

[0019] In some embodiments, storing the preprocessed code submission record in a vectorized database using a vectorized storage method includes:

[0020] According to the vectorization model, converting the preprocessed code submission record into a vector of a preset dimension;

[0021] Generate a hash value according to the code repository name and the preprocessed code submission record;

[0022] According to the hash value, set the record primary key;

[0023] Create source fields;

[0024] According to the code repository name, set the source field;

[0025] According to the record primary key and the source field, the vector of the preset dimension is stored in the vectorized database.

[0026] In some embodiments, the method of performing reasoning processing on the item where the defect is located using a large model based on the target defect information to obtain a group reasoning result includes:

[0027] Generate group definition information according to the business project information, wherein the group definition information includes a plurality of project group aliases;

[0028] Generate a group reasoning prompt word according to the group reasoning guide word, the target defect information, the group definition information and the group reasoning precaution information, wherein the group reasoning guide word includes a group reasoning step, and the group reasoning step includes generating a group where the defect is located according to the target defect information and the group definition information;

[0029] The group reasoning prompt words are input into the large model to obtain the group reasoning results.

[0030] In some embodiments, the performing defect-similar code submission record retrieval processing on the vectorized database according to the group reasoning result to obtain an initial code submission record list includes:

[0031] Generate a code warehouse table according to the multiple code warehouse names, wherein the code warehouse table includes multiple warehouse group aliases, and the warehouse group aliases correspond to the project group aliases;

[0032] Generate a code repository name list according to the group reasoning result and the code repository table;

[0033] According to the preset return quantity, the preset similarity score, the code repository name list and the target defect information, the vectorized database is searched using a preset retrieval module to obtain the initial code submission record list.

[0034] In some embodiments, performing associative submission record reasoning processing using the large model according to the initial code submission record list to obtain a target code submission record list includes:

[0035] Using the target defect information as a defect description;

[0036] Selecting a code submission record from the initial code submission record list as the code submission record to be inferred;

[0037] Generate an associative reasoning prompt word according to the defect description, the submission record of the code to be inferred, the associative reasoning guide word and the associative reasoning precaution information, wherein the associative reasoning guide word includes an associative reasoning step, and the associative reasoning step includes generating an associativity between the defect and the record according to the defect description and the submission record of the code to be inferred;

[0038] Inputting the associative reasoning prompt words into the large model to obtain an associative reasoning result;

[0039] If the relevance reasoning result is that there is a relevance, the submission record of the code to be reasoned is used as the target code submission record;

[0040] Combining multiple target code submission records to obtain the target code submission record list.

[0041] In some embodiments, the method of performing defect reference information and handler reasoning processing using the large model according to the target code submission record list to obtain defect reference information and handler reasoning results includes:

[0042] Generate defect reference information and handler reasoning prompt words according to the target defect information, the target code submission record list, defect reference information and handler reasoning guide words, and defect reference information and handler reasoning precautions information, wherein the defect reference information and handler reasoning guide words include defect reference information and handler reasoning steps;

[0043] The defect reference information and the handler's reasoning prompt words are input into the large model to obtain the defect reference information and the handler's reasoning results.

[0044] In some embodiments, the defect reference information and the handler reasoning step include:

[0045] Determine the defect problem according to the front-end and back-end problem information, where the defect problem includes a front-end problem or a back-end problem;

[0046] Matching the target defect information with the target code submission record list to obtain a matching result;

[0047] According to the matching result, the code submission record with the highest matching degree in the target code submission record list is used as the software defect location result;

[0048] Assign the submitter of the code submission record with the highest matching degree as the software defect assignor;

[0049] Generate defect reference information according to the defect problem, the software defect location result and the software defect assignment handler;

[0050] Generate a defect review result based on the defect reference information.

[0051] On the other hand, an embodiment of the present invention provides a software defect location and allocation device based on a large model, comprising:

[0052] The first module is used to obtain the initial code submission record;

[0053] The second module is used to preprocess the initial code submission record to obtain a preprocessed code submission record;

[0054] A third module is used to store the preprocessed code submission record in a vectorized database using a vectorized storage method;

[0055] The fourth module is used to generate a defect information list of a target project using a preset software interface, wherein the defect information list includes a plurality of items of defect information to be processed, and the defect information to be processed includes a defect title, defect details and a defect module;

[0056] A fifth module is used to select one defect information from the defect information list as target defect information;

[0057] The sixth module is used to perform reasoning processing on the item where the defect is located using the large model according to the target defect information to obtain a group reasoning result;

[0058] The seventh module is used to perform a retrieval process on the defect-similar code submission record of the vectorized database according to the group reasoning result to obtain an initial code submission record list;

[0059] An eighth module is used to perform associative submission record reasoning processing using the large model according to the initial code submission record list to obtain a target code submission record list;

[0060] The ninth module is used to use the large model to perform defect reference information and handler reasoning processing based on the target code submission record list to obtain defect reference information and handler reasoning results, which are used for software defect location and allocation.

[0061] In another aspect, an embodiment of the present invention provides a computer device, comprising:

[0062] at least one processor;

[0063] at least one memory for storing at least one program;

[0064] When the at least one program is executed by the at least one processor, the at least one processor implements the described method.

[0065] The beneficial effects of the present invention are as follows:

[0066] The embodiment of the present invention first obtains the initial code submission record, preprocesses the initial code submission record to obtain the preprocessed code submission record, and then uses the vectorized storage method to store the preprocessed code submission record in a vectorized database, uses a preset software interface to generate a defect information list of the target project, and selects one defect information from the defect information list as the target defect information, and then uses the big model to perform reasoning processing on the project where the defect is located to obtain the group reasoning result, performs defect-similar code submission record retrieval processing on the vectorized database to obtain the initial code submission record list, and uses the big model to perform reasoning processing on the associated submission record to obtain the target code submission record list, and finally uses the big model to perform defect reference information and handler reasoning processing to obtain the defect reference information and handler reasoning results, so that the defect reference information and handler reasoning results can be used to realize software defect location and allocation, thereby improving the accuracy and efficiency.

[0067] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0069] Figure 1 This is a flow chart of a software defect location and allocation method based on a large model according to an embodiment of the present invention;

[0070] Figure 2 A schematic diagram of an overall process of performing defect review according to an embodiment of the present invention;

[0071] Figure 3 This is a structural schematic diagram of a software defect location and allocation device based on a large model according to an embodiment of the present invention;

[0072] Figure 4 The figure is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0073] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.

[0074] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".

[0075] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0076] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0077] Before describing the embodiments of the present application in detail, some nouns and terms involved in the embodiments of the present application are first described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0078] Software Defect: Often called Bug, a software defect is a problem, error, or hidden functional defect in a computer software or program that destroys the normal operation capability. The existence of defects will cause the software product to fail to meet the needs of users to some extent. Definition of defect: From the perspective of the product, a defect is a variety of problems such as errors and faults in the development or maintenance of a software product; from the perspective of the product, a defect is the failure or violation of a certain function that the system needs to implement.

[0079] Vectorized database: refers to a database system that stores and processes data in the form of vectors. The core idea of ​​vectorized database is to represent data as vectors and use the parallel processing capabilities of modern computer hardware to improve data processing speed and efficiency.

[0080] In the related technologies, defect management is a crucial link in the software development process. Traditional defect location and allocation methods mainly rely on manual experience, and have problems such as low efficiency, low accuracy, and poor scalability. The defects of the existing technology are that the defect location process is cumbersome, inefficient, and prone to errors. The existing problems include: (1) Defect location is inaccurate, and it is easy to miss or misjudge. (2) Defect allocation depends on the experience and time of the assignor, which may lead to unfair allocation or inefficiency.

[0081] In view of this, this embodiment uses a large model to infer the location of software defects and defect handlers based on code submission records and defect information lists, shortens the time for software defect location, improves the accuracy of defect location, and realizes an automated process from defect location to allocation, thereby improving accuracy and efficiency. This embodiment implements an efficient, accurate and scalable defect location and intelligent allocation mechanism, which is of great significance for improving software development efficiency and quality.

[0082] The software defect location and allocation method based on a large model provided in the embodiment of the present application relates to the field of computer software technology. The software defect location and allocation method based on a large model provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or it can be configured as a server cluster or a distributed system composed of multiple physical servers, and can also be configured to provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms and other basic cloud computing services. The server can also be a node server in a blockchain network; the software can be an application that implements a software defect location and allocation method based on a large model, etc., but is not limited to the above forms.

[0083] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0084] The following is a detailed explanation of the embodiments of the present application in conjunction with the accompanying drawings:

[0085] Figure 1 is an optional flowchart of the software defect location and allocation method based on a large model provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S109.

[0086] Step S101, obtaining the initial code submission record;

[0087] Step S102: preprocessing the initial code submission record to obtain a preprocessed code submission record;

[0088] Step S103: using a vectorized storage method to store the preprocessed code submission record in a vectorized database;

[0089] Step S104: Generate a defect information list of the target project using a preset software interface, the defect information list including several items of defect information to be processed, the defect information to be processed including defect title, defect details and defect module;

[0090] Step S105, selecting one item of defect information from the defect information list as target defect information;

[0091] Step S106: Based on the target defect information, the large model is used to perform reasoning processing on the item where the defect is located to obtain a group reasoning result;

[0092] Step S107: According to the group reasoning result, the vectorized database is searched for defect-similar code submission records to obtain an initial code submission record list;

[0093] Step S108: Based on the initial code submission record list, the large model is used to perform relevance submission record reasoning processing to obtain a target code submission record list;

[0094] Step S109: Based on the target code submission record list, the large model is used to perform defect reference information and handler reasoning processing to obtain defect reference information and handler reasoning results, which are used for software defect location and allocation.

[0095] Steps S101 to S109 shown in the embodiment of the present application implement software defect location and allocation, and improve accuracy and efficiency.

[0096] In step S101 of some embodiments, the initial code submission record may be obtained through the project code repository. The initial code submission record may also be obtained through other methods, not limited thereto. Exemplarily, all code submission records in the project may be extracted as the initial code submission record.

[0097] In some embodiments, in step S102, preprocessing the initial code submission record to obtain a preprocessed code submission record may include but is not limited to the following steps:

[0098] According to the invalid character information, the initial code submission record is filtered to obtain a first record, wherein the invalid character information includes space character information or line feed character information;

[0099] According to the mapping table, the first record is mapped to obtain a second record, wherein the mapping is used to map the code submitter code to the company personnel name;

[0100] The second record is format converted to obtain a preprocessing code submission record.

[0101] In some embodiments, the initial code submission record can be filtered according to the invalid character information to obtain a first record, wherein the invalid character information includes space character information or line break character information. Then, according to the mapping table, the first record is mapped to obtain a second record, wherein the mapping is used to map the code submitter code to the company personnel name. Exemplarily, the mapping table can be in the form of a hash table, which can be expressed as: {jy: "Luo A", xie: "Xie B"}, in each key-value pair, the code submitter code of the git code repository is used as the key (such as jy), and the company personnel name is used as the value (such as Luo A), and the mapping table can store multiple personnel mapping relationships. More, the code submitter mailbox and the company personnel mailbox can be mapped using a similar mapping table. Finally, the second record is formatted to obtain a preprocessed code submission record. Exemplarily, the format used for conversion may include: "Code submission description: description person; code submitter: company personnel name; submission date: 20XX-XX-XX 18:00", indicating the date and time.

[0102] In some embodiments, in step S103, the preprocessed code submission record is stored in the vectorized database using a vectorized storage method, which may include but is not limited to the following steps:

[0103] According to the vectorization model, the preprocessed code submission record is converted into a vector of a preset dimension;

[0104] Generate a hash value based on the code repository name and preprocessed code submission record;

[0105] According to the hash value, set the record primary key;

[0106] Create source fields;

[0107] Set the source field according to the code repository name;

[0108] According to the record primary key and source field, vectors of preset dimensions are stored in the vectorized database.

[0109] In some embodiments, Milvus' vectorized storage technology can be used to store preprocessed code submission records in a vectorized database for efficient similarity retrieval. The preprocessed code submission records can be first converted into a vector of a preset dimension according to the vectorized model. For example, the vectorized model can be set to BAAI / bge-m3, the preset dimension can be set to 1024 dimensions, and converted into a 1024-dimensional vector, such as [-0.06615690141916275, -0.0182508099824190,......]. Then, a hash value is generated according to the code repository name and the preprocessed code submission record, and the record primary key is set according to the hash value. Then create a source field (source field), and set the source field according to the code repository name for searching different code repositories. Finally, according to the record primary key and source field, the vector of the preset dimension is stored in the vectorized database.

[0110] In some embodiments, in step S104, a preset software interface can be used to generate a defect information list of the target project, wherein the defect information list includes several items of pending defect information, and the pending defect information includes defect title, defect details and defect module. Exemplarily, the pending defect information of the current target project, including defect title, defect details and defect module, etc., can be obtained through the software interface of the Tapd agile software management system, and a defect information list can be output. In the example of pending defect information, the content of the defect title can include [online form] custom input text style, which is not effective for mobile phone area code. The content of the defect details can include the expected result, the mobile phone number and area code follow the input text setting item; the actual result, after the text setting item is entered, the mobile phone number and area code style is not effective, and the mobile phone view is the same. The content of the defect module can include responsive-online form. It can be understood that due to the limitations of the Tapd interface, the request frequency can be set to 50 times per minute in the program code.

[0111] In some embodiments, in step S105, one item of defect information may be selected from the defect information list as target defect information, so that the target defect information can be subsequently inferred using the large model. In addition, the defect information list may be traversed to infer each item of defect information in the defect information list.

[0112] In some embodiments, in step S106, based on the target defect information, the large model is used to perform reasoning processing on the item where the defect is located to obtain a group reasoning result, which may include but is not limited to the following steps:

[0113] Generate group definition information according to the business project information, the group definition information includes multiple project group aliases;

[0114] Generate group reasoning prompt words according to group reasoning guide words, target defect information, group definition information and group reasoning precaution information, wherein the group reasoning guide words include group reasoning steps, and the group reasoning steps include generating a group where the defect is located according to the target defect information and the group definition information;

[0115] Input the group reasoning prompt words into the large model to obtain the group reasoning results.

[0116] In some embodiments, the cloud vendor's large model reasoning technology can be used to infer the project where the defect may be located based on information such as the defect title, defect details and defect module. Group definition information can be generated first based on the business project information, wherein the group definition information includes multiple project group aliases. Exemplarily, the business project information may include Project A: PC-side website building, a project designed for dragging and dropping websites for PC-side websites; Project B: Mobile-side website building, a project designed for dragging and dropping websites for mobile-side websites; Project C: Responsive website building, one-time website building, automatic response matching website projects for PC and mobile terminals, which are divided into computer view, mobile view and iPad view. The group definition information may include project group A: including PC and computer terminals; project group B: including mobile terminals and Mobile; project group C: including responsive, mobile view, computer view, iPad view; default project group: does not include information of project group A, project group B and project group C. The project group alias may be project group A, project group B, project group C or default project group. Then, according to the group reasoning guide words, target defect information, group definition information and group reasoning precaution information, group reasoning prompt words are generated, wherein the group reasoning guide words include group reasoning steps, and the group reasoning steps include generating the group to which the defect belongs according to the target defect information and the group definition information. Exemplarily, the group reasoning guide words can be "You are a classifier, you need to judge which group the defect belongs to based on the target defect information and the group definition information". The group reasoning precaution information can be "Attention: A JSON object must be returned; group: Fill in the inferred grouping here". Furthermore, each defect information in the defect information list can be inferred, and different prompt words can be constructed respectively. Finally, the group reasoning prompt words are input into the large model to obtain the group reasoning result. Exemplarily, the group reasoning result output by the large model can be {group: "Group C"}, indicating that the large model reasoning believes that the defect information contains the information of Group C (responsive, mobile view), so Group C is returned as the group reasoning result.

[0117] In some embodiments, in step S107, based on the group reasoning result, the vectorized database is searched for defect-similar code submission records to obtain an initial code submission record list, which may include but is not limited to the following steps:

[0118] Generate a code repository table based on multiple code repository names. The code repository table includes multiple repository group aliases. There is a corresponding relationship between the repository group aliases and the project group aliases.

[0119] Generate a list of code repository names based on the group reasoning results and the code repository table;

[0120] According to the preset return quantity, preset similarity score, code repository name list and target defect information, the vectorized database is searched using the preset retrieval module to obtain the initial code submission record list.

[0121] In some embodiments, the code submission record most similar to the target defect information can be retrieved from the Milvus vectorization database according to the project group alias of the group reasoning result. A code warehouse table can be generated first according to multiple code warehouse names, wherein the code warehouse table includes multiple warehouse group aliases, and there is a corresponding relationship between the warehouse group aliases and the project group aliases. Exemplarily, the code warehouse table can be represented as {warehouse group A: ["code warehouse 1", "code warehouse 2"]; warehouse group B: ["code warehouse 3", "code warehouse 4"]; warehouse group C: ["code warehouse 5"]; default warehouse group: ["code warehouse 1", "code warehouse 2", "code warehouse 3", "code warehouse 4", "code warehouse 5", "code warehouse 6"]}, the code warehouse table is in the form of a hash table in the program code, the key is the warehouse group alias, and the value is the code warehouse name array. The code warehouse name is consistent with the source field stored in the vectorization database. In the correspondence between the warehouse group alias and the project group alias, warehouse group A can correspond to project group A, warehouse group B corresponds to project group B, and warehouse group C corresponds to project group C. Then, based on the group reasoning results and the code warehouse table, a code warehouse name list is generated. For example, the corresponding code warehouse name list can be found in the code warehouse table according to the project group alias of the group reasoning result. Finally, according to the preset return quantity, preset similarity score, code warehouse name list and target defect information, the preset retrieval module is used to search the vectorized database to obtain the initial code submission record list. For example, the Milvus retrieval module can be called, the preset return quantity topk is set to 10, the preset similarity score score_threshold is set to 0.8, source is the code warehouse name list, the retrieval module retrieves 10 code submission records with a similarity score greater than 0.8, and outputs the initial code submission record list.

[0122] In some embodiments, in step S108, according to the initial code submission record list, the large model is used to perform associative submission record reasoning processing to obtain the target code submission record list, which may include but is not limited to the following steps:

[0123] Using target defect information as defect description;

[0124] Select a code submission record from the initial code submission record list as the code submission record to be reasoned;

[0125] Generate a correlation reasoning prompt word according to the defect description, the submission record of the code to be inferred, the correlation reasoning guide word and the correlation reasoning precaution information, wherein the correlation reasoning guide word includes a correlation reasoning step, and the correlation reasoning step includes generating a correlation between the defect and the record according to the defect description and the submission record of the code to be inferred;

[0126] Input the associative reasoning prompt words into the large model to obtain the associative reasoning results;

[0127] If the result of the correlation reasoning is that there is a correlation, the submission record of the code to be reasoned is used as the target code submission record;

[0128] Combine multiple target code submission records to obtain a target code submission record list.

[0129] In some embodiments, a large model can be used to screen the correlation between code submission records and defects, and output the code submission records with the highest correlation to reduce the illusion problem of the large model. The initial code submission record list can be traversed, and a large model reasoning can be performed on each submission record to screen the relevance of the code submission record. If the large model infers that this submission record is related to the defect description, it is retained, and if it is not related, it is discarded. The target defect information can be first used as the defect description, and a code submission record can be selected from the initial code submission record list as the code submission record to be reasoned. Then, based on the defect description, the code submission record to be reasoned, the correlation reasoning guide words and the correlation reasoning precautions information, the correlation reasoning prompt words are generated, wherein the correlation reasoning guide words include the correlation reasoning steps, and the correlation reasoning steps include generating the correlation between the defect and the record based on the defect description and the code submission record to be reasoned. Exemplarily, the guiding words for relevance reasoning may be "You are a submission record scorer. You need to determine whether the record is related to the defect description based on the defect description and the code submission record to be inferred." The relevance reasoning precaution information may be "Attention: A JSON object must be returned; binary_score: Fill in the inferred binary_score here, and the value range is a yes or no string." Then input the relevance reasoning prompt words into the big model to obtain the relevance reasoning result. Exemplarily, the relevance reasoning result returned by the big model may be {binary_score: "yes"}, where binary_score represents the binary relevance score. If the relevance reasoning result is that there is a correlation, the code submission record to be inferred is used as the target code submission record. After the big model reasoning is performed on each code submission record in the initial code submission record list, multiple target code submission records whose relevance reasoning results show that there is a correlation are combined to obtain a target code submission record list. It is understandable that, based on the relevance reasoning results returned by the large model, the code submission records with binary_score of no can be discarded, as these submission records are considered irrelevant by the large model, and the code submission records with binary_score of yes after the large model reasoning are retained.

[0130] In some embodiments, in step S109, based on the target code submission record list, the defect reference information and the handler reasoning process are performed using the big model to obtain the defect reference information and the handler reasoning result, which may include but is not limited to the following steps:

[0131] Generate defect reference information and handler reasoning prompt words according to target defect information, target code submission record list, defect reference information and handler reasoning guide words, and defect reference information and handler reasoning precaution information, wherein the defect reference information and handler reasoning guide words include defect reference information and handler reasoning steps;

[0132] Input the defect reference information and the processor inference prompt into the large model to obtain the defect reference information and the processor inference result.

[0133] In some embodiments, the target code submission record list can be used as the context, combined with the prompt, and the large model is used to output the defect reference information and infer the processor. Based on the target code submission record list and the target defect information as the context for large model reasoning, the defect reference information and the processor inference prompt can be generated first according to the target defect information, the target code submission record list, the defect reference information and the processor inference guiding words, and the defect reference information and the processor inference precautions information. Among them, the defect reference information and the processor inference guiding words include the defect reference information and the processor inference steps. Exemplarily, the code submission records found in the target code submission record list may include "User group: Front-end: [Fill in the names of the front-end personnel in the project team here] (filled in according to the mapping table of the code submitters); Back-end: [Fill in the names of the back-end personnel in the project team here] (filled in according to the mapping table of the code submitters)". The defect reference information and the processor inference precautions information may be "Attention! The returned data format must be in json format". Then input the defect reference information and the processor inference prompt into the large model to obtain the defect reference information and the processor inference result. Among them, the defect reference information and the processor inference result are used for software defect location and assignment. Exemplarily, the defect reference information and the processor inference result may include {user: "Chen C", indicating one of the user groups, reason: "The problem mainly involves the form style and belongs to a front-end problem. According to the submission record and the user group information, it is initially judged that this problem should be handled by the front-end team. Among them, the record submitted by Chen C is the most relevant to the defect description, and he is a member of the front-end team and is most likely to solve this problem"}. The located software defect can be assigned to "Chen C".

[0134] In some embodiments, the defect reference information and the processor inference steps include:

[0135] Determine the defect problem according to the front-end and back-end problem information, and the defect problem includes front-end problems or back-end problems;

[0136] Match the target defect information with the target code submission record list to obtain a matching result;

[0137] According to the matching result, use the code submission record with the highest matching degree in the target code submission record list as the software defect location result;

[0138] Use the submitter corresponding to the code submission record with the highest matching degree as the software defect assignment processor;

[0139] Generate defect reference information based on defect issues, software defect location results, and software defect assignment handlers;

[0140] Generate defect review results based on defect reference information.

[0141] In some embodiments, the defect problem can be determined based on the front-end and back-end problem information, wherein the defect problem includes a front-end problem or a back-end problem. Exemplarily, it can be determined whether the defect problem belongs to the Web front-end or the Web back-end problem, the front-end problems include interaction, vision, and UI problems, and the back-end problems include logic, security, and error reporting problems. Then the target defect information and the target code submission record list are matched to obtain a matching result, and according to the matching result, the code submission record with the highest matching degree in the target code submission record list is used as the software defect location result, and the submitter corresponding to the code submission record with the highest matching degree is used as the software defect assignment handler. Exemplarily, the found code submission record can be matched with the problem description. If the code submission record has a high matching degree with the problem description, the submission record with the highest matching degree is the problem, and the software defect location result is obtained, and the submitter is the software defect assignment handler. More, if there are multiple code submission records, they are sorted according to the submission time, and the code submission record with the last time is taken as the software defect location result. Then, based on the defect problem, the software defect location result and the software defect assignment handler, defect reference information is generated. For example, the reasoning process can be explained according to the reasoning steps and the returned data logic, and the handler is returned in json format to obtain the defect reference information. The defect reference information can be "user: fill in the inferred handler here; reason: fill in the reasoning reason here". Finally, based on the defect reference information, the defect comment result is generated. For example, the data returned by the big model can be combined with the Tapd comment API to comment on the defect. The comment content can be "Based on intelligent speculation, this defect is suitable for processing by the user field of the data returned by the big model. Reason: reason field information returned by the big model; context: filtered target code submission record list; combined with the Tapd defect API, the software defect assignment handler is changed to the user field returned by the big model".

[0142] In some embodiments, the overall process of performing defect review is as follows: Figure 2As shown, you can first extract data (such as project 1, project 2), and segment or clean the data, then convert the submission record into a vector and store it in the Milvus database (such as submit 1 into vector 1), then retrieve the submission record, sort it by score (such as the similarity of submit 1 is 0.841232, the similarity of submit 2 is 0.821232), filter the large model according to the prompt word (such as filtering out submit 1, submit 2 and submit 3 to retain submit 1 and submit 2), ask questions to the large model, input the defect title, defect description, defect module and prompt word, get the defect comment content, and finally send the defect comment content to Tapd.

[0143] In some embodiments, this embodiment uses the Milvus vector database to store and retrieve code submission records to achieve efficient similarity matching. The cloud vendor's big model is used to perform code repository reasoning and correlation screening to improve the accuracy of defect location. Combined with prompt words, the big model is used to output reference information of defects and inferred handlers to achieve intelligent allocation. The entire process of this embodiment is completed automatically without human intervention, which significantly improves the efficiency of defect handling and has a high degree of automation. Using a big model for in-depth analysis can more accurately locate the defect location and recommend suitable handlers with high accuracy. By comprehensively analyzing code submission records and defect information, the cause of the defect and the treatment plan can be intelligently inferred to provide a reference for developers, with a high degree of intelligence. The model parameters and algorithms can be flexibly adjusted according to different needs to adapt to different development scenarios, with strong scalability.

[0144] The beneficial effects of implementing the embodiments of the present invention include: the embodiments of the present invention first obtain the initial code submission record, preprocess the initial code submission record to obtain the preprocessed code submission record, and then use the vectorized storage method to store the preprocessed code submission record in a vectorized database, use the preset software interface to generate a defect information list of the target project, and select one defect information from the defect information list as the target defect information, and then use the big model to perform reasoning processing on the project where the defect is located to obtain the group reasoning result, perform defect-similar code submission record retrieval processing on the vectorized database to obtain the initial code submission record list, and use the big model to perform association submission record reasoning processing to obtain the target code submission record list, and finally use the big model to perform defect reference information and handler reasoning processing to obtain the defect reference information and handler reasoning results, so that the defect reference information and handler reasoning results can be used to realize software defect location and allocation, thereby improving the accuracy and efficiency.

[0145] like Figure 3 As shown, the embodiment of the present invention also provides a software defect location and allocation device based on a large model, including:

[0146] The first module 801 is used to obtain the initial code submission record;

[0147] The second module 802 is used to pre-process the initial code submission record to obtain a pre-processed code submission record;

[0148] The third module 803 is used to store the preprocessed code submission record into the vectorized database by using the vectorized storage method;

[0149] The fourth module 804 is used to generate a defect information list of the target project using a preset software interface, the defect information list includes a number of defect information to be processed, and the defect information to be processed includes a defect title, defect details and defect module;

[0150] The fifth module 805 is used to select one defect information from the defect information list as target defect information;

[0151] The sixth module 806 is used to perform reasoning processing on the item where the defect is located using the large model according to the target defect information to obtain a group reasoning result;

[0152] The seventh module 807 is used to perform a search process on the defect-similar code submission record of the vectorized database according to the group reasoning result to obtain an initial code submission record list;

[0153] The eighth module 808 is used to perform relevance submission record reasoning processing using the large model according to the initial code submission record list to obtain a target code submission record list;

[0154] The ninth module 809 is used to perform defect reference information and handler reasoning processing using a large model according to a target code submission record list, and obtain defect reference information and handler reasoning results, which are used for software defect location and allocation.

[0155] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0156] like Figure 4 As shown, an embodiment of the present invention further provides a computer device, including:

[0157] at least one processor 901;

[0158] At least one memory 902, used to store at least one program;

[0159] When at least one program is executed by at least one processor, the at least one processor implements Figure 1 The method shown.

[0160] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0161] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. A software defect location and allocation method based on a large model, characterized in that: The following steps are involved: Get the initial code submission record; Preprocessing the initial code submission record to obtain a preprocessed code submission record; Using a vectorized storage method to store the preprocessed code submission record in a vectorized database; Generate a defect information list of a target project using a preset software interface, wherein the defect information list includes a plurality of items of defect information to be processed, and the defect information to be processed includes a defect title, defect details, and a defect module; Selecting one item of defect information from the defect information list as target defect information; According to the target defect information, the large model is used to perform reasoning processing on the item where the defect is located to obtain a group reasoning result; According to the group reasoning result, the vectorized database is searched for defect-similar code submission records to obtain an initial code submission record list; According to the initial code submission record list, the large model is used to perform associative submission record reasoning processing to obtain a target code submission record list; According to the target code submission record list, the large model is used to perform defect reference information and handler reasoning processing to obtain defect reference information and handler reasoning results, which are used for software defect location and allocation.

2. The method according to claim 1, characterized in that: The preprocessing of the initial code submission record to obtain a preprocessed code submission record includes: Filtering the initial code submission record according to the invalid character information to obtain a first record, wherein the invalid character information includes space character information or line feed character information; According to the mapping table, mapping processing is performed on the first record to obtain a second record, wherein the mapping processing is used to map the code submitter code to the company personnel name; The second record is format converted to obtain the preprocessing code submission record.

3. The method according to claim 1, characterized in that The method of storing the preprocessed code submission record in a vectorized database by using a vectorized storage method includes: According to the vectorization model, converting the preprocessed code submission record into a vector of a preset dimension; Generate a hash value according to the code repository name and the preprocessed code submission record; According to the hash value, set the record primary key; Create source fields; According to the code repository name, set the source field; According to the record primary key and the source field, the vector of the preset dimension is stored in the vectorized database.

4. The method according to claim 1, characterized in that According to the target defect information, the large model is used to perform reasoning processing on the item where the defect is located to obtain a group reasoning result, including: Generate group definition information according to the business project information, wherein the group definition information includes a plurality of project group aliases; Generate a group reasoning prompt word according to the group reasoning guide word, the target defect information, the group definition information and the group reasoning precaution information, wherein the group reasoning guide word includes a group reasoning step, and the group reasoning step includes generating a group where the defect is located according to the target defect information and the group definition information; The group reasoning prompt words are input into the large model to obtain the group reasoning results.

5. The method according to claim 4, characterized in that According to the group reasoning result, the vectorized database is searched for defect-similar code submission records to obtain an initial code submission record list, including: Generate a code warehouse table according to the multiple code warehouse names, wherein the code warehouse table includes multiple warehouse group aliases, and the warehouse group aliases correspond to the project group aliases; Generate a code repository name list according to the group reasoning result and the code repository table; According to the preset return quantity, the preset similarity score, the code repository name list and the target defect information, the vectorized database is searched using a preset retrieval module to obtain the initial code submission record list.

6. The method according to claim 1, characterized in that The step of performing relevance submission record reasoning processing using the large model according to the initial code submission record list to obtain a target code submission record list includes: Using the target defect information as a defect description; Selecting a code submission record from the initial code submission record list as the code submission record to be inferred; Generate an associative reasoning prompt word according to the defect description, the submission record of the code to be inferred, the associative reasoning guide word and the associative reasoning precaution information, wherein the associative reasoning guide word includes an associative reasoning step, and the associative reasoning step includes generating an associativity between the defect and the record according to the defect description and the submission record of the code to be inferred; Inputting the associative reasoning prompt words into the large model to obtain an associative reasoning result; If the relevance reasoning result is that there is a relevance, the submission record of the code to be reasoned is used as the target code submission record; Combining multiple target code submission records to obtain the target code submission record list.

7. The method according to claim 1, characterized in that According to the target code submission record list, the large model is used to perform defect reference information and handler reasoning processing to obtain defect reference information and handler reasoning results, including: Generate defect reference information and handler reasoning prompt words according to the target defect information, the target code submission record list, defect reference information and handler reasoning guide words, and defect reference information and handler reasoning precautions information, wherein the defect reference information and handler reasoning guide words include defect reference information and handler reasoning steps; The defect reference information and the handler's reasoning prompt words are input into the large model to obtain the defect reference information and the handler's reasoning results.

8. The method according to claim 7, characterized in that The defect reference information and the reasoning steps of the handler include: Determine the defect problem according to the front-end and back-end problem information, where the defect problem includes a front-end problem or a back-end problem; Matching the target defect information with the target code submission record list to obtain a matching result; According to the matching result, the code submission record with the highest matching degree in the target code submission record list is used as the software defect location result; Assign the submitter of the code submission record with the highest matching degree as the software defect assignor; Generate defect reference information according to the defect problem, the software defect location result and the software defect assignment handler; Generate a defect review result based on the defect reference information.

9. A software defect location and allocation device based on a large model, characterized in that: include: The first module is used to obtain the initial code submission record; The second module is used to preprocess the initial code submission record to obtain a preprocessed code submission record; A third module is used to store the preprocessed code submission record in a vectorized database using a vectorized storage method; The fourth module is used to generate a defect information list of a target project using a preset software interface, wherein the defect information list includes a plurality of items of defect information to be processed, and the defect information to be processed includes a defect title, defect details and a defect module; A fifth module is used to select one defect information from the defect information list as target defect information; The sixth module is used to perform reasoning processing on the item where the defect is located using the large model according to the target defect information to obtain a group reasoning result; The seventh module is used to perform a retrieval process on the defect-similar code submission record of the vectorized database according to the group reasoning result to obtain an initial code submission record list; An eighth module is used to perform associative submission record reasoning processing using the large model according to the initial code submission record list to obtain a target code submission record list; The ninth module is used to use the large model to perform defect reference information and handler reasoning processing based on the target code submission record list to obtain defect reference information and handler reasoning results, which are used for software defect location and allocation.

10. A computer device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 8.