An Automatic Defect Report Monitoring and Synthesis Method for Developer Group Chat

An automated method using deep learning and semantic analysis extracts software error reports from community chat data, enhancing efficiency and reducing costs by accurately identifying and generating error reports from chat content.

CN114610888BActive Publication Date: 2025-07-15INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210272371.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-18
Publication Date
2025-07-15
Estimated Expiration
2042-03-18

AI Technical Summary

Technical Problem

In the prior art, it is difficult for software developers to efficiently obtain software error information mentioned by users in online chats, resulting in high cost and low efficiency of manual sorting.

Method used

Semantic analysis and data mining technology based on deep learning are adopted to decouple chat information through forward neural networks, and context information is obtained using graph neural networks, and information extraction model is constructed in combination with transfer learning to automatically identify and generate software error reports from chat information.

Benefits of technology

It realizes the rapid and accurate generation of software error reports from a large number of chat information, reduces development costs, broadens the access to error reports, and improves the efficiency of software development and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114610888B_ABST
    Figure CN114610888B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic monitoring and synthesis method for defect reports for chatting among developer groups, and its steps include: 1) Collecting online chat data, decoupling the collected chat data, and performing data augmentation on the decoupled data to obtain a conversation decoupled data set after data augmentation; 2) Feeding the conversation decoupled data set into a conversation classification model to classify conversations containing software error information and conversations not containing software error information; 3) Feeding the conversations containing software error information into a software error information extraction model to obtain the category to which each sentence in the conversation belongs, and generating a software error report based on the sentences and their corresponding categories. The present invention realizes the full automation of the process from chat information to the generation of software error reports, can quickly and accurately generate software error reports, reduces the cost of obtaining software error reports in the software development process, broadens the access channels for software error reports, and improves the software development and maintenance efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and particularly relates to an automatic monitoring and synthesis method for defect reports for developer group chats, and a corresponding storage medium and electronic device. Background Art

[0002] In the process of software development iteration, software error reports play a crucial role for software developers to understand and fix current software errors. Software error reports are an important information source for software developers to understand, locate, and reproduce software errors. Currently, software developers mainly obtain software error reports by the way of users actively submitting them. However, this way is limited by the user's own technical level, understanding of the software, and whether they know the feedback channels, etc. Software developers can often only obtain a small number of software error reports. At the same time, users often discuss the software errors or anomalies they encounter and seek solutions through online chats in public communities. And this user habit results in a large amount of software error information existing in the online chat content of the community. However, since the online chat content contains a large amount of irrelevant information, if software developers adopt a manual way to obtain chat error information and organize it into software error reports, it will consume huge human costs and time. If the software error information in the online chat can be automatically identified and automatically integrated into software error reports, this will greatly expand the way for developers to obtain software error reports, improve the software development efficiency, and reduce the software development cost. Summary of the Invention

[0003] In view of the above problems, the present invention provides an automatic monitoring and synthesis method for defect reports for developer group chats, and a corresponding storage medium and electronic device, aiming to solve the problems of quickly and accurately extracting software error information from a large amount of complicated and redundant chat information and generating error reports, broadening the way for software developers to obtain error reports, thereby improving the software development quality and reducing the software development cost. This method combines technologies such as natural language processing, text mining, and deep learning, and trains and optimizes the model on the community chat dialogue database to overcome problems such as the flexibility and non-standardization of software error descriptions in chat information.

[0004] The present invention automatically decouples chat information described in natural language through semantic analysis and data mining techniques based on deep learning, understands the semantic information of user chat conversations, identifies conversations containing software error information, extracts software error information from such conversations, and then automatically generates a software error report. By automatically generating software error reports from user chat information in public communities, the present invention can broaden the channels for software developers to obtain software error reports, enabling software developers to obtain more software error information, so as to help software developers improve software development efficiency and at the same time reduce software maintenance costs.

[0005] The present invention relates to a method for automatically monitoring and synthesizing defect reports for group chats of developers, and its steps include:

[0006] 1) Collect online chat data, use a dialogue decoupling model based on a forward neural network to decouple the above chat information and perform data augmentation based on simple data augmentation techniques (EDA) to obtain a data-augmented dialogue decoupling data set.

[0007] 2) Feed the decoupled dialogue decoupling data set into a dialogue classification model based on a graph neural network to classify conversations containing software error information and conversations not containing software error information.

[0008] 3) Feed the obtained conversations containing software error information into a software error information extraction model based on transfer learning to obtain the category to which each sentence belongs, thereby generating a software error report.

[0009] Furthermore, the steps of using the dialogue decoupling model based on a forward network to decouple the above chat information include:

[0010] 1) Preprocess the original chat information, filter out pictures and expressions, and convert website addresses, codes, email addresses, versions, and HTML elements into five feature tags: [URL], [CODE], [EMAIL], [VERSION], and [HTML].

[0011] 2) The present invention uses a model based on a forward neural network to decouple dialogue data. This model is a dialogue decoupling model proposed by Kummerfeld, which consists of two layers of forward neural networks, each layer having 512 hidden units. We input the preprocessed chat data into the dialogue decoupling model to obtain a decoupled output.

[0012] 3) According to the decoupled output, reconstruct each group of conversations in the chat information, where a group of conversations contains N user speeches with the same theme.

[0013] Furthermore, the process of performing data augmentation based on EDA includes:

[0014] 1) The original conversation contains N user utterances. For each of these utterances, according to the number of words (length) in the utterance, a word replacement or utterance replacement strategy is selected for enhancement. When the length of the utterance is greater than the threshold θ, word replacement is adopted, that is, some words in the user utterance are replaced with synonyms to generate a new user utterance. When the length of the utterance is not greater than the threshold θ, an utterance replacement method is adopted, that is, a user utterance with a length not greater than θ is randomly selected from the conversation dataset to replace the current user utterance.

[0015] 2) Through the above two strategies of word replacement and utterance replacement, for the original conversation containing N user utterances, N new user utterances can be generated. The newly generated N user utterances are combined into a new conversation. Further, the classification process of the dialogue classification model based on the graph neural network for the conversation includes:

[0016] 1) Input the conversation to be predicted into the dialogue classification model.

[0017] 2) Use the pre-trained BERT language model to perform word encoding operations on the words in each sentence of the input conversation to obtain the word vectors corresponding to the words. BERT is a language representation model proposed by Google, consisting of a bidirectional Transformer encoder, and is pre-trained on a large amount of data. The BERT model has been widely used in various natural language processing tasks.

[0018] 3) Input the word vectors of all words in each sentence into the TextCNN model to obtain the sentence vector representation of each sentence.

[0019] 4) Input the sentence vector representation of each sentence into the graph neural network model to obtain the context sentence vector representation of each sentence. The graph neural network is a neural network that directly acts on the graph structure. It can, through the information propagation mechanism, transmit the information of neighboring nodes to adjacent nodes, so that each node in the graph can obtain the information of other nodes. We construct each conversation as a dialogue graph, with each sentence as a node in the dialogue graph, and then use the graph neural network to learn the context representation of each sentence.

[0020] 5) Concatenate the sentence vector and the context sentence vector of each sentence and input them into the sum pooling and max pooling layers to obtain the dialogue-level vector representation.

[0021] 6) Input the dialogue-level vector representation into a fully connected classification layer to obtain the prediction of the model for the category of this conversation.

[0022] Further, the process of constructing the software error information extraction model based on transfer learning includes:

[0023] 1) Obtain a publicly available software error report dataset, and perform text processing operations on the sentences in the software error report dataset, such as converting uppercase letters to lowercase, tokenizing, deleting non-English sentences, and deleting overly long sentences, and manually annotate the sentences.

[0024] 2) Connect a BERT pre-trained language model to a fully connected layer, where the fully connected layer serves as the final output layer, thus constructing an information extraction model. In the information extraction model, the parameters of the first 9 layers of the BERT pre-trained model are frozen, and the parameters of the remaining 3 layers are unfrozen to participate in the following two fine-tuning processes.

[0025] 3) Feed the obtained sentences into the information extraction model for the first fine-tuning to obtain the first fine-tuned information extraction model.

[0026] 4) Annotate the sentences spoken by the conversation initiator in the conversation containing software error information, and filter the annotated sentences according to heuristic rules to obtain a software error information sentence training set.

[0027] 5) Replace the fully connected layer of the first fine-tuned information extraction model with a brand-new fully connected layer. The parameters in the brand-new fully connected layer are randomly initialized. Through this replacement, the information extraction model can learn the information in the software error information sentence training set more quickly. Then use the software error information sentence training set to train the replaced information extraction model, and re-learn the model parameters for the second fine-tuning of transfer learning to obtain a software error information extraction model after two fine-tuning processes based on transfer learning.

[0028] Furthermore, the category to which each sentence belongs includes Observed Behavior (OB), Expected Behavior (EB), Steps to Reproduce (SR), and Others.

[0029] Furthermore, the process of generating a software error report includes:

[0030] 1) According to the results of the information extraction model (i.e., the category to which each sentence belongs), classify the sentences into four categories.

[0031] 2) Input the clustering results into a software error report template to generate a software error report.

[0032] A storage medium stores a computer program, wherein the computer program executes the above method.

[0033] An electronic device includes a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to execute the above method.

[0034] Compared with the prior art, the advantages of the present invention are as follows:

[0035] The present invention makes the first attempt to automatically generate software error messages from chat information.

[0036] The present invention proposes to use a graph neural network to obtain the context information of chat information and construct a dialogue classification model, realizing the efficient identification of whether the dialogue contains software error messages.

[0037] The present invention proposes to use transfer learning and combine it with heuristic rules to construct an information extraction model, realizing the accurate identification of the category to which the sentences in the dialogue belong.

[0038] The present invention does not require manual intervention, overcomes the situation of non-standard and highly flexible language expressions in user descriptions in chat information, and has cross-domain adaptability. It realizes the full process automation from chat information to software error report generation, so as to quickly and accurately generate software error reports from a large amount of chat information, reduce the cost of obtaining software error reports in the software development process, broaden the access to software error reports, and improve the software development and maintenance efficiency.

[0039] The dialogue model in the present invention achieves an average F1 value of 77.74% on the test set, which is 12.96% higher than the baseline. At the same time, the information extraction model achieves an average F1 value of 84.62% in the OB category, 71.46% in the EB category, and 73.13% in the SR category on the test set, and is 9.32%, 12.21%, and 10.91% higher than the baseline respectively. Description of the Drawings

[0040] Figure 1 It is a framework diagram of the software report automatic generation method of the present invention.

[0041] Figure 2 It is a flow chart of training the dialogue classification model of the present invention.

[0042] Figure 3 It is a flow chart of the information extraction model of the present invention. Detailed Embodiments

[0043] Although the specific content, implementation algorithms and drawings of the present invention are disclosed for the purpose of explaining the present invention, the purpose is to help understand the content of the present invention and implement it accordingly. However, those skilled in the art can understand that: without departing from the spirit and scope of the present invention and the appended claims, various substitutions, changes and modifications are possible. The present invention should not be limited to the content disclosed in the best embodiments of this specification and the drawings. The scope of protection required by the present invention shall be subject to the scope defined by the claims.

[0044] The present invention first proposes a method for automatically generating software error reports from chat messages. Through semantic analysis and natural language processing technologies, the present invention uses a forward neural network to automatically decouple chat messages to obtain a set of chat conversations, and then uses a graph neural network to obtain the context information in the chat conversations, thereby obtaining the ability to accurately judge the content of the chat conversations. Further, through the information extraction model obtained by using transfer learning, the software error information content in the chat conversations is extracted, and finally, a software error report is automatically generated through a template. The present invention provides a full-process automation solution for generating software error reports from chat messages. The following is a further description of this aspect through specific implementation examples.

[0045] As Figure 1 shown, it is a method framework diagram for automatically generating software error reports from chat messages according to the present invention. The present invention includes five main steps: dialogue decoupling, data augmentation, training a dialogue classification model, sentence preprocessing, and sentence classification:

[0046] Step 1: Collect and decouple chat messages. First, collect chat messages from the community chat channel and save them locally in the txt file format. Then, through the preprocessing work before chat decoupling, filter out pictures and emojis, and at the same time convert the URLs, codes, email addresses, versions, and HTML elements in the chat into five feature tags: [URL], [CODE], [EMAIL], [VERSION], and [HTML]. In this way, the processed chat message text after preprocessing is obtained. Then, the obtained processed chat message text is sent into a pre-trained forward neural network model, and through model decoupling, a set of chat conversations can be obtained:

[0047] L = {D1, D2, …, D n}

[0048] D = {U1, U2, …, U n}

[0049] where L represents chat messages, which are composed of an array of chat conversations. And D i represents a user conversation, which is composed of several user utterances. And U i represents a user utterance.

[0050] Step 2: Perform data augmentation on the chat conversations. In data augmentation, we adopt two replacement strategies: word replacement and utterance replacement. Define the length of an utterance as L, the threshold of the utterance length as θ, and the enhanced utterance as U′. Then the replacement strategy satisfies the following formula:

[0051]

[0052] When the length L of the speech is greater than θ, the word replacement method is adopted, that is, some words in the user's speech are replaced with synonyms to generate a new user speech; when the length L of the speech is not greater than θ, the speech replacement method is adopted, that is, a user speech with a length not greater than θ is randomly selected from the dialogue dataset to replace the current user speech to generate a new user speech.

[0053] Combine the enhanced speeches into a new user dialogue D aug :

[0054] D aug ={U′1,U′2,…,U′ n}

[0055] After data enhancement, we finally obtain the dialogue decoupling dataset.

[0056] Step 3, dialogue classification model. As Figure 2 shown, the entire dialogue classification model consists of three layers: the sentence encoding layer, the graph-based context encoding layer, and the dialogue encoding and classification layer. For the model training stage: the input is a dialogue D = {u1, u2, …, u n} with known label categories. First, in the sentence encoding layer, the present invention uses the pre-trained deep language model BERT to perform word embedding operations on all sub-words of each sentence u i in the dialogue D to obtain all word vectors of each sentence. Then, the TextCNN model is used to perform convolution operations on the word vectors of each sentence to obtain the sentence vector

[0057] Since there is usually a reply-to relationship between sentences in the same dialogue, the context of a sentence also contains the semantic information of the sentence. To obtain deeper semantic information, the present invention uses a graph neural network to capture the context information of the sentence. Specifically, given a dialogue, based on the reply relationship between the sentences in the dialogue, a dialogue graph G = (V, E, W, T) is constructed. Where V represents the vertex set of the dialogue graph, and each sentence vector in the dialogue serves as each vertex in the dialogue graph E represents the edge set of the dialogue graph. If there is a reply relationship between two sentences in the dialogue, then there is an edge e between the corresponding two vertices in the dialogue graph ij ; W represents the weight of the edge in the dialogue graph. Based on the semantic similarity of adjacent vertices and , the weight w ij of its edge e ij is calculated, is the j-th sentence vector The corresponding vertex; T represents the type of edge in the dialogue graph. For different roles in the dialogue, four different types of edges are considered. After obtaining the dialogue graph G, the present invention uses a two-layer graph neural network to learn the context information of the sentences. The first layer is a basic GNN, and its calculation is as follows:

[0058]

[0059] Where is the vertex vector output by the first-layer GNN network, W1 (1) and W2 (1) are learnable parameters, and N (*,i) represents the set of all vertices pointing to vertex i. The second layer is a relational graph convolutional network (RGCN), and its calculation is as follows:

[0060]

[0061] Where is the vertex vector output by the second-layer RGCN network, and N t (*,i) is the set of all vertices pointing to vertex i under the relationship t, and c i,t is a regularization constant that can be specified in advance. Concatenate the sentence vector and the sentence vector containing context information to obtain the final sentence vector

[0062] Perform Sum pooling and Max pooling on all the obtained sentence vectors to obtain the dialogue-level vector representation:

[0063]

[0064] Finally, input the dialogue-level vector representation into a fully connected classification layer to obtain the prediction of the model for the dialogue category. The present invention uses the Focal Loss function to optimize the model:

[0065]

[0066] Where y k represents the true label, P k represents the predicted probability, α k and γ are adjustable parameters, and k takes the value of 0 or 1. P0 represents the probability of predicting a non-defective dialogue, P1 represents the probability of predicting a defective dialogue, α0 represents the weight when the true label is not the predicted dialogue, and α1 represents the weight when the true label is the predicted dialogue.

[0067] In the model prediction stage, the conversations with unknown label categories are input. Through the calculations of the above three-layer model (sentence encoding layer, graph-based context encoding layer, conversation encoding and classification layer), the probabilities that the conversation belongs to and does not belong to defective conversations are obtained respectively. If the probability that the conversation belongs to a defective conversation is greater than the probability that it does not belong to a defective conversation, then the conversation is predicted as a defective conversation, and vice versa.

[0068] Step 4, sentence preprocessing, involves splitting the user's speech in conversations containing software error information and filtering the user conversations judged to contain software errors using heuristic rules. The rule definitions in the heuristic rules are as follows:

[0069] 1) Delete all sentences that are not from the conversation initiator.

[0070] 2) Delete sentences that meet the conditions of having a conversation length less than five, a stop word ratio exceeding 50% and not containing the five feature tags of [URL], [CODE],

[0071] [EMAIL], [VERSION], [HTML].

[0072] 3) Delete the greetings in the sentences, such as "Hello", "Good afternoon", "Hello everyone", etc.

[0073] The conversation initiator mentioned above refers to the user who starts the topic in a conversation. Through the sentence preprocessing step, we obtain the processed conversation sentence dataset.

[0074] Step 5, sentence classification, mainly classifies the sentences in conversations containing software error information. The classified sentences can be regarded as four categories, namely Observed Behavior (OB), Expected Behavior (EB), Steps to Reproduce (SR) and Others.

[0075] The entire sentence classification step involves the training and application of an information extraction model.

[0076] The training of the information extraction model is as Figure 3 shown. First, collect the open-source dataset of software error reports on the Internet and perform some preprocessing on it. The steps involved are as follows:

[0077] 1) Split the software error reports into sentences.

[0078] 2) Convert the capital letters in the sentences of the software error reports to lowercase letters.

[0079] 3) Tokenize the sentences in the software error reports.

[0080] 4) Delete the non-English sentences in the software error reports.

[0081] 5) Delete overly long sentences in software error reports.

[0082] 6) Manually assign classification labels to sentences in software error reports.

[0083] An information extraction model was constructed by using a pre-trained BERT language model and connecting it to a fully connected layer, where the fully connected layer served as the final output layer. In the information extraction model, the parameters of the first 9 layers of the BERT pre-trained model were frozen, and the parameters of the remaining 3 layers were unfrozen to participate in the following two fine-tuning processes.

[0084] Sentences in the open-source dataset of publicly available software error reports were fed into the information extraction model for the first fine-tuning, and the model with the best performance was saved to obtain the first fine-tuned information extraction model.

[0085] Sentences spoken by the conversation initiator in conversations containing software error information were annotated, and the annotated sentences were filtered according to heuristic rules to obtain a training set of software error information sentences. The heuristic rules are as follows:

[0086] 1) Delete sentences with a length less than 5 that do not contain the five feature tags [URL], [CODE], [EMAIL], [VERSION], and [HTML].

[0087] 2) Delete greeting words or gratitude words in the sentences, such as "thank you", "good afternoon", "hello", etc.

[0088] During the second fine-tuning of the transfer learning of the information extraction model, the last output layer in the information extraction model after the previous first fine-tuning was replaced with an output layer that had not been trained at all, and its parameters were all random. Then, the processed software error information conversation sentence dataset obtained previously was input into the pre-trained information extraction model with the replaced output layer for retraining. Finally, the best information extraction model was saved. During the actual training process, the processed conversation sentence dataset was divided into two parts, one part was the training set and the other part was the test set. The training set was used to train the model, and the test set was used to test the performance of the model.

[0089] During the two training processes of the above information extraction model, the minimum value of the following formula was used as the training objective in the present invention:

[0090]

[0091] where Loss represents the loss during the model training process, P e 、P o 、P sr and P orespectively represent the probabilities that a sentence is judged as an EB label, an OB label, an SR label, and an Other label. And y e , y o , y sr and y o represent the true value situation of a sentence for this label.

[0092] After the sentences in the conversation are extracted by the information extraction model and divided into four different categories. Then, according to the categories of the sentences, clustering operations are performed, and the results are input into the software error report template, and finally a software error report is generated.

[0093] Although specific embodiments of the present invention are disclosed for illustrative purposes, which are intended to help understand the content of the present invention and implement it accordingly, those skilled in the art can understand that: without departing from the spirit and scope of the present invention and the appended claims, various substitutions, changes, and modifications are possible. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection required by the present invention shall be subject to the scope defined by the claims.

Claims

1. An automatic monitoring and synthesis method for defect reports for developer group chats, the steps of which include: 1) Collect online chat data, decouple the collected chat data, and perform data augmentation on the decoupled data to obtain a data-augmented dialogue decoupling dataset; 2) Feed the dialogue decoupling dataset into a dialogue classification model to classify the dialogues containing software error information and the dialogues not containing software error information; 3) Feed the dialogues containing software error information into a software error information extraction model to obtain the category to which each sentence in the dialogue belongs, and generate a software error report based on the sentence and the corresponding category; the categories include observed behavior, expected behavior, reproduction steps, and others.

2. The method according to claim 1, wherein The method for decoupling the collected chat data by a dialogue decoupling model based on a forward neural network is as follows: First, preprocess the collected chat data, filter out pictures and expressions, and convert website addresses, codes, email addresses, versions, and HTML elements into corresponding feature tags [URL], [CODE], [EMAIL], [VERSION], [HTML]; input the preprocessed chat data into the dialogue decoupling model based on a forward neural network to obtain a decoupled output; According to the decoupled output, reconstruct each group of dialogues in the chat information, where each group of dialogues contains several user speeches with the same theme.

3. The method according to claim 2, wherein The method for data augmentation of the decoupled data is as follows: For each user speech in each group of dialogues, if the length of the user speech is greater than the set threshold θ, replace the word w in the user speech with a synonym of the word w in the user speech to generate a new user speech; if the length of the user speech is not greater than the threshold θ, randomly extract a user speech with a length not greater than θ from the dialogue dataset to replace the user speech to generate a new user speech; combine the newly generated user speeches for the same group of dialogues into a new group of dialogues.

4. The method according to claim 1 or 2 or 3, characterized in that, The method for the dialogue classification model to classify the dialogue decoupling dataset is as follows: 1) Input each dialogue in the dialogue decoupling dataset into the dialogue classification model respectively; the dialogue classification model includes a sentence encoding layer and a classification layer; the sentence encoding layer includes a context encoding layer and a dialogue encoding layer; 2) The sentence encoding layer performs word encoding operations on the words of each sentence in the input dialogue to obtain word vectors corresponding to the words; 3) The context encoding layer generates a sentence vector representation corresponding to each sentence according to the word vectors of all the words in each sentence in the dialogue; 4) The dialogue encoding layer generates a context sentence vector representation corresponding to each sentence according to the sentence vector representations of the sentences in the dialogue; Then, splice the sentence vectors and context sentence vectors of each sentence in the dialogue, and perform sum pooling and max pooling processing on the splicing results in sequence to obtain a dialogue-level vector representation; among them, the dialogue encoding layer is a graph neural network, Construct each dialogue as a dialogue graph, each sentence in the dialogue is used as a node in the dialogue graph of the dialogue, and use the graph neural network to learn the context sentence vector representation of each sentence; 5) Input the dialogue-level vector representation into the classification layer to obtain the category of the dialogue.

5. The method according to claim 4, wherein Classify the decoupled dialogue dataset using the trained dialogue classification model; the method for training the dialogue classification model is as follows: the input is a dialogue D = {u1, u2, …, u n}, and the sentence encoding layer performs word embedding operations on all sub-words of each sentence u i in the dialogue D to obtain all word vectors of each sentence; the context encoding layer performs convolution operations on the word vectors of each sentence to obtain the sentence vector of the corresponding sentence where i = 1~n, and n is the total number of sentences in the dialogue D; then, based on the reply relationship between the sentences in the dialogue D, a dialogue graph G = (V, E, W, T) is constructed, where V represents the vertex set of the dialogue graph, and the sentence vector of the i-th sentence in the dialogue is used as the i-th vertex in the dialogue graph E represents the edge set of the dialogue graph. If there is a reply relationship between the i-th sentence and the j-th sentence in the dialogue D, then there is an edge e between the two corresponding vertices in the dialogue graph ij ; W represents the weight of the edge in the dialogue graph, and the weight w of its edge e ij is calculated based on the semantic similarity of the vertices ij ; T represents the type of the edge in the dialogue graph; input the dialogue graph G into the graph neural network to generate the context sentence vector representation of the corresponding sentence; concatenate the sentence vectors and context sentence vectors of each sentence in the dialogue D, and perform sum pooling and max pooling operations on the concatenated results in sequence to obtain the dialogue-level vector representation, input the dialogue-level vector representation into a fully connected classification layer to obtain the predicted category of the dialogue D; use the FocalLoss loss function to optimize the dialogue classification model.

6. The method according to claim 1 or 2 or 3, characterized in that, The method for constructing the software error information extraction model is as follows: 61) Obtain a software error report dataset, convert the sentences in the software error report dataset from uppercase to lowercase, perform word segmentation, delete non-English sentences, and delete sentences exceeding a set length, and then annotate the sentences; 62) Connect a BERT pre-trained language model to a fully connected layer, where the fully connected layer serves as the final output layer, thus constructing an information extraction model; fix the parameters of the first 9 layers of the BERT pre-trained language model, and the parameters of the remaining 3 layers are adjustable; 63) Feed the sentences annotated in step 61) into the information extraction model for the first fine-tuning to obtain the information extraction model after the first fine-tuning; 64) Annotate the sentences spoken by the conversation initiator in the conversation containing software error information, and filter the annotated sentences according to heuristic rules to obtain a software error information sentence training set; 65) Replace the fully connected layer of the information extraction model after the first fine-tuning with a fully connected layer with randomized parameters; Then use the software error information sentence training set to train the replaced information extraction model to obtain a software error information extraction model after two fine-tunings.

7. The method according to claim 6, wherein Each fine-tuning training of the software error message extraction model uses the minimum value as the training objective; Among them, P e , P o , P sr and P o respectively represent the probabilities that a sentence is judged as an EB label, an OB label, an SR label, and an Other label, and y e , y o , y sr and y o respectively represent the true label corresponding to a sentence; the EB label is the observed behavior The OB label is the expected behavior, the SR label is the reproduction steps, and the Other label is others.

8. The method according to claim 1, characterized in that The process of generating a software error report includes: clustering the sentences according to the category to which each sentence belongs; inputting the clustering result into a software error report template to generate a software error report.

9. An electronic device, characterized in that, It includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Software problem report classifying method based on text chaos degree

    CN107273295A

  • Software defect report dispatching method and device based on knowledge graph and semantic role labeling

    CN113138920A