Defect Duplicate Check Implementation Method, Device, Terminal Device and Storage Medium

Through key proprietary word discovery and topic matching, combined with the pre-constructed defect plagiarism check model, automated defect plagiarism check is achieved, solving the problems of low manual plagiarism check efficiency and difficulty in semantic recognition, and improving the efficiency and effectiveness of plagiarism check.

CN114969347BActive Publication Date: 2025-06-13CHINA MERCHANTS BANK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210738950.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2025-06-13
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

The existing defect plagiarism checking methods require manual participation, are inefficient, and cannot perform semantic recognition, especially in the case of short text, which has poor results.

Method used

By obtaining the defect text summary, performing key proprietary word discovery calculations, topic matching and sentence pair combinations are performed based on these keywords, and a pre-constructed defect plagiarism check model is used for plagiarism checking.

Benefits of technology

This method can automate defect plagiarism checking, save manual time, and improve plagiarism checking efficiency and effectiveness through semantic understanding and information refinement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114969347B_ABST
    Figure CN114969347B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, terminal device and storage medium for realizing defect duplicate checking. The method includes: obtaining a defect duplicate checking task, where the defect duplicate checking task includes: a defect text summary to be checked; performing key proper word discovery calculation on the defect text summary to obtain a key proper word calculation result; performing topic matching based on the key proper word calculation result, combining sentence pairs with the matched topic to obtain combined sentence pairs; and based on a pre-constructed defect duplicate checking model, performing duplicate checking evaluation on the sentence pairs to obtain a defect duplicate checking evaluation result. Thus, defect duplicate checking is carried out through a model and an algorithm, which can save the time of manual duplicate checking; moreover, this solution extracts information from the defect text and trains the model, which can refine semantic information from short texts and perform effective duplicate checking, improving the efficiency and effectiveness of defect duplicate checking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of testing, and particularly to a method, device, terminal device and storage medium for realizing duplicate defect checking. Background Art

[0002] With the continuous increase in software complexity, scale and iteration speed, the investment in software testing work is constantly increasing, and the surge in software test cases has also increased the defect management work. For each test task, numerous test defects will be generated. To avoid duplicate defects, the primary task is to check for duplicate defects among these defects. The previous duplicate defect checking has faced the following problems:

[0003] (1) Manual duplicate defect checking is required:

[0004] A common way of duplicate defect checking is to identify by manual. However, for large-scale test tasks, the number of defects is huge. When manually identifying defects, it is easy to forget or miss multiple defects, and it requires repeated viewing and certain experience, consuming a lot of manpower and material resources;

[0005] (2) Semantic recognition cannot be performed:

[0006] For common text matching methods, the main approach is to directly segment the text and then retrieve the text in the database. This method requires the text words to be the same or similar. It cannot recognize texts with different words but the same meaning. Since defects are generally written by different testers, there may be significant differences in grammar expression and word usage. Therefore, it is difficult to check for duplicate defects using this text matching method;

[0007] (3) The defect text is short and the information is scarce:

[0008] Generally, the description of a defect is divided into a defect summary and a defect description. The defect summary is short while the defect description is detailed. If the duplicate defect checking is performed after the defect description is completed, it will waste the time of testers. Therefore, for the duplicate defect checking task, it is generally required to give a prompt when the defect summary is completed. This makes the duplicate defect checking be carried out under the condition of short text and scarce information. Traditional text matching has poor effects in short texts and cannot extract information from short texts for duplicate defect checking. Summary of the Invention

[0009] The main purpose of the embodiments of the present invention is to provide a method, device, terminal device and storage medium for realizing duplicate defect checking, aiming to improve the efficiency and effectiveness of duplicate defect checking.

[0010] To achieve the above object, an embodiment of the present invention provides a method for realizing duplicate defect checking, and the method includes the following steps:

[0011] Obtain a defect duplicate check task, where the defect duplicate check task includes: a defect text summary to be checked for duplicates;

[0012] Perform key proper noun discovery calculation on the defect text summary to obtain a key proper noun calculation result;

[0013] Based on the key proper noun calculation result, perform theme matching, and combine sentence pairs with the matched theme to obtain combined sentence pairs;

[0014] Based on a pre - constructed defect duplicate check model, perform duplicate check evaluation on the sentence pairs to obtain a defect duplicate check evaluation result.

[0015] Optionally, the step of performing theme matching based on the key proper noun calculation result, and combining sentence pairs with the matched theme to obtain combined sentence pairs includes:

[0016] Determine the theme of the key proper nouns in the key proper noun calculation result;

[0017] Match the theme of the key proper nouns in the key proper noun calculation result with the key proper noun classification themes of the pre - stored platform full - volume defect texts;

[0018] Combine sentence pairs with the matched theme to obtain combined sentence pairs.

[0019] Optionally, after the step of combining sentence pairs with the matched theme, the following steps are further included:

[0020] Perform data cleaning on the sentence pairs to obtain cleaned sentence pairs.

[0021] Optionally, the step of performing duplicate check evaluation on the sentence pairs based on a pre - constructed defect duplicate check model to obtain a defect duplicate check evaluation result includes:

[0022] Duplicate the sentence pairs to obtain two copies of the sentence pairs;

[0023] Vectorize one copy of the sentence pairs using a pre - trained weighted word vector model to obtain a weighted vectorization result;

[0024] Input the other copy of the sentence pairs into the pre - trained defect duplicate check model, and through the defect duplicate check model and in combination with the weighted vectorization result, perform duplicate check evaluation on the sentence pairs to obtain a defect duplicate check evaluation result.

[0025] Optionally, before the step of performing key proper noun discovery calculation on the defect text summary to obtain a key proper noun calculation result, the following steps are further included:

[0026] Preprocess the defective text abstract, and the preprocessing methods include one or more of data augmentation and data cleaning.

[0027] Optionally, the step of performing key proper noun discovery calculation on the defective text abstract to obtain the key proper noun calculation result includes:

[0028] Perform new word discovery calculation on the defective text abstract using the left and right information entropy new word discovery algorithm, and screen out the proper nouns in the defective text abstract;

[0029] Calculate the keywords in the defective text abstract using the TFIDF algorithm;

[0030] Based on the proper nouns and keywords, construct a proper keyword table to obtain the key proper noun calculation result.

[0031] Optionally, before the step of performing duplicate check evaluation on the sentence pair based on a pre-trained defective duplicate check model to obtain the defective duplicate check evaluation result, it further includes:

[0032] Construct the defective duplicate check model, specifically including:

[0033] Obtain a defective text data training set, and the training set includes original defective abstract text data;

[0034] Perform key proper noun screening on the original defective abstract text data in the training set, and construct a proper keyword table for the training set according to the screening results;

[0035] Based on the proper keyword table of the training set and a pre-trained text vectorization model, perform weighted vectorization on the defective abstract text data in the training set to obtain defective text data word vectors;

[0036] Based on the defective text data word vectors and the original defective abstract text data, perform model training and fusion to construct the defective duplicate check model.

[0037] Optionally, the step of performing key proper noun screening on the original defective abstract text data in the training set and constructing a proper keyword table for the training set according to the screening results includes:

[0038] Perform new word discovery calculation on the original defective abstract text data in the training set using the left and right information entropy new word discovery algorithm, and screen out the proper nouns in the original defective abstract text data;

[0039] Calculate the keywords in the original defective abstract text data using the TFIDF algorithm;

[0040] Construct a list of specific keywords for the training set based on the proper nouns and keywords in the original defect summary text data.

[0041] Optionally, the steps of training and fusing the model based on the defect text data word vectors and the original defect summary text data to construct the defect duplicate checking model include:

[0042] Input the defect text data word vectors into a pre-created bidirectional LSTM model based on the attention mechanism for training to obtain a first training result;

[0043] Input the original defect summary text data into a pre-selected AlBert pre-training model for training to obtain a second training result;

[0044] Fuse and iteratively train the first training result and the second training result through the XGBoost algorithm to obtain the defect duplicate checking model.

[0045] Optionally, before the step of screening key proper nouns from the original defect summary text data in the training set, it further includes:

[0046] Perform data preprocessing on the defect text data training set, specifically including:

[0047] Perform data augmentation on the defect text data training set to obtain an augmented training set;

[0048] Use common stop words to clean the original defect summary text data in the training set, removing useless and interfering information to obtain a cleaned training set.

[0049] The present invention also proposes a defect duplicate checking implementation device, including:

[0050] An acquisition module for acquiring a defect duplicate checking task, where the defect duplicate checking task includes: a defect text summary to be checked;

[0051] A calculation module for performing key proper word discovery calculation on the defect text summary to obtain a key proper word calculation result;

[0052] A combination module for performing topic matching based on the key proper word calculation result, combining sentence pairs with the matched topic to obtain combined sentence pairs;

[0053] A judgment module for performing duplicate checking judgment on the sentence pairs based on a pre-constructed defect duplicate checking model to obtain a defect duplicate checking judgment result.

[0054] The present invention also provides a terminal device, which includes a memory, a processor, and a defect duplicate checking implementation program stored on the memory and executable on the processor. When the defect duplicate checking implementation program is executed by the processor, the steps of the defect duplicate checking implementation method described above are implemented.

[0055] The present invention also provides a computer-readable storage medium, on which a defect duplicate checking implementation program is stored. When the defect duplicate checking implementation program is executed by a processor, the steps of the defect duplicate checking implementation method described above are implemented.

[0056] The defect duplicate checking implementation method, device, terminal device, and storage medium provided by the embodiments of the present invention obtain a defect duplicate checking task, where the defect duplicate checking task includes: a defect text abstract to be checked; perform key proper word discovery calculation on the defect text abstract to obtain a key proper word calculation result; perform theme matching based on the key proper word calculation result, combine sentence pairs with the matched theme to obtain combined sentence pairs; and perform duplicate checking evaluation on the sentence pairs based on a pre-constructed defect duplicate checking model to obtain a defect duplicate checking evaluation result. Thus, defect duplicate checking can be performed through a model and an algorithm, saving the time of manual duplicate checking; and this solution extracts information from the defect text and trains the model, which can refine semantic information from short texts and perform effective duplicate checking, thereby improving the efficiency and effectiveness of defect duplicate checking. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a schematic diagram of the functional modules of the terminal device to which the defect duplicate checking implementation device of the present invention belongs;

[0058] Figure 2 It is a schematic flowchart of the first embodiment of the defect duplicate checking implementation method of the present invention;

[0059] Figure 3 It is a schematic diagram of the full process of defect duplicate checking in the embodiments of the present invention;

[0060] Figure 4 It is a schematic flowchart of the second embodiment of the defect duplicate checking implementation method of the present invention;

[0061] Figure 5 It is a schematic flowchart of the detailed process of constructing a duplicate checking model in the embodiments of the present invention.

[0062] Figure 6 It is a schematic diagram of the principle of processing text data when constructing a duplicate checking model in the embodiments of the present invention;

[0063] Figure 7 It is a schematic diagram of the full process of constructing a duplicate checking model in the embodiments of the present invention.

[0064] The realization, functional features, and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments

[0065] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0066] The main solution of the embodiment of the present invention is as follows: by obtaining a defect duplicate checking task, the defect duplicate checking task includes: a defect text summary to be checked; performing key proper word discovery calculation on the defect text summary to obtain a key proper word calculation result; performing topic matching based on the key proper word calculation result, combining sentence pairs with the matched topic to obtain combined sentence pairs; and based on a pre-constructed defect duplicate checking model, performing duplicate checking evaluation on the sentence pairs to obtain a defect duplicate checking evaluation result. Thus, defect duplicate checking can be performed through a model and an algorithm, which can save the time of manual duplicate checking; moreover, this solution extracts information from the defect text and trains the model, which can refine semantic information from short texts and perform effective duplicate checking, thereby improving the efficiency and effectiveness of defect duplicate checking.

[0067] Technical terms involved in the embodiment of the present invention:

[0068] Left and right information entropy new word discovery algorithm;

[0069] TFIDF algorithm;

[0070] Word2vec;

[0071] Attention mechanism, Attention Mechanism;

[0072] LSTM model, Long Short-Term Memory network, Long Short-Term Memory;

[0073] AlBert model, deep language model;

[0074] XGBoost, an optimized distributed gradient boosting library.

[0075] The specific explanations are as follows:

[0076] Left and right information entropy new word discovery algorithm: The purpose of the new word discovery algorithm is to discover new words. If the current word segmentation technology is used, sometimes rare words or proper nouns are often mis-segmented, and the improvement measure is to use the new word algorithm to discover new words in the text, and then put the discovered new words into the user-defined dictionary of the word segmentation algorithm, which will increase the accuracy of word segmentation. The following two concepts need to be explained:

[0077] Pointwise Mutual Information - Degree of Cohesion: For example, the formula for pointwise mutual information is: $$\operatorname{PMI}(x,y)=\log_{2}\frac{p(x,y)}{p(x)p(y)}$$

[0078] where $p(x,y)$ represents the probability that two words appear together, and $p(x)$ and $p(y)$ represent the probabilities of each word appearing. For example, in a corpus, the word "deep learning" appears 10 times, "deep" appears 15 times, and "learning" appears 20 times. Since the total number of words in the corpus is a fixed value, the pointwise mutual information of the word "deep learning" with respect to "deep" and "learning" is $\log_{2}\frac{10N}{15\times20}$, where $N$ refers to the total number of words.

[0079] From the above formula, it can be seen that the larger the pointwise mutual information, the more frequently these two words appear together, indicating a greater degree of cohesion between the two words and a greater possibility of them forming a new word.

[0080] Left (Right) Entropy - Degree of Freedom: The formula for left (right) entropy is as follows, which is the formula for information entropy: $$E_{left}(PreW)=-\sum_{\forall Pre\subseteq A}P(PreW)\log_{2}P(PreW)$$

[0081] In summary, the larger the left (right) entropy value, the richer the surrounding words of the word, indicating a greater degree of freedom of the word and a greater possibility of it becoming an independent word.

[0082] TF-IDF Algorithm: TF-IDF is a statistical method used to evaluate the importance of a word or phrase for a document set or a single document in a corpus. The importance of a word increases proportionally with the number of times it appears in a document, but decreases inversely with its frequency of appearance in the corpus. Various forms of TF-IDF weighting are often used by search engines as a measure or rating of the relevance between a document and a user query.

[0083] The main idea of TF-IDF is that if a word or phrase has a high term frequency (TF) in an article and rarely appears in other articles, it is considered to have good category discrimination ability and is suitable for classification.

[0084] TFIDF is actually: TF * IDF, where TF is the Term Frequency and IDF is the Inverse Document Frequency. TF represents the frequency of a term appearing in document d.

[0085] The main idea of IDF is that if the number of documents containing term t is smaller, that is, n is smaller, and the IDF is larger, it indicates that term t has good class discrimination ability. If the number of documents containing term t in a certain class of documents C is m, and the total number of documents containing t in other classes is k, obviously the total number of documents containing t, n = m + k. When m is large, n is also large, and the IDF value obtained according to the IDF formula will be small, indicating that the class discrimination ability of term t is not strong. However, in fact, if a term appears frequently in the documents of a class, it means that the term can well represent the characteristics of the text of this class. Such terms should be given higher weights and selected as the feature words of this class of text to distinguish them from the documents of other classes.

[0086] In a given document, the term frequency (TF) refers to the frequency of a given word appearing in the document. This number is a normalization of the term count to prevent it from biasing towards long documents. (The same word may have a higher term count in a long document than in a short document, regardless of the importance of the word.)

[0087] Word2vec: It originates from NLP (Natural Language Processing). In NLP, the finest-grained unit is the word. Words form sentences, and sentences form paragraphs, chapters, and documents. Therefore, when dealing with NLP problems, words are considered first. For example, to determine the part of speech of a word, whether it is a verb or a noun, using the machine learning approach, there is a series of samples (x, y), where x is the word and y is its part of speech. We need to construct a mapping of f(x) -> y. However, the mathematical model f here (such as neural networks, SVM) only accepts numerical inputs, while the words in NLP are human abstractions and are in symbolic forms (such as Chinese, English, Latin, etc.). Therefore, they need to be converted into numerical forms, or rather - embedded into a mathematical space. This embedding method is called word embedding, and Word2vec is one type of word embedding. In NLP, for f(x) -> y, if x is regarded as a word in a sentence and y is the context word of this word, then f here is the 'language model' that often appears in NLP. The purpose of this model is to determine whether the sample (x, y) conforms to the laws of natural language.

[0088] Word2vec exactly comes from this idea. However, its ultimate goal is not to train f to be perfect. It only cares about the by-products after the model training - the model parameters (specifically referring to the weights of the neural network), and uses these parameters as a certain vectorized representation of the input x. This vector is called the word vector.

[0089] Attention mechanism: The attention to the distribution of input weights was first used in the encoder-decoder. The attention mechanism obtains the input variable of the next layer by taking a weighted average of the hidden states of all time steps of the encoder.

[0090] LSTM model: Long Short-Term Memory (LSTM). Due to its unique design structure, LSTM is suitable for processing and predicting important events with very long intervals and delays in time series.

[0091] LSTM usually performs better than time-recurrent neural networks and Hidden Markov Models (HMMs), such as in continuous unsegmented handwritten recognition. LSTM is also widely used in automatic speech recognition. As a non-linear model, LSTM can be used as a complex non-linear unit to construct larger deep neural networks.

[0092] To minimize the training error, the Gradient Descent method, such as applying the Backpropagation Through Time algorithm, can be used to modify the weights each time according to the error. The main problem of the Gradient Descent method in Recurrent Neural Networks (RNNs) was first discovered in 1991, which is that the error gradient exponentially disappears with the time length between events. When an LSTM block is set, the error also propagates back during the backpropagation calculation, affecting each gate from the output back to the input stage until this value is filtered out. Therefore, the normal backpropagation neural network is an effective method for training the LSTM block to remember long-term values.

[0093] AlBert model, a deep language model, is based on Bert. With the popularity of the Transformer structure, pre-trained models with large corpora and large numbers of parameters have become the mainstream. When actually deploying models such as BERT, it is often necessary to use techniques such as distillation, compression, or other optimization techniques to process the model. The AlBert model achieves better results with fewer parameters. It has achieved state-of-the-art performance on major benchmarks with a 30% reduction in parameters. There are Chinese pre-trained models of different versions of AlBert, including TensorFlow, PyTorch, and Keras.

[0094] XGBoost: The full name of XGBoost is eXtreme Gradient Boosting. XGBoost is an optimized distributed gradient boosting library designed to be efficient, flexible, and portable. It implements machine learning algorithms under the Gradient Boosting framework. XGBoost provides parallel tree boosting (also known as GBDT, GBM), which can quickly and accurately solve many data science problems. The same code runs on major distributed environments (Hadoop, SGE, MPI) and can handle problems beyond billions of examples.

[0095] It is an optimized distributed gradient boosting library designed to be efficient, flexible, and portable. XGBoost is a tool for large-scale parallel boosting trees and is currently the fastest and best open-source boosting tree toolkit. In terms of large-scale industrial data, the distributed version of XGBoost has wide portability and supports running on various distributed environments such as Kubernetes, Hadoop, SGE, MPI, Dask, etc., enabling it to well solve the problems of large-scale industrial data.

[0096] The present invention takes into account that currently, defect duplicate checking requires manual duplicate checking, which has low efficiency and is time-consuming and laborious; moreover, in the current text matching-based duplicate checking, semantic matching cannot be performed, and information extraction cannot be carried out for short text matching, resulting in poor duplicate checking results.

[0097] The present invention provides a solution that can improve the efficiency and effectiveness of defect duplicate checking.

[0098] Specifically, referring to Figure 1 , Figure 1 is a schematic diagram of the functional modules of the terminal device to which the defect duplicate checking implementation device of the present invention belongs. The defect duplicate checking implementation device can be a device independent of the terminal device and capable of data processing, which can be carried on the terminal device in the form of hardware or software. The terminal device can be an intelligent mobile terminal with data processing functions such as a mobile phone or a tablet computer, or a fixed terminal device or a server with data processing functions, etc.

[0099] In this embodiment, the terminal device to which the defect duplicate checking implementation device belongs at least includes an output module 110, a processor 120, a memory 130, and a communication module 140.

[0100] The memory 130 stores an operating system and a defect duplicate checking implementation program; the output module 110 can be a display screen, etc. The communication module 140 can include a WIFI module, a mobile communication module, and a Bluetooth module, etc., and communicates with external devices or servers through the communication module 140.

[0101] Among them, when the defect duplicate checking implementation program in the memory 130 is executed by the processor, the following steps are implemented:

[0102] Obtain a defect duplicate checking task, where the defect duplicate checking task includes: a defect text summary to be checked;

[0103] Perform key proper noun discovery calculation on the defect text summary to obtain a key proper noun calculation result;

[0104] Based on the key proper noun calculation result, perform topic matching, and combine sentence pairs with the matching topic to obtain combined sentence pairs;

[0105] Based on a pre-constructed defect duplicate checking model, perform duplicate checking judgment on the sentence pairs to obtain a defect duplicate checking judgment result.

[0106] Furthermore, when the defect duplicate checking implementation program in the memory 130 is executed by the processor, the following steps are implemented:

[0107] Determine the topic of the key proper nouns in the key proper noun calculation result;

[0108] Match the topic of the key proper nouns in the key proper noun calculation result with the key proper noun classification topics of the pre-stored platform full-volume defect texts;

[0109] Combine sentence pairs with the matching topic to obtain combined sentence pairs.

[0110] Furthermore, when the defect duplicate checking implementation program in the memory 130 is executed by the processor, the following steps are implemented:

[0111] Perform data cleaning on the sentence pairs to obtain cleaned sentence pairs.

[0112] Furthermore, when the defect duplicate checking implementation program in the memory 130 is executed by the processor, the following steps are implemented:

[0113] Duplicate the sentence pairs to obtain two copies of sentence pairs;

[0114] Vectorize one copy of the sentence pairs using a pre-trained weighted word vector model to obtain a weighted vectorization result;

[0115] Input the other copy of the sentence pairs into a pre-trained defect duplicate checking model, and through the defect duplicate checking model and in combination with the weighted vectorization result, perform duplicate checking judgment on the sentence pairs to obtain a defect duplicate checking judgment result.

[0116] Furthermore, when the defect duplicate checking implementation program in the memory 130 is executed by the processor, the following steps are implemented:

[0117] Preprocess the defective text abstract, and the preprocessing methods include one or more of data augmentation and data cleaning.

[0118] Furthermore, when the defective duplicate checking implementation program in the memory 130 is executed by the processor, the following steps are implemented:

[0119] Use the left and right information entropy new word discovery algorithm to perform new word discovery calculation on the defective text abstract, and screen out the proper nouns in the defective text abstract;

[0120] Use the TFIDF algorithm to calculate the keywords in the defective text abstract;

[0121] Based on the proper nouns and keywords, construct a proper keyword table to obtain the calculation result of the key proper words.

[0122] Furthermore, when the defective duplicate checking implementation program in the memory 130 is executed by the processor, the following steps are implemented:

[0123] Construct the defective duplicate checking model, which specifically includes:

[0124] Obtain a defective text data training set, and the training set includes original defective abstract text data;

[0125] Perform key proper noun screening on the original defective abstract text data in the training set, and construct a proper keyword table of the training set according to the screening results;

[0126] Based on the proper keyword table of the training set and a pre-trained text vectorization model, perform weighted vectorization on the defective abstract text data of the training set to obtain defective text data word vectors;

[0127] Based on the defective text data word vectors and the original defective abstract text data, perform model training and fusion to construct the defective duplicate checking model.

[0128] Furthermore, when the defective duplicate checking implementation program in the memory 130 is executed by the processor, the following steps are implemented:

[0129] Use the left and right information entropy new word discovery algorithm to perform new word discovery calculation on the original defective abstract text data in the training set, and screen out the proper nouns in the original defective abstract text data;

[0130] Use the TFIDF algorithm to calculate the keywords in the original defective abstract text data;

[0131] Based on the proper nouns and keywords in the original defective abstract text data, construct the proper keyword table of the training set.

[0132] Further, when the defect duplicate check implementation program in the memory 130 is executed by the processor, the following steps are implemented:

[0133] Input the word vectors of the defect text data into a pre-created bidirectional LSTM model based on the attention mechanism for training to obtain a first training result;

[0134] Input the original defect summary text data into a pre-selected AlBert pre-training model for training to obtain a second training result;

[0135] Fuse and iteratively train the first training result and the second training result through the XGBoost algorithm to obtain the defect duplicate check model.

[0136] Further, when the defect duplicate check implementation program in the memory 130 is executed by the processor, the following steps are implemented:

[0137] Perform data preprocessing on the defect text data training set, specifically including:

[0138] Perform data augmentation on the defect text data training set to obtain an augmented training set;

[0139] Use common stop words to clean the original defect summary text data in the training set to remove useless and interfering information, obtaining a cleaned training set.

[0140] In this embodiment, through the above solution, a defect duplicate check task is obtained. The defect duplicate check task includes: a defect text summary to be checked; performing key proper word discovery calculation on the defect text summary to obtain a key proper word calculation result; performing topic matching based on the key proper word calculation result, combining sentence pairs with the matched topic to obtain combined sentence pairs; and based on a pre-constructed defect duplicate check model, performing duplicate check judgment on the sentence pairs to obtain a defect duplicate check judgment result. Thus, defect duplicate check can be performed through the model and algorithm, saving the time of manual duplicate check; and this solution extracts information from the defect text and trains the model, which can refine semantic information from short texts and perform effective duplicate check, thereby improving the efficiency and effectiveness of defect duplicate check.

[0141] Based on the above terminal device architecture but not limited to the above architecture, an embodiment of the method of the present invention is proposed.

[0142] Refer to Figure 2 , Figure 2 is a schematic flowchart of the first embodiment of the defect duplicate check implementation method of the present invention. The defect duplicate check implementation method includes:

[0143] Step S101, obtain a defect duplicate check task, where the defect duplicate check task includes: a defect text summary to be checked;

[0144] The execution subject of the method in this embodiment can be a defect duplicate checking implementation device, or a defect duplicate checking implementation terminal device or server. This embodiment takes the defect duplicate checking implementation device as an example, and this device can be integrated on terminal devices such as smart phones and tablet computers with data processing functions.

[0145] The solution in this embodiment mainly realizes duplicate checking of test defects.

[0146] This solution abstracts the defect duplicate checking task into a text classification task, that is, pairs the submitted defect text summaries two by two as sentence pairs, and stipulates that the sentence pair is the same defect as 1, and the sentence pair is a different defect as 0. The defect duplicate checking task is transformed into a binary text classification task of using a model to judge all sentence pairs, outputting 1 for the same defect and 0 for different defects.

[0147] Specifically, first, obtain a defect duplicate checking task, and the defect duplicate checking task includes: the defect text summary to be currently checked for duplicates.

[0148] Among them, as an implementation manner, the defect text summary to be checked for duplicates can be input by the user on the system defect registration platform, and the user starts defect duplicate checking after inputting the defect text summary on the system defect registration platform.

[0149] Among them, as another implementation manner, the system can also automatically obtain the defect text summary to be checked for duplicates from external devices or other network devices according to configuration rules, and thus start defect duplicate checking.

[0150] Further, as an implementation manner, after obtaining the defect text summary, the defect text summary can be preprocessed, and the preprocessing methods include: one or more of data augmentation and data cleaning.

[0151] Among them, data augmentation can adopt the following scheme:

[0152] Perform synonym replacement on some words in the text and replace them with other texts with the same meaning for data increment; or use a translation software to translate it into an intermediate language and then translate it back to obtain different text expressions with the same meaning for data increment.

[0153] Data cleaning can use common stop words to clean the data in this article and remove useless and interfering information.

[0154] Step S102, perform key proper word discovery calculation on the defect text summary to obtain a key proper word calculation result;

[0155] Among them, the conventional process needs to compare the input text with all the defective texts in the database. In order to improve performance and reduce the amount of comparison data, in this embodiment, key proper word discovery calculations have been pre-performed in the storage of all platform defective texts and stored by theme according to the calculation results.

[0156] After obtaining the defective text summary to be checked for duplication in this embodiment, key proper word discovery calculations are performed on the defective text summary to obtain key proper word calculation results, and sentence pair combination and data cleaning are performed according to the matching theme.

[0157] Specifically, as an implementation manner, the step of performing key proper word discovery calculations on the defective text summary to obtain key proper word calculation results may include:

[0158] Adopt the left and right information entropy new word discovery algorithm to perform new word discovery calculations on the defective text summary, and screen out the proper nouns in the defective text summary;

[0159] Use the TFIDF algorithm to calculate the keywords in the defective text summary;

[0160] Based on the proper nouns and keywords, construct a proper keyword table to obtain key proper word calculation results.

[0161] Step S103, perform theme matching based on the key proper word calculation results, and perform sentence pair combination according to the matching theme to obtain the combined sentence pairs;

[0162] Specifically, as an implementation manner, first determine the theme of the key proper words in the key proper word calculation results;

[0163] Then, match the theme of the key proper words in the key proper word calculation results with the key proper word classification themes of the pre-stored platform all defective texts;

[0164] Finally, perform sentence pair combination according to the matching theme to obtain the combined sentence pairs.

[0165] Among them, after the step of performing sentence pair combination according to the matching theme, the sentence pairs can be further data-cleaned to obtain the cleaned sentence pairs, and data cleaning can improve the accuracy of data processing.

[0166] Step S104, based on the pre-constructed defective text duplication checking model, perform duplication checking and evaluation on the sentence pairs to obtain defective text duplication checking and evaluation results.

[0167] This embodiment pre-constructs a defective text duplication checking model, which is constructed through training, model fusion, and iterative calculation based on a pre-collected defective text data training set.

[0168] Specifically, as an implementation manner, the step of performing duplicate check judgment on the sentence pair based on the pre-constructed defect duplicate check model to obtain a defect duplicate check judgment result may include:

[0169] First, duplicate the sentence pair to obtain two copies of the sentence pair;

[0170] Vectorize one copy of the sentence pair using a pre-trained weighted word vector model to obtain a weighted vectorization result;

[0171] Input the other copy of the sentence pair into the pre-trained defect duplicate check model, and perform duplicate check judgment on the sentence pair through the defect duplicate check model in combination with the weighted vectorization result to obtain a defect duplicate check judgment result.

[0172] Specifically, after obtaining the sentence pair to be compared, duplicate it into two copies of data. Vectorize one copy using a pre-trained weighted word vector model, and input the other copy in the original format into the pre-trained defect duplicate check model for calculation and judgment. Among them, during the judgment, perform duplicate check judgment on the sentence pair through the defect duplicate check model in combination with the weighted vectorization result to obtain a defect duplicate check judgment result.

[0173] Obtain the defect texts determined to be the same through the model evaluation calculation result, and return the required information to the interface for display, so that the tester can choose whether to continue submitting the defect. Thus, the entire solution process is completed.

[0174] The entire process of performing defect duplicate check in the embodiment of the present invention can be referred to Figure 3 as shown.

[0175] Through the above solution in this embodiment, specifically by obtaining a defect duplicate check task, the defect duplicate check task includes: a defect text summary to be checked; performing key proper word discovery calculation on the defect text summary to obtain a key proper word calculation result; performing theme matching based on the key proper word calculation result, and combining sentence pairs with the matched theme to obtain a combined sentence pair; performing duplicate check judgment on the sentence pair based on a pre-constructed defect duplicate check model to obtain a defect duplicate check judgment result. Thus, defect duplicate check can be performed through the model and algorithm, which can save the time of manual duplicate check; moreover, this solution extracts information from the defect text and performs model training, which can refine semantic information from short texts and perform effective duplicate check, thereby improving the efficiency and effectiveness of defect duplicate check.

[0176] Refer to Figure 4 , Figure 4 which is a schematic flowchart of the second embodiment of the defect duplicate check implementation method of the present invention. As Figure 4 shown, in this embodiment, based on the above Figure 2Based on the embodiments shown, before the step of performing duplicate checking evaluation on the sentence pair based on a pre-trained defect duplicate checking model in step S104 to obtain a defect duplicate checking evaluation result, the following steps are further included:

[0177] Step S100, construct the defect duplicate checking model.

[0178] As Figure 5 shown, the above step S100 may specifically include:

[0179] Step S1001, obtain a defect text data training set, where the training set includes original defect summary text data;

[0180] Among them, the original defect summary text data is the defect summary text known in daily tests, and it constitutes the defect text data training set as sample data.

[0181] This embodiment takes into account that:

[0182] Based on the machine learning related solutions, the primary condition is to have a certain amount of original data for learning and training the model. The characteristics of the original accumulated defect text data are: in one test task, the same defects account for a very small number, and the data for training is unbalanced; the text is short, the grammar is more colloquial, and the information is not clear enough.

[0183] For the above problems, the data processing method proposed in this solution is:

[0184] Data augmentation: perform undersampling, oversampling, and data transformation on the training data set to construct a balanced data set for training;

[0185] Remove stop words and screen key proper nouns: remove useless and repeated words in the text, and screen key proper nouns through algorithms for subsequent weighting;

[0186] Word2Vec word vectorization: use the existing text data to train a word vector model, and add proper nouns for weighted word vectorization.

[0187] Among them, as an implementation method, after obtaining the original defect summary text data, the original defect summary text data can be preprocessed; or after obtaining the defect text data training set, data preprocessing can be performed on the defect text data training set. The specific processing process can be as Figure 6 shown, including:

[0188] Perform data augmentation on the defect text data training set to obtain a data-augmented training set;

[0189] Use common stop words to perform data cleaning on the original defect summary text data in the training set, remove useless and interfering information, and obtain a data-cleaned training set.

[0190] Specifically, the original defective summary text data used as sample data in the defective text data training set is the data statistically checked for duplicate content manually by test management, and is divided into two categories: non-duplicate defective data and duplicate defective data.

[0191] In actual test tasks, the amount of non-duplicate defective data is much larger than that of duplicate defective data. Unbalanced data will cause bias in the training of classification tasks based on the balance threshold.

[0192] Therefore, in this embodiment, undersampling is first performed on the non-duplicate defective data. Undersampling means randomly discarding the category of data with an excessive quantity to reduce the quantity gap between the two categories of data.

[0193] Then, oversampling is performed on the duplicate defective data. Oversampling means repeatedly obtaining the category of data with a smaller quantity to reduce the quantity gap between the two categories of data. However, directly repeatedly obtaining data easily leads to overfitting in training. Therefore, the oversampling used in this solution is to generate data by flipping and transferring sentence pairs based on the characteristics of sentence pairs. For example, for the sentence pair "AA@BB", after flipping, "BB@AA" is considered as newly added data (flipping); secondly, assuming sentence pairs "AA@BB" and "BB@CC", then "AA@CC" is also considered as newly added data (transfer generation).

[0194] In addition to the sampling method, data can also be incremented for the dataset with a smaller quantity through synonym replacement. Part of the vocabulary in the text is replaced with other texts with the same meaning for data increment. Or use a translation software to translate it into an intermediate language and then translate it back to obtain different text expressions with the same meaning for data increment.

[0195] Furthermore, if there are many colloquial expressions and useless information in the defective text data processed by this solution, then after data augmentation, common stop words can be used to clean the text data to remove useless and interfering information, thereby improving the accuracy of subsequent defective duplicate checking.

[0196] Step S1002: Screen the key proper nouns from the original defective summary text data in the training set, and construct a list of proper keywords for the training set according to the screening results;

[0197] After that, the left and right information entropy new word discovery algorithm can be used to calculate new word discovery for the text, screen out the proper nouns in the text data, such as some product names, professional terms, etc., and then use the TFIDF algorithm to calculate the keywords in the text data to construct a list of proper keywords for subsequent weighted calculation.

[0198] Specifically, as an implementation, the step of screening key proper nouns from the original defect summary text data in the training set and constructing a keyword list specific to the training set according to the screening results may include:

[0199] Use the left and right information entropy new word discovery algorithm to perform new word discovery calculation on the original defect summary text data in the training set, and screen out the proper nouns in the original defect summary text data;

[0200] Use the TFIDF algorithm to calculate the keywords in the original defect summary text data;

[0201] Based on the proper nouns and keywords in the original defect summary text data, construct a keyword list specific to the training set.

[0202] Step S1003: Based on the keyword list specific to the training set and a pre-trained text vectorization model, perform weighted vectorization on the defect summary text data in the training set to obtain defect text data word vectors;

[0203] For text model training, it is necessary to import the vectorized text into a deep learning network for calculation. The Word2Vec method adopted in this solution is used to train the text data after data augmentation to obtain a text vectorization model, and then the keyword list specific to the text data is weighted and vectorized according to the above steps to obtain defect text data word vectors.

[0204] Step S1004: Based on the defect text data word vectors and the original defect summary text data, perform model training and fusion to construct the defect duplicate checking model.

[0205] Specifically, first, input the defect text data word vectors into a pre-created bidirectional LSTM model based on the attention mechanism for training to obtain a first training result;

[0206] Then, input the original defect summary text data into a pre-selected AlBert pre-training model for training to obtain a second training result;

[0207] Finally, use the XGBoost algorithm to fuse and iteratively train the first training result and the second training result to obtain the defect duplicate checking model.

[0208] Specifically, the model training method adopted in this solution is to use two machine learning models for training respectively, and then use model fusion to combine the results of the two models to comprehensively obtain the actual judgment result for training. The specific process of model training in this solution can refer to Figure 7 as shown.

[0209] Among them, in model training, the first model structure adopted in this solution is a bidirectional LSTM model structure based on the attention mechanism, namely Bi-LSTM. Its advantage is that through the attention mechanism, it automatically focuses on and weights the important information of the text. The Bi-LSTM structure can gradually learn the main information in the text during training. The "gate" mechanism therein can learn the main information in the text and forget the useless information, while the bidirectional structure enables the model training to enhance text understanding through context learning. In model training, the weighted text data vector obtained in the data processing operation is imported into the Bi-LSTM model for training to obtain the result of determining whether the defects are the same.

[0210] Another model adopted in this solution is the lightweight pre-trained model AlBert. The AlBert pre-trained model adopted in this solution has been learned through a large amount of Chinese data to obtain a pre-trained model with strong versatility. Then, through fine-tuning using some defect texts pre-collected in this solution, the pre-trained model parameters are updated to make it applicable to the task scenario of this solution.

[0211] In model training, the defect text data after data augmentation is directly imported into the fine-tuned AlBert model for training to obtain the result of determining whether the defects are the same.

[0212] Among them, in the training of two different types of models in this solution, multiple models can be fused together through the XGBoost algorithm to improve performance.

[0213] Model fusion essentially gives a greater weight to the examples that were misclassified in the previous training. In subsequent training iterations, the probability of correctly classifying the originally misclassified samples is increased, and different weights are assigned according to different models. Finally, a weighted strong classifier is obtained.

[0214] Through the above steps, continuous iteration is carried out to train a complete classification model for subsequent use.

[0215] In this embodiment, through the above solution, a defect duplicate check model is constructed to obtain a defect duplicate check task. The defect duplicate check task includes: a defect text summary to be checked; calculating key proper words for the defect text summary to obtain a key proper word calculation result; performing topic matching based on the key proper word calculation result, and combining sentence pairs with the matched topic to obtain combined sentence pairs; and based on the pre-constructed defect duplicate check model, performing duplicate check judgment on the sentence pairs to obtain a defect duplicate check judgment result. Thus, defect duplicate check can be performed through the model and algorithm, saving the time of manual duplicate check; moreover, this solution extracts information from the defect text and conducts model training, which can refine semantic information from short texts and perform effective duplicate check, thereby improving the efficiency and effectiveness of defect duplicate check.

[0216] Compared with the prior art, the embodiments of the present invention adopt an algorithm to remove duplicates of defects, solving the cumbersome process of manual duplicate checking, being automatic and efficient; performing customized data augmentation on defect data to solve the problems of unbalanced training data and insufficient data volume; using machine learning semantic understanding for training to solve the problem that conventional methods cannot match at the semantic level; using a key proper noun discovery method to weight the text vectorization and performing topic classification according to the key proper nouns to reduce the number of comparisons; moreover, this solution adopts two different models, Bi-LSTM and AlBert, and performs model fusion through XGBoost, improving the effectiveness of defect duplicate checking.

[0217] In addition, the embodiments of the present invention also propose a defect duplicate checking implementation device, including:

[0218] An acquisition module, configured to acquire a defect duplicate checking task, where the defect duplicate checking task includes: a defect text summary to be checked for duplicates;

[0219] A calculation module, configured to perform key proper noun discovery calculation on the defect text summary to obtain a key proper noun calculation result;

[0220] A combination module, configured to perform topic matching based on the key proper noun calculation result, combine sentence pairs with the matched topics to obtain combined sentence pairs;

[0221] A judgment module, configured to perform duplicate checking judgment on the sentence pairs based on a pre-constructed defect duplicate checking model to obtain a defect duplicate checking judgment result.

[0222] For the implementation principle of defect duplicate checking of the present invention, please refer to the above embodiments and will not be elaborated here.

[0223] In addition, the embodiments of the present invention also propose a terminal device, where the terminal device includes a memory, a processor, and a defect duplicate checking implementation program stored on the memory and executable on the processor. When the defect duplicate checking implementation program is executed by the processor, the steps of the above-mentioned defect duplicate checking implementation method are implemented.

[0224] Since when the defect duplicate checking implementation program is executed by the processor, all the technical solutions of the foregoing all embodiments are adopted, it has at least all the beneficial effects brought by all the technical solutions of the foregoing all embodiments, which will not be elaborated one by one here.

[0225] In addition, the embodiments of the present invention also propose a computer-readable storage medium, on which a defect duplicate checking implementation program is stored. When the defect duplicate checking implementation program is executed by the processor, the steps of the above-mentioned defect duplicate checking implementation method are implemented.

[0226] Since all the technical solutions of the foregoing embodiments are adopted when the defect duplicate checking implementation program is executed by a processor, it has at least all the beneficial effects brought by all the technical solutions of the foregoing embodiments, which will not be elaborated herein one by one.

[0227] Compared with the prior art, a defect duplicate checking implementation method, system, terminal device and storage medium provided by the present invention adopt an algorithm to remove duplicates of defects, solve the cumbersome process of manual duplicate checking, and are automatic and efficient; customize data augmentation for defect data to solve the problems of unbalanced training data and insufficient data volume; adopt a machine learning semantic understanding method for training to solve the problem that conventional methods cannot perform semantic-level matching; adopt a key proper noun discovery method to weight text vectorization and perform topic classification according to key proper nouns to reduce the number of comparisons; furthermore, this solution adopts two different models, Bi-LSTM and AlBert, and performs model fusion through XGBoost, improving the effectiveness of defect duplicate checking.

[0228] It should be noted that in this article, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or method. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or method including that element.

[0229] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0230] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of the present invention.

[0231] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for defect duplicate checking implementation, characterized in that, the method comprises the following steps: Obtain a defect duplicate checking task, where the defect duplicate checking task includes: a defect text summary to be checked; Perform key proper noun discovery calculation on the defect text summary to obtain a key proper noun calculation result; Determine the theme of the key proper nouns in the key proper noun calculation result; Match the theme of the key proper nouns in the key proper noun calculation result with the key proper noun classification themes of the pre-stored platform full-volume defect texts; Perform sentence pair combination based on the matched theme to obtain the combined sentence pairs; Based on a pre-constructed defect duplicate checking model, perform duplicate checking evaluation on the sentence pairs to obtain a defect duplicate checking evaluation result, including: Duplicate the sentence pairs to obtain two copies of the sentence pairs; Vectorize one copy of the sentence pairs using a pre-trained weighted word vector model to obtain a weighted vectorization result; Input the other copy of the sentence pairs into the pre-trained defect duplicate checking model, and through the defect duplicate checking model and in combination with the weighted vectorization result, perform duplicate checking evaluation on the sentence pairs to obtain a defect duplicate checking evaluation result; Before the step of performing duplicate checking evaluation on the sentence pairs based on the pre-trained defect duplicate checking model to obtain a defect duplicate checking evaluation result, it further includes: Construct the defect duplicate checking model, specifically including: Obtain a defect text data training set, where the training set includes original defect summary text data; Screen the original defect summary text data in the training set, and construct a list of specific keywords for the training set according to the screening results; Based on the list of specific keywords of the training set and a pre-trained text vectorization model, perform weighted vectorization on the defect summary text data of the training set to obtain defect text data word vectors; Based on the defect text data word vectors and the original defect summary text data, perform model training and fusion to construct the defect duplicate checking model, specifically including: Input the defect text data word vectors into a pre-created bidirectional LSTM model based on the attention mechanism for training to obtain a first training result; Input the original defect summary text data into a pre-selected AlBert pre-training model for training to obtain a second training result; Fuse and iteratively train the first training result and the second training result through the XGBoost algorithm to obtain the defect duplicate checking model.

2. The method according to claim 1, characterized in that, after the step of performing sentence pair combination based on the matched theme, it further includes: Clean the data of the sentence pairs to obtain the cleaned sentence pairs.

3. The method according to claim 1, characterized in that, before the step of performing key proper noun discovery calculation on the defect text summary to obtain a key proper noun calculation result, it further includes: Preprocess the defect text summary, and the preprocessing methods include one or more of data augmentation and data cleaning.

4. The method according to claim 1, characterized in that, The step of performing key proper noun discovery calculation on the defective text summary to obtain the key proper noun calculation result includes: Performing new word discovery calculation on the defective text summary using the left - right information entropy new word discovery algorithm, and screening out the proper nouns in the defective text summary; Calculating the keywords in the defective text summary using the TFIDF algorithm; Based on the proper nouns and keywords, constructing a list of proprietary keywords to obtain the key proper noun calculation result.

5. The method according to claim 1, wherein, The step of performing key proper noun screening on the original defective summary text data in the training set and constructing the list of proprietary keywords for the training set according to the screening result includes: Performing new word discovery calculation on the original defective summary text data in the training set using the left - right information entropy new word discovery algorithm, and screening out the proper nouns in the original defective summary text data; Calculating the keywords in the original defective summary text data using the TFIDF algorithm; Based on the proper nouns and keywords in the original defective summary text data, constructing the list of proprietary keywords for the training set.

6. The method according to claim 1, wherein, Before the step of performing key proper noun screening on the original defective summary text data in the training set, it further includes: Performing data pre - processing on the defective text data training set, specifically including: Performing data augmentation on the defective text data training set to obtain an augmented training set; Performing data cleaning on the original defective summary text data in the training set using common stop words to remove useless and interfering information, obtaining a data - cleaned training set.

7. A device for realizing defective duplicate checking, wherein, it includes: An acquisition module, configured to acquire a defective duplicate - checking task, and the defective duplicate - checking task includes: a defective text summary to be checked; A calculation module, configured to perform key proper noun discovery calculation on the defective text summary to obtain a key proper noun calculation result; A combination module, configured to determine the theme of the key proper nouns in the key proper noun calculation result, match the theme of the key proper nouns in the key proper noun calculation result with the key proper noun classification theme of the full - volume defective texts of the pre - stored platform, and perform sentence - pair combination with the matched theme to obtain the combined sentence pair; A judgment module, configured to perform duplicate - checking judgment on the sentence pair based on a pre - constructed defective duplicate - checking model to obtain a defective duplicate - checking judgment result. Specifically, it is used to duplicate the sentence pair to obtain two copies of the sentence pair; vectorize one copy of the sentence pair using a pre - trained weighted word vector model to obtain a weighted vectorization result; input the other copy of the sentence pair into the pre - trained defective duplicate - checking model, and perform duplicate - checking judgment on the sentence pair through the defective duplicate - checking model in combination with the weighted vectorization result to obtain a defective duplicate - checking judgment result; A construction module, configured to construct the defective duplicate - checking model, specifically including: Obtaining a defective text data training set, and the training set includes original defective summary text data; Screen the key proper nouns from the original defect summary text data in the training set, and construct a keyword list for the training set according to the screening results; Based on the keyword list for the training set and a pre-trained text vectorization model, perform weighted vectorization on the defect summary text data in the training set to obtain defect text data word vectors; Based on the defect text data word vectors and the original defect summary text data, perform model training and fusion to construct the defect duplicate checking model, specifically including: Input the defect text data word vectors into a pre-created bidirectional LSTM model based on the attention mechanism for training to obtain a first training result; Input the original defect summary text data into a pre-selected AlBert pre-training model for training to obtain a second training result; Through the XGBoost algorithm, fuse and iteratively train the first training result and the second training result to obtain the defect duplicate checking model.

8. A terminal device Characterized in that The terminal device includes a memory, a processor, and a defect duplicate checking implementation program stored on the memory and executable on the processor. When the defect duplicate checking implementation program is executed by the processor, the steps of the defect duplicate checking implementation method according to any one of claims 1-6 are implemented.

9. A computer-readable storage medium Characterized in that A defect duplicate checking implementation program is stored on the computer-readable storage medium. When the defect duplicate checking implementation program is executed by a processor, the steps of the defect duplicate checking implementation method according to any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Article duplicate checking detection method, device and equipment, and storage medium

    CN110472203A