Text data enhancement method and device of FAQ system
By building a training set in the FAQ system to train the Simbert model to generate synonyms, and perform word-level processing and screening, the problem of difficulty in obtaining corpus in the FAQ system is solved and the Q&A effect is improved.
Patent Information
- Application Number
- CN202510335711.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-18
AI Technical Summary
The acquisition of corpus in the FAQ system is high and difficult, and the existing methods are time-consuming and limited in effect, resulting in poor Q&A effects.
The training set is constructed using some original question text and synonyms in the FAQ system's question-and-answer corpus, the Simbert model is trained to generate synonyms, and the synonyms are processed through word-level text enhancement methods, including removing stop words, keyword extraction, substitution of close pronunciations, etc., and finally filtering and adding them to the question-and-answer corpus.
It significantly increases the depth and breadth of the FAQ system's corpus, improves the system's ability to answer questions, solves the problem of difficulty in obtaining corpus, and improves the Q&A effect.
Smart Images

Figure CN120336457A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of question-and-answer systems, and in particular, to a method, apparatus, computer-readable storage medium, and text data enhancement system for text data enhancement of a FAQ system. Background Art
[0002] The question-and-answer process of a question-and-answer system for frequently asked questions usually uses standard questions as a bridge to connect users and answers. By building a database with some pre-organized corpora manually (usually frequently asked question pairs), it is published on the system to provide services for users. The FAQ system has advantages such as high quality and good organization, making the system have a relatively high level of answering questions. However, the FAQ system has the problems of high cost and difficulty in obtaining corpora, mainly in two aspects: 1. The FAQ system needs to collect question-and-answer pairs manually in the early stage, which is time-consuming and inevitable, and the quantity is small; 2. The question-and-answer pairs collected in the early stage of the FAQ system need to be processed to improve the question-and-answer effect. Traditional methods mainly process them manually or partially based on words, and the process is cumbersome and the effect is limited. Summary of the Invention
[0003] The main purpose of this application is to provide a method, apparatus, computer-readable storage medium, and text data enhancement system for text data enhancement of a FAQ system, so as to at least solve the problem of difficulty in obtaining the original question text in the prior art.
[0004] To achieve the above purpose, according to one aspect of this application, a method for text data enhancement of a FAQ system is provided, including: constructing a training set by using part of the original question text in the question-and-answer corpus of the current FAQ system and the synonymous sentences corresponding to the part of the original question text, and training an initial Simbert model by using the training set to obtain a trained Simbert model, where the Simbert model is used for generating synonymous sentences; generating synonymous sentences for all the original question texts by using the trained Simbert model to obtain multiple synonymous sentences corresponding to each of the original question texts; processing all the original question texts and the synonymous sentences corresponding to all the original question texts by using a text enhancement method at the word level to obtain a first target synonymous sentence corresponding to each of the original question texts, where the text enhancement method at the word level includes removing stop words, keyword extraction, near-sound word replacement, synonym replacement, random insertion of words, and random deletion of words; screening each of the first target synonymous sentences to obtain a second target synonymous sentence corresponding to each of the original question texts, where the second target synonymous sentence is a sentence with the same meaning as the original question text and correct grammar; adding all the second target synonymous sentences to the question-and-answer corpus to bind the second target synonymous sentences to the corresponding original question texts.
[0005] Optionally, a training set is constructed using a part of the original question text in the Q&A corpus of the current FAQ system and the synonymous sentences corresponding to the part of the original question text, including: using an Embedding model to convert the part of the original question text in the Q&A corpus of the current FAQ system into question vectors; in the first step, using an ANN algorithm to determine, among the question vectors, a similar question vector with the smallest Euclidean distance from the target question vector, where the similar question vector represents the vector with the closest semantic level to the target question vector, and the target question vector is any one of the question vectors; in the second step, forming a similar vector pair from the target question vector and the similar question vector; in the first repetition step, repeating the first step and the second step to obtain a plurality of similar vector pairs; inputting the pairs of original question texts corresponding to the plurality of similar vector pairs into the Simbert model to train the Simbert model to generate synonymous sentences that meet the requirements.
[0006] Optionally, the training process of the Simbert model includes: training the Simbert model using a plurality of similar vector pairs until the cross-entropy loss function and the contrastive loss function converge to obtain the trained Simbert model, where the cross-entropy loss function is used to calculate the loss between the target question vector and the similar question vector in the similar vector pair, and the contrastive loss function is used to calculate the sum of the losses between the target question vector and other question vectors respectively, and the target question vector is any one of the question vectors.
[0007] Optionally, training the Simbert model using a plurality of similar vector pairs until the cross-entropy loss function and the contrastive loss function converge includes: in the first processing step, separating the target question text corresponding to the target question vector and the similar question text corresponding to the similar question vector in the similar vector pair through an identifier so that the Simbert model can identify the target question text and the similar question text; in the second processing step, inputting the target question text into the encoder of the Simbert model for vectorization processing to obtain a first vector, where the first vector is the vector representation of all the words in the target question text; in the third processing step, inputting the first vector into the decoder of the Simbert model to generate a predicted vocabulary sequence, comparing the predicted vocabulary sequence with the similar question text, and calculating the loss value of the cross-entropy loss function, where the predicted vocabulary sequence is a text composed of words; in the second repetition step, repeating the first processing step, the second processing step, and the third processing step to adjust the parameters of the Simbert model until the cross-entropy loss function converges.
[0008] Optionally, multiple of the similar vector pairs are used to train the Simbert model until the cross-entropy loss function and the contrastive loss function converge. It further includes: concatenating all the problem vectors to obtain a sentence vector matrix D ∈ R b*d , where b is the total number of the problem vectors, d is the dimension of the problem vectors, and R indicates that the elements in the sentence vector matrix are real numbers; normalizing all the problem vectors in the sentence vector matrix with the L2 regularization norm to obtain a regularization matrix Performing an inner product calculation on the regularization matrix to obtain a similarity matrix The elements in the similarity matrix represent the similarity between any two of the problem vectors; determining, according to the problem vector of one of the original problem texts in the original problem text pair, an associated vector of the problem vector of the other original problem text, determining the associated vector as the positive sample vector of the problem vector, and determining the other problem vectors except the associated vector as the negative sample vectors of the problem vector, where the associated vector is a vector that is semantically similar to the problem vector; determining the similarity between each of the positive sample vectors and each of the negative sample vectors according to the similarity matrix; substituting the similarities of each of the positive sample vectors and the similarities of each of the negative sample vectors into the contrastive loss function to calculate the loss value of the contrastive loss function; and adjusting the parameters of the Simbert model until the contrastive loss function converges.
[0009] Optionally, before determining the similarity between each of the positive sample vectors and each of the negative sample vectors according to the similarity matrix, the method further includes: performing a mask process on the diagonal elements of the similarity matrix, where the diagonal elements are the elements corresponding to the diagonal from the upper left corner to the lower right corner in the similarity matrix.
[0010] Optionally, adding all the second target synonymous sentences to the Q&A corpus to bind the second target synonymous sentences to the corresponding original problem texts. The method further includes: when the input question of the FAQ system is the original problem text, using the answer corresponding to the original problem text as the output result of the FAQ system; and when the input question of the FAQ system is the second target synonymous sentence corresponding to the original problem text, using the answer corresponding to the original problem text bound to the second target synonymous sentence as the output result of the FAQ system.
[0011] To achieve the above object, according to one aspect of the present application, there is provided a text data enhancement device for an FAQ system, including: a construction unit for constructing a training set by using a part of the original question texts in the question-answer corpus of the current FAQ system and the synonymous sentences corresponding to the part of the original question texts, and training an initial Simbert model by using the training set to obtain a trained Simbert model, where the Simbert model is used for generating synonymous sentences; a first generation unit for generating synonymous sentences for all the original question texts by using the trained Simbert model to obtain a plurality of the synonymous sentences corresponding to each of the original question texts; a second generation unit for processing all the original question texts and the synonymous sentences corresponding to all the original question texts by using a text enhancement method at the word level to obtain a first target synonymous sentence corresponding to each of the original question texts, where the text enhancement method at the word level includes removing stop words, keyword extraction, homophone replacement, synonym replacement, randomly inserting words, and randomly deleting words; a screening unit for screening each of the first target synonymous sentences to obtain a second target synonymous sentence corresponding to each of the original question texts, where the second target synonymous sentence is a sentence with the same meaning as the original question text and correct grammar; an adding unit for adding all the second target synonymous sentences to the question-answer corpus to bind the second target synonymous sentences to the corresponding original question texts.
[0012] According to another aspect of the present application, there is provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls any one of the methods in the device where the computer-readable storage medium is located.
[0013] According to yet another aspect of the present application, there is provided a text data enhancement system, including: one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the methods.
[0014] Applying the technical solution of the present application in the above-mentioned method for enhancing text data of an FAQ system includes: constructing a training set by using some original question texts in the Q&A corpus of the current FAQ system and synonymous sentences corresponding to the above-mentioned some original question texts, and training an initial Simbert model by using the above-mentioned training set to obtain a trained Simbert model, where the Simbert model is used for generating synonymous sentences; using the above-mentioned trained Simbert model to generate synonymous sentences for all the original question texts to obtain multiple above-mentioned synonymous sentences corresponding to each of the above-mentioned original question texts; using a text enhancement method at the word level to process all the above-mentioned original question texts and the above-mentioned synonymous sentences corresponding to all the above-mentioned original question texts to obtain a first target synonymous sentence corresponding to each of the above-mentioned original question texts, and the above-mentioned text enhancement method at the word level includes removing stop words, keyword extraction, homophone replacement, synonym replacement, random insertion of words, and random deletion of words; screening each of the above-mentioned first target synonymous sentences to obtain a second target synonymous sentence corresponding to each of the above-mentioned original question texts, and the above-mentioned second target synonymous sentence is a sentence with the same meaning as the above-mentioned original question text and correct grammar; adding all the above-mentioned second target synonymous sentences to the above-mentioned Q&A corpus to bind the above-mentioned second target synonymous sentences with the corresponding above-mentioned original question texts. The present application trains the model by constructing a training set for the Simbert model, uses the trained Simbert model to generate synonymous sentences for the original question texts, then performs word-level processing on the original question texts and the synonymous sentences corresponding to the original question texts to obtain the first target synonymous sentences, and finally determines the second target synonymous sentences by screening and adds them to the Q&A corpus, thereby improving the ability of the FAQ system to answer questions, avoiding the poor Q&A effect of the FAQ system due to too few original question texts in the Q&A corpus, and solving the problem of difficult acquisition of original question texts in the prior art. Description of the Drawings
[0015] Figure 1 Fig. shows a hardware structure block diagram of a mobile terminal for implementing a method for enhancing text data of an FAQ system according to an embodiment of the present application;
[0016] Figure 2 Fig. shows a flowchart of a method for enhancing text data of an FAQ system according to an embodiment of the present application;
[0017] Figure 3 Fig. shows a training schematic diagram of a Simbert model for a method for enhancing text data of an FAQ system according to an embodiment of the present application;
[0018] Figure 4 Fig. shows an overall flowchart of a method for enhancing text data of an FAQ system according to an embodiment of the present application;
[0019] Figure 5 The block diagram of a text data enhancement device for an FAQ system provided according to an embodiment of the present application is shown.
[0020] Among them, the above-mentioned drawings include the following reference numerals:
[0021] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. Detailed implementation manners
[0022] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0023] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to describe the embodiments of the present application here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.
[0025] For the convenience of description, some nouns or terms related to the embodiments of the present application are described below:
[0026] FAQ system: The FAQ system is one of the research and application directions with broad development prospects in the fields of artificial intelligence and natural language processing. It is a common question-answering system that usually uses sentences in natural language form for questions and directly returns accurate and reliable answers to users in combination with a corpus to improve the way users obtain information.
[0027] As introduced in the background art, in the prior art, the FAQ system has the problems of high and difficult corpus acquisition cost. To solve this technical problem, the embodiments of the present application provide a method, apparatus, computer-readable storage medium, and text data enhancement system for the text data enhancement of the FAQ system.
[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention.
[0029] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a method of text data enhancement of an FAQ system according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in Figure 1 the processor 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in Figure 1 is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than
[0030] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the text data enhancement method of a FAQ system in an embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0031] In this embodiment, a method for enhancing text data of a FAQ system running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0032] Figure 2 It is a flowchart of a method for enhancing text data of a FAQ system according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:
[0033] Step S201, constructing a training set by using some original question texts in the Q&A corpus of the current FAQ system and synonymous sentences corresponding to the above-mentioned some original question texts, and training the initial Simbert model by applying the above-mentioned training set to obtain a trained Simbert model, where the Simbert model is used for generating synonymous sentences;
[0034] Specifically, a training set is constructed using some of the original question texts in the current FAQ system Q&A corpus and their corresponding synonymous sentences. The initial SimBERT model is trained using this dataset to obtain an optimized SimBERT model, enabling it to efficiently generate synonymous sentences. This optimization process significantly enhances the model's understanding and generation capabilities for domain-specific sentences, directly improving the richness and quality of the question corpus in the FAQ system.
[0035] Step S202: Use the above-trained Simbert model to generate synonymous sentences for all the original question texts, obtaining multiple of the above-mentioned synonymous sentences corresponding to each of the above original question texts.
[0036] Specifically, the trained Simbert model is used to process all the original question texts in the Q&A corpus to generate multiple variant synonymous sentences for each original question text.
[0037] Step S203: Use a word-level text augmentation method to process all the above original question texts and the above-mentioned synonymous sentences corresponding to all the above original question texts, obtaining first target synonymous sentences corresponding to each of the above original question texts. The above word-level text augmentation method includes removing stop words, keyword extraction, near-sound word replacement, synonym replacement, random insertion of words, and random deletion of words.
[0038] Specifically, operations such as removing stop words, keyword extraction, near-sound word replacement, synonym replacement, random insertion of words, and random deletion of words are performed on all the original question texts and the synonymous sentences derived from them, thereby obtaining more relevant sentences based on the original question texts and the synonymous sentences corresponding to the original question texts, and further expanding the original question texts.
[0039] Step S204: Screen each of the above first target synonymous sentences to obtain second target synonymous sentences corresponding to each of the above original question texts. The above second target synonymous sentences are sentences with the same meaning as the above original question texts and correct grammar.
[0040] Specifically, each first target synonymous sentence generated after the word-level text augmentation method is carefully screened manually to ensure that the retained second target synonymous sentences not only match the meaning of the original question text but also have a correct grammatical structure. This screening process effectively eliminates sentences that may have semantic deviations or grammar errors during the data augmentation process, thus ensuring the high quality of the augmented corpus.
[0041] Step S205: Add all the above second target synonymous sentences to the above Q&A corpus to bind the above second target synonymous sentences to the corresponding above original question texts.
[0042] Specifically, all the second target synonymous sentences that have been screened and verified to ensure accurate meaning and correct grammar are added to the Q&A corpus of the FAQ system to achieve precise binding with the original question text. This integration step significantly increases the depth and breadth of the corpus, enabling the system to understand and respond to various questions raised by users more flexibly and comprehensively.
[0043] Through this embodiment, in the above text data enhancement method for a FAQ system, a training set is constructed using some of the original question texts in the Q&A corpus of the current FAQ system and the synonymous sentences corresponding to the above-mentioned some original question texts, and the initial Simbert model is trained using the above training set to obtain a trained Simbert model for generating synonymous sentences; the trained Simbert model is used to generate synonymous sentences for all the original question texts to obtain multiple synonymous sentences corresponding to each of the above original question texts; a word-level text enhancement method is used to process all the above original question texts and the synonymous sentences corresponding to all the above original question texts to obtain the first target synonymous sentences corresponding to each of the above original question texts, and the above word-level text enhancement method includes removing stop words, keyword extraction, homophone replacement, synonym replacement, random insertion of words, and random deletion of words; each of the above first target synonymous sentences is screened to obtain the second target synonymous sentences corresponding to each of the above original question texts, and the above second target synonymous sentences are sentences with the same meaning as the above original question text and correct grammar; all the above second target synonymous sentences are added to the above Q&A corpus to bind the above second target synonymous sentences with the corresponding above original question texts. This application trains the model by constructing a training set for the Simbert model, uses the trained Simbert model to generate synonymous sentences for the original question texts, then performs word-level processing on the original question texts and the synonymous sentences corresponding to the original question texts to obtain the first target synonymous sentences, and finally determines the second target synonymous sentences through screening and adds them to the Q&A corpus, thereby improving the ability of the FAQ system to answer questions, avoiding the poor Q&A effect of the FAQ system due to too few original question texts in the Q&A corpus, and solving the problem of difficult acquisition of original question texts in the prior art.
[0044] In order to obtain the training set of the Simbert model, in an optional implementation manner, a training set is constructed using some of the original question texts in the Q&A corpus of the current FAQ system and the synonymous sentences corresponding to the above-mentioned some original question texts, and the above step S201 includes:
[0045] Step S2011, using the Embedding model to convert the above-mentioned some original question texts in the Q&A corpus of the current FAQ system into question vectors;
[0046] Specifically, using the Embedding model, a selected part of the original question text in the current Q&A corpus of the FAQ system is converted into the corresponding question vector representation. This conversion process maps the text information into a multi-dimensional vector space, effectively capturing the semantic features of the questions.
[0047] Step S2012, the first step, use the ANN algorithm to determine the similar question vector with the smallest Euclidean distance from the target question vector among the above question vectors. The above similar question vector represents the vector with the closest semantic level to the above target question vector, and the above target question vector is any one of the above question vectors.
[0048] Specifically, determine the similar question vector with the shortest Euclidean distance from the target question vector through the ANN approximate nearest neighbor search algorithm. Among them, the similar question vector represents the vector representation that is closest to the target vector at the semantic level. Here, the target question vector can be any selected question vector in the Q&A corpus. This step can find the similar question that is semantically most relevant to the target question in the high-dimensional vector space.
[0049] Step S2013, the second step, form a similar vector pair with the above target question vector and the above similar question vector.
[0050] Specifically, pair the target question vector with the similar question vector to construct a similar vector pair. The effect of this construction process is that it can provide the Simbert model with question pairs with related similar semantics to prepare for the subsequent training of the Simbert model.
[0051] Step S2014, the first repeated step, repeat the above first step and the above second step to obtain multiple above similar vector pairs.
[0052] Specifically, loop through the operations of the first step and the second step, that is, continuously use the ANN algorithm to match the similar question vector closest to the selected target question vector, and then combine these vectors to generate multiple pairs of similar vector pairs.
[0053] Step S2015, input the original question text pairs corresponding to the multiple above similar vector pairs into the above Simbert model to train the above Simbert model to generate the above synonymous sentences that meet the requirements.
[0054] Specifically, the original question text pairs mapped back by multiple similar vector pairs are used as training materials and input into the Simbert model, aiming to train the Simbert model to produce synonymous sentences that meet the standard requirements. The effect of this process is that through targeted training, the Simbert model can more accurately understand and generate synonymous sentences with the same semantics as the original questions but different expressions, thereby significantly enhancing the ability of the FAQ system to handle synonymous questions and improving the overall user experience.
[0055] In order to train the above Simbert model so that the Simbert model has the ability to generate synonymous sentences, in an alternative embodiment, the training process of the above Simbert model includes:
[0056] Step S301, training the above Simbert model with multiple above-mentioned similar vector pairs until the cross-entropy loss function and the contrast loss function converge, to obtain the above-mentioned trained Simbert model. The cross-entropy loss function is used to calculate the loss between the target question vector and the similar question vector in the above-mentioned similar vector pair, and the contrast loss function is used to calculate the sum of the losses between the target question vector and other above-mentioned question vectors respectively. The target question vector is any one of the above-mentioned question vectors.
[0057] Specifically, the Simbert model is iteratively trained with multiple groups of similar vector pairs until the values of the cross-entropy loss function and the contrast loss function reach stable convergence, so as to obtain an optimized Simbert model through training, enabling the Simbert model to accurately understand the semantics of questions and generate high-quality synonymous sentences.
[0058] In order to make the cross-entropy loss function converge, in an alternative embodiment, the above Simbert model is trained with multiple above-mentioned similar vector pairs until the above cross-entropy loss function and the above contrast loss function converge. The above step S301 includes:
[0059] Step S30101, the first processing step, separating the target question text corresponding to the target question vector and the similar question text corresponding to the similar question vector in the above-mentioned similar vector pair through an identifier, so that the above Simbert model can identify the above target question text and the above similar question text;
[0060] Specifically, a specific identifier is used to divide and identify the target question text and the similar question text in the similar vector pair. This operation aims to enable the Simbert model to accurately distinguish and identify the target question text and its similar question text.
[0061] Step S30102, the second processing step, input the above target problem text into the encoder of the above Simbert model for vectorization processing to obtain a first vector, where the first vector is the vector representation of all the words in the above target problem text;
[0062] Specifically, the target problem text is sent to the encoder module of the Simbert model for processing to generate a first vector. The first vector is essentially the vectorized representation of all the words in the target problem text. Through vectorization, the model can capture the semantic relationships and context information between words, which helps the model understand and generate synonymous sentences with similar semantics but different expressions from the original problem.
[0063] Step S30103, the third processing step, input the above first vector into the decoder of the above Simbert model to generate a predicted word sequence, compare the above predicted word sequence with the above similar problem text, and calculate the loss value of the cross-entropy loss function. The above predicted word sequence is a text composed of words;
[0064] Specifically, the obtained first vector, that is, the vector representation of the words in the target problem text, is sent to the decoder part of the Simbert model to prompt the model to generate a series of predicted words to form a so-called predicted word sequence. This sequence is essentially a text composed of words. Then, this predicted text is compared and analyzed with the similar problem text, and the difference between the two is quantified by calculating the cross-entropy loss function, which is used as a feedback signal for model training.
[0065] Step S30104, the second repetition step, repeat the above first processing step, the above second processing step, and the above third processing step, and adjust the parameters of the above Simbert model until the above cross-entropy loss function converges.
[0066] Specifically, repeat the first processing step, the second processing step, and the third processing step. In this iterative process, the parameters of the Simbert model will be dynamically adjusted according to the feedback of the calculated cross-entropy loss function until the value of this loss function reaches stability, that is, convergence is achieved.
[0067] To make the contrast loss function converge, in an optional implementation manner, multiple above similar vector pairs are used to train the above Simbert model until the above cross-entropy loss function and the above contrast loss function converge. The above step S301 further includes:
[0068] Step S30105, concatenate all the above problem vectors to obtain a sentence vector matrix D ∈ R b*d , where b is the total number of the above problem vectors, d is the dimension of the above problem vectors, and R indicates that the elements in the above sentence vector matrix are real numbers;
[0069] Specifically, a concatenation operation is performed on the problem vectors to construct a sentence vector matrix. The dimension of this matrix is determined by the total number b of problem vectors and the dimension d of a single vector. This operation integrates the scattered vector information into a compact matrix form, facilitating batch processing and optimized calculations by the model.
[0070] Step S30106: Normalize all the above-mentioned problem vectors in the above-mentioned sentence vector matrix with the L2 regularization norm to obtain a regular matrix.
[0071] Specifically, perform the normalization operation of the L2 regularization norm on each vector element in the constructed sentence vector matrix, thereby generating a regular matrix. This step standardizes the lengths of the individual vectors, avoids the influence of vector magnitudes on the model learning process, and ensures that each vector in the matrix has the same weight and comparison benchmark. Through normalization, the Simbert model becomes more accurate when calculating similarities.
[0072] Step S30107: Perform an inner product calculation on the above-mentioned regular matrix to obtain a similarity matrix. The elements in the above-mentioned similarity matrix represent the similarity between any two of the above-mentioned problem vectors.
[0073] Specifically, perform an inner product operation on the obtained regular matrix, thereby constructing a similarity matrix. Each element in the matrix specifically reflects the semantic similarity degree between any two problem vectors. Through this step, the semantic relevance between problem texts is quantified, enabling us to intuitively evaluate and compare the similarities between different problem vectors numerically.
[0074] Step S30108: Determine the associated vector of one of the above-mentioned problem vectors in the above-mentioned original problem text pair as the associated vector of the other above-mentioned problem vector. Determine the above-mentioned associated vector as the positive sample vector of the above-mentioned problem vector, and determine the other above-mentioned problem vectors except the above-mentioned associated vector as the negative sample vectors of the above-mentioned problem vector. The above-mentioned associated vector is a vector that is semantically similar to the above-mentioned problem vector.
[0075] Specifically, according to one text in the original question text pair, its corresponding question vector is determined and identified as the associated vector of the question vector corresponding to another original question text. This associated vector is determined as the positive sample vector, and all other question vectors are determined as negative sample vectors. By marking vectors with similar semantics as positive samples, the model can, when generating synonymous sentences, preferentially imitate and learn the expression patterns of these positive sample vectors, thereby ensuring that the generated sentences are semantically closely related to the original sentence, while negative samples prompt the model to learn how to distinguish and avoid generating semantically unrelated sentences.
[0076] Step S30109, determine the similarity between each of the above positive sample vectors and each of the above negative sample vectors according to the above similarity matrix;
[0077] Specifically, according to the similarity matrix, precisely quantify the similarity of the positive sample vector and the negative sample vector respectively relative to the target question vector, so that the model can, based on these specific values, learn to distinguish and identify text features that are semantically similar and different.
[0078] Step S30110, substitute the similarity of each of the above positive sample vectors and the similarity of each of the above negative sample vectors into the contrast loss function to calculate the loss value of the above contrast loss function;
[0079] Specifically, input the similarity of each positive sample vector and the similarity of each negative sample vector into the contrast loss function respectively to calculate the specific loss value of the contrast loss function.
[0080] Step S30111, adjust the parameters of the above Simbert model until the above contrast loss function converges.
[0081] Specifically, adjust the various parameters of the SimBERT model until the contrast loss function reaches the convergence state, so that SimBERT can more accurately understand and generate synonymous sentences related to the semantics of the original question, reducing the output of irrelevant or semantically deviated sentences, and providing a rich and high-quality corpus for the FAQ system.
[0082] To make the calculation of similarity more accurate, before determining the similarity between each of the above positive sample vectors and each of the above negative sample vectors according to the above similarity matrix, the above method further includes:
[0083] Step S401, perform mask processing on the diagonal elements of the above similarity matrix, and the diagonal elements are the elements corresponding to the diagonal from the upper left corner to the lower right corner of the above similarity matrix.
[0084] Specifically, the diagonal elements from the upper left to the lower right in the similarity matrix are masked. The goal is to prevent the model from comparing the vector of the same sentence with itself, thereby avoiding redundancy in similarity calculation and ensuring that the model can focus on distinguishing the semantic similarity between different sentences rather than being misled by self-repeating results.
[0085] In order to enable the FAQ system to give correct output results according to the input question after expanding the Q&A corpus of the FAQ system, in an alternative embodiment, all the above-mentioned second-target synonymous sentences are added to the above-mentioned Q&A corpus to bind the above-mentioned second-target synonymous sentences to the corresponding above-mentioned original question text. The method further includes:
[0086] Step S501, when the input question of the above-mentioned FAQ system is the above-mentioned original question text, the answer corresponding to the above-mentioned original question text is used as the output result of the above-mentioned FAQ system;
[0087] Specifically, when the query received by the FAQ system is exactly the above-mentioned original question text, the system directly returns the preset answer matching the above-mentioned original question text as the output result, ensuring that the FAQ system can quickly and accurately answer the queries for the existing questions in the Q&A corpus.
[0088] Step S502, when the input question of the above-mentioned FAQ system is the above-mentioned second-target synonymous sentence corresponding to the above-mentioned original question text, the answer corresponding to the above-mentioned original question text bound to the above-mentioned second-target synonymous sentence is used as the output result of the above-mentioned FAQ system.
[0089] Specifically, when the received query is the second-target synonymous sentence corresponding to the original question text, the system outputs the preset answer corresponding to the original question text bound to the second-target synonymous sentence. That is, even if the question proposed by the user varies in expression, the system can, through the binding of synonymous sentences, locate the original question text and provide the answer corresponding to the original question text, significantly improving the user's query experience.
[0090] In a specific embodiment, taking the original question text: "The photo in the online verification system does not match the photo on the second-generation resident ID card. Can it be optimized?" as an example, the specific implementation process of the present invention is described.
[0091] Original statement: "The photo in the online verification system does not match the photo on the second-generation resident ID card. Can it be optimized?"
[0092] I. Based on sentence-level text augmentation, the Simbert model generates synonymous sentences. Taking 5 sentences as an example:
[0093] Synonymous sentence 1: "The photo in the online verification system does not match the photo on the second-generation resident ID card. What should I do?"
[0094] Synonymous Sentence 2: How to handle the situation where the photo in the online verification system does not match the photo on the resident ID card?
[0095] Synonymous Sentence 3: The identity information in the instruction document of the online verification system does not match the associated photo when querying the resident ID card. What solutions are there?
[0096] Synonymous Sentence 4: Can the situation where the photo in the online verification system does not match the photo on the second-generation ID card be optimized?
[0097] Synonymous Sentence 5: Hello, regarding the handling of the problem that "the photo in the system verification system does not match the photo on the second-generation resident ID card. Can it be optimized?"
[0098] Second, the generated synonymous sentences are sequentially subjected to word-based text enhancement. If the input synonymous sentence is "The photo in the online verification system does not match the photo on the second-generation resident ID card. What should I do?", the result after processing by the word-level text enhancement method is:
[0099] Result of removing stop words enhancement: The photo in the online verification system does not match the photo on the second-generation resident ID card;
[0100] Result of keyword enhancement: Resident ID card, solution, verification, query, and networking;
[0101] Result of near-homophone / synonym enhancement: The photo in the connectivity verification system does not match the photo on the second-generation resident ID card. What should I do;
[0102] Result of random insertion / deletion enhancement: The photo in the online verification instruction system file does not match the photo on the second-generation resident ID card.
[0103] Finally, the results processed by the word-level text enhancement method are manually screened, and the screened results are added to the Q&A corpus of the FAQ system.
[0104] Figure 3 Shows a Simbert model training schematic diagram of a text data enhancement method for an FAQ system provided according to an embodiment of the present application, as Figure 3 shown:
[0105] During the training of the SimBERT model, different sentences of the similar sentence pairs are separated by the [SEP] identifier, and the Attention Mask method is used to use bidirectional attention between each token in the first half of the [SEP] identifier statement and unidirectional attention between each token in the second half, so that the model has the NLG ability. During the training process, Simbert concatenates all the [CLS] sentence vectors within each batch to form a sentence vector matrix, and performs L_2 regularization to obtain a regular matrix The final similarity matrix is obtained after the inner product operation. Among them, the diagonal part is masked. The loss function of the final Simbert model is the combined loss function obtained by adding the cross-entropy loss function and the contrast loss function in the similar sentence classifier.
[0106] Figure 4 The overall flowchart of a method for enhancing text data of an FAQ system according to an embodiment of the present application is shown, as Figure 4 shown:
[0107] The overall process of a method for enhancing text data of an FAQ system mainly consists of two parts: The part within the dashed box 1 is a method for enhancing text data at the sentence level, which generates n synonymous sentences from 1 original question corpus input through a deep learning model. The part within the dashed box 2 is to sequentially perform stop word removal, keyword extraction, near-sound / synonym replacement, and random word insertion / deletion on the generated n synonymous sentences (including the original sentence), generating a total of n*m text enhancement samples. This process can be implemented through public tools such as Jieba, and there is no need to manually collect and organize a large amount of data. If the problem raw materials of the artificial data consist of N, after being processed by the method of this article, a total of m*n*N enhancement results are obtained, and then they are stored in the database after manual modification. The enhanced results stored in the database are used to improve the Q&A effect of the FAQ system on the one hand and to improve the effect of the synonym training model on the other hand.
[0108] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0109] The embodiment of the present application also provides a device for enhancing text data of an FAQ system. It should be noted that a device for enhancing text data of an FAQ system in the embodiment of the present application can be used to execute the method for enhancing text data of an FAQ system provided in the embodiment of the present application. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0110] The following introduces a device for enhancing text data of an FAQ system provided by an embodiment of the present application.
[0111] Figure 5 is a structural block diagram of a device for enhancing text data of an FAQ system according to an embodiment of the present application. As Figure 5 shown, the device includes:
[0112] The construction unit 10 is used to construct a training set by using some original question texts in the Q&A corpus of the current FAQ system and the synonymous sentences corresponding to the above-mentioned some original question texts, and train the initial Simbert model by applying the above training set to obtain a trained Simbert model, and the Simbert model is used for generating synonymous sentences.
[0113] Specifically, a training set is constructed by using some original question texts in the Q&A corpus of the current FAQ system and their corresponding synonymous sentences, and the initial SimBERT model is trained by using this data set, so as to obtain an optimized SimBERT model, so that it can efficiently generate synonymous sentences. This optimization process significantly enhances the model's understanding and generation ability of sentences in a specific field, and directly improves the richness and quality of the question corpus in the FAQ system.
[0114] The first generation unit 20 is used to generate synonymous sentences for all the original question texts by using the above-mentioned trained Simbert model, and obtain multiple above-mentioned synonymous sentences corresponding to each of the above original question texts.
[0115] Specifically, the trained Simbert model is used to process all the original question texts in the Q&A corpus to generate multiple variant synonymous sentences for each original question text.
[0116] The second generation unit 30 is used to process all the above-mentioned original question texts and the above-mentioned synonymous sentences corresponding to all the above-mentioned original question texts by using a text enhancement method at the word level to obtain a first target synonymous sentence corresponding to each of the above original question texts, and the text enhancement method at the word level includes removing stop words, keyword extraction, near-sound word replacement, synonym replacement, random insertion of words and random deletion of words.
[0117] Specifically, operations such as removing stop words, keyword extraction, near-sound word replacement, synonym replacement, random insertion of words and random deletion of words are performed on all the original question texts and the synonymous sentences derived therefrom, so as to obtain more relevant sentences based on the original question texts and the synonymous sentences corresponding to the original question texts, and further expand the original question texts.
[0118] The screening unit 40 is used to screen each of the above first target synonymous sentences to obtain a second target synonymous sentence corresponding to each of the above original question texts, and the second target synonymous sentence is a sentence with the same meaning as the above original question text and correct grammar.
[0119] Specifically, each of the first target synonymous sentences generated after processing by the word-level text enhancement method is carefully screened manually to ensure that the retained second target synonymous sentences not only match the meaning of the original question text but also have correct grammar structures. This screening process effectively eliminates sentences that may have semantic deviations or grammar errors during the data enhancement process, thus ensuring the high quality of the enhanced corpus.
[0120] An adding unit 50 is added for adding all the above-mentioned second target synonymous sentences to the above-mentioned Q&A corpus to bind the above-mentioned second target synonymous sentences to the corresponding above-mentioned original question text.
[0121] Specifically, all the second target synonymous sentences that have been screened and verified to ensure accurate meaning and correct grammar are added to the Q&A corpus of the FAQ system to achieve precise binding with the original question text. This integration step significantly increases the depth and breadth of the corpus, enabling the system to more flexibly and comprehensively understand and respond to various questions raised by users.
[0122] Through this embodiment, a construction unit is configured to construct a training set by using some of the original question texts in the current Q&A corpus of the FAQ system and the synonymous sentences corresponding to the above-mentioned some original question texts, and train an initial Simbert model by applying the above-mentioned training set to obtain a trained Simbert model, where the Simbert model is used for generating synonymous sentences; a first generation unit is configured to generate synonymous sentences for all the original question texts by using the above-mentioned trained Simbert model to obtain multiple above-mentioned synonymous sentences corresponding to each of the above-mentioned original question texts; a second generation unit is configured to process all the above-mentioned original question texts and the above-mentioned synonymous sentences corresponding to all the above-mentioned original question texts by using a text enhancement method at the word level to obtain first target synonymous sentences corresponding to each of the above-mentioned original question texts, and the above-mentioned text enhancement method at the word level includes removing stop words, keyword extraction, homophone replacement, synonym replacement, randomly inserting words, and randomly deleting words; a screening unit is configured to screen each of the above-mentioned first target synonymous sentences to obtain second target synonymous sentences corresponding to each of the above-mentioned original question texts, and the above-mentioned second target synonymous sentences are sentences with the same meaning as the above-mentioned original question texts and correct grammar; an adding unit is configured to add all the above-mentioned second target synonymous sentences to the above-mentioned Q&A corpus to bind the above-mentioned second target synonymous sentences to the corresponding above-mentioned original question texts. Through constructing a training set of the Simbert model to train the model, generating synonymous sentences for the original question texts by using the trained Simbert model, then performing word-level processing on the original question texts and the synonymous sentences corresponding to the original question texts to obtain first target synonymous sentences, and finally determining second target synonymous sentences through screening and adding them to the Q&A corpus, the ability of the FAQ system to answer questions is improved, the problem that the Q&A effect of the FAQ system is poor due to too few original question texts in the Q&A corpus is avoided, and the problem of difficult acquisition of original question texts in the prior art is solved.
[0123] In an alternative embodiment, in order to obtain a training set of the Simbert model, a training set is constructed by using some of the original question texts in the current Q&A corpus of the FAQ system and the synonymous sentences corresponding to the above-mentioned some original question texts, and the above-mentioned construction unit includes:
[0124] A first construction module is configured to convert the above-mentioned some original question texts in the current Q&A corpus of the above-mentioned FAQ system into question vectors by using an Embedding model;
[0125] Specifically, by using the Embedding model, some selected original question texts in the current Q&A corpus of the FAQ system are converted into corresponding question vector representations. This conversion process maps the text information into a multi-dimensional vector space, effectively capturing the semantic features of the questions.
[0126] A second construction module for performing a first step of determining, using an ANN algorithm, a similar problem vector with the smallest Euclidean distance from the target problem vector among each of the above-mentioned problem vectors, where the similar problem vector represents a vector with the closest semantic level to the target problem vector, and the target problem vector is any one of the above-mentioned problem vectors;
[0127] Specifically, an ANN approximate nearest neighbor search algorithm is used to determine a similar problem vector with the shortest Euclidean distance from the target problem vector. Here, the similar problem vector represents a vector representation that is closest to the target vector at the semantic level. Here, the target problem vector can be any selected problem vector in the Q&A corpus. This step can find a similar problem that is semantically most relevant to the target problem in a high-dimensional vector space.
[0128] A third construction module for performing a second step of forming a similar vector pair from the above-mentioned target problem vector and the above-mentioned similar problem vector;
[0129] Specifically, the target problem vector is paired with the similar problem vector to construct a similar vector pair. The effect of this construction process is that it can provide the Simbert model with pairs of problems with related and similar semantics to prepare for subsequent training of the Simbert model.
[0130] A fourth construction module for performing a first repetition step of repeating the above-mentioned first step and the above-mentioned second step to obtain multiple above-mentioned similar vector pairs;
[0131] Specifically, the operations of the first step and the second step are repeatedly executed, that is, the ANN algorithm is continuously used to match the similar problem vector closest to the selected target problem vector, and then these vectors are combined to generate multiple pairs of similar vector pairs.
[0132] A fifth construction module for inputting the original problem text pairs corresponding to the multiple above-mentioned similar vector pairs into the above-mentioned Simbert model to train the above-mentioned Simbert model to generate the above-mentioned synonymous sentences that meet the requirements.
[0133] Specifically, the original problem text pairs mapped back from multiple similar vector pairs are used as training materials and input into the Simbert model, aiming to train the Simbert model to be able to produce synonymous sentences that meet the standard requirements. The effect of this process is that through targeted training, the Simbert model can more accurately understand and generate synonymous sentences with the same semantics as the original problem but different expressions, thus significantly enhancing the ability of the FAQ system to handle synonymous questions and improving the overall user experience.
[0134] To train the above Simbert model so that the Simbert model has the ability to generate synonymous sentences, in an alternative embodiment, in the training process of the above Simbert model, the device includes:
[0135] A training unit for training the above Simbert model using multiple pairs of the above similarity vectors until the cross-entropy loss function and the contrastive loss function converge, to obtain the above trained Simbert model. The cross-entropy loss function is used to calculate the loss between the above target problem vector and the above similar problem vector in the above similarity vector pair, and the contrastive loss function is used to calculate the sum of the losses between the target problem vector and each of the other above problem vectors. The target problem vector is any one of the above problem vectors.
[0136] Specifically, the Simbert model is iteratively trained using multiple pairs of similarity vectors until the values of the cross-entropy loss function and the contrastive loss function reach stable convergence, thereby obtaining an optimized Simbert model through training, enabling the Simbert model to accurately understand the problem semantics and generate high-quality synonymous sentences.
[0137] To make the cross-entropy loss function converge, in an alternative embodiment, the above Simbert model is trained using multiple pairs of the above similarity vectors until the above cross-entropy loss function and the above contrastive loss function converge. The training unit includes:
[0138] A first training module for performing a first processing step of separating the target problem text corresponding to the above target problem vector and the similar problem text corresponding to the above similar problem vector in the above similarity vector pair through an identifier, so that the above Simbert model can identify the above target problem text and the above similar problem text;
[0139] Specifically, a specific identifier is used to divide and identify the target problem text and the similar problem text in the similarity vector pair. This operation aims to enable the Simbert model to accurately distinguish and identify the target problem text and its similar problem text.
[0140] A second training module for performing a second processing step of inputting the above target problem text into the encoder of the above Simbert model for vectorization processing to obtain a first vector, where the first vector is the vector representation of all the words in the above target problem text;
[0141] Specifically, the target problem text is sent to the encoder module of the Simbert model for processing to generate a first vector. The first vector is essentially the vectorized representation of all the words in the target problem text. Through vectorization, the model can capture the semantic relationships and context information between words, which helps the model understand and generate synonymous sentences with similar semantics but different expressions from the original problem.
[0142] A third training module for performing a third processing step of inputting the first vector into the decoder of the Simbert model to generate a predicted vocabulary sequence, comparing the predicted vocabulary sequence with the similar problem text, and calculating the loss value of the cross-entropy loss function. The predicted vocabulary sequence is a text composed of words.
[0143] Specifically, the obtained first vector, that is, the vectorized representation of the words in the target problem text, is sent to the decoder part of the Simbert model to prompt the model to generate a series of predicted words to form a so-called predicted vocabulary sequence. This sequence is essentially a text composed of a string of words. Then, this predicted text is compared and analyzed with the similar problem text, and the difference between the two is quantified by calculating the cross-entropy loss function, which is used as a feedback signal for model training.
[0144] A fourth training module for performing a second repetition step of repeating the first processing step, the second processing step, and the third processing step to adjust the parameters of the Simbert model until the cross-entropy loss function converges.
[0145] Specifically, the first processing step, the second processing step, and the third processing step are repeated. During this iterative process, the parameters of the Simbert model will be dynamically adjusted according to the feedback of the calculated cross-entropy loss function until the value of this loss function reaches stability, that is, convergence is achieved.
[0146] To make the contrast loss function converge, in an optional implementation, multiple pairs of the above similar vectors are used to train the Simbert model until the cross-entropy loss function and the contrast loss function converge. The training unit further includes:
[0147] A fifth training module for concatenating all the above problem vectors to obtain a sentence vector matrix D ∈ R b*d , where b is the total number of the above problem vectors, d is the dimension of the above problem vectors, and R indicates that the elements in the sentence vector matrix are real numbers.
[0148] Specifically, a concatenation operation is performed on the problem vectors to construct a sentence vector matrix. The dimension of this matrix is determined by the total number b of problem vectors and the dimension d of a single vector. This operation integrates the scattered vector information into a compact matrix form, facilitating batch processing and optimized calculations by the model.
[0149] The sixth training module is used to normalize all the above-mentioned problem vectors in the above-mentioned sentence vector matrix with the L2 regularization norm to obtain a regular matrix
[0150] Specifically, a normalization operation of the L2 regularization norm is performed on each vector element in the constructed sentence vector matrix, thereby generating a regular matrix. This step standardizes the lengths of the individual vectors, avoids the influence of vector magnitudes on the model learning process, and ensures that each vector in the matrix has the same weight and comparison benchmark. Through normalization, the Simbert model is made more accurate when performing similarity calculations.
[0151] The seventh training module is used to perform an inner product calculation on the above-mentioned regular matrix to obtain a similarity matrix The elements in the above-mentioned similarity matrix represent the similarity between any two of the above-mentioned problem vectors;
[0152] Specifically, an inner product operation is performed on the obtained regular matrix, thereby constructing a similarity matrix. Each element in the matrix specifically reflects the semantic similarity degree between any two problem vectors. Through this step, the semantic relevance between problem texts is quantified, enabling us to numerically and intuitively evaluate and compare the similarities between different problem vectors.
[0153] The eighth training module is used to determine, based on one of the above-mentioned original problem texts in the above-mentioned original problem text pair, the associated vector of the above-mentioned problem vector of the other above-mentioned original problem text, determine the above-mentioned associated vector as the positive sample vector of the above-mentioned problem vector, and determine the other above-mentioned problem vectors except the above-mentioned associated vector as the negative sample vectors of the above-mentioned problem vector. The above-mentioned associated vector is a vector that is semantically similar to the above-mentioned problem vector;
[0154] Specifically, according to one text in the original problem text pair, its corresponding problem vector is determined and identified as the associated vector of the problem vector corresponding to the other original problem text. This associated vector is determined as the positive sample vector, and all other problem vectors are determined as negative sample vectors. By marking semantically similar vectors as positive samples, the model can, when generating synonymous sentences, preferentially imitate and learn the expression patterns of these positive sample vectors, thereby ensuring that the generated sentences are semantically closely related to the original sentence, while the negative samples prompt the model to learn how to distinguish and avoid generating semantically unrelated sentences.
[0155] The ninth training module is used to determine the similarity between each of the above positive sample vectors and each of the above negative sample vectors according to the above similarity matrix;
[0156] Specifically, according to the similarity matrix, the similarities of the positive sample vector and the negative sample vector relative to the target problem vector are accurately quantified, enabling the model to learn to distinguish and identify semantically similar and different text features based on these specific values.
[0157] The tenth training module is used to calculate the loss value of the above contrast loss function by substituting the similarities of each of the above positive sample vectors and the similarities of each of the above negative sample vectors into the contrast loss function;
[0158] Specifically, the similarity of each positive sample vector and the similarity of each negative sample vector are respectively input into the contrast loss function to calculate the specific loss value of the contrast loss function.
[0159] The eleventh training module is used to adjust the parameters of the above Simbert model until the above contrast loss function converges.
[0160] Specifically, the parameters of the SimBERT model are adjusted until the contrast loss function reaches the convergence state, enabling SimBERT to more accurately understand and generate synonymous sentences related to the semantics of the original question, reducing the output of irrelevant or semantically deviated sentences, and providing a rich and high-quality corpus for the FAQ system.
[0161] To make the calculation of similarity more accurate, before determining the similarity between each of the above positive sample vectors and each of the above negative sample vectors according to the above similarity matrix, the above device further includes:
[0162] The first masking unit is used to perform mask processing on the diagonal elements of the above similarity matrix, and the above diagonal elements are the elements corresponding to the diagonal from the upper left corner to the lower right corner of the above similarity matrix.
[0163] Specifically, mask processing is performed on the diagonal elements from the upper left to the lower right of the similarity matrix, aiming to prevent the model from comparing the vector of the same sentence with itself, thus avoiding redundancy in similarity calculation, and ensuring that the model can focus on distinguishing the semantic similarity between different sentences rather than being misled by self-repeated results.
[0164] In order to enable the FAQ system to issue correct output results according to the input question after expanding the Q&A corpus of the FAQ system, in an optional implementation manner, all the above second target synonymous sentences are added to the above Q&A corpus to bind the above second target synonymous sentences to the corresponding above original question text, and the above device further includes:
[0165] The first control unit is configured to, when the input question of the above FAQ system is the above original question text, use the answer corresponding to the above original question text as the output result of the above FAQ system;
[0166] Specifically, when the query received by the FAQ system is exactly the above original question text, the system directly returns the preset answer that matches the above original question text as the output result, ensuring that the FAQ system can quickly and accurately answer queries for questions already existing in the Q&A corpus.
[0167] The second control unit is configured to, when the input question of the above FAQ system is the above second target synonymous sentence corresponding to the above original question text, use the answer corresponding to the above original question text bound to the above second target synonymous sentence as the output result of the above FAQ system.
[0168] Specifically, when the received query is the second target synonymous sentence corresponding to the original question text, the system outputs the preset answer corresponding to the original question text bound to the second target synonymous sentence. That is, even if the question proposed by the user varies in expression, the system can, through the binding of synonymous sentences, locate the original question text and provide the answer corresponding to the original question text, significantly improving the user's query experience.
[0169] The above text data enhancement device for a FAQ system includes a processor and a memory. The above construction unit, first generation unit, second generation unit, screening unit, addition unit, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions. The above modules are all located in the same processor; or, the above respective modules are separately located in different processors in any combination form.
[0170] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the original question text of the FAQ system can be obtained to expand the Q&A corpus.
[0171] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in forms such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0172] An embodiment of the present invention provides a computer-readable storage medium. The above computer-readable storage medium includes a stored program, wherein when the above program runs, it controls the device where the above computer-readable storage medium is located to execute the above text data enhancement method for a FAQ system.
[0173] Specifically, a method for text data enhancement of an FAQ system includes:
[0174] Step S201: Construct a training set using some original question texts in the Q&A corpus of the current FAQ system and synonymous sentences corresponding to the above-mentioned some original question texts, and use the above training set to train an initial Simbert model to obtain a trained Simbert model, where the Simbert model is used for generating synonymous sentences;
[0175] Step S202: Use the above-trained Simbert model to generate synonymous sentences for all the original question texts, and obtain multiple above-mentioned synonymous sentences corresponding to each of the above original question texts;
[0176] Step S203: Process all the above original question texts and the above synonymous sentences corresponding to all the above original question texts using a word-level text enhancement method to obtain first target synonymous sentences corresponding to each of the above original question texts, and the above word-level text enhancement method includes removing stop words, keyword extraction, near-sound word replacement, synonym replacement, randomly inserting words, and randomly deleting words;
[0177] Step S204: Screen each of the above first target synonymous sentences to obtain second target synonymous sentences corresponding to each of the above original question texts, and the above second target synonymous sentences are sentences with the same meaning as the above original question texts and correct grammar;
[0178] Step S205: Add all the above second target synonymous sentences to the above Q&A corpus to bind the above second target synonymous sentences to the corresponding above original question texts.
[0179] An embodiment of the present invention provides a processor, where the above processor is used to run a program, and when the above program runs, it executes the above method for text data enhancement of an FAQ system.
[0180] Specifically, a method for text data enhancement of an FAQ system includes:
[0181] Step S201: Construct a training set using some original question texts in the Q&A corpus of the current FAQ system and synonymous sentences corresponding to the above-mentioned some original question texts, and use the above training set to train an initial Simbert model to obtain a trained Simbert model, where the Simbert model is used for generating synonymous sentences;
[0182] Step S202: Use the above-trained Simbert model to generate synonymous sentences for all the original question texts, and obtain multiple above-mentioned synonymous sentences corresponding to each of the above original question texts;
[0183] Step S203: Process all the above-mentioned original question texts and the above-mentioned synonymous sentences corresponding to all the above-mentioned original question texts by using a text enhancement method at the word level to obtain first target synonymous sentences corresponding to each of the above-mentioned original question texts. The text enhancement method at the word level includes removing stop words, keyword extraction, homophone replacement, synonym replacement, random word insertion, and random word deletion;
[0184] Step S204: Screen each of the above-mentioned first target synonymous sentences to obtain second target synonymous sentences corresponding to each of the above-mentioned original question texts. The above-mentioned second target synonymous sentences are sentences with the same meaning as the above-mentioned original question texts and correct grammar;
[0185] Step S205: Add all the above-mentioned second target synonymous sentences to the above-mentioned Q&A corpus to bind the above-mentioned second target synonymous sentences to the corresponding above-mentioned original question texts.
[0186] The embodiment of the present application also provides a text data enhancement system, including: one or more processors, a memory, and one or more programs, wherein the above-mentioned one or more programs are stored in the above-mentioned memory and are configured to be executed by the above-mentioned one or more processors, including executing any one of the above-mentioned methods in the text data enhancement method of the above-mentioned FAQ system.
[0187] Specifically, a text data enhancement method for an FAQ system includes:
[0188] Step S201: Construct a training set by using some original question texts in the Q&A corpus of the current FAQ system and the synonymous sentences corresponding to the above-mentioned some original question texts, and apply the above-mentioned training set to train an initial Simbert model to obtain a trained Simbert model. The Simbert model is used for generating synonymous sentences;
[0189] Step S202: Use the above-mentioned trained Simbert model to generate synonymous sentences for all the original question texts to obtain multiple above-mentioned synonymous sentences corresponding to each of the above-mentioned original question texts;
[0190] Step S203: Process all the above-mentioned original question texts and the above-mentioned synonymous sentences corresponding to all the above-mentioned original question texts by using a text enhancement method at the word level to obtain first target synonymous sentences corresponding to each of the above-mentioned original question texts. The text enhancement method at the word level includes removing stop words, keyword extraction, homophone replacement, synonym replacement, random word insertion, and random word deletion;
[0191] Step S204, screen each of the above first target synonymous sentences to obtain second target synonymous sentences corresponding to each of the above original question texts, where the second target synonymous sentences are sentences that have the same meaning as the above original question texts and are grammatically correct;
[0192] Step S205, add all the above second target synonymous sentences to the above Q&A corpus to bind the above second target synonymous sentences to the corresponding above original question texts.
[0193] Obviously, those skilled in the art should understand that each module or each step of the above-mentioned present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0194] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0195] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0196] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction means that implements the function specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the block or blocks.
[0197] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the block or blocks.
[0198] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0199] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory. Memory is an example of computer-readable media.
[0200] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0201] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0202] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0203] 1), A text data enhancement method for a FAQ system of the present application constructs a training set by using some original question texts in the Q&A corpus of the current FAQ system and synonymous sentences corresponding to the above-mentioned some original question texts, and trains an initial Simbert model by using the above training set to obtain a trained Simbert model, where the Simbert model is used for generating synonymous sentences; generates synonymous sentences for all the original question texts by using the above-trained Simbert model to obtain multiple above-mentioned synonymous sentences corresponding to each of the above original question texts; processes all the above original question texts and the above synonymous sentences corresponding to all the above original question texts by using a text enhancement method at the word level to obtain a first target synonymous sentence corresponding to each of the above original question texts, and the above text enhancement method at the word level includes removing stop words, keyword extraction, homophone replacement, synonym replacement, randomly inserting words and randomly deleting words; screens each of the above first target synonymous sentences to obtain a second target synonymous sentence corresponding to each of the above original question texts, and the above second target synonymous sentence is a sentence with the same meaning as the above original question text and correct grammar; adds all the above second target synonymous sentences to the above Q&A corpus to bind the above second target synonymous sentences to the corresponding above original question texts. The present application trains the model by constructing a training set of the Simbert model, generates synonymous sentences for the original question texts by using the trained Simbert model, then performs word-level processing on the original question texts and the synonymous sentences corresponding to the original question texts to obtain the first target synonymous sentences, and finally determines the second target synonymous sentences by screening and adds them to the Q&A corpus, thereby improving the ability of the FAQ system to answer questions, avoiding the poor Q&A effect of the FAQ system due to too few original question texts in the Q&A corpus, and solving the problem of difficult acquisition of original question texts in the prior art.
[0204] 2) A text data enhancement device for a FAQ system of the present application. The construction unit is used to construct a training set by using some original question texts in the Q&A corpus of the current FAQ system and synonymous sentences corresponding to the above-mentioned some original question texts, and apply the above training set to train an initial Simbert model to obtain a trained Simbert model, which is used for generating synonymous sentences; The first generation unit is used to generate synonymous sentences for all the original question texts by using the above-mentioned trained Simbert model to obtain multiple above-mentioned synonymous sentences corresponding to each of the above original question texts; The second generation unit is used to process all the above original question texts and the above synonymous sentences corresponding to all the above original question texts by using a text enhancement method at the word level to obtain a first target synonymous sentence corresponding to each of the above original question texts. The above text enhancement method at the word level includes removing stop words, keyword extraction, near-sound word replacement, synonym replacement, random insertion of words and random deletion of words; The screening unit is used to screen each of the above first target synonymous sentences to obtain a second target synonymous sentence corresponding to each of the above original question texts. The above second target synonymous sentence is a sentence with the same meaning as the above original question text and correct grammar; The addition unit is used to add all the above second target synonymous sentences to the above Q&A corpus to bind the above second target synonymous sentences to the corresponding above original question texts. The present application trains the model by constructing a training set of the Simbert model, uses the trained Simbert model to generate synonymous sentences for the original question texts, then performs word-level processing on the original question texts and the synonymous sentences corresponding to the original question texts to obtain the first target synonymous sentences, and finally determines the second target synonymous sentences through screening and adds them to the Q&A corpus, thereby improving the ability of the FAQ system to answer questions, avoiding the poor Q&A effect of the FAQ system due to too few original question texts in the Q&A corpus, and solving the problem of difficult acquisition of original question texts in the prior art.
[0205] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for text data enhancement of an FAQ system, characterized in that Including: Construct a training set by using a part of the original question texts in the Q&A corpus of the current FAQ system and the synonymous sentences corresponding to the part of the original question texts, and train the initial Simbert model with the training set to obtain a trained Simbert model, where the Simbert model is used for generating synonymous sentences; Use the trained Simbert model to generate synonymous sentences for all the original question texts, and obtain multiple synonymous sentences corresponding to each of the original question texts; Process all the original question texts and the synonymous sentences corresponding to all the original question texts by using a word-level text enhancement method to obtain first target synonymous sentences corresponding to each of the original question texts, and the word-level text enhancement method includes removing stop words, keyword extraction, near-sound word replacement, synonym replacement, randomly inserting words, and randomly deleting words; Screen each of the first target synonymous sentences to obtain second target synonymous sentences corresponding to each of the original question texts, and the second target synonymous sentences are sentences with the same meaning as the original question texts and correct grammar; Add all the second target synonymous sentences to the Q&A corpus to bind the second target synonymous sentences with the corresponding original question texts.
2. The method according to claim 1, wherein Constructing a training set by using a part of the original question texts in the Q&A corpus of the current FAQ system and the synonymous sentences corresponding to the part of the original question texts includes: Use an Embedding model to convert the part of the original question texts in the Q&A corpus of the current FAQ system into question vectors; In the first step, use the ANN algorithm to determine, among the question vectors, a similar question vector with the smallest Euclidean distance from the target question vector, where the similar question vector represents the vector with the closest semantic level to the target question vector, and the target question vector is any one of the question vectors; In the second step, form a similar vector pair with the target question vector and the similar question vector; In the first repeated step, repeat the first step and the second step at least once in sequence to obtain multiple similar vector pairs; Input the pairs of original question texts corresponding to the multiple similar vector pairs into the Simbert model to train the Simbert model to generate synonymous sentences that meet the requirements.
3. The method according to claim 2, wherein The training process of the Simbert model includes: Train the Simbert model with multiple similar vector pairs until the cross-entropy loss function and the contrastive loss function converge to obtain the trained Simbert model. The cross-entropy loss function is used to calculate the loss between the target question vector and the similar question vector in the similar vector pair, and the contrastive loss function is used to calculate the sum of the losses between the target question vector and other question vectors respectively, where the target question vector is any one of the question vectors.
4. The method according to claim 3, wherein Training the Simbert model with multiple similar vector pairs until the cross-entropy loss function and the contrastive loss function converge includes: The first processing step is to separate the target problem text corresponding to the target problem vector and the similar problem text corresponding to the similar problem vector in the pair of similar vectors by an identifier, so that the Simbert model can identify the target problem text and the similar problem text; The second processing step is to input the target problem text into the encoder of the Simbert model for vectorization processing to obtain a first vector, where the first vector is the vector representation of all the words in the target problem text; The third processing step is to input the first vector into the decoder of the Simbert model to generate a predicted vocabulary sequence, compare the predicted vocabulary sequence with the similar problem text, and calculate the loss value of the cross-entropy loss function, where the predicted vocabulary sequence is a text composed of words; The second repetition step is to repeat the first processing step, the second processing step, and the third processing step at least once in sequence to adjust the parameters of the Simbert model until the cross-entropy loss function converges.
5. The method according to claim 3, wherein When training the Simbert model with multiple pairs of similar vectors until the cross-entropy loss function and the contrast loss function converge, it further includes: Concatenate all the said problem vectors to obtain a sentence vector matrix D ∈ R b*d , where b is the total number of the said problem vectors, d is the dimension of the said problem vectors, and R indicates that the elements in the said sentence vector matrix are real numbers; Normalize all the problem vectors in the sentence vector matrix with the L2 regularization norm to obtain a regular matrix Perform an inner product calculation on the regular matrix to obtain a similarity matrix The elements in the similarity matrix represent the similarity between any two of the problem vectors; Determine the associated vector of the problem vector of one of the original problem texts in the pair of original problem texts as the problem vector of the other original problem text, determine the associated vector as the positive sample vector of the problem vector, and determine the other problem vectors except the associated vector as the negative sample vectors of the problem vector, where the associated vector is a vector that is semantically similar to the problem vector; Determine the similarity of each positive sample vector and each negative sample vector according to the similarity matrix; Substitute the similarities of each positive sample vector and each negative sample vector into the contrast loss function to calculate the loss value of the contrast loss function; Adjust the parameters of the Simbert model until the contrast loss function converges.
6. The method according to claim 5, characterized in that Before determining the similarity of each positive sample vector and each negative sample vector according to the similarity matrix, the method further includes: Perform a mask process on the diagonal elements of the similarity matrix, where the diagonal elements are the elements corresponding to the diagonal from the upper left corner to the lower right corner of the similarity matrix.
7. The method according to any one of claims 1 to 6, characterized in that, Add all the second target synonymous sentences to the Q&A corpus to bind the second target synonymous sentences to the corresponding original problem texts. The method further includes: When the input question of the FAQ system is the original problem text, use the answer corresponding to the original problem text as the output result of the FAQ system; When the input question of the FAQ system is the second target synonymous sentence corresponding to the original problem text, use the answer corresponding to the original problem text bound to the second target synonymous sentence as the output result of the FAQ system.
8. A text data enhancement device for an FAQ system, characterized in that, It includes: A construction unit is used to construct a training set by using a part of the original question texts in the Q&A corpus of the current FAQ system and the synonymous sentences corresponding to the part of the original question texts, and use the training set to train an initial Simbert model to obtain a trained Simbert model, where the Simbert model is used for generating synonymous sentences; A first generation unit is used to generate synonymous sentences for all the original question texts by using the trained Simbert model to obtain a plurality of the synonymous sentences corresponding to each of the original question texts; A second generation unit is used to process all the original question texts and the synonymous sentences corresponding to all the original question texts by using a word-level text enhancement method to obtain a first target synonymous sentence corresponding to each of the original question texts, and the word-level text enhancement method includes removing stop words, keyword extraction, near-sound word replacement, synonym replacement, randomly inserting words and randomly deleting words; A screening unit is used to screen each of the first target synonymous sentences to obtain a second target synonymous sentence corresponding to each of the original question texts, and the second target synonymous sentence is a sentence that has the same meaning as the original question text and is grammatically correct; An adding unit is used to add all the second target synonymous sentences to the Q&A corpus to bind the second target synonymous sentences to the corresponding original question texts.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 7.
10. A text data enhancement system, characterized in that, Comprising: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing the method according to any one of claims 1 to 7.
Citation Information
Cited By
FAQ knowledge base closed-loop optimization method based on million-level user message data
CN122064792A