Text screening method and device, electronic device and storage medium

By calculating the semantic similarity between data items in triplets and calculating consistency index, the accuracy and consistency problems when extracting question-and-answer pairs in unsupervised scenarios are solved, and automated checksum is achieved to improve the extraction quality.

CN119938888BActive Publication Date: 2025-06-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510418587.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-27
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

When extracting Q&A in an unsupervised scenario, there are problems such as inaccurate extraction results and incorrect answers, which leads to inconsistent quality of Q&A and requires manual inspection, which is time-consuming and labor-intensive and subjective.

Method used

By calculating the semantic similarity between data items in the triplet, compute the consistency index, and filtering the triplets based on the preset threshold, we obtain triplets with consistency index greater than the preset threshold.

Benefits of technology

It realizes automated verification of the quality of Q&A pairs, reduces manual intervention, improves the accuracy and consistency of Q&A pair extraction, and improves the reliability and practicality of Q&A pairs in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938888B_ABST
    Figure CN119938888B_ABST
Patent Text Reader

Abstract

The present application discloses a method and device for screening text, an electronic device, and a storage medium, relating to the technical field of data processing. The method includes calculating a consistency index based on the semantic similarity between data items in a triple, and screening the triple using a preset threshold. Therefore, it is possible to solve the technical problems such as inaccurate extraction results and wrong answers in the extraction of question-and-answer pairs in the prior art in an unsupervised scenario, and achieve the technical effects of automatically verifying the quality of question-and-answer pairs, reducing manual intervention, and improving the extraction accuracy and consistency of question-and-answer pairs. Through the quantitative evaluation of the consistency index, it is possible to effectively eliminate question-and-answer pairs with semantic inconsistencies, thereby improving the reliability and practicality of question-and-answer pairs in actual applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technologies, and in particular, to a method and device for screening texts, an electronic device, and a storage medium. Background Art

[0002] With the development of artificial intelligence technologies, especially the wide application of large language models (LLMs), automatically extracting question-and-answer pairs from text corpora has become an important research direction in the field of natural language processing.

[0003] Currently, when extracting question-and-answer pairs in an unsupervised scenario, problems such as inaccurate extraction results and incorrect answers are often faced, resulting in potential quality issues in the actual application of question-and-answer pairs. To verify the quality of the extracted question-and-answer pairs, full-scale manual inspection is usually required, which is not only time-consuming and laborious but also leads to inconsistent quality of the inspected question-and-answer pairs due to human subjectivity. Summary of the Invention

[0004] This application provides a method, device, electronic device, and storage medium for screening texts to at least solve the problem of inconsistent quality of inspected question-and-answer pairs caused by human subjectivity in related technologies.

[0005] This application provides a method for screening texts, including:

[0006] Obtaining a triple, and calculating a consistency index of the triple based on the semantic similarity between data items in the triple; wherein, the data items in the triple include the original corpus, the question, and the answer; the consistency index is used to quantify the semantic consistency of the data items in the triple;

[0007] Screening the triple based on a preset threshold and the consistency index to obtain a triple with a consistency index greater than the preset threshold.

[0008] This application further provides a device for screening texts, including:

[0009] An obtaining unit, configured to obtain a triple and calculate a consistency index of the triple based on the semantic similarity between data items in the triple; wherein, the data items in the triple include the original corpus, the question, and the answer; the consistency index is used to quantify the semantic consistency of the data items in the triple;

[0010] A screening unit, configured to screen the triple based on a preset threshold and the consistency index to obtain a triple with a consistency index greater than the preset threshold.

[0011] This application further provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above text screening methods when executing the computer program.

[0012] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned text screening methods are implemented.

[0013] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the above-mentioned text screening methods are implemented.

[0014] Through the present application, since a method for calculating a consistency index based on the semantic similarity between data items in a triple is adopted, and the triple is screened by using a preset threshold, it is possible to solve the technical problems such as inaccurate extraction results and wrong answers in the prior art when extracting question-and-answer pairs in an unsupervised scenario, and achieve the technical effects of automatically verifying the quality of question-and-answer pairs, reducing manual intervention, and improving the accuracy and consistency of question-and-answer pair extraction. Through the quantitative evaluation of the consistency index, it is possible to effectively eliminate question-and-answer pairs with inconsistent semantics, thereby improving the reliability and practicality of question-and-answer pairs in actual applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a schematic flowchart of a method for screening a text provided by an embodiment of the present application;

[0017] Figure 2 It is a schematic flowchart of a text screening process provided by the present application;

[0018] Figure 3 It is a schematic structural diagram of a text screening device provided by an embodiment of the present application;

[0019] Figure 4 It is a schematic structural diagram of a text screening device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0021] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and not to describe a specific order or sequence.

[0022] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.

[0023] An embodiment of this application provides a method for screening texts. In combination with the execution process of the text screening method, the method will be described in detail.

[0024] Figure 1 It is a schematic flowchart of a method for screening texts provided by an embodiment of this disclosure.

[0025] As Figure 1 shown, the method includes the following steps:

[0026] Step 101, obtain a triple, and calculate the consistency index of the triple based on the semantic similarity between the data items in the triple; wherein, the data items in the triple include the original corpus, the question, and the answer; the consistency index is used to quantify the semantic consistency of the data items in the triple;

[0027] The process of calculating the consistency index based on the semantic similarity between triples (original corpus, question, answer) involves converting each data item into a vector form. In some embodiments, a deep learning model or a pre-trained language model can be used to capture the semantic relationships between them, and then through calculation steps to determine the matching degree of these data items at the semantic level. This process includes, but is not limited to, performing detailed preprocessing on the text, such as word segmentation, stop word removal, part-of-speech tagging, and using advanced semantic encoding techniques to extract features, and then combining the similarity scores between the original corpus and the question, the original corpus and the answer, and the question and the answer through weighted average or other fusion strategies to obtain a comprehensive consistency index.

[0028] The consistency index not only considers the direct similarity between individual data items but also their relative importance in the overall context. Through comparison and normalization processing, a quantitative index that can reflect the semantic consistency between the data items in the triple is finally obtained. This index plays an important role in evaluating the accuracy of the question-and-answer system and improving the answer generation strategy.

[0029] Step 102: Filter the triples based on a preset threshold and the consistency index to obtain the triples with a consistency index greater than the preset threshold.

[0030] After calculating the consistency index for each triple, filter the triples according to a preset threshold to identify and retain those combinations of data items with high semantic consistency. Specifically, first, set a threshold for the consistency index, which is determined based on pre-experiments on the performance requirements of the question-and-answer system and represents the acceptable minimum level of semantic consistency. Then, iterate through all the triples for which the consistency index has been calculated and compare the consistency index of each triple with the preset threshold. Select the triples whose consistency index is greater than or equal to the preset threshold. These triples are considered to be semantically more consistent and reliable, ensuring that the information provided to the user is based on highly consistent and trustworthy data sources, thereby enhancing the overall user experience and the reliability of the system.

[0031] In some embodiments, before calculating the consistency index of the triples based on the semantic similarity between the data items in the triples, the following steps are further included:

[0032] Perform embedding processing on the triples to obtain a first semantic vector corresponding to the original corpus, a second semantic vector corresponding to the question, and a third semantic vector corresponding to the answer;

[0033] Please refer to Figure 2 , Figure 2 which is a schematic diagram of a screening process for a text provided by this application; as Figure 2 shown, perform embedding on the triples to obtain a vector representation; the embedding representation is to convert the original corpus, question, and answer in the triples into semantic vector representations. In some embodiments, it can be implemented through pre-trained language models such as BERT, Word2Vec, GloVe, etc. Specifically, the embodiments of this application do not limit the way of embedding representation.

[0034] Perform embedding processing on the original corpus to obtain a first semantic vector.

[0035] Perform embedding processing on the question to obtain a second semantic vector.

[0036] Perform embedding processing on the answer to obtain a third semantic vector.

[0037] Calculating the consistency index of the triples based on the semantic similarity between the data items in the triples includes:

[0038] Calculate the semantic similarity between every two data items in the triples based on the semantic vectors to generate a semantic similarity array;

[0039] Calculate the consistency metric according to the semantic similarity array.

[0040] In some embodiments, when calculating semantic similarity, cosine similarity, Euclidean distance, or other similarity measurement methods can be used to calculate the semantic similarity between any two data items in the triple. The embodiments of the present application do not limit this.

[0041] Calculate the semantic similarity between the original corpus (the first semantic vector) and the question (the second semantic vector); calculate the semantic similarity between the original corpus (the first semantic vector) and the answer (the third semantic vector); calculate the semantic similarity between the question (the second semantic vector) and the answer (the third semantic vector).

[0042] Please continue to refer to Figure 2 , and form a semantic similarity array with the calculated similarity values. For example, in the form: [(original corpus, question), (original corpus, answer), (question, answer)].

[0043] In some embodiments, calculating the semantic similarity between any two data items in the triple based on the semantic vector includes:

[0044] Calculate the first semantic similarity between the first semantic vector and the second semantic vector;

[0045] Calculate the second semantic similarity between the second semantic vector and the third semantic vector;

[0046] Calculate the third semantic similarity between the first semantic vector and the third semantic vector;

[0047] Obtain a semantic similarity array according to the first semantic similarity, the second semantic similarity, and the third semantic similarity. In some embodiments, the semantic similarity array can be represented as (S1, S2, S3); where S1 is the first semantic similarity, S2 is the second semantic similarity, and S3 is the third semantic similarity.

[0048] Calculate the first semantic similarity: Calculate the semantic similarity between the original corpus (the first semantic vector) and the question (the second semantic vector).

[0049] Calculate the second semantic similarity: This refers to calculating the semantic similarity between the question (the second semantic vector) and the answer (the third semantic vector).

[0050] Calculate the third semantic similarity: This refers to calculating the semantic similarity between the original corpus (the first semantic vector) and the answer (the third semantic vector).

[0051] In some embodiments, a similarity measurement method (such as cosine similarity) can be used to calculate the similarity between two vectors, and the calculation result is recorded as the first semantic similarity.

[0052] Combine the first semantic similarity, the second semantic similarity, and the third semantic similarity calculated above into an array. The form of the array can be [the first semantic similarity, the second semantic similarity, the third semantic similarity].

[0053] In some embodiments, calculating the consistency index based on the semantic similarity array includes:

[0054] Calculate the range factor, the deviation factor, and the coverage factor according to the semantic similarity array;

[0055] Please continue to refer to Figure 2 , calculate three consistency factors. In some embodiments, when calculating the range factor, the deviation factor, and the coverage factor, the calculation can be performed according to the following steps:

[0056] Based on the semantic similarity array, calculate the range factor using the maximum semantic similarity and the minimum semantic similarity;

[0057] In some implementations, when calculating the range factor, the following steps can also be performed:

[0058] Divide the maximum eigenvalue by the minimum eigenvalue, and then square the result to calculate the range factor.

[0059] First, extract the maximum and minimum semantic similarities from the semantic similarity array. The range factor is obtained by squaring the ratio of the maximum semantic similarity to the minimum semantic similarity. Specifically, divide the maximum semantic similarity by the minimum semantic similarity, and then square the result to calculate the range factor. The calculation formula is:

[0060]

[0061] where κ(S) is the range factor

[0062] The range factor is used to represent the overall stability of the semantic similarity array. The smaller its value, the more stable the distribution of semantic similarities in the array.

[0063] In some implementations, the calculation of the range factor can also be based on eigenvalues. Specifically, regard the semantic similarity array as a matrix, and calculate its maximum eigenvalue and minimum eigenvalue. Then divide the maximum eigenvalue by the minimum eigenvalue, and square the result to calculate the range factor. Eigenvalues are important concepts in matrix analysis, representing the scaling ratio of the matrix in a specific direction. By calculating eigenvalues, the stability of the semantic similarity array can be more accurately reflected.

[0064] Through the above steps, the calculation of the range factor is completed, providing basic data support for the subsequent calculation of the consistency index.

[0065] Calculate the deviation factor based on the minimum semantic similarity value in the semantic similarity array;

[0066] In some implementations, when calculating the deviation factor, the following steps can also be performed: Use the difference between the third constant and the minimum semantic similarity value as the exponent of a variable positive number to calculate the deviation factor; where the variable constant is a positive number greater than 1.

[0067] Extract the minimum semantic similarity value from the semantic similarity array. The deviation factor is calculated by the difference between a variable positive number greater than 1 and the minimum semantic similarity value. Specifically, multiply the variable positive number by 1 minus the minimum semantic similarity value to obtain the deviation factor, and the calculation formula is:

[0068]

[0069] where δ(S) is the deviation factor.

[0070] The deviation factor is used to penalize low similarity values in the semantic similarity array. The larger its value, the lower the semantic similarity in the array, which has a negative impact on the semantic consistency of the triples.

[0071] In some implementations, the calculation of the deviation factor can also be based on an exponential function. Specifically, use the difference between the third constant and the minimum semantic similarity value as the exponent of a variable positive number to calculate the deviation factor. Among them, the third constant is a preset fixed value, and the variable positive number is a positive number greater than 1. Through the calculation of the exponential function, the influence of low similarity values can be further amplified, so as to more significantly penalize the low values in the semantic similarity array.

[0072] Through the above steps, the calculation of the deviation factor is completed, providing basic data support for the subsequent calculation of the consistency index.

[0073] Calculate the coverage factor based on the third semantic similarity and the first semantic similarity in the semantic similarity array.

[0074] In some implementations, when calculating the coverage factor, the following steps can also be performed:

[0075] Use the difference between the third semantic similarity and the first semantic similarity as the exponent of the natural constant, and then subtract the fourth constant to calculate the coverage factor.

[0076] Extract the third semantic similarity and the first semantic similarity from the semantic similarity array. The coverage factor is calculated by the difference between the third semantic similarity and the first semantic similarity. Specifically, subtract the first semantic similarity from the third semantic similarity to obtain the difference, then use this difference as the exponent of the natural constant (i.e., the base e of the natural logarithm), and then subtract the fourth constant to calculate the coverage factor, and the calculation formula is:

[0077]

[0078] Among them, C is the coverage factor, and e is the base of the natural logarithm. S3 is the third semantic similarity between the original corpus (the first semantic vector) and the answer (the third semantic vector); S1 is the first semantic similarity between the original corpus (the first semantic vector) and the question (the second semantic vector).

[0079] When the order of the semantic vectors in the semantic similarity array is different, the calculation formula of the coverage factor needs to be adjusted according to the order of the semantic vectors to ensure that the exponent of the natural constant is the difference between the semantic similarity between the original corpus and the answer and the semantic similarity between the original corpus and the question.

[0080] The coverage factor is used to ensure that the semantic similarity between the answer and the original corpus is more consistent. When its value is positive, it indicates a high semantic consistency between the answer and the original corpus; when it is negative, it indicates a semantic deviation.

[0081] In some implementations, the calculation of the coverage factor can be further optimized. Specifically, the difference between the third semantic similarity and the first semantic similarity is used as the exponent of the natural constant, and then a fourth constant is subtracted to calculate the coverage factor. Among them, the fourth constant is a preset fixed value used to adjust the benchmark of the coverage factor. Through the exponential operation of the natural constant, the change of the semantic similarity difference can be more sensitively reflected, so as to more accurately evaluate the consistency between the answer and the original corpus.

[0082] Through the above steps, the calculation of the coverage factor is completed, providing basic data support for the calculation of subsequent consistency indicators.

[0083] Calculate the consistency indicator based on the range factor, deviation factor and coverage factor.

[0084] In some embodiments, the calculation of the meander-shaped consistency indicator can also be performed according to the following steps:

[0085] Multiply the first weight coefficient by the sum of the reciprocals of the range factor and the deviation factor, and then add the product of the second weight coefficient and the coverage factor to calculate the consistency indicator.

[0086] Obtain the calculation results of the range factor, deviation factor and coverage factor. The consistency indicator is calculated by multiplying the first weight coefficient by the sum of the reciprocals of the range factor and the deviation factor, and then adding the product of the second weight coefficient and the coverage factor. Specifically, the calculation formula of the consistency indicator is:

[0087] CI = α·(1 / κ(S) + 1 / δ(S)) + β·C

[0088] Among them, CI is the consistency index, α is the first constant, β is the second constant, and the first constant and the second constant can be equal.

[0089] Through this formula, the range factor, deviation factor, and coverage factor are combined to quantify the semantic consistency of data items in the triple.

[0090] In some embodiments, the calculation of the consistency index can be further optimized. Specifically, the first weight coefficient is multiplied by the sum of the reciprocals of the range factor and the deviation factor, and then added to the product of the second weight coefficient and the coverage factor to calculate the consistency index. Among them, the first weight coefficient and the second weight coefficient are used to adjust the contribution ratio of the range factor, deviation factor, and coverage factor in the consistency index. By adjusting the weights, it can more flexibly adapt to the semantic consistency evaluation requirements in different scenarios.

[0091] Through the above steps, the calculation of the consistency index is completed, providing basic data support for the subsequent quality verification of question and answer extraction.

[0092] In some embodiments, based on a preset threshold and the consistency index, the triples are screened, and the triples with a consistency index greater than the preset threshold include:

[0093] When the consistency index is lower than the preset threshold, the triple is determined to be invalid text.

[0094] The preset threshold is used to determine whether the semantic consistency of the triple meets the requirements. The calculated consistency index is compared with the preset threshold. If the consistency index is greater than or equal to the preset threshold, the triple is determined to be valid text and retained for the subsequent question and answer extraction task. If the consistency index is lower than the preset threshold, the triple is determined to be invalid text and excluded. Through this screening process, triples with low semantic consistency can be effectively excluded, thereby improving the accuracy and reliability of question and answer extraction.

[0095] In some embodiments, the preset threshold is proportional to the first constant and the second constant. The setting of the preset threshold can be adjusted according to the specific application scenario, requirements, and the numerical values of the first constant and the second constant. For example, in a scenario with higher requirements, a higher threshold can be set to screen out triples with higher semantic consistency; while in a scenario with relatively loose requirements, the threshold can be appropriately lowered to retain more triples; when the first constant and the second constant are set larger, the preset threshold will be synchronously increased; when the first constant and the second constant are set smaller, the preset threshold will be synchronously decreased. By combining the preset threshold with the consistency index, the automatic screening of triples is achieved, providing strong support for the improvement of the quality of question and answer extraction.

[0096] The following uses two embodiments to clearly illustrate the method for screening texts provided by this application.

[0097] The following is illustrated with a positive example. Embodiment 1: Suppose there is the following ABC triple:

[0098] A: Text corpus (original input corpus): "Due to the sharp increase in greenhouse gas emissions, climate change has occurred, and the global temperature is rising."

[0099] B: Question (question extracted by a third-party system such as a large language model) "Why is the global temperature rising?"

[0100] C: Answer (answer extracted or organized by a third-party system such as a large language model): "The global temperature is rising due to the sharp increase in greenhouse gas emissions."

[0101] First, use a general embedding model to embed the triple to obtain the corresponding vector encoding, and then calculate the cosine similarity in turn to obtain:

[0102] Similarity array S = [0.708, 0.743, 0.948]

[0103] Then calculate the three factors required by the consistency index formula:

[0104] Range factor: ≈ 1.7928

[0105] Deviation factor (take K = 20): ≈ 2.3982

[0106] Coverage factor: C = 2.718 ^ (0.948 - 0.743 - 1) ≈ 0.2275, where e is taken as the approximate value 2.718

[0107] Finally, apply the consistency index calculation formula, and set the weights to 1:

[0108] CI = 1·(1 / κ(S) + 1 / δ(S)) + 1·C

[0109] Substitute the factor values, so:

[0110] CI = (1 / 1.7928 + 1 / 2.3982) + 0.2275 = 1.2023

[0111] The CI index is greater than 1, indicating that this Q&A triple has significant rationality and is a high-quality Q&A extraction result.

[0112] The following is illustrated with a negative example. Embodiment 2: Suppose there is the following ABC triple:

[0113] A: Text corpus (original input corpus): "Due to climate change, the global temperature is rising."

[0114] B: Question (question extracted by a third-party system such as a large language model): "What is the weather condition in Area A?"

[0115] C: Answer (answer extracted or organized by a third-party system such as a large language model): "Quantum computers use qubits to perform complex computational tasks."

[0116] First, use a general embedding model to embed the triples to obtain the corresponding vector encodings, and then calculate the cosine similarity in sequence to get:

[0117] Similarity array S = [0.368, 0.175, 0.350]

[0118] Then calculate the three factors required for the consistency index formula:

[0119] Range factor: ≈ 4.4220

[0120] Deviation factor (taking K = 20): ≈ 11.8399

[0121] Coverage factor: C = 2.718 ^ (0.350 - 0.175 - 1) ≈ 0.1912, where e is approximated as 2.718

[0122] Finally, apply the consistency index calculation formula, setting the weights to 1:

[0123] CI = 1 · (1 / κ(S) + 1 / δ(S)) - 1 + 1 · C

[0124] Substitute the factor values, so:

[0125] CI = (1 / 4.4220 + 1 / 11.8399) + 0.1912 = 0.5018

[0126] The CI index is significantly less than 1, indicating that the Q&A triple has a large semantic inconsistency and is an unqualified Q&A extraction result.

[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0128] The embodiments of the present application also provide a text screening device. Figure 3The following is a schematic structural diagram of a text screening device provided by an embodiment of the present application. As shown in Figure 3 the figure, it includes:

[0129] An acquisition unit 21, configured to acquire a triple, and calculate a consistency index of the triple based on the semantic similarity between data items in the triple; wherein, the data items in the triple include the original corpus, the question, and the answer; the consistency index is used to quantify the semantic consistency of the data items in the triple;

[0130] A screening unit 22, configured to screen the triple based on a preset threshold and the consistency index, and obtain a triple whose consistency index is greater than the preset threshold.

[0131] Further, in a possible implementation manner of an embodiment of the present application, as shown in Figure 4, the device further includes:

[0132] A processing unit 23, configured to perform embedding processing on the triple before the acquisition unit 21 calculates the consistency index of the triple based on the semantic similarity between data items in the triple, so as to obtain a first semantic vector corresponding to the original corpus, a second semantic vector corresponding to the question, and a third semantic vector corresponding to the answer;

[0133] The acquisition unit 21 is further configured to:

[0134] Calculate the semantic similarity between pairwise data items in the triple based on the semantic vectors, and generate a semantic similarity array;

[0135] Calculate a consistency index according to the semantic similarity array.

[0136] Further, in a possible implementation manner of an embodiment of the present application, the acquisition unit 21 is further configured to:

[0137] Calculate a first semantic similarity between the first semantic vector and the second semantic vector;

[0138] Calculate a second semantic similarity between the second semantic vector and the third semantic vector;

[0139] Calculate a third semantic similarity between the first semantic vector and the third semantic vector;

[0140] Obtain a semantic similarity array according to the first semantic similarity, the second semantic similarity, and the third semantic similarity.

[0141] Further, in a possible implementation manner of an embodiment of the present application, the acquisition unit 21 is further configured to:

[0142] Calculate a range factor, a deviation factor, and a coverage factor according to the semantic similarity array;

[0143] Calculate a consistency index based on the range factor, the deviation factor, and the coverage factor.

[0144] Further, in a possible implementation manner of the embodiment of the present application, the obtaining unit 21 is further configured to:

[0145] Calculate a range factor based on the maximum semantic similarity value and the minimum semantic similarity value in the semantic similarity array;

[0146] Calculate a deviation factor based on the minimum semantic similarity value in the semantic similarity array;

[0147] Calculate a coverage factor based on the third semantic similarity and the first semantic similarity in the semantic similarity array.

[0148] Further, in a possible implementation manner of the embodiment of the present application, the calculation unit is further configured to:

[0149] Multiply the first weight coefficient by the sum of the reciprocals of the range factor and calculate the consistency index by adding the product of the second weight coefficient and the coverage factor.

[0150] Further, in a possible implementation manner of the embodiment of the present application, the obtaining unit 21 is further configured to:

[0151] Divide the maximum eigenvalue by the minimum eigenvalue and square the result to calculate the range factor.

[0152] Further, in a possible implementation manner of the embodiment of the present application, the obtaining unit 21 is further configured to:

[0153] Use the difference between the third constant and the minimum semantic similarity as the exponent of a variable positive number to calculate the deviation factor; where the variable constant is a positive number greater than 1.

[0154] Further, in a possible implementation manner of the embodiment of the present application, the obtaining unit 21 is further configured to:

[0155] Use the difference between the third semantic similarity and the first semantic similarity as the exponent of the natural constant and subtract the fourth constant to calculate the coverage factor.

[0156] Further, in a possible implementation manner of the embodiment of the present application, the preset threshold is proportional to the first constant and the second constant.

[0157] Further, in a possible implementation manner of the embodiment of the present application, the screening unit 22 is further configured to:

[0158] When the consistency index is lower than the preset threshold, it is determined that the triple is invalid text.

[0159] For the description of the features in the corresponding embodiment of the text screening device, reference may be made to the relevant description in the corresponding embodiment of the text screening method, which will not be elaborated here one by one.

[0160] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-described embodiments of the text screening method.

[0161] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described embodiments of the text screening method when running.

[0162] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc and other media that can store computer programs.

[0163] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and the steps in any of the above-described embodiments of the text screening method are implemented when the computer program is executed by a processor.

[0164] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and the steps in any of the above-described embodiments of the text screening method are implemented when the computer program is executed by a processor.

[0165] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0166] The above has introduced in detail a method and device for screening a kind of text, an electronic device, and a storage medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A text screening method, characterized in that: include: Acquire a triple, and calculate a consistency index of the triple based on the semantic similarity between the data items in the triple; wherein the data items in the triple include the original corpus, the question and the answer; and the consistency index is used to quantify the semantic consistency of the data items in the triple; Screening the triples based on a preset threshold and the consistency index to obtain the triples whose consistency index is greater than the preset threshold; Before calculating the consistency index of the triples based on the semantic similarity between the data items in the triples, the method further includes: Embedding the triples to obtain a first semantic vector corresponding to the original corpus, a second semantic vector corresponding to the question, and a third semantic vector corresponding to the answer; The calculating the consistency index of the triplet based on the semantic similarity between the data items in the triplet comprises: Based on the semantic vector, the semantic similarity between the two data items in the triple is calculated to generate a semantic similarity array; Calculating a consistency index according to the semantic similarity array; The step of calculating the consistency index according to the semantic similarity array includes: Calculate the extreme difference factor, deviation factor and coverage factor according to the semantic similarity array; Calculate the consistency index based on the range factor, the deviation factor and the coverage factor; The step of calculating the range factor, the deviation factor and the coverage factor according to the semantic similarity array includes: Calculate the range factor based on the maximum semantic similarity value and the minimum semantic similarity value in the semantic similarity array; Calculate the deviation factor based on the minimum semantic similarity value in the semantic similarity array; The coverage factor is calculated based on the third semantic similarity and the first semantic similarity in the semantic similarity array.

2. The text screening method according to claim 1, characterized in that: The calculating the semantic similarity between any two data items in the triple based on the semantic vector comprises: Calculating a first semantic similarity between the first semantic vector and the second semantic vector; Calculating a second semantic similarity between the second semantic vector and the third semantic vector; Calculating a third semantic similarity between the first semantic vector and the third semantic vector; The semantic similarity array is obtained according to the first semantic similarity, the second semantic similarity and the third semantic similarity.

3. The text screening method according to claim 1, characterized in that: The calculation of the consistency index based on the range factor, the range factor and the coverage factor includes: The consistency index is calculated by multiplying the first weight coefficient by the reciprocal of the range factor and the sum of the reciprocal of the range factor, and then adding the product of the second weight coefficient and the coverage factor.

4. The text screening method according to claim 1, characterized in that: The calculating of the range factor based on the maximum semantic similarity value and the minimum semantic similarity value in the semantic similarity array includes: The range factor is calculated by dividing the maximum eigenvalue by the minimum eigenvalue and squaring the result.

5. The text screening method according to claim 1, characterized in that: The calculating of the deviation factor based on the minimum semantic similarity value in the semantic similarity array comprises: The deviation factor is obtained by calculating the difference between the third constant and the minimum value of the semantic similarity as the exponent of a variable positive number; wherein the variable constant is a positive number greater than 1.

6. The text screening method according to claim 1, characterized in that: The calculation of the coverage factor based on the third semantic similarity and the first semantic similarity in the semantic similarity array includes: The coverage factor is calculated by taking the difference between the third semantic similarity and the first semantic similarity as the exponent of the natural constant and subtracting a fourth constant.

7. The text screening method according to claim 3, characterized in that: The preset threshold is proportional to the first constant and the second constant.

8. The method for screening text according to any one of claims 1 to 7, characterized in that: The screening of the triples based on the preset threshold and the consistency index to obtain the triples whose consistency index is greater than the preset threshold comprises: When the consistency index is lower than a preset threshold, the triple is determined to be invalid text.

9. A text screening device, characterized in that: include: An acquisition unit, used to acquire a triple, and calculate a consistency index of the triple based on the semantic similarity between the data items in the triple; wherein the data items in the triple include the original corpus, the question and the answer; and the consistency index is used to quantify the semantic consistency of the data items in the triple; A screening unit, configured to screen the triples based on a preset threshold and the consistency index to obtain the triples whose consistency index is greater than the preset threshold; Wherein, the device further comprises: a processing unit, configured to perform embedding processing on the triples to obtain a first semantic vector corresponding to the original corpus, a second semantic vector corresponding to the question, and a third semantic vector corresponding to the answer before the obtaining unit calculates the consistency index of the triples based on the semantic similarity between the data items in the triples; The acquisition unit is also used for: Based on the semantic vector, the semantic similarity between the two data items in the triple is calculated to generate a semantic similarity array; Calculating a consistency index according to the semantic similarity array; Wherein, the acquisition unit is also used for: Calculate the extreme difference factor, deviation factor and coverage factor according to the semantic similarity array; Calculate the consistency index based on the range factor, the deviation factor and the coverage factor; Wherein, the acquisition unit is further used for: Calculate the range factor based on the maximum semantic similarity value and the minimum semantic similarity value in the semantic similarity array; Calculate the deviation factor based on the minimum semantic similarity value in the semantic similarity array; The coverage factor is calculated based on the third semantic similarity and the first semantic similarity in the semantic similarity array.

10. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the text screening method as claimed in any one of claims 1 to 8 when executing the computer program.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the text screening method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the text screening method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Medical question and answer pair quality inspection method and device, computer equipment and storage medium

    CN111444724A

  • Intelligent question answering system construction method and system based on LLM large language model

    CN119691140A