Text screening method and device, electronic equipment and storage medium
By calculating the semantic similarity between data items in triplets and filtering triplets with consistency index greater than the threshold, the accuracy and consistency problems when extracting question-and-answer pairs in unsupervised scenarios are solved, and the effect of automated checksums to improve the extraction quality is achieved.
Patent Information
- Application Number
- CN202510418587.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
When extracting Q&A pairs in unsupervised scenarios, there are problems such as inaccurate extraction results and incorrect answers in the prior art, resulting in inconsistent quality of Q&A pairs, which requires manual inspection, which is time-consuming and labor-intensive and subjective differences.
By obtaining the semantic similarity between data items (original corpus, questions, and answers) in the triplet, compute the consistency index, and filter the triplets based on the preset threshold, we obtain triplets with consistency index greater than the threshold.
It realizes automated verification of the quality of Q&A pairs, reduces manual intervention, improves the accuracy and consistency of the extraction of Q&A pairs, and improves the reliability and practicality of Q&A pairs in practical applications.
Smart Images

Figure CN119938888A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a text screening method and device, electronic equipment and storage medium. Background Technology
[0002] With the development of artificial intelligence technology, especially the widespread application of large-scale language models (LLMs), automatically extracting question-answer pairs from text corpora has become an important research direction in the field of natural language processing.
[0003] Currently, when extracting question-answer pairs in unsupervised scenarios, problems such as inaccurate extraction results and wrong answers are often encountered, resulting in quality risks in practical applications of question-answer pairs. In order to verify the quality of the extracted question-answer pairs, full manual inspection is usually required, which is not only time-consuming and laborious, but also leads to inconsistent quality of the question-answer pairs after inspection due to manual subjectivity. SUMMARY OF THE INVENTION
[0004] This application provides a text screening method, device, electronic device and storage medium to at least solve the problem in the related art that human subjectivity leads to inconsistent quality of question and answer pairs after inspection.
[0005] This application provides a text screening method, including: Get triples, and calculate the consistency index of triples based on the semantic similarity between the data items in the triples; the data items in the triples include the original corpus, questions and answers; the consistency index is used to quantify the semantic consistency of the data items in the triples; Based on the preset threshold and consistency index, the triples are screened to obtain the triples whose consistency index is greater than the preset threshold.
[0006] The present application also provides a text screening device, including: The acquisition unit is used to acquire triples and calculate the consistency index of the triples based on the semantic similarity between the data items in the triples; wherein the data items in the triples include the original corpus, the question and the answer; the consistency index is used to quantify the semantic consistency of the data items in the triples; The screening unit is used to screen the triples based on the preset threshold and consistency index to obtain the triples whose consistency index is greater than the preset threshold.
[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; a processor for implementing the steps of any of the above-mentioned text screening methods when executing the computer program.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned text screening methods are implemented.
[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned text screening methods when executed by a processor.
[0010] Through this application, since a method of calculating the consistency index based on the semantic similarity between data items in a triple is adopted, and the triples are screened using a preset threshold, the technical problems of inaccurate extraction results and wrong answers in the extraction of question-answer pairs in unsupervised scenarios in the prior art can be solved, and the technical effects of automatically verifying the quality of question-answer pairs, reducing manual intervention, and improving the accuracy and consistency of question-answer pair extraction can be achieved. Through the quantitative evaluation of the consistency index, semantically inconsistent question-answer pairs can be effectively eliminated, thereby improving the reliability and practicality of question-answer pairs in practical applications. Brief Description of the Figures
[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0012] Figure 1 A flow chart of a text screening method provided in an embodiment of the present application; Figure 2 A schematic diagram of a text screening process provided in this application; Figure 3 A schematic diagram of the structure of a text screening device provided in an embodiment of the present application; Figure 4 This is a schematic diagram of the structure of a text screening device provided in an embodiment of the present application. Specific implementation method
[0013] The following will combine the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of this application.
[0014] It should be noted that in the description of this application, the terms "include", "comprises" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second" and the like in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0015] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0016] The embodiment of the present application provides a text screening method, and the method is described in detail in combination with the execution process of the text screening method.
[0017] Figure 1 This is a flowchart of a text screening method provided by an embodiment of the present disclosure.
[0018] If Figure 1 As shown, the method comprises the following steps: Step 101, obtaining triples, and calculating the consistency index of the triples based on the semantic similarity between the data items in the triples; wherein the data items in the triples include the original corpus, the question and the answer; the consistency index is used to quantify the semantic consistency of the data items in the triples; The process of calculating the consistency index based on the semantic similarity between triples (original corpus, question, answer) involves converting each data item into a vector form. In some embodiments, the semantic relationship between them can be captured by a deep learning model or a pre-trained language model, and then the degree of matching of these data items at the semantic level is determined through calculation steps. This process includes but is not limited to detailed preprocessing of the text, such as word segmentation, stop word removal, part-of-speech tagging, and the use of advanced semantic encoding technology to extract features, and then combining the similarity scores between the original corpus and the question, the original corpus and the answer, and the question and the answer through weighted averaging or other fusion strategies to derive a comprehensive consistency index.
[0019] The consistency index not only considers the direct similarity between the data items, but also their relative importance in the overall context. Through comparison and normalization, a quantitative index that can reflect the semantic consistency between the data items in the triple is finally obtained. This index plays an important role in evaluating the accuracy of the question-answering system and improving the answer generation strategy.
[0020] Step 102, screening the triples based on the preset threshold and consistency index, and obtaining the triples whose consistency index is greater than the preset threshold.
[0021] After calculating the consistency index of each triple, the triples are screened according to the preset threshold to identify and retain those data item combinations with high semantic consistency. Specifically, first, a consistency index threshold is set. The threshold is determined based on the preliminary experiment of the performance requirements of the question-answering system and represents the minimum acceptable level of semantic consistency; then, all triples with calculated consistency indexes are traversed, and the consistency index of each triple is compared with the preset threshold; the triples with a consistency index greater than or equal to the preset threshold are screened out. These triples are considered to be semantically more consistent and reliable, ensuring that the information provided to users is based on highly consistent and reliable data sources, thereby improving the overall user experience and system reliability.
[0022] In some embodiments, before calculating the consistency index of the triples based on the semantic similarity between the data items in the triples, the following steps are also included: Embed the triples to obtain the first semantic vector corresponding to the original corpus, the second semantic vector corresponding to the question, and the third semantic vector corresponding to the answer; Please refer to Figure 2 , Figure 2 is a schematic diagram of a text screening process provided by this application; such as Figure 2 As shown in the figure, the triples are embedded to obtain vector representation; the embedding representation is to convert the original corpus, questions and answers in the triples into semantic vector representation. In some embodiments, it can be implemented by pre-trained language models such as BERT, Word2Vec, GloVe, etc. The specific embodiment of this application does not limit the way of embedding representation.
[0023] Embed the original corpus to obtain the first semantic vector.
[0024] Embed the question to get the second semantic vector.
[0025] Embed the answer to obtain the third semantic vector.
[0026] Based on the semantic similarity between the data items in the triple, the consistency index of the triple is calculated including: Based on the semantic vector, calculate the semantic similarity between the two data items in the triple and generate a semantic similarity array; Calculate the consistency index based on the semantic similarity array.
[0027] In some embodiments, when calculating semantic similarity, cosine similarity, Euclidean distance or other similarity measurement methods can be used to calculate the semantic similarity between two data items in a triple. This embodiment of the present application is not limited to this.
[0028] Calculate the semantic similarity between the original corpus (first semantic vector) and the question (second semantic vector); calculate the semantic similarity between the original corpus (first semantic vector) and the answer (third semantic vector); calculate the semantic similarity between the question (second semantic vector) and the answer (third semantic vector).
[0029] Please continue reading Figure 2 , the similarity values calculated above are combined into a semantic similarity array, for example, in the form of: [(original corpus, question), (original corpus, answer), (question, answer)].
[0030] In some embodiments, based on the semantic vector, calculating the semantic similarity between two data items in the triple includes: Calculate the first semantic similarity between the first semantic vector and the second semantic vector; Calculate the second semantic similarity between the second semantic vector and the third semantic vector; Calculate the third semantic similarity between the first semantic vector and the third semantic vector; According to the first semantic similarity, the second semantic similarity and the third semantic similarity, a semantic similarity array is obtained. In some embodiments, the semantic similarity array can be represented as (S1, S2, S3); wherein S1 is the first semantic similarity, S2 is the second semantic similarity, and S3 is the third semantic similarity.
[0031] Calculate the first semantic similarity: Calculate the semantic similarity between the original corpus (first semantic vector) and the question (second semantic vector).
[0032] Calculate the second semantic similarity: This refers to calculating the semantic similarity between the question (second semantic vector) and the answer (third semantic vector).
[0033] Calculate the third semantic similarity: This refers to calculating the semantic similarity between the original corpus (first semantic vector) and the answer (third semantic vector).
[0034] In some embodiments, a similarity measurement method (such as cosine similarity) may be used to calculate the similarity between two vectors, and the calculation result is recorded as the first semantic similarity.
[0035] Combine the first semantic similarity, the second semantic similarity and the third semantic similarity calculated above into an array. The array may be in the form of [first semantic similarity, second semantic similarity, third semantic similarity].
[0036] In some embodiments, calculating the consistency index according to the semantic similarity array includes: Calculate the extreme difference factor, deviation factor and coverage factor based on the semantic similarity array; Please continue reading Figure 3 , calculate three consistency factors. In some embodiments, when calculating the range factor, deviation factor and coverage factor, the calculation can be performed according to the following steps: Calculate the range factor based on the maximum semantic similarity and the minimum semantic similarity in the semantic similarity array; In some implementations, when calculating the range factor, the following steps may also be performed: Divide the maximum eigenvalue by the minimum eigenvalue, then square the result to calculate the range factor.
[0037] First, extract the maximum and minimum values of semantic similarity from the semantic similarity array. The range factor is obtained by the square of the ratio of the maximum semantic similarity to the minimum semantic similarity. Specifically, divide the maximum semantic similarity by the minimum semantic similarity, then square the result to calculate the range factor. The calculation formula is:
[0038] where κ(S) is the range factor The range factor is used to indicate the overall stability of the semantic similarity array. The smaller its value, the more stable the semantic similarity distribution in the array.
[0039] In some implementations, the calculation of the range factor can also be performed based on the eigenvalue. Specifically, the semantic similarity array is regarded as a matrix, and its maximum eigenvalue and minimum eigenvalue are calculated. Then the maximum eigenvalue is divided by the minimum eigenvalue, and the result is squared to calculate the range factor. Eigenvalue is an important concept in matrix analysis, which indicates the scaling ratio of the matrix in a specific direction. By calculating the eigenvalue, the stability of the semantic similarity array can be more accurately reflected.
[0040] Through the above steps, the calculation of the range factor is completed, providing basic data support for the calculation of subsequent consistency indicators.
[0041] Calculate the deviation factor based on the minimum semantic similarity value in the semantic similarity array; In some implementations, when calculating the deviation factor, the following steps may also be performed: the difference between the third constant and the minimum value of the semantic similarity is used as the exponent of a variable positive number to calculate the deviation factor; wherein the variable constant is a positive number greater than 1.
[0042] Extract the minimum value of semantic similarity from the semantic similarity array. The deviation factor is calculated by the difference between a variable positive number greater than 1 and the minimum value of semantic similarity. Specifically, the deviation factor is obtained by multiplying the variable positive number by 1 minus the minimum value of semantic similarity. The calculation formula is:
[0043] Wherein, δ(S) is the deviation factor.
[0044] The deviation factor is used to penalize low similarity values in the semantic similarity array. The larger its value, the lower the semantic similarity in the array, which has a negative impact on the semantic consistency of the triple.
[0045] In some implementations, the calculation of the deviation factor can also be performed based on an exponential function. Specifically, the difference between the third constant and the minimum value of the semantic similarity is used as the exponent of the variable positive number to calculate the deviation factor. Among them, the third constant is a preset fixed value, and the variable positive number is a positive number greater than 1. Through the calculation of the exponential function, the influence of low similarity values can be further amplified, thereby more significantly penalizing low values in the semantic similarity array.
[0046] Through the above steps, the calculation of the deviation factor is completed, providing basic data support for the calculation of subsequent consistency indicators.
[0047] Calculate the coverage factor based on the third semantic similarity and the first semantic similarity in the semantic similarity array.
[0048] In some implementations, when calculating the coverage factor, the following steps may also be performed: The difference between the third semantic similarity and the first semantic similarity is used as the exponent of the natural constant, and then the fourth constant is subtracted to calculate the coverage factor.
[0049] The third semantic similarity and the first semantic similarity are extracted from the semantic similarity array. The coverage factor is calculated by the difference between the third semantic similarity and the first semantic similarity. Specifically, the third semantic similarity is subtracted from the first semantic similarity to obtain the difference, which is then used as the exponent of a natural constant (i.e., the base of the natural logarithm, e), and then subtracted from the fourth constant to calculate the coverage factor. The calculation formula is:
[0050] Where C is the coverage factor, e is the base of the natural logarithm. S3 is the third semantic similarity between the original corpus (first semantic vector) and the answer (third semantic vector); S1 is the first semantic similarity between the original corpus (first semantic vector) and the question (second semantic vector).
[0051] When the order of semantic vectors in the semantic similarity array is different, the calculation formula of the coverage factor needs to be adjusted according to the order of semantic vectors to ensure that the exponent of the natural constant is the difference between the semantic similarity between the original corpus and the answer and the semantic similarity between the original corpus and the question.
[0052] The coverage factor is used to ensure that the semantic similarity between the answer and the original corpus is more consistent. When its value is positive, it indicates a high semantic consistency between the answer and the original corpus; when it is negative, it indicates a semantic deviation.
[0053] In some implementations, the calculation of the coverage factor can be further optimized. Specifically, the difference between the third semantic similarity and the first semantic similarity is used as the exponent of the natural constant, and then the fourth constant is subtracted to calculate the coverage factor. Here, the fourth constant is a preset fixed value used to adjust the benchmark of the coverage factor. Through the exponential operation of the natural constant, the change in the difference of semantic similarity can be more sensitively reflected, so as to more accurately evaluate the consistency between the answer and the original corpus.
[0054] Through the above steps, the calculation of the coverage factor is completed, providing basic data support for the subsequent calculation of the consistency index.
[0055] Calculate the consistency index based on the range factor, deviation factor, and coverage factor.
[0056] In some embodiments, the calculation of the loop-shaped consistency index can also be performed according to the following steps: Multiply the first weight coefficient by the sum of the reciprocals of the range factor and the deviation factor, and then add the product of the second weight coefficient and the coverage factor to calculate the consistency index.
[0057] Obtain the calculation results of the range factor, deviation factor, and coverage factor. The consistency index is calculated by multiplying the first weight coefficient by the sum of the reciprocals of the range factor and the deviation factor, and then adding the product of the second weight coefficient and the coverage factor. Specifically, the calculation formula of the consistency index is: CI = α·(1 / κ(S) + 1 / δ(S)) + β·C Where CI is the consistency index, α is the first constant, β is the second constant, and the first constant and the second constant can be equal.
[0058] Through this formula, the range factor, deviation factor, and coverage factor are combined to quantify the semantic consistency of the data items in the triple.
[0059] In some embodiments, the calculation of the consistency index can be further optimized. Specifically, the first weight coefficient is multiplied by the sum of the inverse of the range factor and the inverse of the deviation factor, and then the product of the second weight coefficient and the coverage factor is added to calculate the consistency index. The first weight coefficient and the second weight coefficient are used to adjust the contribution ratio of the range factor, the deviation factor, and the coverage factor in the consistency index. By adjusting the weights, it is possible to more flexibly adapt to the semantic consistency evaluation needs in different scenarios.
[0060] Through the above steps, the calculation of consistency indicators is completed, providing basic data support for the subsequent question and answer extraction quality verification.
[0061] In some embodiments, the triples are screened based on a preset threshold and a consistency index, and the triples whose consistency index is greater than the preset threshold include: When the consistency index is lower than the preset threshold, the triple is determined to be invalid text.
[0062] The preset threshold is used to determine whether the semantic consistency of the triple meets the requirements. The calculated consistency index is compared with the preset threshold. If the consistency index is greater than or equal to the preset threshold, the triple is determined to be valid text and retained for subsequent question and answer extraction tasks. If the consistency index is lower than the preset threshold, the triple is determined to be invalid text and is removed. Through this screening process, triplets with low semantic consistency can be effectively removed, thereby improving the accuracy and reliability of question and answer extraction.
[0063] In some embodiments, the preset threshold is proportional to the first constant and the second constant. The setting of the preset threshold can be adjusted according to the specific application scenario, requirements and the numerical values of the first constant and the second constant. For example, in scenarios with higher requirements, a higher threshold can be set to filter out triples with higher semantic consistency; in scenarios with relatively loose requirements, the threshold can be appropriately lowered to retain more triples; when the first constant and the second constant are set larger, the preset threshold will be increased synchronously; when the first constant and the second constant are set smaller, the preset threshold will be reduced synchronously. By combining the preset threshold with the consistency index, automatic screening of triples is achieved, providing strong support for improving the quality of question and answer extraction.
[0064] The following two examples are used to clearly illustrate the text screening method provided by this application.
[0065] The following is a positive example. Example 1: Assume the following ABC triple: A: Text corpus (original input corpus): "Global temperatures are rising due to a surge in greenhouse gas emissions, causing climate change." B: Questions (questions extracted by third-party systems such as large language models) "Why is global temperature rising?" C: Answer (extracted or organized by a third-party system such as a large language model): "The rise in global temperatures is due to a surge in greenhouse gas emissions." First, use the general embedding model to embed the triples to get the corresponding vector encoding, and then calculate the cosine similarity in turn to get: Similarity array S = [ 0.708, 0.743, 0.948] Then calculate the three factors needed for the consistency index formula: Range Factor: ≈ 1.7928 Deviation factor (take K=20): ≈ 2.3982 Coverage factor: C = 2.7180.948-0.743 -1 ≈ 0.2275, where e is taken as an approximate value of 2.718 Finally, apply the consistency index calculation formula and set the weights to 1: CI = 1·(1 / κ(S) + 1 / δ(S)) + 1·C Substitute the factor values, so: CI = (1 / 1.7928 + 1 / 2.3982) + 0.2275 = 1.2023 The CI index is greater than 1, indicating that the question-answer triplet has significant rationality and is a high-quality question-answer extraction result.
[0066] The following is an explanation using a negative example. Example 2: Assume the following ABC triple: A: Text corpus (original input corpus): "Global temperatures are rising due to climate change." B: Question (extracted by a third-party system such as a large language model): "What is the weather condition in place A?" C: Answer (extracted or organized by a third-party system such as a large language model): "Quantum computers use qubits to perform complex computing tasks." First, use the general embedding model to embed the triples to get the corresponding vector encoding, and then calculate the cosine similarity in turn to get: Similarity array S = [0.368, 0.175, 0.350] Then calculate the three factors needed for the consistency index formula: Range Factor: ≈ 4.4220 Deviation factor (take K=20): ≈ 11.8399 Coverage factor: C =2.7180.350-0.175 -1 ≈ 0.1912, where e is taken as an approximate value of 2.718 Finally, apply the consistency index calculation formula and set the weights to 1: CI = 1·(1 / κ(S) + 1 / δ(S))-1 + 1·C Substitute the factor values, so: CI = (1 / 4.4220 + 1 / 11.8399) + 0.1912 = 0.5018 The CI index is significantly less than 1, indicating that the question-answer triple has a large semantic inconsistency and is an unqualified question-answer extraction result.
[0067] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, or by hardware, but in many cases the former is a better implementation method.
[0068] The embodiment of the present application also provides a text screening device, Figure 3 is a structural schematic diagram of a text screening device provided in an embodiment of the present application, such as shown, including: The acquisition unit 21 is used to acquire triples, and calculate the consistency index of the triples based on the semantic similarity between the data items in the triples; wherein the data items in the triples include the original corpus, the question and the answer; the consistency index is used to quantify the semantic consistency of the data items in the triples; The screening unit 22 is used to screen the triples based on the preset threshold and the consistency index to obtain the triples whose consistency index is greater than the preset threshold.
[0069] Further, in a possible implementation of the embodiment of the present application, as shown in FIG4, the device further includes: The processing unit 23 is used to embed the triples before the acquisition unit 21 calculates the consistency index of the triples based on the semantic similarity between the data items in the triples, so as to obtain the first semantic vector corresponding to the original corpus, the second semantic vector corresponding to the question and the third semantic vector corresponding to the answer; The acquisition unit 21 is further used for: Based on the semantic vector, calculate the semantic similarity between the two data items in the triple and generate a semantic similarity array; Calculate the consistency index based on the semantic similarity array.
[0070] Further, in a possible implementation of the embodiment of the present application, the acquisition unit 21 is also used for: Calculate the first semantic similarity between the first semantic vector and the second semantic vector; Calculate the second semantic similarity between the second semantic vector and the third semantic vector; Calculate the third semantic similarity between the first semantic vector and the third semantic vector; According to the first semantic similarity, the second semantic similarity and the third semantic similarity, a semantic similarity array is obtained.
[0071] Further, in a possible implementation of the embodiment of the present application, the acquisition unit 21 is also used for: Calculate the extreme difference factor, deviation factor and coverage factor based on the semantic similarity array; Calculate consistency index based on range factor, deviation factor and coverage factor.
[0072] Further, in a possible implementation of the embodiment of the present application, the acquisition unit 21 is also used for: Calculate the range factor based on the maximum semantic similarity and the minimum semantic similarity in the semantic similarity array; Calculate the deviation factor based on the minimum semantic similarity value in the semantic similarity array; Calculate the coverage factor based on the third semantic similarity and the first semantic similarity in the semantic similarity array.
[0073] Further, in a possible implementation of the embodiment of the present application, the computing unit is also used for: The consistency index is calculated by multiplying the first weight coefficient by the reciprocal of the range factor and the sum of the reciprocals of the range factor, and then adding the product of the second weight coefficient and the coverage factor.
[0074] Further, in a possible implementation of the embodiment of the present application, the acquisition unit 21 is also used for: Divide the maximum eigenvalue by the minimum eigenvalue, then square the result to calculate the range factor.
[0075] Further, in a possible implementation of the embodiment of the present application, the acquisition unit 21 is also used for: The difference between the third constant and the minimum value of semantic similarity is used as the exponent of the variable positive number to calculate the deviation factor; wherein the variable constant is a positive number greater than 1.
[0076] Further, in a possible implementation of the embodiment of the present application, the acquisition unit 21 is also used for: The difference between the third semantic similarity and the first semantic similarity is used as the exponent of the natural constant, and then the fourth constant is subtracted to calculate the coverage factor.
[0077] Furthermore, in a possible implementation of the embodiment of the present application, the preset threshold is proportional to the first constant and the second constant.
[0078] Further, in a possible implementation of the embodiment of the present application, the screening unit 22 is also used for: When the consistency index is lower than the preset threshold, the triple is determined to be invalid text.
[0079] For the description of the features in the embodiment corresponding to the text screening device, please refer to the relevant description of the embodiment corresponding to the text screening method, which will not be repeated here.
[0080] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned text screening method embodiments.
[0081] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned text screening method embodiments when it is run.
[0082] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0083] The embodiment of the present application also provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned text screening method embodiments are implemented.
[0084] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned text screening method embodiments are implemented.
[0085] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0086] The above has introduced in detail a method and apparatus for screening a text, an electronic device, and a storage medium provided by this application. Specific examples have been used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A text screening method, characterized in that: include: Acquire a triple, and calculate a consistency index of the triple based on the semantic similarity between the data items in the triple; wherein the data items in the triple include the original corpus, the question and the answer; and the consistency index is used to quantify the semantic consistency of the data items in the triple; The triples are screened based on a preset threshold and the consistency index to obtain the triples whose consistency index is greater than the preset threshold.
2. The text screening method according to claim 1, characterized in that: Before calculating the consistency index of the triples based on the semantic similarity between the data items in the triples, the method further includes: Embedding the triples to obtain a first semantic vector corresponding to the original corpus, a second semantic vector corresponding to the question, and a third semantic vector corresponding to the answer; The calculating the consistency index of the triplet based on the semantic similarity between the data items in the triplet comprises: Based on the semantic vector, the semantic similarity between the two data items in the triple is calculated to generate a semantic similarity array; A consistency index is calculated based on the semantic similarity array.
3. The text screening method according to claim 2, characterized in that: The calculating the semantic similarity between any two data items in the triple based on the semantic vector comprises: Calculating a first semantic similarity between the first semantic vector and the second semantic vector; Calculating a second semantic similarity between the second semantic vector and the third semantic vector; Calculating a third semantic similarity between the first semantic vector and the third semantic vector; The semantic similarity array is obtained according to the first semantic similarity, the second semantic similarity and the third semantic similarity.
4. The text screening method according to claim 2, characterized in that: The calculating the consistency index according to the semantic similarity array comprises: Calculate the extreme difference factor, deviation factor and coverage factor according to the semantic similarity array; The consistency index is calculated based on the range factor, the deviation factor and the coverage factor.
5. The text screening method according to claim 4, characterized in that: The calculating of the extreme difference factor, the deviation factor and the coverage factor according to the semantic similarity array comprises: Calculate the range factor based on the maximum semantic similarity value and the minimum semantic similarity value in the semantic similarity array; Calculate the deviation factor based on the minimum semantic similarity value in the semantic similarity array; The coverage factor is calculated based on the third semantic similarity and the first semantic similarity in the semantic similarity array.
6. The text screening method according to claim 4, characterized in that: The calculation of the consistency index based on the range factor, the range factor and the coverage factor includes: The consistency index is calculated by multiplying the first weight coefficient by the reciprocal of the range factor and the sum of the reciprocal of the range factor, and then adding the product of the second weight coefficient and the coverage factor.
7. The text screening method according to claim 5, characterized in that: The calculating of the range factor based on the maximum semantic similarity value and the minimum semantic similarity value in the semantic similarity array includes: The range factor is calculated by dividing the maximum eigenvalue by the minimum eigenvalue and squaring the result.
8. The text screening method according to claim 5, characterized in that: The calculating of the deviation factor based on the minimum semantic similarity value in the semantic similarity array comprises: The deviation factor is obtained by calculating the difference between the third constant and the minimum value of the semantic similarity as the exponent of a variable positive number; wherein the variable constant is a positive number greater than 1.
9. The text screening method according to claim 5, characterized in that: The calculation of the coverage factor based on the third semantic similarity and the first semantic similarity in the semantic similarity array includes: The coverage factor is calculated by taking the difference between the third semantic similarity and the first semantic similarity as the exponent of the natural constant and subtracting a fourth constant.
10. The text screening method according to claim 6, characterized in that: The preset threshold is proportional to the first constant and the second constant.
11. The method for screening text according to any one of claims 1 to 10, characterized in that: The screening of the triples based on the preset threshold and the consistency index to obtain the triples whose consistency index is greater than the preset threshold comprises: When the consistency index is lower than a preset threshold, the triple is determined to be invalid text.
12. A text screening device, characterized in that: include: An acquisition unit, used to acquire a triple, and calculate a consistency index of the triple based on the semantic similarity between the data items in the triple; wherein the data items in the triple include the original corpus, the question and the answer; and the consistency index is used to quantify the semantic consistency of the data items in the triple; A screening unit is used to screen the triples based on a preset threshold and the consistency index to obtain the triples whose consistency index is greater than the preset threshold.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the text screening method as claimed in any one of claims 1 to 11 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the text screening method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the text screening method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Medical question and answer pair quality inspection method and device, computer equipment and storage medium
CN111444724A
Intelligent question answering system construction method and system based on LLM large language model
CN119691140A
Semantic consistency model for determining a semantic consistency of contents of at least two screenshots
US20240273381A1