Vector similarity search-based encyclopedia entry cross contradiction inspection method and system, electronic equipment and storage medium

By using the BERT-wwm model to convert sentences into vectors, calculating the similarity between sentences, identifying and analyzing the information contradiction between encyclopedia entries, the problem of difficulty in accurately detecting sentence semantic relationships in the prior art is solved, and the accuracy and consistency of encyclopedia content is improved.

CN119990098APending Publication Date: 2025-05-13MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510159384.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When the prior art detects information conflicts between encyclopedia entries, it is difficult to accurately understand and compare the deep semantic content of sentences, resulting in limited detection effect.

Method used

Using a semantic model based on BERT-wwm, the cleaned sentences are converted into vectors, the cosine similarity method is used to calculate the similarity between sentence vector representations, the similarity threshold is set to identify potential counter-information sentences, and the conflicts of actual information are analyzed and determined.

Benefits of technology

By deeply understanding the overall meaning of the sentence, the information conflict between entries can be more accurately identified and corrected, and the accuracy and consistency of the encyclopedia content can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990098A_ABST
    Figure CN119990098A_ABST
Patent Text Reader

Abstract

The invention discloses an encyclopedia entry cross contradiction check method and system based on vector similarity search, electronic equipment and a storage medium, and the method comprises the following steps: extracting text contents of all entries from an encyclopedia database, segmenting and cleaning the text contents in each entry according to sentences, and obtaining cleaned sentences; the cleaned sentences are converted into vectors through a BERT-wwm semantic model, and sentence vector representation of the sentences is obtained; calculating the similarity between the sentence vector representations by using a similarity method; setting a similarity threshold, and when the similarity exceeds the similarity threshold, taking the corresponding sentence pair as a potential contradiction, analyzing the sentence pair, and determining whether there is a conflict of actual information; and summarizing all contradiction sentences and related information entries to generate a detailed report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information retrieval technology, and in particular relates to a method, system, electronic device and storage medium for checking cross-contradiction of encyclopedia entries based on vector similarity search. Background Art

[0002] In the digital age, knowledge-based databases such as encyclopedias have become an important way for people to obtain information. In encyclopedia databases, the content of each entry should maintain consistency and accuracy, but due to the background and opinions of the compilers, or the lack of synchronization of data updates, there may be information conflicts between different entries. At present, common solutions include manual review and automatic inspection methods based on traditional text matching. Traditional methods assist compilers in comparing and correcting the consistency of encyclopedia entry content by checking the content of sentences with similar descriptions in different encyclopedia entries. However, these methods face multiple technical challenges.

[0003] Traditional text similarity algorithms, such as the search engine-based text similarity algorithm BM25, usually focus on keyword matching, but are not good at understanding the deep semantic content of sentences. In addition, sentences vary in length, information density, and contextual dependence. It is difficult to accurately determine the true semantic relationship between sentences simply by relying on keyword matching. Although long text vector similarity algorithms, such as the BERT-wwm-based method, can better capture semantic information, their effect is still limited when dealing with shorter or more semantically complex sentences. Summary of the invention

[0004] The present invention aims to solve the deficiencies of the prior art and provides the following solutions:

[0005] The cross-contradiction checking method of encyclopedia entries based on vector similarity search includes the following steps:

[0006] Extracting text content of all entries from the encyclopedia database, segmenting and cleaning the text content in each entry into sentences, and obtaining cleaned sentences;

[0007] The cleaned sentence is converted into a vector using a BERT-wwm semantic model to obtain a sentence vector representation of the sentence;

[0008] Calculate the similarity between each of the sentence vector representations using a similarity method;

[0009] A similarity threshold is set, and when the similarity exceeds the similarity threshold, the corresponding sentence pair is regarded as a potential contradiction, and the sentence pair is analyzed to determine whether there is an actual information conflict;

[0010] Summarize all conflicting sentences and related information items and generate a detailed report.

[0011] Preferably, the method for obtaining the sentence vector representation includes:

[0012] v=M(x)

[0013] Among them, v represents the sentence vector representation, x represents the cleaned sentence, and M represents the pre-trained BERT-wwm model.

[0014] Preferably, the method for obtaining the similarity includes: using a cosine similarity method to calculate the similarity of two sentences:

[0015]

[0016] in, and They represent the sentence vector representations of the two sentences respectively, and |·| represents the norm.

[0017] Preferably, the detailed report includes: the specific content of the contradictory sentence, the items involved and the suggested corrective measures.

[0018] The present invention also provides an encyclopedia entry cross-contradiction checking system based on vector similarity search, the system applying any of the above-mentioned methods, including: a text processing module, a vector conversion module, a similarity calculation module, a contradiction analysis module and a reporting module;

[0019] The text processing module is used to extract the text content of all entries from the encyclopedia database, segment and clean the text content in each entry according to sentences, and obtain cleaned sentences;

[0020] The vector conversion module converts the cleaned sentence into a vector using the BERT-wwm semantic model to obtain a sentence vector representation of the sentence;

[0021] The similarity calculation module calculates the similarity between each of the sentence vector representations using a similarity method;

[0022] The contradiction analysis module is used to set a similarity threshold. When the similarity exceeds the similarity threshold, the corresponding sentence pair is regarded as a potential contradiction, and the sentence pair is analyzed to determine whether there is an actual information conflict.

[0023] The reporting module is used to summarize all conflicting sentences and related information items and generate a detailed report. Preferably, the workflow of the vector conversion module includes:

[0024] v=M(x)

[0025] Among them, v represents the sentence vector representation, x represents the cleaned sentence, and M represents the pre-trained BERT-wwm model.

[0026] Preferably, the process flow of the similarity calculation module includes:

[0027] Use the cosine similarity method to calculate the similarity between two sentences:

[0028]

[0029] in, and They represent the sentence vector representations of the two sentences respectively, and |·| represents the norm.

[0030] Preferably, the detailed report includes: the specific content of the contradictory sentences, the items involved and the suggested corrective measures.

[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the above-mentioned encyclopedia entry cross-contradiction checking method based on vector similarity search is implemented.

[0032] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed, the above-mentioned encyclopedia entry cross-contradiction checking method based on vector similarity search is implemented.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] (1) By using a semantic model based on BERT-wwm, the present invention can more accurately understand and compare the deep semantic information of sentences, which not only helps to detect subtle differences in semantics, but also effectively identifies and corrects information contradictions between entries, thereby improving the accuracy and consistency of encyclopedia content. Without relying on specific words, the overall meaning of the sentence is deeply understood, thereby more accurately identifying and comparing semantic content, solving the problem that traditional keyword matching methods cannot deeply explore semantics.

[0035] (2) The BERT-wwm model of the present invention can capture the implicit meaning and contextual relationship in a sentence, so that similarity comparison and contradiction detection can be effectively performed even in sentences with complex semantics or small amounts of information. Since the BERT-wwm model is vectorized based on the context of the entire sentence rather than individual words, it can effectively process sentences of different lengths and maintain efficient detection performance regardless of whether the sentence is short or long. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] To more clearly illustrate the technical solution of the present invention, the accompanying drawings required in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0037] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention;

[0038] Figure 2 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention.

[0039] Description of the reference numerals:

[0040] 1010, processor; 1020, memory; 1030, input / output interface; 1040, communication interface; 1050, bus. Detailed implementation manners

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0042] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0043] Embodiment 1

[0044] In this embodiment, as Figure 1 shown, the method for cross-contradiction checking of encyclopedia entries based on vector similarity search includes the following steps:

[0045] S1. Extract the text content of all entries from the encyclopedia database, split and clean the text content in each entry according to sentences to obtain the cleaned sentences.

[0046] In this embodiment, obtaining the text content of all entries from the encyclopedia database is a large-scale text collection. Splitting the text of each entry by sentence can process and compare them in units of sentences in subsequent steps. Then clean the text, remove meaningless symbols (such as punctuation marks, special characters) and stop words (such as "de", "shi") to reduce noise and improve the efficiency and accuracy of subsequent analysis.

[0047] S2. Use the BERT-wwm semantic model to convert the cleaned sentence into a vector and obtain the sentence vector representation of the sentence.

[0048] In this embodiment, the semantic model BERT-wwm is used to convert each sentence into its vector representation. The converted vector captures the semantic information of the sentence, so that subsequent comparison can be made based on vector similarity. The vector representation of each sentence is then stored, and subsequent similarity calculation and contradiction detection are performed. The method for obtaining the sentence vector representation includes:

[0049] v=M(x)

[0050] Among them, v represents the sentence vector representation, x represents the cleaned sentence, which is in the form of natural language text, and M represents the pre-trained BERT-wwm model; when the sentence x is input, it can convert the sentence into a vector v of fixed dimension according to the parameters and structure inside the model. This vector can represent the semantic features of the sentence to a certain extent.

[0051] Specifically, the BERT-wwm model is called, the text sentence x is input into the model, the model outputs the semantic vector of sentence x as the vector representation of the sentence, and the vector representation of the sentence is returned as the output of the function. It can be understood that in this step, each sentence is converted into a numerical vector, which is vectorized using a pre-trained semantic model such as BERT. Sentence vectorization can convert natural language text into mathematical vectors, so that computers can process and compare text semantics; this conversion can use the pre-trained BERT semantic model to encode text information into dense vectors while retaining semantic information, which is beneficial for subsequent similarity calculations and contradiction detection.

[0052] S3. Use the similarity method to calculate the similarity between each sentence vector representation.

[0053] The method of obtaining similarity includes: using the cosine similarity method to calculate the similarity of two sentences:

[0054]

[0055] in, and The sentence vector representations of the two sentences are obtained by the BERT-wwm sentence vectorization function in step S2. Their dimensions are the same, and the value of each dimension represents the representation strength of the sentence on a certain semantic feature; Representation vector and The dot product of , which reflects the degree of similarity between two vectors in direction. The larger the dot product value, the closer the two vectors are in direction. |·| represents the norm, that is, the length of the vector. Furthermore, for the vector representation of each sentence, the cosine similarity algorithm is used to calculate its similarity with all other sentence vectors, and the weight of the similarity calculation is adjusted according to the length of the sentence. Among them, long sentences focus more on the results of the semantic model, while short sentences focus more on the similarity algorithm, which can more accurately measure the similarity between sentences.

[0056] S4. Set a similarity threshold. When the similarity exceeds the similarity threshold, treat the corresponding sentence pair as a potential contradiction, and analyze the sentence pair to determine whether there is an actual information conflict.

[0057] In this embodiment, a threshold is set, and sentence pairs exceeding the threshold are considered to be potentially contradictory. Potentially contradictory sentence pairs are analyzed to determine whether there is indeed an information conflict, and the reality of the contradiction is confirmed through manual intervention and judgment; the similarity threshold is set to t, which is a value determined based on experience or experiment. If the cosine similarity of two sentence vectors is Then these two sentence pairs are regarded as potential contradictions. The main goal of this step is to identify potential contradictory sentence pairs based on the results of similarity calculation. The threshold setting can be adjusted according to specific application requirements to help the system control the sensitivity of contradiction detection. By setting appropriate thresholds, the system can reduce false positives or false negatives and improve the accuracy and efficiency of contradiction detection. This control capability makes the contradiction detection system more flexible and adaptable to the needs of different scenarios, thereby improving the practicality and reliability of the system.

[0058] S5. Summarize all conflicting sentences and related information items and generate a detailed report.

[0059] The detailed report includes: the specific content of the conflicting sentences, the items involved and the recommended corrections.

[0060] In this embodiment, feedback from users and editors is also collected to understand the effectiveness and accuracy of the detection method, and similarity thresholds and weights are adjusted based on the feedback to improve the accuracy and efficiency of future detections;

[0061] Parameter adjustment function expression:

[0062] θ new =F(θ old , f)

[0063] Among them, θ represents adjustable parameters, including the similarity threshold or the weight in adjusting the similarity calculation weight according to the sentence length, and f represents the feedback information of the user and the editor. The feedback information includes the evaluation of the accuracy of the detection result and the analysis of the detected conflict situation. By analyzing this feedback information, the function F can reasonably adjust the parameter θ to improve the accuracy and efficiency of the detection method.

[0064] Embodiment 2

[0065] In this embodiment, an encyclopedia entry cross-conflict checking system based on vector similarity search includes: a text processing module, a vector conversion module, a similarity calculation module, a conflict analysis module, and a reporting module.

[0066] The text processing module is used to extract the text content of all entries from the encyclopedia database, split the text content in each entry into sentences and clean them to obtain the cleaned sentences.

[0067] In this embodiment, obtaining the text content of all entries from the encyclopedia database is a large-scale text collection. Splitting the text of each entry into sentences can be processed and compared sentence by sentence in subsequent steps, and then cleaning the text to remove meaningless symbols (such as punctuation marks, special characters) and stop words (such as "de", "shi") to reduce noise and improve the efficiency and accuracy of subsequent analysis.

[0068] The vector conversion module uses the BERT-wwm semantic model to convert the cleaned sentences into vectors to obtain the sentence vector representation of the sentences.

[0069] In this embodiment, using the semantic model BERT-wwm, each sentence is converted into its vector representation. The converted vector captures the semantic information of the sentence, enabling subsequent comparison based on vector similarity. Then, the vector representation of each sentence is stored for subsequent similarity calculation and conflict detection. The workflow of the vector conversion module includes:

[0070] v = M(x)

[0071] Among them, v represents the sentence vector representation, x represents the cleaned sentence, which is in the form of natural language text, and M represents the pre-trained BERT-wwm model; when the input sentence is x, it can convert the sentence into a vector v with a fixed dimension according to the parameters and structure inside the model. This vector can represent the semantic features of the sentence to a certain extent.

[0072] Specifically, the BERT-wwm model is called, the text sentence x is input into the model, the model outputs the semantic vector of sentence x as the vector representation of the sentence, and the vector representation of the sentence is returned as the output of the function. It can be understood that in this step, each sentence is converted into a numerical vector, which is vectorized using a pre-trained semantic model such as BERT. Sentence vectorization can convert natural language text into mathematical vectors, so that computers can process and compare text semantics; this conversion can use the pre-trained BERT semantic model to encode text information into dense vectors while retaining semantic information, which is beneficial for subsequent similarity calculations and contradiction detection.

[0073] The similarity calculation module uses the similarity method to calculate the similarity between each sentence vector representation.

[0074] The workflow of the similarity calculation module includes: using the cosine similarity method to calculate the similarity of two sentences:

[0075]

[0076] in, and The sentence vector representations of the two sentences are obtained by the BERT-wwm sentence vectorization function in step S2. Their dimensions are the same, and the value of each dimension represents the representation strength of the sentence on a certain semantic feature; Representation vector and The dot product of , which reflects the degree of similarity between two vectors in direction. The larger the dot product value, the closer the two vectors are in direction. |·| represents the norm, that is, the length of the vector. Furthermore, for the vector representation of each sentence, the cosine similarity algorithm is used to calculate its similarity with all other sentence vectors, and the weight of the similarity calculation is adjusted according to the length of the sentence. Among them, long sentences focus more on the results of the semantic model, while short sentences focus more on the similarity algorithm, which can more accurately measure the similarity between sentences.

[0077] The contradiction analysis module is used to set a similarity threshold. When the similarity exceeds the similarity threshold, the corresponding sentence pair is regarded as a potential contradiction, and the sentence pair is analyzed to determine whether there is an actual information conflict.

[0078] In this embodiment, a threshold is set, and sentence pairs exceeding the threshold are considered to be potentially contradictory. Potentially contradictory sentence pairs are analyzed to determine whether there is indeed an information conflict, and the reality of the contradiction is confirmed through manual intervention and judgment; the similarity threshold is set to t, which is a value determined based on experience or experiment. If the cosine similarity of two sentence vectors is Then these two sentence pairs are regarded as potential contradictions. The main goal of this step is to identify potential contradictory sentence pairs based on the results of similarity calculation. The threshold setting can be adjusted according to specific application requirements to help the system control the sensitivity of contradiction detection. By setting appropriate thresholds, the system can reduce false positives or false negatives and improve the accuracy and efficiency of contradiction detection. This control capability makes the contradiction detection system more flexible and adaptable to the needs of different scenarios, thereby improving the practicality and reliability of the system.

[0079] The report module is used to summarize all conflicting sentences and related information items and generate a detailed report.

[0080] The detailed report includes: the specific content of the conflicting sentences, the items involved and the recommended corrections.

[0081] In this embodiment, feedback from users and editors is also collected to understand the effectiveness and accuracy of the detection method, and similarity thresholds and weights are adjusted based on the feedback to improve the accuracy and efficiency of future detections;

[0082] Parameter adjustment function expression:

[0083] θ new =F(θ old , f)

[0084] Among them, θ represents an adjustable parameter, including a similarity threshold or adjusting the weight in the similarity calculation weight according to the sentence length, and f represents feedback information from users and editors. The feedback information includes an evaluation of the accuracy of the detection results and an analysis of the detected contradictions. By analyzing these feedback information, function F can reasonably adjust the parameter θ to improve the accuracy and efficiency of the detection method.

[0085] Embodiment 3

[0086] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the encyclopedia entry cross-contradiction checking method based on vector similarity search described in any of the above embodiments is implemented.

[0087] Figure 2 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.

[0088] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0089] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0090] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0091] The communication interface 1040 is used to connect a communication module (not shown in the figure) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB (Universal Serial Bus), network cable, etc.), or through a wireless mode (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).

[0092] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0093] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0094] The system of the above embodiment is used to implement the corresponding encyclopedia entry cross-contradiction checking method based on vector similarity search in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0095] Embodiment 4

[0096] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the encyclopedia entry cross-conflict checking method based on vector similarity search as described in any of the above embodiments.

[0097] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0098] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the encyclopedia entry cross-contradiction checking method based on vector similarity search as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0099] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.

[0100] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the known power / ground connections to the integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it is apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0101] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0102] Therefore, the units of each example described in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present application.

[0103] The embodiments of the present disclosure are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A method for checking cross-contradiction of encyclopedia entries based on vector similarity search, characterized in that: The following steps are involved: Extracting text content of all entries from the encyclopedia database, segmenting and cleaning the text content in each entry into sentences, and obtaining cleaned sentences; The cleaned sentence is converted into a vector using a BERT-wwm semantic model to obtain a sentence vector representation of the sentence; Calculate the similarity between each of the sentence vector representations using a similarity method; A similarity threshold is set, and when the similarity exceeds the similarity threshold, the corresponding sentence pair is regarded as a potential contradiction, and the sentence pair is analyzed to determine whether there is an actual information conflict; Summarize all conflicting sentences and related information items and generate a detailed report.

2. The method for checking cross-contradiction of encyclopedia entries based on vector similarity search according to claim 1, characterized in that: The method for obtaining the sentence vector representation includes: v=M(x) Among them, v represents the sentence vector representation, x represents the cleaned sentence, and M represents the pre-trained BERT-wwm model.

3. The method for checking cross-contradiction of encyclopedia entries based on vector similarity search according to claim 1, characterized in that: The method for obtaining the similarity includes: using the cosine similarity method to calculate the similarity of two sentences: in, and They represent the sentence vector representations of the two sentences respectively, and |·| represents the norm.

4. The method for checking cross-contradiction of encyclopedia entries based on vector similarity search according to claim 1, characterized in that: The detailed report includes: the specific content of the conflicting sentences, the items involved and the recommended corrective measures.

5. A system for checking cross-contradiction of encyclopedia entries based on vector similarity search, wherein the system applies the method according to any one of claims 1 to 4, and is characterized in that: include: Text processing module, vector conversion module, similarity calculation module, contradiction analysis module and reporting module; The text processing module is used to extract the text content of all entries from the encyclopedia database, segment and clean the text content in each entry according to sentences, and obtain cleaned sentences; The vector conversion module converts the cleaned sentence into a vector using the BERT-wwm semantic model to obtain a sentence vector representation of the sentence; The similarity calculation module calculates the similarity between each of the sentence vector representations using a similarity method; The contradiction analysis module is used to set a similarity threshold. When the similarity exceeds the similarity threshold, the corresponding sentence pair is regarded as a potential contradiction, and the sentence pair is analyzed to determine whether there is an actual information conflict. The report module is used to summarize all conflicting sentences and related information items and generate a detailed report.

6. The encyclopedia entry cross-contradiction checking system based on vector similarity search according to claim 5, characterized in that: The workflow of the vector conversion module includes: v=M(x) Among them, v represents the sentence vector representation, x represents the cleaned sentence, and M represents the pre-trained BERT-wwm model.

7. The encyclopedia entry cross-contradiction checking system based on vector similarity search according to claim 5, characterized in that: The workflow of the similarity calculation module includes: Use the cosine similarity method to calculate the similarity between two sentences: in, and They represent the sentence vector representations of the two sentences respectively, and |·| represents the norm.

8. The encyclopedia entry cross-contradiction checking system based on vector similarity search according to claim 5, characterized in that: The detailed report includes: the specific content of the conflicting sentences, the items involved and the recommended corrective measures.

9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for checking cross-contradiction of encyclopedia entries based on vector similarity search as claimed in any one of claims 1 to 4 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, it implements the encyclopedia entry cross-contradiction checking method based on vector similarity search as described in any one of claims 1 to 4.