Text similarity evaluation method, device and system of coal mine knowledge base
Patent Information
- Application Number
- CN202410969597.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-07-18
AI Technical Summary
然而,随着知识库的传播和共享,知识库中的文件也面临着高相似度的问题,导致知识库中信息质量低,难以保证数据的真实性和原创性
[0039]本公开的实施例提供的技术方案可以包括以下有益效果:通过对待评估的第一文本进行分词处理,得到多个第一分词,以及对煤矿知识库中待比对的第二文本进行分词处理,得到多个第二分词,针对多个分词中的每个第一分词,确定第一分词在第一文本中的重要程度值;确定第一分词与多个第二分词中每个第二分词的至少一个相似度指标值,针对多个第二分词中的每个第二分词,根据第一分词对应的重要程度值和第二分词对应的至少一个相似度指标值,确定第一分词的相似度评估值,对多个第一分词的相似度评估值进行融合处理,得到第一文本的融合评估值,从而根据融合评估值评估第一文本与第二文本的相似程度,进而能够对煤矿知识库中相似度较高的文本进行识别,提高文本的真实性和原创性。
Smart Images

Figure CN118797364B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer processing technology for natural language, and in particular to a method, apparatus and system for evaluating text similarity in a coal mine knowledge base. Background Technology
[0002] In related technologies, with the rapid development of information technology, knowledge base applications and management systems have become important tools for academic research and technological development. These systems store a large number of research reports, technical documents, and data, providing researchers with a convenient platform for resource sharing and knowledge management. However, with the dissemination and sharing of knowledge bases, the files within them also face the problem of high similarity, resulting in low information quality and making it difficult to guarantee the authenticity and originality of the data. Summary of the Invention
[0003] To overcome the problems existing in related technologies, this disclosure provides a method, apparatus and system for evaluating the text similarity of a coal mine knowledge base.
[0004] According to a first aspect of the present disclosure, a method for evaluating the text similarity of a coal mine knowledge base is provided, comprising:
[0005] The first text to be evaluated is segmented to obtain multiple first segments, and the second text to be compared in the coal mine knowledge base is segmented to obtain multiple second segments.
[0006] For each first word in the plurality of word segments, determine the importance value of the first word in the first text;
[0007] Determine at least one similarity index value between the first word segment and each of the plurality of second words;
[0008] For each of the plurality of second word segments, a similarity evaluation value for the first word segment is determined based on the importance value corresponding to the first word segment and at least one similarity index value corresponding to the second word segment.
[0009] The similarity evaluation values of the multiple first word segments are fused to obtain the fused evaluation value of the first text, so as to evaluate the similarity between the first text and the second text based on the fused evaluation value.
[0010] In some embodiments of this disclosure, the fusion processing of the similarity evaluation values of the plurality of first word segments to obtain the fusion evaluation value of the first text includes:
[0011] For each first word segment, the average of multiple similarity evaluation values corresponding to the first word segment is determined; the multiple similarity evaluation values correspond one-to-one with the multiple second words.
[0012] The average values corresponding to each of the multiple first word segments are fused to obtain the fusion evaluation value of the first text.
[0013] In some embodiments of this disclosure, the step of fusing the average values corresponding to the plurality of first word segments to obtain the fusion evaluation value of the first text includes:
[0014] The average values corresponding to the multiple first word segments are fused using the following formula to obtain the fusion evaluation value:
[0015]
[0016] Where x is any first word segment, and Target[x] is an array including all first words in the first text.
[0017] In some embodiments of this disclosure, determining the similarity evaluation value of the first segment based on the importance value corresponding to the first segment and at least one similarity index value corresponding to the second segment for each of the plurality of second segment words includes:
[0018] For each of the plurality of second word segments, the importance value corresponding to the first word and at least one similarity index value corresponding to the second word are weighted and summed to obtain the similarity evaluation value of the first word segment.
[0019] In some embodiments of this disclosure, evaluating the similarity between the first text and the second text based on the fusion evaluation value includes:
[0020] Based on multiple preset fusion evaluation threshold ranges, determine the fusion evaluation value to which the fusion evaluation threshold range belongs;
[0021] Determine the preset level corresponding to the fusion evaluation threshold range;
[0022] If the preset level meets the preset conditions, the first text is determined to be similar to the second text.
[0023] In some embodiments of this disclosure, before the weighted summation of the importance value corresponding to the first segmented word and at least one similarity index value corresponding to the second segmented word to obtain the similarity evaluation value of the first segmented word, the method further includes:
[0024] Each first word segment and its corresponding importance value are stored in a first hash table as key-value pairs.
[0025] For each of the at least one similarity index values, the similarity index value and the corresponding first word segment are stored in a corresponding second hash table in the form of key-value pairs; the second hash table stores similarity index values of the same type.
[0026] The step of determining the similarity evaluation value of the first segment based on the importance value corresponding to the first segment and at least one similarity index value corresponding to the second segment includes:
[0027] Search the first hash table for the importance value corresponding to the first word segmentation;
[0028] Find the similarity index value corresponding to the first word in each of at least one second hash table;
[0029] The similarity evaluation value of the first word is determined based on the importance value and the similarity index value corresponding to the first word in each second hash table.
[0030] According to a second aspect of the present disclosure, a text similarity evaluation apparatus for a coal mine knowledge base is provided, comprising:
[0031] The word segmentation unit is used to segment the first text to be evaluated to obtain multiple first words, and to segment the second text to be compared to obtain multiple second words; the second text is used to perform similarity comparison with the first text.
[0032] The first determining unit is configured to determine the importance value of each of the plurality of word segments in the first text;
[0033] The second determining unit is used to determine at least one similarity index value between the first word segment and each of the plurality of second words segmented;
[0034] The third determining unit is used to determine the similarity evaluation value of the first segment for each of the plurality of second segmented words, based on the importance value corresponding to the first segmented word and at least one similarity index value corresponding to the second segmented word.
[0035] An evaluation unit is used to fuse the similarity evaluation values of the plurality of first word segments to obtain a fused evaluation value of the first text, so as to evaluate the similarity between the first text and the second text based on the fused evaluation value.
[0036] According to a third aspect of the present disclosure, an electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as described in any one of the first aspects.
[0037] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of the first aspects.
[0038] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method as described in any one of the first aspects.
[0039] The technical solution provided by the embodiments of this disclosure may include the following beneficial effects: by performing word segmentation on the first text to be evaluated to obtain multiple first words, and by performing word segmentation on the second text to be compared in the coal mine knowledge base to obtain multiple second words, for each of the multiple first words, the importance value of the first word in the first text is determined; at least one similarity index value between the first word and each of the multiple second words is determined; for each of the multiple second words, a similarity evaluation value of the first word is determined based on the importance value corresponding to the first word and the at least one similarity index value corresponding to the second word; the similarity evaluation values of the multiple first words are fused to obtain a fused evaluation value of the first text, thereby evaluating the similarity between the first text and the second text based on the fused evaluation value, and thus being able to identify texts with high similarity in the coal mine knowledge base, improving the authenticity and originality of the text.
[0040] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0042] Figure 1 This is a flowchart illustrating a text similarity assessment method for a coal mine knowledge base according to an exemplary embodiment.
[0043] Figure 2 This is a block diagram illustrating a text similarity assessment device for a coal mine knowledge base according to an exemplary embodiment.
[0044] Figure 3This is a block diagram illustrating an apparatus for text similarity assessment of a coal mine knowledge base according to an exemplary embodiment. Detailed Implementation
[0045] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.
[0046] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. The singular forms “a” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0047] It should be understood that although the terms first, second, third, etc., may be used to describe various information in embodiments of this disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of embodiments of this disclosure, and similarly, second information may also be referred to as first information. Depending on the context, the words “if” and “suppose” as used herein may be interpreted as “when”, “when”, or “in response to a determination”.
[0048] Furthermore, various forms of processes shown in the embodiments of this disclosure can be used to reorder, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0049] In related technologies, with the rapid development of information technology, knowledge base applications and management systems have become important tools for academic research and technological development. These systems store a large number of research reports, technical documents, and data, providing researchers with a convenient platform for resource sharing and knowledge management. However, with the dissemination and sharing of knowledge bases, the files within them also face the problem of high similarity, resulting in low information quality and making it difficult to guarantee the authenticity and originality of the data.
[0050] To address the aforementioned issues, this disclosure provides a method, apparatus, and system for evaluating text similarity in a coal mine knowledge base. The method involves segmenting a first text to be evaluated into multiple first segments, and segmenting a second text in the coal mine knowledge base into multiple second segments. For each first segment, an importance value is determined within the first text. At least one similarity index value is determined between the first segment and each of the multiple second segments. For each second segment, a similarity evaluation value is determined based on the importance value of the first segment and the at least one similarity index value of the second segment. The similarity evaluation values of the multiple first segments are then fused to obtain a fused evaluation value for the first text. This fused evaluation value is used to assess the similarity between the first and second texts, thereby enabling the identification of highly similar texts in the coal mine knowledge base and improving the authenticity and originality of the text.
[0051] Figure 1 This is a flowchart illustrating a text similarity assessment method, apparatus, and system for a coal mine knowledge base according to an exemplary embodiment, such as... Figure 1 As shown, it should be noted that the text similarity evaluation method, apparatus, and system for coal mine knowledge bases in this disclosure are applied in the text similarity evaluation method, apparatus, and system for coal mine knowledge bases. For example... Figure 1 As shown, the method may include the following steps:
[0052] Step 101: Perform word segmentation on the first text to be evaluated to obtain multiple first words, and perform word segmentation on the second text to be compared in the coal mine knowledge base to obtain multiple second words.
[0053] It is understandable that the first text is the text that needs to be evaluated for similarity, and the second text is the object that the first text is compared with.
[0054] In one embodiment, the first text can be text from a coal mine knowledge base, or it can be other text.
[0055] In this embodiment of the disclosure, the first text to be evaluated and the second text to be compared are segmented to obtain multiple first segments corresponding to the first text and multiple second segments corresponding to the second text.
[0056] It should be noted that both the first and second participles can be words or phrases.
[0057] In some embodiments of this disclosure, step 101 may specifically include the following steps:
[0058] Step a1: Perform word segmentation on the first text to be evaluated to obtain the first word segmentation result.
[0059] Step a2: removing the words with functional part-of-speech from the first word segmentation result, and determining other words except the words with functional part-of-speech as first words.
[0060] In the embodiment of the present disclosure, after obtaining the first word segmentation result, part-of-speech tagging may be performed on each word in the first word segmentation result, so as to distinguish verbs, nouns, attributives and functional words. According to the tagged part-of-speech, words that are functional words such as "de" (a structural particle), "shi" (the copula "is") are removed from the first word segmentation result, and other words except the words with functional part-of-speech are determined as first words, so as to reduce interference and improve the text information rate.
[0061] Step a3: performing word segmentation processing on the second text to be evaluated to obtain a second word segmentation result.
[0062] Step a4: removing the words with functional part-of-speech from the second word segmentation result, and determining other words except the words with functional part-of-speech as second words.
[0063] In the embodiment of the present disclosure, after obtaining the second word segmentation result, part-of-speech tagging may be performed on each word in the second word segmentation result, so as to distinguish verbs, nouns, attributives and functional words. According to the tagged part-of-speech, words that are functional words such as "de" (a structural particle), "shi" (the copula "is") are removed from the second word segmentation result, and other words except the words with functional part-of-speech are determined as second words, so as to reduce interference and improve the text information rate.
[0064] Step 102: for each first word among the plurality of segmented words, determining an importance value of the first word in the first text.
[0065] In some embodiments of the present disclosure, TF-IDF (Term Frequency-Inverse Document Frequency) can be used to calculate the importance value of the first word in the first text.
[0066] It should be noted that TF-IDF is a common weighting technology used for information retrieval and text mining. TF-IDF is a statistical method used to evaluate the importance of a word to one document in a document set or a corpus. The importance of a word increases in direct proportion to the number of times it appears in the document, but decreases in inverse proportion to the frequency of its appearance in the corpus. Various forms of TF-IDF weighting are often applied by search engines as a measure or rating of the correlation between a document and a user query.
[0067] Step 103: determining at least one similarity index value between the first word and each second word among the plurality of second words.
[0068] In some embodiments of this disclosure, at least one similarity index value includes one or more of the following:
[0069] Edit distance index, cosine similarity index, Jaccard similarity coefficient.
[0070] In one embodiment, the edit distance metric can be used to calculate the minimum number of edit operations required between two strings, including insertion, deletion, and replacement. For example, converting "kitten" to "sitting" requires at least 3 edit operations.
[0071] In one embodiment, cosine similarity represents text as vectors and calculates the cosine similarity between two vectors. Cosine similarity is a method for measuring the cosine of the angle between two non-zero vectors. It is commonly used to compare document similarity. Cosine similarity values range from -1 to 1, where 1 represents identical vectors, 0 represents dissimilar vectors, and -1 represents identical vectors.
[0072] It should be noted that the first text includes multiple first words, and each first word can be compared with all the second words separately. That is, any first word can be compared with all the second words separately. The similarity can be compared by calculating the edit distance index, the cosine similarity index, or the Jaccard similarity coefficient, or any one or more of these methods.
[0073] Step 104: For each of the multiple second word segments, determine the similarity evaluation value of the first word segment based on the importance value corresponding to the first word segment and at least one similarity index value corresponding to the second word segment.
[0074] In some embodiments of this disclosure, step 104 may specifically include the following steps: for each of the plurality of second word segments, a weighted sum of the importance value corresponding to the first word segment and at least one similarity index value corresponding to the second word segment is performed to obtain the similarity evaluation value of the first word segment.
[0075] In one embodiment, the similarity evaluation value V(n) of the first word segment can be calculated using the following formula:
[0076] V(n)=(tf_hash(Test[n])*P1+d_hash(Test[n])*P2+cos_hash(Test[n])*P3+jac_hash(Test[n])*P4) / (P1+P2+P3+P4)
[0077] Wherein, tf_hash(Test[n]) is the importance value corresponding to the second word n in the first hash table, d_hash(Test[n]) is the first similarity evaluation value in the second hash table corresponding to the edit distance index value, cos_hash(Test[n]) is the second similarity evaluation value in the second hash table corresponding to the cosine similarity index value, jac_hash(Test[n]) is the third similarity evaluation value in the second hash table corresponding to the Jaccard similarity coefficient, P1 is the weight value corresponding to the importance value, P2 is the weight value corresponding to the first similarity evaluation value, P3 is the weight value corresponding to the second similarity evaluation value, and P4 is the weight value corresponding to the third similarity evaluation value.
[0078] In some embodiments of this disclosure, prior to step 104, the method may further include the following steps:
[0079] Each first word segment and its corresponding importance value are stored as key-value pairs in the first hash table.
[0080] For each similarity index value among at least one similarity index value, the similarity index value and the corresponding first word segment are stored in the corresponding second hash table in the form of key-value pairs; the second hash table stores similarity index values of the same type;
[0081] Based on the importance value corresponding to the first word segment and at least one similarity index value corresponding to the second word segment, determine the similarity evaluation value of the first word segment, including:
[0082] Find the importance value corresponding to the first word in the first hash table;
[0083] Find the similarity index value corresponding to the first word in each of at least one second hash table;
[0084] The similarity evaluation value of the first word is determined based on the importance value and the similarity index value corresponding to the first word in each second hash table.
[0085] It is understandable that since the first text contains multiple first-segment words, each first-segment word corresponds to multiple importance values. In addition, since the second text contains multiple second-segment words, each second-segment word corresponding to a first-segment word corresponds to multiple similarity index values.
[0086] To efficiently locate the importance value and similarity index value corresponding to the current first word segment when calculating similarity evaluation values, the importance value of each first word segment and its corresponding value can be stored as key-value pairs in a first hash table for each similarity index value among at least one similarity index value. Additionally, the similarity index value and the corresponding first word segment can be stored as key-value pairs in a corresponding second hash table. The second hash table stores similarity index values of the same type.
[0087] In one embodiment, when it is necessary to calculate the similarity evaluation value of the first word segment, the importance value corresponding to the first word segment can be found in the first hash table, and the similarity index value corresponding to the first word segment can be found in each of the at least one second hash table. Based on the importance value and the similarity index value corresponding to the first word segment in each second hash table, the similarity evaluation value of the first word segment is determined.
[0088] Step 105: The similarity evaluation values of multiple first word segments are fused to obtain the fused evaluation value of the first text, so as to evaluate the similarity between the first text and the second text based on the fused evaluation value.
[0089] It is understandable that, since the first text includes multiple first words, after obtaining the similarity evaluation value of each first word, the multiple similarity evaluation values can be fused to obtain the fused evaluation value of the first text, so as to evaluate the similarity between the first text and the second text based on the fused evaluation value.
[0090] In some embodiments of this disclosure, the fusion processing of the similarity evaluation values of multiple first word segments in step 105 to obtain the fusion evaluation value of the first text may specifically include the following steps:
[0091] Step a1: For each first word segment, determine the average of multiple similarity evaluation values corresponding to the first word segment.
[0092] Multiple similarity assessment values correspond one-to-one with multiple second-word segments.
[0093] In one embodiment, the above evaluation value R can be calculated using the following formula:
[0094]
[0095] Where V(n) is the similarity evaluation value of the first word segmentation.
[0096] Step a2: The average values corresponding to the multiple first word segments are fused to obtain the fusion evaluation value of the first text.
[0097] In some embodiments of this disclosure, step a2 may specifically include the following steps:
[0098] The average values of multiple first-segment words are fused using the following formula to obtain the fusion evaluation value:
[0099]
[0100] Where x is any first word segment, and Target[x] is an array including all first words in the first text.
[0101] In some embodiments of this disclosure, step 105, which assesses the similarity between the first text and the second text based on the fusion evaluation value, may specifically include the following steps:
[0102] Step b1: Determine the fusion evaluation threshold range to which the fusion evaluation value belongs based on multiple preset fusion evaluation threshold ranges.
[0103] Step b2: Determine the preset level corresponding to the fusion evaluation threshold range.
[0104] Step b3: If the preset level meets the preset conditions, determine that the first text and the second text are similar.
[0105] Understandably, the fusion evaluation threshold range can be pre-defined according to actual needs, the level of the fusion evaluation value can be determined according to the fusion evaluation threshold range, and the similarity between the first text and the second text can be judged based on the level, thereby further improving the judgment efficiency.
[0106] For example, assessment levels can be preset as shown in Table 1:
[0107] Table 1 Similarity Rating Evaluation Table
[0108] Test Results [0,0.2) [0.2,0.45) [0.45,0.75) [0.75,1.0]
[0109] When the similarity level is determined to be 3, a warning message is output for the text, indicating a suspected high similarity declaration. When the similarity level is determined to be 4, the knowledge base determines that the first text is highly similar to the second text and identifies the first text as non-original text.
[0110] According to the text similarity assessment method for a coal mine knowledge base proposed in this disclosure, the method involves segmenting a first text to be assessed to obtain multiple first words, and segmenting a second text to be compared in the coal mine knowledge base to obtain multiple second words. For each first word, the importance value of the first word in the first text is determined. At least one similarity index value between the first word and each of the multiple second words is determined. For each second word, a similarity assessment value is determined based on the importance value corresponding to the first word and the at least one similarity index value corresponding to the second word. The similarity assessment values of the multiple first words are then fused to obtain a fused assessment value for the first text. The similarity between the first text and the second text is assessed based on the fused assessment value, thereby enabling the identification of texts with high similarity in the coal mine knowledge base and improving the authenticity and originality of the text.
[0111] Figure 2 This is a block diagram of a text similarity assessment device for a coal mine knowledge base, according to an exemplary embodiment. (Refer to...) Figure 2 The device includes a word segmentation unit 201, a first determination unit 202, a second determination unit 203, a third determination unit 204, and an evaluation unit 205.
[0112] The word segmentation unit 201 is used to perform word segmentation on the first text to be evaluated to obtain multiple first words, and to perform word segmentation on the second text to be compared to obtain multiple second words; the second text is used to perform similarity comparison with the first text.
[0113] The first determining unit 202 is used to determine the importance value of each first segment in the first text for each of the multiple segmented words;
[0114] The second determining unit 203 is used to determine at least one similarity index value between the first word segment and each of the plurality of second words segmented;
[0115] The third determining unit 204 is used to determine the similarity evaluation value of the first segment for each of the multiple second segment words, based on the importance value corresponding to the first segment word and at least one similarity index value corresponding to the second segment word.
[0116] Evaluation unit 205 is used to fuse the similarity evaluation values of multiple first word segments to obtain the fused evaluation value of the first text, so as to evaluate the similarity between the first text and the second text based on the fused evaluation value.
[0117] In some embodiments of this disclosure, the evaluation unit 205 may specifically be used for:
[0118] For each first word segment, determine the average of multiple similarity evaluation values corresponding to the first word segment; multiple similarity evaluation values correspond one-to-one with multiple second words;
[0119] The average values of the first word segments are fused to obtain the fusion evaluation value of the first text.
[0120] In some embodiments of this disclosure, the evaluation unit 205 may specifically be used for:
[0121] The average values of multiple first-segment words are fused using the following formula to obtain the fusion evaluation value:
[0122]
[0123] Where x is any first word segment, and Target[x] is an array including all first words in the first text.
[0124] In some embodiments of this disclosure, the third determining unit 204 may specifically be used to: for each of the plurality of second word segments, perform a weighted summation of the importance value corresponding to the first word segment and at least one similarity index value corresponding to the second word segment to obtain a similarity evaluation value of the first word segment.
[0125] In some embodiments of this disclosure, the evaluation unit 205 may specifically be used for:
[0126] Based on multiple preset fusion evaluation threshold ranges, determine the fusion evaluation value to which the fusion evaluation threshold range belongs;
[0127] Determine the preset level corresponding to the relevant fusion evaluation threshold range;
[0128] If the preset conditions are met at the preset level, the first text and the second text are determined to be similar.
[0129] In some embodiments of this disclosure, the apparatus may further include:
[0130] The first storage unit is used to store each first word segment and its corresponding importance value in the first hash table in the form of key-value pairs;
[0131] The second storage unit is used to store the similarity index value and the corresponding first word segment in the form of key-value pairs into the corresponding second hash table for each of the at least one similarity index values; the second hash table stores similarity index values of the same type;
[0132] The third determining unit 204 can be specifically used for:
[0133] Find the importance value corresponding to the first word in the first hash table;
[0134] Find the similarity index value corresponding to the first word in each of at least one second hash table;
[0135] The similarity evaluation value of the first word is determined based on the importance value and the similarity index value corresponding to the first word in each second hash table.
[0136] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0137] According to the text similarity assessment method for a coal mine knowledge base proposed in this disclosure, the method involves segmenting a first text to be assessed to obtain multiple first words, and segmenting a second text to be compared in the coal mine knowledge base to obtain multiple second words. For each first word, the importance value of the first word in the first text is determined. At least one similarity index value between the first word and each of the multiple second words is determined. For each second word, a similarity assessment value is determined based on the importance value corresponding to the first word and the at least one similarity index value corresponding to the second word. The similarity assessment values of the multiple first words are then fused to obtain a fused assessment value for the first text. The similarity between the first text and the second text is assessed based on the fused assessment value, thereby enabling the identification of texts with high similarity in the coal mine knowledge base and improving the authenticity and originality of the text.
[0138] Figure 3 This is a block diagram illustrating an apparatus for text similarity evaluation of a coal mine knowledge base according to an exemplary embodiment. For example, apparatus 300 may be an electronic device, such as a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0139] Reference Figure 3 The device 300 may include one or more of the following components: a processing component 302, a memory 304, a power component 306, a multimedia component 308, an audio component 310, an input / output (I / O) interface 312, a sensor component 314, and a communication component 316.
[0140] Processing component 302 typically controls the overall operation of device 300, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 302 may include one or more processors 320 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 302 may include one or more modules to facilitate interaction between processing component 302 and other components. For example, processing component 302 may include a multimedia module to facilitate interaction between multimedia component 308 and processing component 302.
[0141] Memory 304 is configured to store various types of data to support the operation of device 300. Examples of this data include instructions for any application or method operating on device 300, contact data, phonebook data, messages, pictures, videos, etc. Memory 304 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0142] The power supply component 306 provides power to the various components of the device 300. The power supply component 306 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 300.
[0143] Multimedia component 308 includes a screen that provides an output interface between the device 300 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 308 includes a front-facing camera and / or a rear-facing camera. When the device 300 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0144] Audio component 310 is configured to output and / or input audio signals. For example, audio component 310 includes a microphone (MIC) configured to receive external audio signals when device 300 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 304 or transmitted via communication component 316. In some embodiments, audio component 310 also includes a speaker for outputting audio signals.
[0145] I / O interface 312 provides an interface between processing component 302 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0146] Sensor assembly 314 includes one or more sensors for providing status assessments of various aspects of device 300. For example, sensor assembly 314 may detect the on / off state of device 300, the relative positioning of components such as the display and keypad of device 300, changes in the position of device 300 or a component of device 300, the presence or absence of user contact with device 300, the orientation or acceleration / deceleration of device 300, and temperature changes of device 300. Sensor assembly 314 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 314 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 314 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0147] Communication component 316 is configured to facilitate wired or wireless communication between device 300 and other devices. Device 300 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 316 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 316 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0148] In an exemplary embodiment, the apparatus 300 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0149] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 304 including instructions, which can be executed by a processor 320 of the device 300 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0150] In an exemplary embodiment, a computer program product is also provided, including a computer program that implements the above-described method when executed by the processor 320 of the device 300.
[0151] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.
[0152] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for evaluating text similarity in a coal mine knowledge base, characterized in that, include: The first text to be evaluated is segmented to obtain multiple first segments, and the second text to be compared in the coal mine knowledge base is segmented to obtain multiple second segments. For each of the plurality of first word segments, determine the importance value of the first word segment in the first text; Determine at least one similarity index value between the first word segment and each of the plurality of second words; For each of the plurality of second word segments, a similarity evaluation value for the first word segment is determined based on the importance value corresponding to the first word segment and at least one similarity index value corresponding to the second word segment, including: For each of the plurality of second word segments, the importance value corresponding to the first word and at least one similarity index value corresponding to the second word are weighted and summed to obtain the similarity evaluation value of the first word segment. The similarity evaluation values of the multiple first word segments are fused to obtain a fused evaluation value for the first text. The similarity between the first text and the second text is then evaluated based on this fused evaluation value, including: For each first word segment, the average of multiple similarity evaluation values corresponding to the first word segment is determined; the multiple similarity evaluation values correspond one-to-one with the multiple second words. The average values corresponding to each of the multiple first word segments are fused to obtain the fusion evaluation value of the first text; Before performing a weighted summation of the importance value corresponding to the first word segment and at least one similarity index value corresponding to the second word segment to obtain the similarity evaluation value of the first word segment, the method further includes: Each first word segment and its corresponding importance value are stored in a first hash table as key-value pairs. For each of the at least one similarity index values, the similarity index value and the corresponding first word segment are stored in a corresponding second hash table in the form of key-value pairs; the second hash table stores similarity index values of the same type. The step of determining the similarity evaluation value of the first segment based on the importance value corresponding to the first segment and at least one similarity index value corresponding to the second segment includes: Search the first hash table for the importance value corresponding to the first word segmentation; Find the similarity index value corresponding to the first word in each of at least one second hash table; The similarity evaluation value of the first word is determined based on the importance value and the similarity index value corresponding to the first word in each second hash table.
2. The text similarity evaluation method for a coal mine knowledge base according to claim 1, characterized in that, The step of fusing the average values corresponding to the multiple first word segments to obtain the fusion evaluation value of the first text includes: The average values corresponding to the multiple first word segments are fused using the following formula to obtain the fusion evaluation value: Value= Where x is any first word segment, and Target[x] is an array including all first words in the first text.
3. The text similarity evaluation method for a coal mine knowledge base according to claim 1, characterized in that, The step of evaluating the similarity between the first text and the second text based on the fusion evaluation value includes: Based on multiple preset fusion evaluation threshold ranges, determine the fusion evaluation value to which the fusion evaluation threshold range belongs; Determine the preset level corresponding to the fusion evaluation threshold range; If the preset level meets the preset conditions, the first text is determined to be similar to the second text.
4. A text similarity evaluation device for a coal mine knowledge base, characterized in that, The apparatus implements the method as described in claim 1, the apparatus comprising: The word segmentation unit is used to segment the first text to be evaluated to obtain multiple first words, and to segment the second text to be compared to obtain multiple second words; the second text is used to perform similarity comparison with the first text. The first determining unit is configured to determine the importance value of each of the plurality of first word segments in the first text; The second determining unit is used to determine at least one similarity index value between the first word segment and each of the plurality of second words segmented; The third determining unit is used to determine the similarity evaluation value of the first segment for each of the plurality of second segmented words, based on the importance value corresponding to the first segmented word and at least one similarity index value corresponding to the second segmented word. An evaluation unit is used to fuse the similarity evaluation values of the plurality of first word segments to obtain a fused evaluation value of the first text, so as to evaluate the similarity between the first text and the second text based on the fused evaluation value.
5. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as described in any one of claims 1 to 3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 3.
7. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Text similarity assessment method and device
CN105488023A
Innovative evaluation method based on science and technology big data text content
CN114154498A