A text similarity determination method and device, electronic equipment and storage medium
By identifying and removing quoted texts in the text to be processed, and using word distribution information and similarity algorithms to determine the similarity of the text to be checked for duplicates, the problem of quoted texts affecting inaccurate duplicate checking results is solved, and more accurate duplicate checking results are achieved.
Patent Information
- Application Number
- CN202311108038.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-30
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2043-08-30
AI Technical Summary
When there are quoted texts in the text to be checked for duplicates, the existing technology leads to inaccurate duplicate checking results.
By obtaining the word distribution information of the text to be processed and the reference text, identifying and eliminating the quoted text, the text to be checked for duplicates and the text to be compared are obtained, and their similarity is determined using a similarity algorithm.
Effectively reduce the impact of quoted text on duplicate checking results and improve the accuracy of duplicate checking results.
Smart Images

Figure CN119538900B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a text similarity determination method and device, electronic equipment and a storage medium. BACKGROUND
[0002] At present, the text duplication detection technology is widely used in scientific research institutions and periodical magazines.
[0003] When detecting the text, the to-be-detected text is compared with the text in the text library one by one to determine the similarity of the to-be-detected text. However, in general, when the to-be-detected text includes a quoted text, the quoted text in the to-be-detected text will seriously affect the detection result of the to-be-detected text, resulting in an inaccurate text detection result.
[0004] In order to solve the above problems, the method of text duplication detection needs to be improved. SUMMARY
[0005] The present application provides a text similarity determination method and device, electronic equipment and a storage medium to solve the problem of inaccurate detection result of the to-be-detected text caused by the quoted text when the to-be-detected text includes a quoted text.
[0006] In a first aspect, the present application provides a text similarity determination method, comprising:
[0007] obtaining a to-be-processed text and at least one reference text associated with the to-be-processed text;
[0008] For each reference text, determining a quoted text in the to-be-processed text according to the first word distribution information of the to-be-processed text and the second word distribution information of the current reference text;
[0009] eliminating the quoted text from the to-be-processed text to obtain a to-be-detected text, and eliminating the associated text corresponding to the quoted text from the current reference text to obtain a to-be-compared text;
[0010] determining the text similarity between the to-be-detected text and the to-be-compared text based on at least one similarity algorithm.
[0011] In a second aspect, the present application further provides a text similarity determination device, comprising:
[0012] a text acquisition module for acquiring a to-be-processed text and at least one reference text associated with the to-be-processed text;
[0013] The text determining module is configured to determine, for each reference text, a quoted text in the text to be processed according to the first word distribution information of the text to be processed and second word distribution information of a current reference text;
[0014] The text eliminating module is configured to eliminate the quoted text from the text to be processed to obtain a text to be searched, and eliminate associated text corresponding to the quoted text from the current reference text to obtain a text to be compared;
[0015] The similarity determining module is configured to determine a text similarity between the text to be searched and the text to be compared based on at least one similarity algorithm.
[0016] In a third aspect, an electronic device is provided, and the electronic device comprises:
[0017] at least one processor; and
[0018] a memory connected with the at least one processor; wherein
[0019] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the method for determining the text similarity according to any one of the embodiments of the present application.
[0020] In a fourth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores computer instructions for enabling a processor to implement the method for determining the text similarity according to any one of the embodiments of the present application when the processor executes the computer instructions.
[0021] The technical scheme of the embodiment of the present application is to obtain a to-be-processed text and at least one reference text associated with the to-be-processed text. In actual application, after obtaining the to-be-processed text, at least one text is retrieved from a pre-constructed text library as a reference text for duplicate detection of the to-be-processed text. Further, for each reference text, the reference text in the to-be-processed text is determined according to the first word distribution information of the to-be-processed text and the second word distribution information of the current reference text. In the technical scheme, in order to reduce the influence of the reference content in the to-be-processed text on the duplicate detection result, after the to-be-processed text is divided by word segmentation, at least one first text segmentation is obtained, and the word distribution information corresponding to each first text segmentation is determined according to the position distribution information of each first text segmentation in the to-be-processed text, so as to determine the word distribution information corresponding to the to-be-processed text according to the word distribution information of each first text segmentation. Similarly, the same text processing is performed on the current reference text to determine the second word distribution information corresponding to the current reference text. Further, based on the similarity of the first word distribution information and the second word distribution information, the reference text is determined from the to-be-processed text, and the associated text corresponding to the reference text in the current reference text is determined. The reference text is removed from the to-be-processed text to obtain a to-be-detected text, and the associated text corresponding to the reference text is removed from the current reference text to obtain a to-be-compared text. Based on at least one similarity algorithm, the text similarity of the to-be-detected text and the to-be-compared text is determined. The advantage of this setting is that when the to-be-compared text is used for duplicate detection of the to-be-detected text, the influence of the reference content on the duplicate detection result of the to-be-processed text can be effectively reduced, and a more accurate duplicate detection result can be obtained. The problem that the duplicate detection result of the to-be-detected text is inaccurate due to the reference text is solved. By removing the reference text and performing duplicate detection on the text after removing the reference text, the effect of obtaining a more accurate duplicate detection result is achieved.
[0022] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0024] Figure 1 is a flowchart of a text similarity determination method provided by the first embodiment of the present application;
[0025] Figure 2 is a flow chart of a text similarity determination method according to the second embodiment of the present application;
[0026] Figure 3 is a flow chart of a text similarity determination method according to the second embodiment of the present application;
[0027] Figure 4 is a structural schematic diagram of a text similarity determination device according to the third embodiment of the present application;
[0028] Figure 5 is a structural schematic diagram of an electronic device implementing a text similarity determination method according to the present application. DETAILED DESCRIPTION
[0029] In order to make the technical personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor should be within the scope of protection of the present application.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0031] Embodiment one
[0032] Figure 1 A flow chart of a text similarity determination method is provided for the first embodiment of the present application. The present embodiment can be applied to obtain a to-be-processed text and at least one reference text, determine a reference text in the to-be-processed text and an associated text corresponding to the reference text in the reference text, obtain a to-be-duplicated text by removing the reference text from the to-be-processed text, obtain a to-be-compared text by removing the associated text from the reference text, perform text duplication on the to-be-duplicated text based on the to-be-compared text, and obtain a text similarity between the to-be-duplicated text and the to-be-compared text. The similarity between the to-be-compared text and the to-be-duplicated text is represented based on the text similarity. The method can be executed by a text similarity determination device, which can be realized in the form of hardware and / or software, and can be configured in a computing device capable of executing the text similarity determination method.
[0033] As Figure 1As shown, the method comprises:
[0034] S110, obtaining a to-be-processed text and at least one reference text associated with the to-be-processed text.
[0035] The to-be-processed text can be understood as a text that needs to be checked for duplication and does not have any modification operation, including addition, deletion or modification, etc. For example, the to-be-processed text can be a thesis or a journal, etc. that needs to be checked for duplication. The reference text can be understood as a text used for checking the to-be-processed text for duplication.
[0036] For example, taking a duplication checking system providing a thesis duplication checking function as an example, the system usually has a text library for checking for duplication, which includes a large number of published or uploaded thesis texts or journal texts. After obtaining the to-be-processed text that needs to be checked for duplication, at least one thesis text or journal text in the text library is called as a reference text.
[0037] S120, for each reference text, determining a quoted text in the to-be-processed text according to the first word distribution information of the to-be-processed text and the second word distribution information of the current reference text.
[0038] In actual application, before checking the to-be-processed text for duplication, the to-be-processed text needs to be segmented to obtain at least one first text segment. Taking one of the first text segments as an example, the first word distribution information can be understood as the position distribution information of the first text segment in the to-be-processed text, such as the first text segment being located in the first paragraph, the main text and the last paragraph of the to-be-processed text. Correspondingly, when checking the to-be-processed text based on the reference text, the same operation needs to be performed on the reference text, that is, the reference text is prohibited from being segmented to obtain at least one second text segment, and the position distribution information of the second text segment in the reference text is taken as the second word distribution information. The quoted text can be understood as the text content in the to-be-processed text quoted from the reference text. For example, the text content in the reference text is quoted in the to-be-processed text as an example, argumentation or argument support, etc. At this time, the text content quoted from the reference text becomes the quoted text. It can be understood that the quoted text in the to-be-processed text can be completely consistent with the quoted content in the reference text, or can have some differences with the quoted content in the reference text, such as changes in noun names, changes in noun positions and changes in text sentence orders, etc.
[0039] Specifically, in order to accurately perform the duplicate checking on the to-be-processed text, the to-be-processed text needs to be checked based on at least one reference text. For each reference text, taking one of the reference texts as a current reference text as an example, the to-be-processed text and the current reference text are respectively subjected to word segmentation processing, to obtain at least one first text segment in the to-be-processed text and at least one second text segment in the current reference text. Further, according to the distribution of each first text segment in the to-be-processed text, first word distribution information corresponding to the to-be-processed text is determined, and according to the distribution of each second text segment in the current reference text, second word distribution information corresponding to the current reference text is determined, so as to determine the quoted text in the to-be-processed text according to the first word distribution information and the second word distribution information.
[0040] It should be noted that the process of determining the first word distribution information corresponding to the to-be-processed text is similar to that of determining the second word distribution information in the current reference text. Therefore, the first word distribution information is taken as an example for description.
[0041] For example, when performing the duplicate checking on the to-be-processed text, the to-be-processed text is first subjected to word segmentation processing to obtain at least one first text segment, for example, the Jieba word segmentation processing method can be used. For each first text segment obtained by segmentation, the score vector of each first text segment in the to-be-processed text can be calculated based on the term frequency-inverse document frequency algorithm (i.e., TF-IDF), so as to measure the importance of the corresponding first text segment in the to-be-processed text according to the score vector corresponding to each first text segment. Among them, the higher the score of the first text segment, the higher the importance of the first text segment in the to-be-processed text. Further, according to the importance of each first text segment in the to-be-processed text, the word distribution information corresponding to each first text segment is determined.
[0042] Taking one of the first text segments as an example, the position information of the first text segment in the to-be-processed text is located to determine the first word distribution information of the first text segment in the to-be-processed text. In order to determine whether the content associated with the first text segment in the to-be-processed text is a quoted text, when the current reference text contains the first text segment, the first text segment is taken as a second text segment, and the second word distribution information of the second text segment in the current reference text is located. If the similarity between the first word distribution information and the second word distribution information is high, it can be determined that the text content associated with the first text segment is the quoted text quoted from the current reference text by the to-be-processed text.
[0043] S130, the quoted text is removed from the to-be-processed text to obtain a to-be-checked text, and the associated text corresponding to the quoted text is removed from the current reference text to obtain a to-be-compared text.
[0044] The to-be-checked text can be understood as the to-be-processed text after removing the cited text. The associated text can be understood as the text content in the current reference text corresponding to the cited text. The to-be-compared text can be understood as the current reference text after removing the associated text.
[0045] In actual application, the to-be-processed text may cite the text information in other articles in the writing process to enrich the text content of the to-be-processed text. However, for the cited text, the to-be-processed text should be removed to obtain the to-be-checked text when the to-be-processed text is checked, and the to-be-checked text is checked based on the to-be-checked text to obtain a more accurate check result. Correspondingly, when the to-be-processed text is checked based on the current reference text, if the current reference text contains the associated text corresponding to the cited text, the associated text should also be removed from the current reference text to obtain the to-be-compared text. Further, the to-be-checked text is checked based on the to-be-compared text.
[0046] The advantage of such a setting is that the to-be-processed text inevitably needs to cite the content of other articles in the writing process. If the to-be-processed text is directly checked, the check result of the to-be-processed text will be biased high, that is, there is a problem of inaccurate check result. In the present scheme, the cited text is removed from the to-be-processed text to obtain the to-be-checked text, and the to-be-checked text is checked, which can effectively reduce the influence of the cited text on the to-be-processed text in the checking process, so as to make the check result of the to-be-processed text more accurate.
[0047] It should be noted that the so-called removal of the cited text from the to-be-processed text and the removal of the associated text from the current reference text in the present technical solution is for the purpose of more conveniently understanding the checking process of the present technical solution, rather than actually removing the cited text and the associated text. For example, after determining the cited text and the associated text, they can be specially marked, so that when the text information carrying the special mark is detected during the text checking, the influence weight of the corresponding content on the text is set to the lowest, so as to reduce the influence of the cited text on the to-be-processed text.
[0048] S140, determining a text similarity between the to-be-checked text and the to-be-compared text based on at least one similarity algorithm.
[0049] The text similarity can be understood as an index for representing the similarity between the to-be-checked text and the to-be-compared text. The higher the text similarity, the higher the similarity between the to-be-checked text and the to-be-compared text, and vice versa. The similarity algorithm can be a cosine similarity algorithm, a text similarity algorithm, and a distance similarity algorithm.
[0050] Optionally, the text similarity between the to-be-duplicated text and the to-be-compared text is determined based on at least one similarity algorithm, including: performing vectorization processing on the to-be-duplicated text and the to-be-compared text to obtain a first text vector corresponding to the to-be-duplicated text and a second text vector corresponding to the to-be-compared text; determining a vector similarity between the first text vector and the second text vector based on at least one similarity algorithm, and taking the vector similarity as the text similarity between the to-be-duplicated text and the to-be-compared text.
[0051] Specifically, the to-be-duplicated text is subjected to vectorization processing to obtain a first text vector corresponding to the to-be-duplicated text. The to-be-duplicated text includes at least one text subsequence, and the first text vector includes a first subsequence vector corresponding to each text subsequence. Similarly, the to-be-compared text is subjected to vectorization processing to obtain a second text vector. The to-be-compared text includes at least one text subsequence, and the second text vector includes a second subsequence vector corresponding to each text subsequence. The first text vector and the second text vector are subjected to similarity calculation based on at least one similarity algorithm to obtain a vector similarity, and the vector similarity is taken as the text similarity between the to-be-duplicated text and the to-be-compared text, i.e., the text similarity between the to-be-processed text and the current reference text.
[0052] The technical scheme of the embodiment of the present application obtains the to-be-processed text and at least one reference text associated with the to-be-processed text. In actual application, after the to-be-processed text is obtained, at least one text is called from a pre-constructed text library as a reference text for duplicate detection on the to-be-processed text. Further, for each reference text, the reference text in the to-be-processed text is determined according to the first word distribution information of the to-be-processed text and the second word distribution information of the current reference text. In the technical scheme, in order to reduce the influence of the reference content in the to-be-processed text on the duplicate detection result, after the to-be-processed text is divided by word segmentation, at least one first text segmentation is obtained, and the word distribution information corresponding to each first text segmentation is determined according to the position distribution information of each first text segmentation in the to-be-processed text, so as to determine the word distribution information corresponding to the to-be-processed text according to the word distribution information of each first text segmentation. Similarly, the same text processing is performed on the current reference text to determine the second word distribution information corresponding to the current reference text. Further, based on the similarity of the first word distribution information and the second word distribution information, the reference text is determined from the to-be-processed text, and the associated text corresponding to the reference text in the current reference text is determined. The reference text is removed from the to-be-processed text to obtain a to-be-detected text, and the associated text corresponding to the reference text is removed from the current reference text to obtain a to-be-compared text. Based on at least one similarity algorithm, the text similarity of the to-be-detected text and the to-be-compared text is determined. The advantage of this setting is that when the to-be-compared text is used for duplicate detection on the to-be-detected text, the influence of the reference content on the duplicate detection result of the to-be-processed text can be effectively reduced, and a more accurate duplicate detection result is obtained. The problem that the duplicate detection result of the to-be-detected text is inaccurate due to the reference text is solved. By removing the reference text and performing duplicate detection on the text after removing the reference text, the effect of obtaining a more accurate duplicate detection result is achieved.
[0053] Embodiment two
[0054] Figure 2 A flowchart of a text similarity determination method provided for the second embodiment of the present application. Optionally, the first word distribution information of the to-be-processed text and the second word distribution information of the current reference text are refined.
[0055] As Figure 2 shown, the method comprises:
[0056] S210, obtaining a to-be-processed text and at least one reference text associated with the to-be-processed text.
[0057] S220, performing a jiba word segmentation processing on the target text to obtain at least one to-be-screened segmentation.
[0058] The target text is the to-be-processed text or the current reference text.
[0059] It should be noted that when the current reference text is used to perform the duplicate detection on the to-be-processed text, similar operations are needed to be performed on the current reference text and the to-be-processed text. Therefore, in this embodiment, the target text is taken as the to-be-processed text or the current reference text as an example to illustrate, so as to make the technical solution more clearly understood.
[0060] The to-be-screened word segmentation can be understood as word segmentation of the to-be-processed text. If the target text is the to-be-processed text, the to-be-screened word segmentation can be understood as the first word segmentation text in the to-be-processed text. If the target text is the current reference text, the to-be-screened word segmentation can be understood as the second word segmentation text in the current reference text.
[0061] In this embodiment, the target text is taken as the to-be-processed text for illustration.
[0062] Specifically, after obtaining the to-be-processed text, the to-be-processed text is preprocessed. For example, stop words such as "of", "is", "in" and the like are removed. Such words usually have a very high frequency of occurrence in the text, but do not contain specific meanings, and can be regarded as noise in the text. Therefore, such words need to be removed when the to-be-processed text is preprocessed. Further, the to-be-processed text is segmented by using the jieba word segmentation to obtain at least one to-be-screened word segmentation.
[0063] In S230, a score attribute corresponding to each to-be-screened word segmentation is determined based on a term frequency-inverse document frequency algorithm, and at least one to-be-used word segmentation is determined from the at least one to-be-screened word segmentation according to the score attribute.
[0064] The to-be-used word segmentation includes a high-frequency word segmentation or a professional term word segmentation. In this technical solution, the higher the score attribute corresponding to the to-be-screened word segmentation, the more it can be used to represent the text content of the to-be-processed text. Therefore, the to-be-screened word segmentation with a higher score attribute is taken as the to-be-used word segmentation.
[0065] The term frequency-inverse document frequency (TF-IDF) is a weighting technique used in information retrieval and text mining, which can be used to evaluate the importance of a word to a document set or a corpus. The importance of a word is proportional to the number of times it appears in a document, but inversely proportional to the frequency of its appearance in the corpus. In short, if a word appears more frequently in an article and rarely appears in other articles, it is considered that the word or phrase has good class discrimination ability, and the score attribute of the word or phrase in the article is higher.
[0066] Specifically, after obtaining at least one to-be-screened word, in order to facilitate the calculation of the score attribute corresponding to each to-be-screened word, the to-be-screened words are vectorized, and the score attribute corresponding to each to-be-screened word is calculated based on the TF-IDF algorithm. The score attribute corresponding to each to-be-screened word is obtained. It can be understood that the score attribute can be used to represent the importance of the corresponding to-be-screened word in the to-be-processed text. The higher the score attribute, the higher the importance of the corresponding to-be-screened word in the to-be-processed text. Further, based on the score attribute of each to-be-screened word, at least one to-be-used word is screened from at least one to-be-screened word, for example, the to-be-screened words are sorted according to the score attribute, and the top 50% of the to-be-screened words are selected as the to-be-used words.
[0067] S240, determining target word distribution information corresponding to the target text according to the word distribution information of the at least one to-be-used word.
[0068] The target word distribution information includes first word distribution information corresponding to the to-be-processed text or second word distribution information corresponding to the current reference text. In this embodiment, the target word distribution information refers to the first word distribution information in the to-be-processed text.
[0069] In actual application, after obtaining at least one to-be-used word, whether there is a reference text in the to-be-processed text can be determined according to the distribution information of each to-be-used word in the to-be-processed text. Specifically, in order to reduce the influence of the reference text in the to-be-processed text on the duplicate detection result of the to-be-processed text, the word distribution information of each to-be-used word is determined to determine the first word distribution information corresponding to the to-be-processed text.
[0070] Taking the word distribution information corresponding to one of the to-be-used words as an example, the target word distribution information corresponding to the target text is determined according to the word distribution information of the at least one to-be-used word, which includes: for each to-be-used word, determining a local text window corresponding to the current word, and determining a nearest neighbor window corresponding to the local text window; performing linear normalization on the window distance between the local text window and the nearest neighbor window to obtain the word distribution information corresponding to the current word; and determining the target word distribution information corresponding to the target text according to the word distribution information of the at least one to-be-used word.
[0071] The local text window refers to a text window containing the current word. The local text window has a certain window length, which can be determined according to the number of words that can be contained in the window, for example, the window length is set to 11, which means that the number of words that can be contained in the local text window corresponding to the current word is 11. The nearest neighbor window can be understood as a text window containing the current word in the to-be-processed text and having the closest window distance from the local text window.
[0072] In a specific example, Figure 3 As shown, the window length of the local text window corresponding to the current segmentation is pre-set to L = 11. That is, the local distribution is measured using the 10 segmentations on both sides of the current segmentation. During the text duplication check, the local text window corresponding to the current segmentation is used to determine the similarities between the texts. For example, for the same quoted part in different texts, since the narrative order of the quoted underwear Rong may be different, the window similarity between the local text window and the nearest neighbor window can be used to determine the same aspect description content in the text. At the same time, the similarity of words can be measured when there are specific fields in the text.
[0073] Specifically, after obtaining the local text window corresponding to the current word segmentation, determining the nearest neighbor window corresponding to the local text window includes: determining at least one to-be-compared window corresponding to the local text window; for each to-be-compared window, determining the to-be-compared window distance between the current to-be-compared window and the local window based on the difference in score attributes between the to-be-used word segmentation contained in the current to-be-compared window and the current word segmentation; and determining the to-be-compared window corresponding to the smallest to-be-compared window distance among the at least one to-be-compared window as the nearest neighbor window corresponding to the local text window.
[0074] The window to be compared can be understood as the text window containing the current segmentation except the local text window corresponding to the current segmentation. The window distance to be compared can be understood as the window distance between the window to be compared and the local text window.
[0075] For example, if the current word segmentation is at the first paragraph of the text to be processed, the local text window containing the current word segmentation will be used as the local text window, and the other text windows in the text to be processed that contain the current word segmentation, except for the local window, will be used as the windows to be compared. That is to say, in the text to be processed, the current word segmentation appears multiple times, and accordingly, the number of text windows containing the current word segmentation is also large. In order to distinguish the text windows, one of the text windows containing the current word segmentation will be used as the local text window, and the remaining text windows containing the current word segmentation will be used as the windows to be compared. It should be noted that in addition to the current word segmentation, the local text window and the window to be compared may also include other word segmentations to be used.
[0076] Specifically, when determining the nearest neighbor window corresponding to the local text window, it can be determined in the following manner:
[0077] After obtaining the text window corresponding to each to-be-used word in the to-be-processed text, the nearest neighbor window corresponding to the local text window of the current word is determined. Specifically, the window distance between each to-be-compared window (i.e., the text window corresponding to the to-be-used word) and the local distribution window is compared in sequence. For the to-be-used words between the to-be-compared window and the local text window, the distance between the same to-be-used words is 0. For different to-be-used words between the to-be-compared window and the local text window, the correspondence is determined according to the sequence from left to right of the window, and the difference between the score attributes of the two to-be-used words is taken as the to-be-compared window distance between the two windows. After determining the to-be-compared window distance between each to-be-compared window and the local text window, the to-be-compared window corresponding to the smallest to-be-compared window distance is determined as the nearest neighbor window.
[0078] Specifically, after obtaining the nearest neighbor window of each to-be-used word corresponding to the local text window, there is a to-be-compared window distance for each pair of nearest neighbor windows. The distance d i between the local text window D j of the i-th to-be-used word and the local text window D ij of the nearest neighbor word j corresponding thereto is:
[0079]
[0080] wherein d ij represents the to-be-compared window distance between the local text window and the nearest neighbor window, TF-IDF iq represents the score attribute corresponding to the i-th to-be-used word, TF-IDF jq represents the score attribute corresponding to the nearest neighbor word j.
[0081] On this basis, the to-be-compared window distance between the nearest neighbor windows corresponding to the current word is linearly normalized to obtain the word distribution information corresponding to the current word:
[0082]
[0083] wherein, represents the word distribution information corresponding to the current word, d ij represents the to-be-compared window distance between the local text window and the nearest neighbor window, and Norm represents the norm symbol.
[0084] Based on the above operation, the position information of the local text window corresponding to the current word in the to-be-processed text can be determined, and further the word distribution information of the current word in the to-be-processed text can be determined. Based on this, the first word distribution information corresponding to the to-be-processed text can be determined according to the word distribution information of each to-be-used word.
[0085] S250, determining the quoted text in the text to be processed.
[0086] In practical applications, in order to eliminate the quoted text in the text to be processed, it is necessary to accurately determine the quoted text in the text to be processed. Optionally, determining the quoted text in the text to be processed comprises: determining at least one nearest neighbor curve associated with the local text window; performing text segmentation on the target text according to the number of at least one nearest neighbor curve to obtain at least one text subsequence; and determining the quoted text in the text to be processed based on the at least one text subsequence.
[0087] The target text is the text to be processed. The text subsequence can be understood as a subsequence obtained by dividing the text to be processed into words, phrases or paragraphs.
[0088] In practical applications, by using the window distribution information corresponding to the text window corresponding to each to-be-used word segmentation, the number of nearest neighbor curves corresponding to the to-be-used word segmentation can be corrected, and the text to be processed can be segmented according to the number of nearest neighbor curves.
[0089] Specifically, referring again to Figure 3 After obtaining the nearest neighbor relationship of the local text window corresponding to each to-be-used word segmentation, the text can be segmented by the nearest neighbor curve between the local text windows. In the corresponding relationship between the nearest neighbor windows, there are a plurality of corresponding nearest neighbor relationships above each to-be-used word segmentation, that is, there are a plurality of nearest neighbor relationships between the local text windows on both sides of the local text window of each to-be-used word segmentation, and the number of nearest neighbor relationships on both sides can be used as a text segmentation standard for the text to be processed. It can be understood that the more the number of nearest neighbor relationships on both sides of a to-be-used word segmentation, the more similar the range of the to-be-used word segmentation, and the less the number of nearest neighbor relationships on both sides, the less the content on both sides of the to-be-used word segmentation is associated. Then, the number of corresponding nearest neighbor windows on both sides of each to-be-used word segmentation can be used to measure the segmentation degree of each to-be-used word segmentation. The number of nearest neighbor relationships N i on both sides of each to-be-used word segmentation is linearly normalized to obtain the segmentable degree ξ i of each to-be-used word segmentation. For the to-be-used word segmentation with a segmentation degree higher than 0.5, the segmentation result is obtained as a segmentation point, that is, at least one text subsequence is obtained.
[0090] Optionally, based on the at least one text subsequence, the cited text in the to-be-processed text is determined, including: for each text subsequence, determining a score attribute of the current subsequence in the local text window corresponding to the at least one to-be-used word, and determining an average score attribute corresponding to the current subsequence; if the average score attribute is less than a preset score attribute, the current subsequence is determined as the cited text.
[0091] After obtaining all the segmentation results, for each segmented text subsequence, the average score of the entire text subsequence is obtained through the average TF-IDF score (i.e., the average score attribute corresponding to the plurality of to-be-used words) of each local text window in the to-be-used word, and then the text subsequence with low importance is determined through the average score attribute of each text subsequence, and is removed. The average score attribute of the mth text subsequence in the total text subsequence is normalized by the score of the total text subsequence, and the importance of each text subsequence ε m :
[0092]
[0093] wherein, TF-IDF wn represents the score attribute of the nth to-be-used word in the wth local text window in the mth text subsequence, and Norm represents a norm symbol.
[0094] In actual application, at least one text subsequence is obtained after the to-be-processed text is divided, and at least one local text window corresponding to the to-be-used word is included in each text subsequence. Further, taking one of the text subsequences as the current subsequence as an example, the number of local text windows in the current subsequence is determined, and the corresponding local text window score attribute is determined according to the score attribute of the to-be-used word in each local text window. Further, the average score attribute corresponding to all local text windows in the current subsequence is determined, and if the average score attribute is less than a preset score attribute, the current subsequence is determined as the cited text.
[0095] The advantage of such setting is that by determining the number of nearest neighbor relationships of the text windows on both sides of each local text window to divide the to-be-processed text, the to-be-processed text can be divided into different text subsequences. In the subsequent comparison process with the text information in the text library, the similarity is determined according to each text subsequence in the reference text, and the similarity is used as an optimization factor in the cosine similarity measurement process. The higher the similarity, the more important the local window of the divided text subsequence, so that in the subsequent process, the text similarity between the to-be-processed text and the reference text is determined according to the text similarity between the to-be-processed text and the reference text.
[0096] It should be noted that in actual operation, the text processing operation on the to-be-processed text is similar to the text processing operation on the current reference text, and accordingly, the cited text corresponding to the to-be-processed text can be obtained, and the associated text in the current reference text corresponding to the cited text can be determined.
[0097] S260, the cited text is removed from the to-be-processed text to obtain a to-be-duplicated text, and the associated text corresponding to the cited text is removed from the current reference text to obtain a to-be-compared text.
[0098] S270, based on at least one similarity algorithm, the text similarity of the to-be-duplicated text and the to-be-compared text is determined.
[0099] On the basis of the above examples, each text subsequence contains one or more local text windows corresponding to the to-be-used word segmentation, and for each text subsequence, the more the text subsequence exists in the current reference text, the higher the optimization factor of the to-be-used word segmentation contained in the text subsequence in the cosine similarity measurement.
[0100] Exemplarily, the optimization factor of the mth text subsequence in the to-be-duplicated text can be represented by the following formula:
[0101]
[0102] wherein, N' md represents the number of text subsequences in the current reference text with a similarity higher than the preset similarity to the mth text subsequence, N d represents the number of subsequences in the current reference text including the length of the mth text subsequence.
[0103] After obtaining the cosine similarity optimization factor, the cosine similarity of the feature vectors of the to-be-duplicated text and the to-be-compared text is obtained by multiplying the score attribute of each dimension (i.e., the to-be-used word segmentation) and η, and then calculating the feature vector angle to obtain the accurate cosine similarity. The obtained cosine similarity is the text similarity of the to-be-duplicated text and the current reference text, that is, the text similarity of the to-be-processed text and the current reference text.
[0104] The technical scheme of the embodiment of the present application obtains the to-be-processed text and at least one reference text associated with the to-be-processed text, in actual application, after obtaining the to-be-processed text, at least one text is called from the pre-constructed text library as a reference text, which is used for duplicate detection on the to-be-processed text. Further, for each reference text, according to the first word distribution information of the to-be-processed text and the second word distribution information of the current reference text, the reference text in the to-be-processed text is determined. In the technical scheme, in order to reduce the influence of the reference content in the to-be-processed text on the duplicate detection result, after the to-be-processed text is divided by word segmentation, at least one first text segmentation is obtained, and the word distribution information corresponding to each first text segmentation is determined according to the position distribution information of each first text segmentation in the to-be-processed text, so as to determine the word distribution information corresponding to the to-be-processed text according to the word distribution information of each first text segmentation. Similarly, the same text processing is performed on the current reference text to determine the second word distribution information corresponding to the current reference text. Further, based on the similarity of the first word distribution information and the second word distribution information, the reference text is determined from the to-be-processed text, and the associated text corresponding to the reference text in the current reference text is determined. The reference text is removed from the to-be-processed text to obtain a to-be-detected text, and the associated text corresponding to the reference text is removed from the current reference text to obtain a to-be-compared text. Based on at least one similarity algorithm, the text similarity of the to-be-detected text and the to-be-compared text is determined. The advantage of this setting is that when the to-be-compared text is used for duplicate detection on the to-be-detected text, the influence of the reference content on the duplicate detection result of the to-be-processed text can be effectively reduced, and a more accurate duplicate detection result can be obtained. The problem that the duplicate detection result of the to-be-detected text is inaccurate due to the reference text is solved. By removing the reference text and performing duplicate detection on the text after removing the reference text, the effect of obtaining a more accurate duplicate detection result is achieved.
[0105] Embodiment three
[0106] Figure 4 A structure diagram of a text similarity determination device provided by the third embodiment of the present application is shown in FIG. 3. Figure 4 As shown in the figure, the device comprises a text acquisition module 310, a text determination module 320, a text removal module 330 and a similarity determination module 340.
[0107] The text acquisition module 310 is configured to acquire the to-be-processed text and at least one reference text associated with the to-be-processed text.
[0108] The text determination module 320 is configured to determine the reference text in the to-be-processed text according to the first word distribution information of the to-be-processed text and the second word distribution information of the current reference text for each reference text.
[0109] The text elimination module 330 is configured to eliminate the reference text from the to-be-de-duplication text, and eliminate the associated text corresponding to the reference text from the current reference text to obtain the to-be-compared text;
[0110] The similarity determination module 340 is configured to determine the text similarity between the to-be-de-duplication text and the to-be-compared text based on at least one similarity algorithm.
[0111] The technical scheme of the embodiment of the application obtains the to-be-processed text and at least one reference text associated with the to-be-processed text. In actual application, after obtaining the to-be-processed text, at least one text is retrieved from a pre-constructed text library as a reference text for de-duplication detection of the to-be-processed text. Further, for each reference text, the reference text in the to-be-processed text is determined according to the first word distribution information of the to-be-processed text and the second word distribution information of the current reference text. In the technical scheme, in order to reduce the influence of the reference content in the to-be-processed text on the de-duplication result, after the to-be-processed text is divided by word segmentation, at least one first text segmentation is obtained, and the word distribution information corresponding to each first text segmentation is determined according to the position distribution information of each first text segmentation in the to-be-processed text, so as to determine the to-be-processed text corresponding to the word distribution information according to the word distribution information of each first text segmentation. Similarly, the same text processing is performed on the current reference text to determine the second word distribution information corresponding to the current reference text. Further, the reference text is determined from the to-be-processed text based on the similarity of the first word distribution information and the second word distribution information, and the associated text corresponding to the reference text in the current reference text is determined. The reference text is eliminated from the to-be-processed text to obtain the to-be-de-duplication text, and the associated text corresponding to the reference text is eliminated from the current reference text to obtain the to-be-compared text. The text similarity between the to-be-de-duplication text and the to-be-compared text is determined based on at least one similarity algorithm. The advantage of this setting is that when the to-be-compared text is used to de-duplicate the to-be-de-duplication text, the influence of the reference content on the de-duplication result of the to-be-processed text can be effectively reduced, and a more accurate de-duplication result can be obtained. The problem that the de-duplication result of the to-be-de-duplication text is inaccurate due to the reference text is solved. By eliminating the reference text and de-duplicating the text after eliminating the reference text, a more accurate de-duplication result is obtained.
[0112] Optionally, the text acquisition module comprises: a to-be-screened segmentation determination submodule configured to perform a jieba word segmentation processing on a target text to obtain at least one to-be-screened segmentation; wherein the target text is the to-be-processed text or the current reference text.
[0113] The to-be-used word determination sub-module is configured to determine a score attribute corresponding to each to-be-screened word based on a term frequency-inverse document frequency algorithm, and determine at least one to-be-used word from the at least one to-be-screened word according to the score attribute; wherein the to-be-used word includes a high-frequency word or a professional term word;
[0114] The word distribution information determination sub-module is configured to determine target word distribution information corresponding to the target text according to the word distribution information of the at least one to-be-used word; wherein the target word distribution information includes first word distribution information corresponding to the to-be-processed text or second word distribution information corresponding to the current reference text.
[0115] Optionally, the word distribution information determination sub-module includes a nearest neighbor window determination unit configured to determine a local text window corresponding to the current word for each to-be-used word, and determine a nearest neighbor window corresponding to the local text window.
[0116] The normalization processing unit is configured to perform linear normalization processing on a window distance between the local text window and the nearest neighbor window to obtain word distribution information corresponding to the current word.
[0117] The word distribution information determination unit is configured to determine target word distribution information corresponding to the target text according to the word distribution information corresponding to the at least one to-be-used word.
[0118] Optionally, the nearest neighbor window determination unit includes a to-be-compared window determination sub-unit configured to determine at least one to-be-compared window corresponding to the local text window.
[0119] The window distance determination sub-unit is configured to determine a to-be-compared window distance between the current to-be-compared window and the local window according to a difference in score attribute between the to-be-used word contained in the current to-be-compared window and the current word for each to-be-compared window.
[0120] The nearest neighbor window determination sub-unit is configured to determine, from the at least one to-be-compared window, a to-be-compared window corresponding to the smallest to-be-compared window distance as the nearest neighbor window corresponding to the local text window.
[0121] Optionally, the text determination module includes a nearest neighbor curve determination sub-module configured to determine at least one nearest neighbor curve associated with the local text window.
[0122] The text sub-sequence determination sub-module is configured to perform text segmentation on the target text according to the number of the at least one nearest neighbor curve to obtain at least one text sub-sequence.
[0123] The reference text determination sub-module is configured to determine a reference text in the to-be-processed text based on the at least one text sub-sequence.
[0124] Optionally, the cited text determining sub-module comprises: a score attribute determining unit configured to determine, for each text sub-sequence, a score attribute of the current sub-sequence in a local text window corresponding to at least one to-be-used segmentation, and determine an average score attribute corresponding to the current sub-sequence.
[0125] The cited text determining unit is configured to determine the current sub-sequence as the cited text if the average score attribute is less than a preset score attribute.
[0126] Optionally, the similarity determining module comprises: a text vector determining sub-module configured to perform vectorization processing on the to-be-duplicated text and the to-be-compared text to obtain a first text vector corresponding to the to-be-duplicated text and a second text vector corresponding to the to-be-compared text.
[0127] The similarity determining sub-module is configured to determine a vector similarity between the first text vector and the second text vector based on at least one similarity algorithm, and take the vector similarity as the text similarity between the to-be-duplicated text and the to-be-compared text.
[0128] The text similarity determining apparatus provided in the embodiments of the present application can perform the text similarity determining method provided in any of the embodiments of the present application, and has the function modules and beneficial effects corresponding to the performing method.
[0129] Embodiment Four
[0130] Figure 5 A structural schematic diagram of an electronic device 10 of an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.
[0131] As Figure 5As shown, the electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., connected to the at least one processor 11 in communication. The memory stores a computer program executable by the at least one processor 11, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded into the random access memory (RAM) 13 from the storage unit 18. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0132] Various components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, a speaker, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0133] The processor 11 can be various general and / or special-purpose processing components having processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the method of determining text similarity.
[0134] In some embodiments, the method of determining text similarity can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method of determining text similarity described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the method of determining text similarity by any other appropriate means, such as by means of firmware.
[0135] The various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, specially designed application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0136] Computer programs implementing the method of determining text similarity of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program running on the processor implements the functions / operations specified in the flowcharts and / or the block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package and partially on a remote machine or entirely on a remote machine or server.
[0137] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0138] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0139] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0140] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0141] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the present disclosure are achieved, and the present disclosure is not limited herein.
[0142] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and principles of the disclosure. Accordingly, the disclosure is not limited to the specific embodiments described above, but only by the scope of the appended claims.
Claims
1. A method for determining text similarity, characterized in that: include: Acquire a text to be processed and at least one reference text associated with the text to be processed; For each reference text, determining a reference text in the text to be processed according to the first word distribution information of the text to be processed and the second word distribution information of the current reference text; Eliminating the reference text from the to-be-processed text to obtain a to-be-checked text, and eliminating associated text corresponding to the reference text from the current reference text to obtain a to-be-compared text; Determining the text similarity between the text to be checked for duplicates and the text to be compared based on at least one similarity algorithm; Wherein, the method according to the first word distribution information of the text to be processed and the second word distribution information of the current reference text includes: performing a jieba word segmentation process on the target text to obtain at least one word to be screened; wherein, the target text is the text to be processed or the current reference text; based on the word frequency-inverse file frequency algorithm, determining the score attributes corresponding to each word to be screened, and determining at least one word to be used from at least one word to be screened according to each score attribute; wherein, the word to be used includes a high-frequency word or a professional term word; according to the word distribution information of at least one word to be used, determining the target word distribution information corresponding to the target text; wherein, the target word distribution information includes the first word distribution information corresponding to the text to be processed or the second word distribution information corresponding to the current reference text; The step of determining target word distribution information corresponding to the target text based on word distribution information of at least one of the to-be-used segmentations includes: for each to-be-used segmentation, determining a local text window corresponding to the current segmentation, and determining a nearest neighbor window corresponding to the local text window; performing linear normalization processing on the window distance between the local text window and the nearest neighbor window to obtain word distribution information corresponding to the current segmentation; and determining target word distribution information corresponding to the target text based on word distribution information corresponding to at least one of the to-be-used segmentations. The determining of the nearest neighbor window corresponding to the local text window comprises: determining at least one to-be-compared window corresponding to the local text window; for each to-be-compared window, determining a to-be-compared window distance between the current to-be-compared window and the local text window based on a difference in a score attribute between a to-be-compared word contained in the current to-be-compared window and the current word; and determining the to-be-compared window corresponding to the smallest to-be-compared window distance among at least one to-be-compared window as the nearest neighbor window corresponding to the local text window; The method of determining the quoted text in the text to be processed includes: determining at least one nearest neighbor curve associated with the local text window; performing text segmentation on the target text according to the number of at least one nearest neighbor curve to obtain at least one text subsequence; and determining the quoted text in the text to be processed based on at least one text subsequence.
2. The method according to claim 1, characterized in that The step of determining a reference text in the text to be processed based on at least one of the text subsequences includes: For each text subsequence, determining a score attribute of the current subsequence in at least one local text window corresponding to a to-be-used word segmentation, and determining an average score attribute corresponding to the current subsequence; If the average score attribute is less than a preset score attribute, the current subsequence is determined to be a quoted text.
3. The method according to claim 1, characterized in that The determining of the text similarity between the text to be checked for duplicates and the text to be compared based on at least one similarity algorithm includes: Performing vectorization processing on the text to be checked for duplicates and the text to be compared to obtain a first text vector corresponding to the text to be checked for duplicates and a second text vector corresponding to the text to be compared; Based on at least one similarity algorithm, a vector similarity between the first text vector and the second text vector is determined, and the vector similarity is used as the text similarity between the text to be checked for duplicates and the text to be compared.
4. A device for determining text similarity, characterized in that: include: A text acquisition module, configured to acquire a text to be processed and at least one reference text associated with the text to be processed; a text determination module, configured to determine, for each reference text, a reference text in the text to be processed based on the first word distribution information of the text to be processed and the second word distribution information of the current reference text; A text elimination module is used to eliminate the reference text from the text to be processed to obtain a text to be checked for duplicates, and to eliminate associated text corresponding to the reference text from the current reference text to obtain a text to be compared; A similarity determination module, configured to determine the text similarity between the text to be checked for duplicates and the text to be compared based on at least one similarity algorithm; The text acquisition module includes: a to-be-screened segmentation determination submodule, configured to perform segmentation processing on a target text to obtain at least one to-be-screened segmentation; wherein the target text is the to-be-processed text or the current reference text; a to-be-used segmentation determination submodule, configured to determine, based on a word frequency-inverse file frequency algorithm, a score attribute corresponding to each to-be-screened segmentation, and determine at least one to-be-used segmentation from at least one to-be-screened segmentation according to each score attribute; wherein the to-be-used segmentation includes high-frequency segmentations or professional term segmentations; a word distribution information determination submodule, configured to determine target word distribution information corresponding to the target text based on word distribution information of at least one to-be-used segmentation; wherein the target word distribution information includes first word distribution information corresponding to the to-be-processed text or second word distribution information corresponding to the current reference text; The word distribution information determination submodule includes: a nearest neighbor window determination unit, for determining, for each word segment to be used, a local text window corresponding to the current word segment, and determining a nearest neighbor window corresponding to the local text window; a normalization processing unit, for performing linear normalization processing on the window distance between the local text window and the nearest neighbor window to obtain word distribution information corresponding to the current word segment; a word distribution information determination unit, for determining target word distribution information corresponding to the target text based on word distribution information corresponding to at least one word segment to be used; The nearest neighbor window determination unit includes: a to-be-compared window determination subunit, configured to determine at least one to-be-compared window corresponding to the local text window; a window distance determination subunit, configured to determine, for each to-be-compared window, a to-be-compared window distance between the current to-be-compared window and the local text window based on a difference in score attributes between a to-be-compared word contained in the current to-be-compared window and the current word; and a nearest neighbor window determination subunit, configured to determine, among at least one of the to-be-compared windows, a to-be-compared window corresponding to the smallest to-be-compared window distance as the nearest neighbor window corresponding to the local text window. Among them, the text determination module includes: a nearest neighbor curve determination submodule, which is used to determine at least one nearest neighbor curve associated with the local text window; a text subsequence determination submodule, which is used to perform text segmentation on the target text according to the number of at least one nearest neighbor curve to obtain at least one text subsequence; and a reference text determination submodule, which is used to determine the reference text in the text to be processed based on at least one text subsequence.
5. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can perform the method for determining text similarity according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for determining text similarity according to any one of claims 1 to 3 when executed.
Citation Information
Patent Citations
Reference-based literature searching method
CN110232120A
Paper query method and device, equipment and storage medium
CN114818701A