Text similarity recognition method and device, electronic equipment, storage medium and product

By obtaining multi-dimensional features of the text, including word, sentence and overall paragraph dimensional features, the problem of low accuracy in text similarity recognition in the existing technology is solved, and more efficient and accurate text similarity recognition is achieved.

CN120671657APending Publication Date: 2025-09-19MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410318340.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The existing text similarity recognition methods are mainly based on text vector representation and only consider the semantic level, resulting in low accuracy.

Method used

By obtaining the multidimensional features of the text to be processed, including word dimension features, sentence dimension features and overall paragraph dimension features, the similarity recognition results between texts are determined, and multidimensional feature extraction is used instead of vector representation.

Benefits of technology

The accuracy and recognition rate of text similarity recognition are improved, the multi-dimensional feature extraction speed is faster, and the recognition accuracy and efficiency are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671657A_ABST
    Figure CN120671657A_ABST
Patent Text Reader

Abstract

The invention provides a text similarity recognition method and device, electronic equipment, a storage medium and a product. The method comprises the steps of obtaining a first text and a second text to be processed; multi-dimensional features between the first text and the second text are determined, the multi-dimensional features comprise at least one of word dimension features, sentence dimension features and overall chapter dimension features, and the different dimension features are used for representing similarity degrees of the texts on different dimension levels; and according to the multi-dimensional features, determining a similarity recognition result of the first text and the second text, so that the accuracy and efficiency of similarity recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a text similarity recognition method, device, electronic device, storage medium, and product. Background Art

[0002] Currently, text similarity recognition is required in many scenarios, such as text classification and search. In related technologies, different texts are mainly represented by vectors, and the similarity between texts is determined based on the text vectors of the texts. However, this method is based on text vectors and only considers the semantic level of representing the entire text, so its accuracy is low. Summary of the Invention

[0003] The present disclosure provides a text similarity recognition method, device, electronic device, storage medium and product.

[0004] In a first aspect, the present disclosure provides a method for identifying text similarity, the method comprising:

[0005] Get the first text and the second text to be processed;

[0006] Determining multidimensional features between the first text and the second text, wherein the multidimensional features include at least one of the following: word dimension features, sentence dimension features, and overall text dimension features, and different dimension features are used to represent the degree of similarity between the texts at different dimensional levels;

[0007] Determine similar recognition results of the first text and the second text based on the multi-dimensional features.

[0008] In a second aspect, the present disclosure provides a text similarity recognition device, the device comprising:

[0009] An acquisition module, configured to acquire a first text and a second text to be processed;

[0010] a determination module, configured to determine multidimensional features between the first text and the second text, wherein the multidimensional features include at least one of the following: a word dimension feature, a sentence dimension feature, and an overall paragraph dimension feature, wherein different dimension features are used to represent the degree of similarity between the texts at different dimensional levels;

[0011] The recognition module is configured to determine similar recognition results of the first text and the second text based on the multi-dimensional features.

[0012] In a third aspect, the present disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor so that the at least one processor can execute the above-mentioned text similarity recognition method.

[0013] In a fourth aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-mentioned text similarity recognition method when executed by a processor.

[0014] In a fifth aspect, the present disclosure provides a computer program product comprising a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned text similarity recognition method.

[0015] In the embodiment of the present disclosure, the features between the first text and the second text are determined from multiple dimensions such as words, sentences, and the entire text, and then the similarity recognition results of the first text and the second text are determined through multi-dimensional features, taking into account multi-level and multi-dimensional information, thereby improving the accuracy of similarity recognition.

[0016] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent to those skilled in the art by describing detailed example embodiments with reference to the accompanying drawings. In the accompanying drawings:

[0018] Figure 1 A schematic diagram of an application scenario of a text similarity recognition method provided by an embodiment of the present disclosure;

[0019] Figure 2 A flowchart of a text similarity recognition method provided by an embodiment of the present disclosure;

[0020] Figure 3 A schematic diagram of a training process of a recognition model provided in an embodiment of the present disclosure;

[0021] Figure 4 A block diagram of a text similarity recognition device provided by an embodiment of the present disclosure;

[0022] Figure 5 A block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] To enable those skilled in the art to better understand the technical solutions of the present disclosure, exemplary embodiments of the present disclosure are described below in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0024] In the absence of conflict, the various embodiments of the present disclosure and the various features therein may be combined with each other.

[0025] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0026] The terms used herein are only used to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprising" and / or "made of" are used in this specification, the presence of the features, wholes, steps, operations, elements and / or components is specified, but the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof is not excluded. Similar words such as "connected" or "connected" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect.

[0027] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined as such herein.

[0028] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution complies with relevant national laws and regulations (for example, the "Information Security Technology Personal Information Security Specification", etc.). For example: corresponding prescribed measures are taken to control access to personal information; the display of personal information is subject to prescribed restrictions; the purpose of using personal information does not exceed the scope of direct or reasonable connection; when using personal information, clear identity reference is eliminated to avoid precise positioning of specific individuals.

[0029] At present, text similarity recognition is required in many scenarios. The text similarity recognition method in related technologies mainly represents different texts by vectors separately. For example, based on the encoding model, a word vector for each word in the text is generated, and a text vector is obtained based on the word vector of each word in the text. Then, based on the text vectors of different texts, the similarity between different texts is determined by methods such as cosine similarity calculation or Euclidean distance calculation. However, this method only considers the semantic level of representing the entire text, and has low accuracy.

[0030] Therefore, in response to this, an embodiment of the present disclosure provides a text similarity recognition method, which obtains a first text and a second text to be processed, determines the multidimensional features between the first text and the second text, and then determines the similarity recognition results of the first text and the second text based on the multidimensional features, wherein the multidimensional features include at least one of the following: word dimension features, sentence dimension features and overall chapter dimension features. In this way, the features of the text can be extracted from multiple dimensions such as words, sentences, and overall chapters, and the similarity of the text can be determined, thereby improving the accuracy of text similarity recognition. In addition, the multidimensional features do not require vector representation, the feature extraction speed is faster, and the recognition rate is improved.

[0031] Figure 1 The following schematically illustrates an application scenario diagram of the text similarity recognition method and apparatus provided by an embodiment of the present disclosure.

[0032] like Figure 1 As shown, an application scenario of an embodiment of the present disclosure may include a terminal device 101, a network 103, and a server 102. The network 103 is used as a medium for providing a communication link between the terminal device 101 and the server 102. The network 103 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0033] The user can use the terminal device 101 to interact with the server 102 via the network 103 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0034] The terminal device 101 may be any electronic device having a display screen and supporting web browsing, including but not limited to a smart phone, a tablet computer, a portable computer, a desktop computer, and the like.

[0035] The server 102 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the terminal device 101. The background management server may analyze and process received user requests and other data, and feed back the processing results (e.g., web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0036] It should be noted that the text similarity recognition method and apparatus provided in the embodiments of the present disclosure can be executed by the server 102. Accordingly, the text similarity recognition method and apparatus provided in the embodiments of the present disclosure can be set in the server 102. The text similarity recognition method and apparatus provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 102 and can communicate with the terminal device 101 and / or the server 102. Accordingly, the text similarity recognition method and apparatus provided in the embodiments of the present disclosure can also be set in a server or server cluster that is different from the server 102 and can communicate with the terminal device 101 and / or the server 102.

[0037] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0038] Figure 2 A flowchart of a text similarity recognition method provided by the embodiment of the present disclosure, referring to Figure 2 , the method comprising:

[0039] S201: Acquire a first text and a second text to be processed.

[0040] In the embodiments of the present disclosure, similarity recognition can be performed on text pairs in any scenario, for example, similarity recognition between letter texts, similarity recognition between article or paper texts, similarity recognition between contract texts, etc., without limitation.

[0041] S202: Determine multidimensional features between the first text and the second text, wherein the multidimensional features include at least one of the following: word dimension features, sentence dimension features, and overall paragraph dimension features, and different dimensional features are used to represent the degree of similarity between the texts at different dimensional levels.

[0042] S203: Determine similar recognition results of the first text and the second text based on the multi-dimensional features.

[0043] For step S203, the present disclosure provides a possible implementation method, including: based on a recognition model, taking multidimensional features as input, and performing similarity recognition on the first text and the second text according to the value of each feature in the multidimensional features, to obtain similarity recognition results of the first text and the second text; wherein the recognition model is obtained after training based on a training text pair set, and each training text pair in the training text pair set includes multidimensional features between the first training text and the second training text, as well as similarity labels.

[0044] In the embodiment of the present disclosure, a recognition model applicable to the application scenario can be pre-trained for a set of training text pairs in different application scenarios. The recognition model is used to identify whether the texts are similar. For example, the multidimensional features of the first text and the second text include 6 features. The values ​​of the 6 features can be converted into an array format as a feature value array. The feature value array is input into the recognition model, and then the recognition model analyzes and recognizes the feature value array to obtain a similar recognition result of the first text and the second text, that is, the recognition model can output a result of whether the first text and the second text are similar or dissimilar.

[0045] Furthermore, in the embodiment of the present disclosure, for the array after the values ​​of the multidimensional features are converted, the order of the features in the array is not restricted. Preferably, in actual applications, the order of the features in the number is the same as the order during the recognition model training.

[0046] In addition, in order to further improve the recognition accuracy of the model, the values ​​of the multidimensional features can be normalized before entering the recognition model. This can reduce the impact of excessive or insufficient weights of a single feature on similar recognition accuracy.

[0047] Furthermore, in the embodiments of the present disclosure, several possible application scenarios are provided. After determining similar recognition results between the first text and the second text based on multi-dimensional features, different subsequent processing can be performed for different application scenarios. Specifically:

[0048] In a possible embodiment, when the first text and the second text are letter texts, similar results are screened out from multiple letter texts and identified as similar letter texts, and the source of the screened letter files is determined, and it is judged whether the source is an abnormal source. When the source is determined to be an abnormal source, an alarm is issued.

[0049] For example, in the financial field, users may send letters to related companies based on third-party agents, which may cause certain troubles to the companies. Therefore, it is very important to identify whether the letters received are sent by third-party agents from a large number of letters, and then to carry out targeted processing. Since the letters sent by third-party agents have certain similarities, similar letters can be identified from a large number of received letters based on the text similarity recognition method in the embodiment of the present disclosure, and then the sources of these similar letters can be determined. When it is determined to be an abnormal source, for example, the abnormal source is a third-party agent, an alarm can be issued, and the relevant personnel or departments can track and handle it.

[0050] In a possible embodiment, it is determined whether there is plagiarism between the first text and the second text based on the similarity recognition result.

[0051] For example, in a plagiarism detection scenario, when it is determined that the first text and the second text are not similar, it is determined that there is no plagiarism between the first text and the second text, or when it is determined that the first text and the second text are similar, it is determined that there is plagiarism between the first text and the second text. In addition, when it is determined that plagiarism exists, the identified similar or repeated parts can also be output and displayed to facilitate users to obtain similar content more clearly and conveniently.

[0052] In the embodiment of the present disclosure, the multidimensional features can be calculated based on a pre-constructed calculation formula. The multidimensional features representing the similarity between the text pairs can be pre-constructed, and then the features between the first text and the second text can be extracted based on the calculation formula corresponding to each feature. Specifically, with respect to the above step S202, the present disclosure provides a possible implementation method for determining the multidimensional features between the first text and the second text, including:

[0053] 1) Segmenting the first text and the second text to obtain a first segmentation and the number of first segmentations of the first text, and a second segmentation and the number of second segmentations of the second text.

[0054] Furthermore, in the embodiment of the present disclosure, deduplication processing can be performed after word segmentation, and the number of occurrences of repeated word segmentations can be recorded, which can facilitate subsequent statistics and calculations and improve efficiency.

[0055] For example, the first text textA and the second text textB can be segmented and deduplicated based on the n-gram method to obtain the first segmentation corresponding to textA as textA_words, the number of first segmentations is M, and the second segmentation corresponding to textB is textB_words, the number of second segmentations is N.

[0056] There is no restriction on the length n of the segmentation in the n-gram method. For example, the value of n is 8, which can be set according to needs and actual experience.

[0057] 2) Determine the common words and the number of common words between the first text and the second text based on the first participle and the second participle.

[0058] For example, taking the intersection of the first participle textA_words and the second participle textB_words, we can get the common words as all_words = [k1, k2, ..., k m ], the total number of words is m.

[0059] 3) Segment the first text and the second text to obtain the first sentences and the number of first sentences contained in the first text, and the second sentences and the number of second sentences contained in the second text.

[0060] For example, the sentence can be divided into sentences according to the punctuation marks, and the first sentence contained in the first text is obtained as textA_s, and the number of the first sentences is T A The second sentence contained in the second text is textB_s, and the number of second sentences is T B .

[0061] 4) Obtain the total number of first characters included in the first text and the total number of second characters included in the second text.

[0062] For example, the first text contains a total of Q first characters A , the second text contains a total of Q second characters B .

[0063] The total number of the first characters and the total number of the second characters may preferably be the number of words.

[0064] 5) Determine multidimensional features between the first text and the second text based on the multidimensional information, wherein the multidimensional information includes at least two of the following: a first participle, the number of first participles, a second participle, the number of second participles, a common word, the number of common words, a first sentence, the number of first sentences, a second sentence, the number of second sentences, a total number of first characters, and a total number of second characters.

[0065] In this way, in the embodiment of the present disclosure, feature calculation can be performed based on the multi-dimensional information of the first text and the second text, thereby obtaining the required multi-dimensional features and improving the accuracy of similarity recognition.

[0066] The above step S202 is described in detail below with respect to the calculation formulas of each dimensional feature.

[0067] 1. In one possible embodiment, the multidimensional feature includes a first word feature representing a word dimensional feature. Determining the multidimensional feature between the first text and the second text based on the multidimensional information includes:

[0068] 1) Determine a first absolute difference between the first number of segmented words and the second number of segmented words, and determine a first product of the first number of segmented words and the second number of segmented words.

[0069] For example, the number of first segmentations is M and the number of second segmentations is N. Then, the first absolute difference between the number of first segmentations and the number of second segmentations is |MN|, and the first product is M*N.

[0070] 2) Obtain a second product between the first absolute difference and the number of common words.

[0071] For example, the number of common words is m, and the second product between the first absolute difference and the number of common words is m*|MN|.

[0072] 3) Determining a first word feature between the first text and the second text based on a ratio between the second product and the first product, wherein the value of the first word feature is negatively correlated with the degree of similarity at the word level.

[0073] For example, the first word feature is represented by a, and the calculation formula of the first word feature is:

[0074] In the embodiment of the present disclosure, the first word feature a can represent the absolute value of the difference between the proportions of common words in the vocabulary textA_words and textB_words. Therefore, the smaller the value of the first word feature, the more common words are contained between the first text and the second text, and the greater the similarity between the first text and the second text.

[0075] 2. In one possible embodiment, the multidimensional feature includes a second word feature representing a word dimensional feature. Determining the multidimensional feature between the first text and the second text based on the multidimensional information includes:

[0076] 1) Determine the first occurrence number of the common word in the first text and the second occurrence number of the common word in the second text.

[0077] In an embodiment of the present disclosure, after the first text and the second text are segmented, multiple identical segmented words may appear, and the number of occurrences of each segmented word can be recorded, thereby obtaining the first number of occurrences of the common word in the first text and the second number of occurrences in the second text. In the case where there are multiple common words, the first number of occurrences and the second number of occurrences of each common word can be obtained respectively, and then the sum of the first number of occurrences of the multiple common words and the sum of the second number of occurrences of the multiple common words can be calculated.

[0078] For example, if the number of common words is m, the first occurrence of m common words in the first text textA is text_A m , the second occurrence number of m common words in the second text textB is text_B m .

[0079] 2) Determine a second absolute difference between the first number of occurrences and the second number of occurrences, and determine a maximum value between the first number of occurrences and the second number of occurrences.

[0080] For example, the second absolute difference between the first and second occurrences is |text_A m -text_B m |, the maximum value between the first and second occurrences is max(text_A m ,text_B m ).

[0081] 3) Determine a second word feature between the first text and the second text based on the proportion of the second absolute difference in the maximum value, wherein the value of the second word feature is negatively correlated with the degree of similarity at the word level.

[0082] For example, the second word feature is represented by b, and the calculation formula of the second word feature is:

[0083] In the embodiment of the present disclosure, the second word feature b may indicate that the absolute difference in the number of occurrences of the common word in textA and textB is within max(text_A m ,text_B m ), the smaller the proportion, the closer the number of times the common words appear in the first text and the second text, and therefore the greater the similarity between the first text and the second text.

[0084] 3. In one possible embodiment, the multidimensional feature includes a first sentence feature representing a sentence dimensional feature. Determining the multidimensional feature between the first text and the second text based on the multidimensional information includes:

[0085] 1) Determining the total number of first sentences in the first sentence that meet a threshold condition, and the total number of second sentences in the second sentence that meet a threshold condition, wherein the threshold condition represents that the ratio of the sum of the lengths of the character strings occupied by the common words in the sentences to the total length of the character strings of the sentences is greater than or equal to a preset threshold.

[0086] For example, the threshold condition is Among them, l s Indicates the total string length of a sentence s, Z m It represents the sum of the lengths of the character strings occupied by the m common words in the sentence, and α is a preset threshold, where the value of α is, for example, 0.8. There is no specific limitation and it can be set according to experience and needs.

[0087] Furthermore, in the embodiment of the present disclosure, all first sentences in the first text can be traversed to determine whether each first sentence meets the threshold condition, and the total number of first sentences meeting the threshold condition S is obtained. A Similarly, traverse all the second sentences in the second text and obtain the total number of second sentences that meet the threshold condition S B .

[0088] 2) Determine a first sentence feature between the first text and the second text based on the absolute difference between the total number of the first sentences and the total number of the second sentences, wherein the value of the first sentence feature is negatively correlated with the degree of similarity at the sentence level.

[0089] For example, if the feature of the first sentence is c, the calculation formula of the feature of the second sentence is: c = |S A -S B |.

[0090] In an embodiment of the present disclosure, the first sentence feature can represent the absolute value of the difference in the number of sentences that meet the threshold condition in the first text and the second text. The smaller the value of the first sentence feature, the more similar sentences there are in the first text and the second text, and therefore the greater the similarity between the first text and the second text.

[0091] 4. In one possible embodiment, the multidimensional feature includes a second sentence feature representing a sentence dimensional feature. Determining the multidimensional feature between the first text and the second text based on the multidimensional information includes:

[0092] 1) Determining a first maximum number of consecutive sentences that meet a threshold condition in a first sentence, and a second maximum number of consecutive sentences that meet a threshold condition in a second sentence, wherein the threshold condition represents a ratio of the sum of the lengths of character strings occupied by common words in the sentences to the total length of the character strings of the sentences being greater than or equal to a preset threshold.

[0093] For example, if there are three consecutive first sentences in the first text that all meet the threshold condition, then the number of consecutive sentences is 3. Thus, by traversing all the first sentences in the first text in sequence, the first maximum number of consecutive sentences that meet the threshold condition can be determined. The first maximum number of consecutive sentences is recorded as L(S A ), similarly, traverse all second sentences in the second text and determine the second maximum number of consecutive sentences L (S B ).

[0094] 2) Obtaining a first difference between the first maximum number of consecutive sentences and the second maximum number of consecutive sentences.

[0095] For example, the first difference is L(S A )-L(S B ).

[0096] 3) Obtaining a second difference between the first number of sentences and the second number of sentences, wherein the second difference is used to represent the first length adjustment factor.

[0097] For example, the first sentence number is T A , the second number of sentences is T B , then the second difference is T A -T B .

[0098] 4) Determining a second sentence feature between the first text and the second text based on the absolute value of the product of the first difference and the second difference, wherein the value of the second sentence feature is negatively correlated with the degree of similarity at the sentence level.

[0099] For example, the feature of the second sentence is represented by d, and the calculation formula of the feature of the second sentence can be:

[0100] d=|(T A -T B )*(L(S A )-L(S B ))|

[0101] In the embodiment of the present disclosure, the second sentence feature d represents the absolute value of the product of the difference between the maximum number of consecutive sentences that meet the threshold condition in the first text and the second text and the difference between the total number of sentences. The smaller the value of the second sentence feature, the more maximum consecutive similar sentences there are in the first text and the second text, and the greater the similarity between the first text and the second text at the sentence level. A -T B ) can represent the first length adjustment factor, for example, in (L(S A )-L(S B )) is known, (T A -T B) is smaller, the proportion of the maximum consecutive similar sentences in the first text and the second text is more balanced, and therefore the degree of similarity is greater. In this way, in the embodiment of the present disclosure, based on the first length adjustment factor, the negative impact of the length difference between the first text and the second text can be reduced, and the accuracy of similarity recognition between texts with large length differences can be improved.

[0102] 5. In one possible embodiment, the multidimensional features include a first passage feature representing the overall passage dimensional features. Determining the multidimensional features between the first text and the second text based on the multidimensional information includes:

[0103] 1) Determine the total number of a first word count of the common words in the first text and the total number of a second word count of the common words in the second text.

[0104] For example, the number of common words is m, the number of times each common word appears in the first text is ml, and the number of characters contained in each common word is m2. Then the total number of the first characters occupied by all common words in the first text is m*m1*m2. Similarly, the total number of the second characters occupied by the common words in the second text can be determined.

[0105] 2) Obtaining the absolute value of the difference between the proportion of the number of words of the common words in the first text and the second text according to the first word total, the first character total, the second character total, and the second word total.

[0106] For example, the sum of the first word is recorded as A m , the sum of the second word number is recorded as B m , the first text contains the first character total Q A , the second text contains a total of Q second characters B , then the absolute value of the difference between the proportion of common words in the first text and the second text can be expressed as:

[0107] 3) Determine the minimum value and the maximum value between the first total number of characters and the second total number of characters, and obtain a ratio between the maximum value and the minimum value, wherein the ratio between the maximum value and the minimum value is used to represent the second length adjustment factor.

[0108] For example, the minimum value between the first character total and the second character total is expressed as min(Q A , Q B ), the maximum value is expressed as max(Q A , Q B ), then the ratio between the maximum and minimum values ​​is:

[0109] 4) Determine the first chapter feature between the first text and the second text based on the ratio and the absolute value of the difference between the proportion of the common words in the first text and the second text, wherein the value of the first chapter feature is negatively correlated with the degree of similarity at the overall chapter level.

[0110] For example, the first chapter feature is expressed as e, and the calculation formula of the first chapter feature can be:

[0111]

[0112] In the embodiment of the present disclosure, the first chapter feature may represent the absolute value of the difference between the proportion of the common words in the first text and the second text. The smaller the value of the first chapter feature, the greater the similarity between the first text and the second text at the overall chapter level. A second length regulation factor can be characterized, e.g. Under known conditions, The smaller the A m and B m The proportion in the first text and the second text is more balanced, and therefore the degree of similarity is greater. In the embodiment of the present disclosure, based on the second length adjustment factor, the similarity evaluation between texts of different lengths can also be balanced, and it can also have better adaptability to texts with large length differences, thereby improving the accuracy of similarity recognition.

[0113] 6. In one possible embodiment, the multidimensional features include a second passage feature representing the overall passage dimension. Determining the multidimensional features between the first text and the second text based on the multidimensional information includes:

[0114] 1) Determine a common string in a first text and a second text, and determine that the length of the common string is a first number within a first range, a second number within a second range, and a third number within a third range.

[0115] For example, for common strings in the first text and the second text, the first number of times the common strings appear in the first range is β1, the second number of times in the second range is β2, and the third number in the third range is β3, where, for example, the first range is [100,), the second range is [50, 100), and the third range is [10, 50), which is not limited in the embodiments of the present disclosure.

[0116] 2) Determine the second chapter feature between the first text and the second text based on the first number and the first weight, the second number and the second weight, and the third number and the third weight, wherein the value of the second chapter feature is positively correlated with the degree of similarity at the overall chapter level.

[0117] For example, the second chapter feature is expressed as f, and the calculation formula of the second chapter feature can be:

[0118] f=β1*γ1+β2*γ2+β3*γ3

[0119] Among them, γ1 is the first weight, γ2 is the second weight, and γ3 is the third weight, which respectively represent the weights of the contributions of common character strings of different lengths to the similarity. For example, the value of γ1 is 0.5, the value of γ2 is 0.3, and the value of γ3 is 0.2. There is no specific limitation and it can be set according to experience and needs. In the embodiment of the present disclosure, the second chapter feature can represent the sum of the contributions of common character strings of different lengths in the first text and the second text to the similarity. The larger the value of the second chapter feature, the greater the similarity between the first text and the second text at the overall chapter level.

[0120] It should also be noted that in the embodiment of the present disclosure, the calculation formulas for the above-mentioned several different dimensional features are only possible examples and are not specifically limited.

[0121] In the embodiment of the present disclosure, the first text and the second text can be processed by word segmentation, sentence segmentation, etc., and the value of each feature can be calculated by using the calculation formula corresponding to the multi-dimensional features. In this way, the features used for similarity recognition between the first text and the second text at different dimensions and levels can be obtained, rather than only considering a single semantic level, thereby improving the accuracy of similarity recognition, and the calculation is simple and fast, thereby improving efficiency.

[0122] Furthermore, since in the embodiments of the present disclosure, feature extraction is performed on the first text and the second text from different dimensions of words, sentences, and entire chapters instead of semantic vector encoding, it is possible to obtain information such as common words and similar sentences between the first text and the second text during the feature extraction process. Based on this, the present disclosure also provides a possible implementation method, which includes, after determining the similarity recognition results of the first text and the second text based on multi-dimensional features, displaying the common words between the first text and the second text in a preset manner when the similarity recognition results of the first text and the second text are similar.

[0123] Among them, the preset method is, for example, highlighting or displaying in different colors in the first text and the second text, or displaying in a comment box in the form of an annotation, or outputting separately, etc., which is not limited in the embodiments of the present disclosure.

[0124] In this way, in the embodiment of the present disclosure, when similarity is determined, common words can be displayed in a preset manner. In addition to displaying common words, sentences that meet threshold conditions can also be displayed, and there is no restriction on this. Users can easily know the similar parts, and it is also convenient for manual similarity review to further improve accuracy.

[0125] The following uses a specific application scenario to illustrate the training process of the recognition model in the embodiment of the present disclosure. Figure 3 As shown in FIG, a schematic diagram of the training process of the recognition model provided by the embodiment of the present disclosure is shown in FIG. Figure 3 Shown, including:

[0126] S301: Obtain a text pair set.

[0127] In the embodiment of the present disclosure, in order to improve the accuracy of text similarity recognition in different application scenarios, multiple texts in the application scenario can be obtained for the required application scenario, and they can be combined in pairs to obtain multiple text pairs, that is, a text pair set.

[0128] S302: Determine the multidimensional features and similarity labels between each text pair in the text pair set.

[0129] For example, for each text pair in the text pair set, similarity labels of the text pairs can be determined by manual labeling or other methods, such as setting label 1 to indicate similarity and setting label 0 to indicate dissimilarity.

[0130] In addition, in the embodiment of the present disclosure, a calculation formula for multidimensional features is pre-constructed. Taking the multidimensional features including the above-mentioned first word feature, second word feature, first sentence feature, second sentence feature, first chapter feature and second chapter feature as an example, the values ​​of the six features between each text pair can be calculated according to the calculation formulas corresponding to the multidimensional features, and then the training text pair set can be obtained based on the multidimensional features and similar labels between multiple text pairs. For example, as shown in Table 1, it is an example of the training text pair set in the embodiment of the present disclosure.

[0131] Table 1.

[0132]

[0133] S303: Normalization processing.

[0134] In the embodiment of the present disclosure, normalizing the values ​​of each feature can reduce the impact of excessive or insufficient weights of individual features on the accuracy of the recognition model. There is no restriction on the normalization method.

[0135] S304: Recognition model training.

[0136] In an embodiment of the present disclosure, the value of each feature in the multidimensional features of each text in the text pair set is converted into an array format as a feature value array X, each feature value array corresponds to a similarity label, and can also be split into a training text pair set and a test text pair set. Training is performed based on the training text pair set. Specifically, X corresponding to each training text is input into the recognition model to obtain a predicted similar recognition result. The recognition model is trained based on the predicted similar recognition result and the similarity label until a preset number of iterations is reached or a loss function converges. The loss function can represent a loss function between the predicted similar recognition result and the similarity label.

[0137] Among them, the loss function of the recognition model can use a normalized exponential function (softmax), a log-likelihood loss function, etc., which is not limited in the embodiments of the present disclosure, and the network structure of the recognition model is not limited. Through training, the model hyperparameters are continuously adjusted, the loss function is transformed, the amount of data is increased, the machine learning model is replaced, and other tuning methods are used to continuously optimize the trained recognition model. After the recognition model is trained, it can also be tested based on the test text set to improve the accuracy of similarity recognition of the recognition model.

[0138] S305: Perform similarity recognition based on the recognition model.

[0139] That is, in the embodiment of the present disclosure, after the recognition model training and testing are completed, the recognition model can be used for similarity recognition. For any two texts, multi-dimensional features are calculated and then input into the recognition model, and the recognition model can output similarity or dissimilarity.

[0140] In this way, in the embodiment of the present disclosure, similarity recognition between texts can be achieved through word segmentation, sentence segmentation, multi-dimensional feature construction, normalization processing, recognition model training, etc., and features at different levels and dimensions of words, sentences, and entire chapters can be constructed, thereby improving the accuracy of similarity recognition, and making calculations faster, thereby improving the efficiency and processing speed of similarity recognition. It is also possible to locate similar parts between texts for easy display, further improving recognition performance.

[0141] It is understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this disclosure will not go into details. It is understood by those skilled in the art that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.

[0142] In addition, the present disclosure also provides a text similarity recognition device, an electronic device, and a computer-readable storage medium, all of which can be used to implement any text similarity recognition method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method section and will not be repeated here.

[0143] Figure 4 This is a block diagram of a text similarity recognition device provided by an embodiment of the present disclosure. Figure 4 The present disclosure provides a text similarity recognition device, which includes:

[0144] An acquisition module 41 is configured to acquire a first text and a second text to be processed;

[0145] a determination module 42 configured to determine multidimensional features between the first text and the second text, wherein the multidimensional features include at least one of the following: a word dimension feature, a sentence dimension feature, and an overall text dimension feature, wherein different dimension features are used to represent the degree of similarity between the texts at different dimensional levels;

[0146] The recognition module 43 is configured to determine similar recognition results between the first text and the second text based on the multi-dimensional features.

[0147] In a possible embodiment, when determining the multidimensional feature between the first text and the second text, the determination module 42 is configured to:

[0148] Performing word segmentation on the first text and the second text to obtain a first word segmentation and the number of first word segmentations of the first text, and a second word segmentation and the number of second word segmentations of the second text;

[0149] Determining, based on the first segmented word and the second segmented word, common words and the number of common words between the first text and the second text;

[0150] Performing sentence segmentation processing on the first text and the second text to obtain first sentences and the number of first sentences contained in the first text, and second sentences and the number of second sentences contained in the second text;

[0151] Obtaining a total number of first characters included in the first text and a total number of second characters included in the second text;

[0152] Determine multidimensional features between the first text and the second text based on multidimensional information, wherein the multidimensional information includes at least two of the following: the first participle, the number of the first participles, the second participle, the number of the second participles, the common words, the number of the common words, the first sentence, the number of the first sentences, the second sentence, the number of the second sentences, the total number of the first characters, and the total number of the second characters.

[0153] In a possible embodiment, the multidimensional feature includes a first word feature representing a word dimensional feature. When determining the multidimensional feature between the first text and the second text based on the multidimensional information, the determination module 42 is configured to:

[0154] Determining a first absolute difference between the first number of participles and the second number of participles, and determining a first product of the first number of participles and the second number of participles;

[0155] obtaining a second product between the first absolute difference and the number of common words;

[0156] A first word feature between the first text and the second text is determined based on a ratio between the second product and the first product, wherein a value of the first word feature is negatively correlated with a similarity degree at a word level.

[0157] In a possible embodiment, the multidimensional feature includes a second word feature representing a word dimensional feature. When determining the multidimensional feature between the first text and the second text based on the multidimensional information, the determination module 42 is configured to:

[0158] Determining a first occurrence number of the common word in the first text, and a second occurrence number of the common word in the second text;

[0159] determining a second absolute difference between the first number of occurrences and the second number of occurrences, and determining a maximum value between the first number of occurrences and the second number of occurrences;

[0160] A second word feature between the first text and the second text is determined based on the proportion of the second absolute difference in the maximum value, wherein the value of the second word feature is negatively correlated with the degree of similarity at the word level.

[0161] In a possible embodiment, the multidimensional feature includes a first sentence feature representing a sentence dimensional feature. When determining the multidimensional feature between the first text and the second text based on the multidimensional information, the determination module 42 is configured to:

[0162] Determining a total number of first sentences in the first sentence that meet a threshold condition, and a total number of second sentences in the second sentence that meet a threshold condition, wherein the threshold condition represents that a ratio of a sum of the lengths of character strings occupied by common words in the sentences to a total character string length of the sentences is greater than or equal to a preset threshold;

[0163] A first sentence feature between the first text and the second text is determined based on an absolute difference between the total number of the first sentences and the total number of the second sentences, wherein a value of the first sentence feature is negatively correlated with a degree of similarity at the sentence level.

[0164] In a possible embodiment, the multidimensional feature includes a second sentence feature representing a sentence dimensional feature. When determining the multidimensional feature between the first text and the second text based on the multidimensional information, the determination module 42 is configured to:

[0165] Determining a first maximum number of consecutive sentences in the first sentence that meets a threshold condition, and a second maximum number of consecutive sentences in the second sentence that meets a threshold condition, wherein the threshold condition represents that a ratio of a sum of the lengths of character strings occupied by common words in the sentence to a total character string length of the sentence is greater than or equal to a preset threshold;

[0166] obtaining a first difference between the first maximum number of consecutive sentences and the second maximum number of consecutive sentences;

[0167] Obtaining a second difference between the first number of sentences and the second number of sentences, wherein the second difference is used to represent a first length adjustment factor;

[0168] A second sentence feature between the first text and the second text is determined based on the absolute value of the product of the first difference and the second difference, wherein the value of the second sentence feature is negatively correlated with the similarity at the sentence level.

[0169] In a possible embodiment, the multidimensional feature includes a first passage feature representing a dimensional feature of the entire passage. When determining the multidimensional feature between the first text and the second text based on the multidimensional information, the determination module 42 is configured to:

[0170] Determining a first total number of characters occupied by the common words in the first text, and a second total number of characters occupied by the common words in the second text;

[0171] Obtaining an absolute value of a difference between a proportion of the number of words of the common words in the first text and the second text according to the first word total, the first character total, the second character total, and the second word total;

[0172] determining a minimum value and a maximum value between the first total number of characters and the second total number of characters, and obtaining a ratio between the maximum value and the minimum value, wherein the ratio between the maximum value and the minimum value is used to represent a second length adjustment factor;

[0173] Based on the ratio and the absolute value of the difference between the proportion of the common words in the first text and the second text, a first chapter feature between the first text and the second text is determined, wherein the value of the first chapter feature is negatively correlated with the degree of similarity at the overall chapter level.

[0174] In a possible embodiment, the multidimensional feature includes a second passage feature representing a dimensional feature of the entire passage. When determining the multidimensional feature between the first text and the second text based on the multidimensional information, the determination module 42 is configured to:

[0175] Determine a common string in the first text and the second text, and determine that the length of the common string is a first number within a first range, a second number within a second range, and a third number within a third range;

[0176] A second chapter feature between the first text and the second text is determined based on the first number and the first weight, the second number and the second weight, and the third number and the third weight, wherein the value of the second chapter feature is positively correlated with the degree of similarity at the overall chapter level.

[0177] Each module in the above-mentioned text similarity recognition device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0178] Figure 5 This is a block diagram of an electronic device provided by an embodiment of the present disclosure. Figure 5 An embodiment of the present disclosure provides an electronic device, which includes: at least one processor 501; at least one memory 502, and one or more I / O interfaces 503 connected between the processor 501 and the memory 502; wherein the memory 502 stores one or more computer programs that can be executed by the at least one processor 501, and the one or more computer programs are executed by the at least one processor 501 so that the at least one processor 501 can execute the above-mentioned text similarity recognition method.

[0179] Each module in the above-mentioned electronic device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0180] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-mentioned text similarity recognition method. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0181] An embodiment of the present disclosure also provides a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned text similarity recognition method.

[0182] It will be understood by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable storage medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium).

[0183] As is well known to those skilled in the art, the term computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information (such as computer-readable program instructions, data structures, program modules or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technology, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those skilled in the art, communication media typically contains computer-readable program instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0184] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0185] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0186] The computer program product described herein may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0187] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0188] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0189] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0190] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0191] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the present disclosure as set forth in the appended claims.

Claims

1. A text similarity recognition method, characterized in that: include: Get the first text and the second text to be processed; Determining multidimensional features between the first text and the second text, wherein the multidimensional features include at least one of the following: word dimension features, sentence dimension features, and overall text dimension features, and different dimension features are used to represent the degree of similarity between the texts at different dimensional levels; Determine similar recognition results of the first text and the second text based on the multi-dimensional features.

2. The method according to claim 1, characterized in that The determining of the multidimensional features between the first text and the second text includes: Performing word segmentation on the first text and the second text to obtain a first word segmentation and the number of first word segmentations of the first text, and a second word segmentation and the number of second word segmentations of the second text; Determining, based on the first segmented word and the second segmented word, common words and the number of common words between the first text and the second text; Performing sentence segmentation processing on the first text and the second text to obtain first sentences and the number of first sentences contained in the first text, and second sentences and the number of second sentences contained in the second text; Obtaining a total number of first characters included in the first text and a total number of second characters included in the second text; Determine multidimensional features between the first text and the second text based on multidimensional information, wherein the multidimensional information includes at least two of the following: the first participle, the number of the first participles, the second participle, the number of the second participles, the common words, the number of the common words, the first sentence, the number of the first sentences, the second sentence, the number of the second sentences, the total number of the first characters, and the total number of the second characters.

3. The method according to claim 2, characterized in that The multidimensional feature includes a first word feature representing a word dimensional feature, and determining the multidimensional feature between the first text and the second text based on the multidimensional information includes: Determining a first absolute difference between the first number of participles and the second number of participles, and determining a first product of the first number of participles and the second number of participles; obtaining a second product between the first absolute difference and the number of common words; A first word feature between the first text and the second text is determined based on a ratio between the second product and the first product, wherein a value of the first word feature is negatively correlated with a similarity degree at a word level.

4. The method according to claim 2, characterized in that The multidimensional feature includes a second word feature representing a word dimensional feature, and determining the multidimensional feature between the first text and the second text based on the multidimensional information includes: Determining a first occurrence number of the common word in the first text, and a second occurrence number of the common word in the second text; determining a second absolute difference between the first number of occurrences and the second number of occurrences, and determining a maximum value between the first number of occurrences and the second number of occurrences; A second word feature between the first text and the second text is determined based on the proportion of the second absolute difference in the maximum value, wherein the value of the second word feature is negatively correlated with the degree of similarity at the word level.

5. The method according to claim 2, characterized in that The multidimensional feature includes a first sentence feature representing a sentence dimensional feature, and determining the multidimensional feature between the first text and the second text based on the multidimensional information includes: Determining a total number of first sentences in the first sentence that meet a threshold condition, and a total number of second sentences in the second sentence that meet a threshold condition, wherein the threshold condition represents that a ratio of a sum of the lengths of character strings occupied by common words in the sentences to a total character string length of the sentences is greater than or equal to a preset threshold; A first sentence feature between the first text and the second text is determined based on an absolute difference between the total number of the first sentences and the total number of the second sentences, wherein a value of the first sentence feature is negatively correlated with a degree of similarity at the sentence level.

6. The method according to claim 2, characterized in that The multidimensional feature includes a second sentence feature representing a sentence dimensional feature, and determining the multidimensional feature between the first text and the second text based on the multidimensional information includes: Determining a first maximum number of consecutive sentences in the first sentence that meets a threshold condition, and a second maximum number of consecutive sentences in the second sentence that meets a threshold condition, wherein the threshold condition represents that a ratio of a sum of the lengths of character strings occupied by common words in the sentence to a total character string length of the sentence is greater than or equal to a preset threshold; obtaining a first difference between the first maximum number of consecutive sentences and the second maximum number of consecutive sentences; Obtaining a second difference between the first number of sentences and the second number of sentences, wherein the second difference is used to represent a first length adjustment factor; A second sentence feature between the first text and the second text is determined based on the absolute value of the product of the first difference and the second difference, wherein the value of the second sentence feature is negatively correlated with the similarity at the sentence level.

7. The method according to claim 2, characterized in that The multidimensional features include a first passage feature representing the overall passage dimensional features, and determining the multidimensional features between the first text and the second text based on the multidimensional information includes: Determining a first total number of characters occupied by the common words in the first text, and a second total number of characters occupied by the common words in the second text; According to the first word sum, the first total number of characters, the second total number of characters and the second word sum, Obtaining the absolute value of the difference between the proportion of the common words in the first text and the second text; determining a minimum value and a maximum value between the first total number of characters and the second total number of characters, and obtaining a ratio between the maximum value and the minimum value, wherein the ratio between the maximum value and the minimum value is used to represent a second length adjustment factor; Based on the ratio and the absolute value of the difference between the proportion of the common words in the first text and the second text, a first chapter feature between the first text and the second text is determined, wherein the value of the first chapter feature is negatively correlated with the degree of similarity at the overall chapter level.

8. The method according to claim 2, characterized in that The multidimensional features include a second passage feature representing the dimensional features of the entire passage. The determining of the multidimensional features between the first text and the second text based on the multidimensional information includes: Determine a common string in the first text and the second text, and determine that the length of the common string is a first number within a first range, a second number within a second range, and a third number within a third range; A second chapter feature between the first text and the second text is determined based on the first number and the first weight, the second number and the second weight, and the third number and the third weight, wherein the value of the second chapter feature is positively correlated with the degree of similarity at the overall chapter level.

9. A text similarity recognition device, characterized in that: include: An acquisition module, configured to acquire a first text and a second text to be processed; a determination module, configured to determine multidimensional features between the first text and the second text, wherein the multidimensional features include at least one of the following: a word dimension feature, a sentence dimension feature, and an overall paragraph dimension feature, wherein different dimension features are used to represent the degree of similarity between the texts at different dimensional levels; The recognition module is configured to determine similar recognition results of the first text and the second text based on the multi-dimensional features.

10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor. The one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the text similarity recognition method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the text similarity recognition method according to any one of claims 1 to 8.

12. A computer program product, characterized in that The invention comprises a computer-readable code or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the text similarity recognition method according to any one of claims 1 to 8.