Text similarity measurement method, device, equipment, storage medium and program product

By constructing the joint probability distribution of text strings and sampling, and calculating the distance matrix between text and sampled strings, the problems of calculation efficiency and information loss in massive text batch processing in the prior art are solved, and efficient text similarity measurement is achieved.

CN115640523BActive Publication Date: 2025-05-27DOUYIN VISION CO LTD

Patent Information

Application Number
CN202211274116.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-18
Publication Date
2025-05-27
Estimated Expiration
2042-10-18

AI Technical Summary

Technical Problem

The existing text similarity measurement methods have problems of computational efficiency and information loss in batch processing of massive texts, especially word frequency calculation, planning solution, statistical models and word vector encoding technologies are difficult to effectively apply in big data and long text processing.

Method used

By obtaining two text strings, constructing their joint probability distribution, and sampling the distribution to get the sampled string. Then, the distance matrix between the two text strings and the sampled strings is calculated separately, and the text similarity is determined based on these matrices, and the dimensionality reduction calculation of the text similarity measurement method is realized, which improves the calculation efficiency and reduces information loss while reducing the dimensionality.

Benefits of technology

The calculation efficiency of text similarity measurement method is improved, especially in the processing of massive text data, which significantly reduces the computational complexity and information loss, and achieves a more efficient similarity measurement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115640523B_ABST
    Figure CN115640523B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, apparatus, device, storage medium, and program product for text similarity measurement. The method includes: obtaining a first text string and a second text string; constructing a joint probability distribution of the two based on the first text string and the second text string, and sampling the joint probability distribution to obtain a sampled string; calculating the distance from the first text string to the sampled string to obtain a first distance matrix, and calculating the distance from the second text string to the sampled string to obtain a second distance matrix; determining the similarity between the first text string and the second text string based on the first distance matrix and the second distance matrix. The embodiments of the present disclosure achieve dimensionality reduction calculation of the text similarity method, improving the calculation efficiency. In addition, since the sampled string contains information of the two text strings, information loss is reduced while dimensionality reduction is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer processing technologies, and in particular, to a method, apparatus, device, storage medium, and program product for measuring text similarity. Background Art

[0002] With the development of emerging technologies such as the Internet, Internet of Things, artificial intelligence, and big data, a large amount of text data has emerged in various fields. Text similarity measurement methods are applied in more and more scenarios. For example, in search engines, understanding search content and establishing indexes of web page links; evaluating repetitions, plagiarism, homogenization, etc. of various information flow articles. Text similarity measurement methods are involved in all of them.

[0003] Existing text similarity measurement methods include word frequency calculation, linear programming, statistical models, word vector encoding, neural network inference, etc., but there are varying degrees of deficiencies for batch processing of a large amount of text. The word frequency calculation method is too rough and is currently only used for the analysis and characterization of extremely small amounts of data. The linear programming algorithm is difficult to apply to big data and large-scale text processing due to its high computational cost. Statistical models are difficult to independently calculate the similarity between two texts because their statistical probabilities are derived from full-scale processed text. Word vector encoding techniques all require a large amount of text corpus for pre-training, need to calculate word encodings from the full-scale document, cannot independently calculate the similarity between two texts, and the computational consumption is also very high. Summary of the Invention

[0004] To solve the above technical problems, embodiments of the present disclosure provide a method, apparatus, device, storage medium, and program product for measuring text similarity, which realizes the dimensionality reduction calculation of the text similarity measurement method and improves the calculation efficiency. In addition, since the sampled string contains the information of the two text strings, information loss is reduced while dimensionality reduction is achieved.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for measuring text similarity, the method including:

[0006] Obtain a first text string and a second text string;

[0007] Based on the first text string and the second text string, construct a joint probability distribution of the two, and sample the joint probability distribution to obtain a sampled string;

[0008] Calculate the distance from the first text string to the sampled string to obtain a first distance matrix, and calculate the distance from the second text string to the sampled string to obtain a second distance matrix;

[0009] Determine the similarity between the first text string and the second text string based on the first distance matrix and the second distance matrix.

[0010] In a second aspect, an embodiment of the present disclosure provides a text similarity measurement device, which includes:

[0011] A text string acquisition module, configured to acquire a first text string and a second text string;

[0012] A sampled string determination module, configured to construct a joint probability distribution of the two based on the first text string and the second text string, and sample the joint probability distribution to obtain a sampled string;

[0013] A distance matrix calculation module, configured to calculate the distance from the first text string to the sampled string to obtain a first distance matrix, and calculate the distance from the second text string to the sampled string to obtain a second distance matrix;

[0014] A text similarity measurement module, configured to determine the similarity between the first text string and the second text string based on the first distance matrix and the second distance matrix.

[0015] In a third aspect, an embodiment of the present disclosure provides an electronic device, which includes:

[0016] One or more processors;

[0017] A storage device, configured to store one or more programs;

[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the text similarity measurement method as described in any one of the above first aspects.

[0019] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the text similarity measurement method as described in any one of the above first aspects.

[0020] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, which includes a computer program or instruction, and when the computer program or instruction is executed by a processor, it implements the text similarity measurement method as described in any one of the above first aspects.

[0021] Embodiments of the present disclosure provide a method, apparatus, device, storage medium, and program product for text similarity measurement. The method includes: obtaining a first text string and a second text string; constructing a joint probability distribution of the two based on the first text string and the second text string, and sampling the joint probability distribution to obtain a sampled string; calculating the distance from the first text string to the sampled string to obtain a first distance matrix, and calculating the distance from the second text string to the sampled string to obtain a second distance matrix; determining the similarity between the first text string and the second text string based on the first distance matrix and the second distance matrix. By sampling the joint probability distribution of two text strings, obtaining a sampled string, calculating the distance matrices between the two text strings and the sampled string respectively, and then calculating the similarity of the two distance matrices, the embodiments of the present disclosure achieve dimensionality reduction calculation of the text similarity method and improve the calculation efficiency. In addition, since the sampled string contains information of the two text strings, information loss is reduced while dimensionality reduction is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and that the original elements and elements are not necessarily drawn to scale.

[0023] Figure 1 is a flowchart of a method for text similarity measurement in an embodiment of the present disclosure;

[0024] Figure 2 is a schematic diagram of a joint probability distribution in an embodiment of the present disclosure;

[0025] Figure 3a is a schematic diagram of initialization of an edit distance in an embodiment of the present disclosure;

[0026] Figure 3b is a schematic diagram of the actual calculation process of an edit distance in an embodiment of the present disclosure;

[0027] Figure 4a is a schematic diagram of the calculation process of an edit distance in an embodiment of the present disclosure;

[0028] Figure 4b is a schematic diagram of the calculation process of an edit distance in an embodiment of the present disclosure;

[0029] Figure 5a is a schematic diagram of the complete calculation process of an edit distance in an embodiment of the present disclosure;

[0030] Figure 5bIt is a schematic diagram of a complete edit distance calculation process in an embodiment of the present disclosure;

[0031] Figure 6 It is a schematic diagram of vector extraction of an edit distance matrix in an embodiment of the present disclosure;

[0032] Figure 7 It is a schematic flowchart of text similarity measurement in an embodiment of the present disclosure

[0033] Figure 8 It is a schematic diagram of a vector calculation process in an embodiment of the present disclosure;

[0034] Figure 9 It is a schematic structural diagram of a text similarity measurement device in an embodiment of the present disclosure;

[0035] Figure 10 It is a schematic structural diagram of an electronic device in an embodiment of the present disclosure. Detailed implementation manners

[0036] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0037] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0038] The term "including" and its variations used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0039] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.

[0040] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0041] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0042] Before further describing the embodiments of the present disclosure in detail, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are subject to the following interpretations.

[0043] In order to solve various problems in text similarity measurement, many measurement methods have been proposed, which have undergone technical evolutions such as word frequency calculation, planning solution, statistical model, word vector encoding, and neural network reasoning.

[0044] The earliest method for calculating text similarity was word frequency statistics, which was mainly measured by counting the number of identical characters in two texts. For example, the string "abc" and the string "abd" have the same character "ab". The more identical characters there are, the more similar the two strings are. Later, in order to further evaluate the impact of string order on similarity based on statistical quantity, the string arrangement number was introduced. When the word frequency is the same, the smaller the difference in arrangement number, the higher the similarity.

[0045] Statistical models such as TF-IDF, which consists of two parts, TF and IDF, count the occurrence ratio of a certain character or word in a single sentence or document, and then count the occurrence ratio of the character or word in all sentences or documents to obtain the probability value of the character or word appearance, thereby obtaining a probability value vector for the sentence or document, and calculating the vector similarity between the two probability value vectors as the similarity of the two documents.

[0046] TF-IDF is an early word vector encoding method. However, due to the difficulty in obtaining semantic similarity from statistical values, with the development of neural networks, many word vector encoding technologies based on neural networks have emerged. The most commonly used ones are CBOW (Continuous Bag-of-Words) and Skip-Gram. CBOW infers the word vector of a specific word through the word vectors of context-related words. The calculation method of Skip-Gram is just the opposite. It calculates the context word vector corresponding to a specific word through the word vector of a specific word. Both calculation methods can obtain context semantics to a certain extent, but because there are more context inferences, the calculation consumption is relatively large.

[0047] The current main text similarity technologies include word frequency calculation, planning solution, statistical model, word vector encoding, neural network inference, etc. However, for the batch processing of massive texts, they all have deficiencies to varying degrees.

[0048] Due to its coarseness, the word frequency calculation method is currently only used for the analysis and characterization of extremely small amounts of data. The edit distance algorithm is difficult to apply to big data and massive text processing because of its high computational cost. Since the statistical probability of the statistical model comes from the full-scale processed text, it is difficult to independently calculate the similarity between two texts. In addition, almost all current word vector encoding technologies require a large amount of text corpus for pre-training, need to calculate word encoding from the full-scale document, cannot independently calculate the similarity between two texts, and the computational consumption is also very high.

[0049] In addition, to cope with massive data, many data preprocessing methods have emerged, such as data dimensionality reduction. Data dimensionality reduction is to map the data in a high-dimensional space to a low-dimensional space on the basis of minimizing information loss, so as to improve the data processing efficiency.

[0050] The current existing data dimensionality reduction technologies mainly include factor analysis, autoencoders, topic models, local embedding, etc., but they all require a large amount of text corpus training, and the calculation of two independent texts and the batch processing effect of massive texts are not good.

[0051] In the embodiments of the present disclosure, the basic algorithm relied on by the text similarity measurement method is the edit distance algorithm proposed by the Russian scientist Vladimir I. Levenshtein in 1965, also called the Levenshtein distance. Because of its intuitive and easy-to-interpret nature and good effect on the similarity measurement of strings, this algorithm has undergone some optimizations and is still widely used today. There is even a dedicated third-party calculation package in coding languages such as Python. This algorithm needs to be solved based on dynamic programming or recursion, with relatively large computational consumption, and it is difficult to apply to text big data and long text similarity measurement.

[0052] In the edit distance algorithm, the power level is relatively high. For a small amount of data, the time complexity and space complexity do not have a great impact. However, for big data and long texts, the resources required are huge. This also leads to the difficulty of widely promoting this method in the current calculation of text big data.

[0053] To solve the above technical problems, in the embodiments of the present disclosure, a text similarity measurement method is provided. By selecting a sampling string, the distance matrices between two text strings and the sampling string are calculated respectively, and then the similarity between the two text strings is calculated according to the two distance matrices, turning the single-stage calculation into a two-stage calculation, improving the calculation efficiency, especially being particularly obvious in massive text big data.

[0054] From the perspective of application value, the text similarity measurement method provided by the embodiments of the present disclosure can be widely applied to fields such as precision medicine, quantitative finance, and intelligent speech. It has very significant application value for web link analysis, cryptographic coding, genome calculation, speech proofreading, etc.

[0055] The following will introduce in detail the text similarity measurement method provided by the embodiments of the present disclosure with reference to the accompanying drawings.

[0056] Figure 1 It is a flowchart of a text similarity measurement method in the embodiments of the present disclosure. This embodiment is applicable to the situation of calculating the similarity between two texts. This method can be executed by a text similarity measurement device, and the text similarity measurement device can be implemented in software and / or hardware.

[0057] As Figure 1 shown, the text similarity measurement method provided by the embodiments of the present disclosure mainly includes steps S101-S104.

[0058] S101. Obtain a first text string and a second text string.

[0059] In the embodiments of the present disclosure, a text string is a form of written language and can be a combination of multiple characters. The characters can include one or more of the following: English characters, Chinese characters, punctuation marks, Roman characters, Greek characters, or other special characters. The first text string and the second text string refer to two text strings whose similarity needs to be measured.

[0060] In an embodiment of the present disclosure, after detecting a trigger event for text similarity measurement, a first text string and a second text string can be obtained.

[0061] In an embodiment of the present disclosure, the above trigger event can be an event of receiving a user input text string; for example: after the user inputs a text string in a browser or an intelligent question and answer system, and the terminal device receives this text string, it can be considered that the trigger event for text similarity measurement has been detected. At this time, the terminal can use the received user input text string as the first text string and obtain any text string from the database as the second text string.

[0062] In one embodiment of the present disclosure, the triggering event here may also be an event of receiving a text similarity measurement instruction; for example, when a user wants to count the similarity between any two text strings in the terminal database, the user can input a similarity measurement instruction to the terminal, and the similarity measurement instruction may be a click instruction, a press instruction, a voice instruction, etc. When the terminal receives this similarity measurement instruction, it can be considered that the triggering event for text similarity measurement has been detected. At this time, the terminal can arbitrarily select two text strings from the database as the first text string and the second text string.

[0063] In one embodiment of the present disclosure, the triggering event here may also be an event of receiving a completion instruction for a target task; for example: a security risk is found in network interface A, and the terminal device receives the program code corresponding to network interface A, then it can be considered that the triggering event for text similarity measurement has been detected. At this time, the terminal can use the program code corresponding to network interface A received as the first text string, and arbitrarily obtain the program code corresponding to another network interface other than network interface A as the second text string.

[0064] S102. Construct a joint probability distribution of the first text string and the second text string, and sample the joint probability distribution to obtain a sampled string.

[0065] In the embodiments of the present disclosure, the characters included in the sampled string are all characters included in the first text string and / or the second text string.

[0066] In the embodiments of the present disclosure, constructing the joint probability distribution of the first text string and the second text string and sampling based on the joint probability distribution is the optimal solution for simultaneously retaining the information of the two strings. Among them, the above joint probability distribution may be the joint probability between a normal distribution and a normal distribution, or the joint probability distribution between a normal distribution and an exponential distribution, which is not specifically limited in the embodiments of the present disclosure.

[0067] As Figure 2 shown, the joint probability distribution diagram between a normal distribution and a normal distribution and the schematic diagram of the joint probability distribution between a normal distribution and an exponential distribution are respectively shown. Random sampling in the joint probability distribution can retain the original text information to the greatest extent, and at this time, the random sampling has no additional computational overhead for big data calculation and can perform massive data calculations.

[0068] The sampled string obtained by sampling the data of the joint probability distribution contains the information of both the first text string and the second text string, reducing information loss, and sampling also reduces data noise.

[0069] In one embodiment of the present disclosure, sampling the joint probability distribution to obtain a sampled string includes: randomly downsampling the joint probability distribution based on a preset sampling ratio to obtain the sampled string, where the preset sampling ratio is inversely proportional to the length of the text string.

[0070] Among them, the preset sampling ratio is a preset ratio of the sampled string obtained by sampling from the joint probability distribution. Preferably, the maximum value of the preset sampling ratio is one-fourth. When the preset sampling ratio is one-fourth, it is possible to reduce information loss as much as possible while performing dimensionality reduction calculations. Specifically, in the embodiments of the present disclosure, at most one-fourth sampling of the joint probability distribution is required.

[0071] Furthermore, the preset sampling ratio is inversely proportional to the length of the text string. In other words, the longer the text string, the smaller the above preset sampling ratio, and the better the performance optimization effect for big data. It should be noted that as long as the preset sampling ratio does not cause the center position of the joint probability distribution to shift, it can be as small as possible, so that dimensionality reduction can be performed to the greatest extent.

[0072] S103. Calculate the distance from the first text string to the sampled string to obtain a first distance matrix, and calculate the distance from the second text string to the sampled string to obtain a second distance matrix.

[0073] In the embodiments of the present disclosure, the first distance matrix and the second distance matrix can be any distance matrix for measuring text similarity. It should be noted that different distance matrices are obtained by using different similarity measurement strategies, and different methods are used to generate distance representation vectors.

[0074] In one embodiment of the present disclosure, the edit distance calculation algorithm is used to calculate the first distance matrix and the second distance matrix. The first distance matrix is used to represent the edit distance from the first text string to the sampled string; the second distance matrix is used to represent the edit distance from the second text string to the sampled string.

[0075] Among them, the most representative in the planning solution is the edit distance, also called the Levenshtein Distance. The measurement method is to determine at least how many times of processing are required to change one string into another string. Here, the processing includes deleting a string, inserting a string, and replacing a string. The fewer the number of processing times, the higher the similarity.

[0076] Among them, the edit distance represents the minimum number of transformation times required to transform one string into another string. Here, the transformations mainly include deletion, insertion, and replacement operations. If |a| and |b| represent the lengths of strings a and b respectively, the edit distance between string a and string b, lev_ab(|a|, |b|), can be expressed as:

[0077]

[0078] Among them, the first line in otherwise in the above formula represents the deletion operation, the second line represents the insertion operation, and the third line represents the replacement operation. In the embodiments of the present disclosure, taking the calculation of the similarity between the strings " / tts_sync" and "tts / sync / " using the edit distance algorithm as an example, the calculation process of the edit distance and the existing problems are introduced in detail through FIGS. 3a, 3b, 4a, 4b, 5a, and 5b.

[0079] As Figure 3a shown, the initialization result of the distance is given, that is, when ifmin(i, j) = 0 in the above edit distance calculation formula, the part of taking max(i, j). Therefore, the sequence of 0-9 in the second row and the sequence of 0-9 in the second column can be obtained. As Figure 3b shown, the actual calculation of the edit distance begins, that is, the part in otherwise in the above edit distance calculation formula. The value in the 3rd row and 3rd column is the minimum of the following three values. The first value: the value in the 2nd row and 3rd column + 1 is equal to 2; the second value: the value in the 3rd row and 2nd column + 1 is equal to 2; the third value: because " / " and "t" are not the same, the value is the value in the 2nd row and 2nd column + 1 which is equal to 1. The minimum of the above three values is 1. That is, the value in the 3rd row and 3rd column is 1.

[0080] Figure 4a And Figure 4b further gives the calculation process of the edit distance. As Figure 4a shown, the value in the 6th row and 3rd column is 3. Because both characters are " / ", the number in the diagonal 5th row and 2nd column does not change. At this time, compared with the value in the 5th row and 3rd column + 1 and the value in the 6th row and 2nd column + 1, the minimum of them is 3. Figure 4b shown, the value in the 4th row and 5th column is 1, which is less than the value in the 3rd row and 5th column which is 2. This further verifies the problem solved by the edit distance algorithm, that is, the distance measured at the (i, j) position is the distance between the first i characters of one string and the first j characters of another string. Obviously, the number of edit times from " / tt" to "tt" is significantly less than the number of times from " / tt" to "t".

[0081] Figure 5a gives the intermediate process of the deduction, Figure 5bThe entire calculation process is given, where the bold numbers are the change paths of the minimum distance. Finally, the edit distance between the strings " / tts_sync" and "tts / sync / " is calculated to be 3, that is, 3 conversions are required. Intuitively, it is necessary to delete the " / " at the starting position, replace the "_" with " / ", and insert a " / " at the ending position, for a total of 3 operations.

[0082] Combined with the deduction process, it is found that there are two solutions to the edit distance algorithm, namely recursive solution and dynamic programming solution. Any of the above solution methods can be used to calculate the distance from the first text string to the sampled string to obtain a first distance matrix, and calculate the distance from the second text string to the sampled string to obtain a second distance matrix.

[0083] S104. Determine the similarity between the first text string and the second text string based on the first distance matrix and the second distance matrix.

[0084] In an embodiment of the present disclosure, calculate the similarity between the first distance matrix and the second distance matrix, and use the similarity between the two distance matrices as the similarity between the first text string and the second text string.

[0085] Optionally, the similarity between the two matrices is calculated by means of Euclidean distance or cosine similarity. It should be noted that other matrix similarity calculation methods can also be used to calculate the similarity between the first distance matrix and the second distance matrix, which is not specifically limited in this embodiment.

[0086] In an embodiment of the present disclosure, the eigenvector calculation method is used to calculate the eigenvector of the first distance matrix, denoted as the first eigenvector, and the eigenvector calculation method is used to calculate the eigenvector of the second distance matrix, denoted as the second eigenvector. Calculate the similarity between the two eigenvectors as the similarity between the first text string and the second text string. Among them, the eigenvector calculation can be realized by matrix decomposition. This is not specifically limited in this embodiment.

[0087] In one embodiment of the present disclosure, determining the similarity between the first text string and the second text string based on the first distance matrix and the second distance matrix includes: extracting features from the first distance matrix to obtain a first distance representation vector, where the first distance representation vector includes a column representation vector, a diagonal representation vector, and a row representation vector corresponding to the first distance matrix; extracting features from the second distance matrix to obtain a second distance representation vector, where the second distance representation vector includes a column representation vector, a diagonal representation vector, and a row representation vector corresponding to the second distance matrix; calculating the vector similarity between the first distance representation vector and the second distance representation vector as the similarity between the first string and the second text string.

[0088] Among them, the first distance representation vector can be understood as a vector that can be extracted from the first distance matrix and can represent the first distance matrix for matrix similarity calculation. The second distance representation vector can be understood as a vector that can be extracted from the second distance matrix and can represent the second distance matrix for matrix similarity calculation.

[0089] In one embodiment of the present disclosure, the feature extraction includes: extracting the last row sequence in the distance matrix as the row representation vector; extracting the last column sequence in the distance matrix as the column representation vector; extracting the diagonal sequence in the distance matrix as the diagonal representation vector.

[0090] The text similarity measurement method provided by the embodiments of the present disclosure starts from the design idea of the edit distance. The distance values at the back (right and bottom) in the distance matrix depend on the previous distance values. Theoretically, selecting the changes in the last row, the last column, and the diagonal sequence in the distance matrix can represent the similarity information of the two text strings. Taking the text strings " / tts_sync" and "tts / sync / " as an example, the shortest path on the diagonal of the distance matrix, the last column in the distance matrix, and the distance changes in the last row of the distance matrix can be extracted to form a matrix representation vector. As Figure 6 shown, extracting the shortest path (1, 1, 1, 1, 2, 2, 2, 2, 2, 3) on the diagonal of the distance matrix, the last column (9, 8, 7, 6, 6, 5, 4, 3, 2, 3) in the distance matrix, and the last row (9, 8, 8, 8, 7, 7, 6, 5, 4, 3) in the distance matrix gives Figure 6 the distance representation vector of the left distance matrix in. In the embodiments of the present disclosure, for the extraction of the distance representation vector of the distance matrix, the similarity calculation of the text is directly converted into the numerical calculation of the vector, potentially completing the encoding of the text string to the word vector.

[0091] In the embodiments of the present disclosure, distance representation vectors are respectively extracted from two distance matrices, and the similarity between the two distance representation vectors is calculated as the similarity between the first text string and the second text string. Since the calculation method of vector similarity has a smaller computational complexity than that of matrix similarity, the technical solution provided by the embodiments of the present disclosure can further reduce the computational complexity and improve the computational efficiency.

[0092] In one embodiment of the present disclosure, the Euclidean distance or cosine similarity is used to determine the vector similarity.

[0093] Specifically, the Euclidean distance between the first distance representation vector and the second distance representation vector; the cosine angle between the first distance representation vector and the second distance representation vector.

[0094] The embodiments of the present disclosure provide a method, apparatus, device, storage medium, and program product for text similarity measurement. The method includes: obtaining a first text string and a second text string; constructing a joint probability distribution of the two based on the first text string and the second text string, and sampling the joint probability distribution to obtain a sampled string; calculating the distance from the first text string to the sampled string to obtain a first distance matrix, and calculating the distance from the second text string to the sampled string to obtain a second distance matrix; determining the similarity between the first text string and the second text string based on the first distance matrix and the second distance matrix. By sampling the joint probability distribution of two text strings, obtaining a sampled string, calculating the distance matrices of the two text strings and the sampled string respectively, and then calculating the similarity of the two distance matrices, the embodiments of the present disclosure realize the dimensionality reduction calculation of the text similarity method and improve the computational efficiency. In addition, since the sampled string contains the information of the two text strings, information loss is reduced while dimensionality reduction is achieved.

[0095] In the embodiments of the present disclosure, a schematic diagram of the architecture for text similarity measurement is provided, as Figure 7 shown. First, a joint probability distribution is made for the two text strings a and text string b to be calculated, and then the joint probability distribution is downsampled to obtain a sampled string. Then, the edit distance between text string a and text string b and the sampled string is calculated respectively. The calculation process will generate a distance matrix as Figure 5b shown. The distance representation vectors are extracted from the distance matrix, and finally, a numerical operation is performed on the distance representation vectors to obtain the similarity between the two distance representation vectors, which is the similarity between text string a and text string b.

[0096] Figure 8The process of calculating the edit distance between text string a and text string b and a sampling string respectively. The first text string a and the second text string b calculate distance matrices with the sampling string respectively, extract row, column and diagonal vectors as distance representation vectors, and then calculate the numerical similarity of the distance representation vectors to obtain the first text string a and the second text string b. Among them, the two processes of calculating the distance matrices between the first text string a and the second text string b and the sampling string respectively can be executed in parallel, further improving the batch processing performance for big data.

[0097] A text similarity measurement method proposed by the present disclosure splits a single-stage high-dimensional calculation process into a two-stage low-dimensional calculation process, effectively improving the similarity measurement efficiency of massive texts. Its main technical effects are: making batch processing of massive text similarity possible, reducing the dimension of data through the distributed sampling and representation vector extraction processes, while ensuring the information content as much as possible and improving the calculation efficiency. Compared with the traditional edit distance calculation method, the complexity is reduced by at least more than 70%, and the longer the text, the more significant the performance improvement effect. For text big data, the performance is expected to be improved by more than one order of magnitude. It can measure the similarity of any two strings or character sequences, and generate word vectors between pairwise specific comparison strings during the calculation process, without additional requirements for text segmentation and semantics. It does not require a large amount of corpus for supervised or unsupervised pre-training, and there is no additional computational overhead except for the actual batch processing process. Thanks to the downsampling process of the text distribution, the influence of data noise on similarity calculation is also reduced to a certain extent.

[0098] A text similarity measurement method proposed by the present disclosure is a general big data processing technology, which can be migrated between different application scenarios with almost zero cost, and can be directly deployed into big data engines such as MapReduce, Spark, Storm, etc. through code.

[0099] In an embodiment of the present disclosure, an application scenario of the text similarity measurement method is provided. Specifically, the method further includes: obtaining the request response data transmitted by the first network interface as the first text string; obtaining the request response data transmitted by the second network interface as the second text string; and determining whether there are similar security risks between the first network interface and the second network interface according to the similarity between the first text string and the second text string.

[0100] The text similarity measurement method provided by the embodiments of the present disclosure can be applied to many scenarios of text similarity measurement under big data. For example, in the field of network security, the text similarity measurement method provided by the embodiments of the present disclosure can be directly used to measure the similarity of data returned by different network interfaces, so as to discover more similar security risks.

[0101] Among them, the first network interface is the network interface where a security risk is detected, and the second network interface is any other network interface except the first network interface. For example: a security risk is detected in the first network interface A, and the request response data of the first network interface A is obtained as the first text string. Any other network interface except the first network interface is obtained as the second network interface B, and the request response data of the second network interface B is obtained as the second text string. The text similarity measurement method provided in the above embodiment is used to calculate the similarity between the first text string and the second text string. If the similarity between the first text string and the second text string exceeds a set value, it means that the second network interface B has request and return parameters similar to those of the first network interface A, and it can be considered to a large extent that the network interface B may also have a similar security risk. When faced with hundreds of millions of network interface traffic, the text similarity measurement method provided by the embodiments of the present disclosure can quickly complete the measurement task. Among them, the set value can be set according to the actual situation, for example: 80% or 90%.

[0102] In an embodiment of the present disclosure, an application scenario of the text similarity measurement method is further provided. Specifically, any test result obtained from white-box security testing is obtained as the first text string; the uniform resource locator system URL interface information is obtained as the second text string; according to the similarity between the first text string and the second text string, the URL interface information corresponding to the test result is determined.

[0103] In the embodiments of the present disclosure, there are many automated tools for risk scanning in the field of network security, and there are also many manual test validations. The automated tools for risk scanning include: white-box security testing that identifies potential vulnerability risks based on the data flow dependence of the application program source code; black-box security testing that imitates hackers at various skill levels to initiate data requests from the outside; and security vulnerabilities discovered manually. In order to improve the efficiency between automated tools and between automated tools and humans as much as possible. At least the test results in each dimension need to be connected in the URL interface dimension. However, white-box security testing is based on the source code of the application program repository, and it is difficult to directly obtain the interface information in the URL dimension.

[0104] In the embodiments of the present disclosure, the above test result refers to the test result of any one of the information such as the repository, route, file, source code, etc. scanned by the white-box security test. As the first text string, information such as the path, request body, response body, etc. in the URL dimension is obtained as the second text string, and the text similarity measurement method provided in the above embodiments is used to calculate the similarity between the first text string and the second text string. If the similarity between the first text string and the second text string exceeds the set value, it can be considered that the URL dimension information as the second text string has a corresponding relationship with the test result, and a similarity association relationship between the two is established.

[0105] Figure 9 It is a schematic structural diagram of a text similarity measurement device in the embodiments of the present disclosure. This embodiment is applicable to the situation of calculating the similarity between two texts, and this text similarity measurement device can be implemented in software and / or hardware.

[0106] As Figure 9 shown, the text similarity measurement device 90 provided in the embodiments of the present disclosure mainly includes: a text string acquisition module 91, a sampling string determination module 92, a distance matrix calculation module 93, and a text similarity measurement module 94.

[0107] Among them, the text string acquisition module 91 is used to acquire the first text string and the second text string; the sampling string determination module 92 is used to construct the joint probability distribution of the two based on the first text string and the second text string, and sample the joint probability distribution to obtain a sampling string; the distance matrix calculation module 93 is used to calculate the distance from the first text string to the sampling string to obtain a first distance matrix, and calculate the distance from the second text string to the sampling string to obtain a second distance matrix; the text similarity measurement module 94 is used to determine the similarity between the first text string and the second text string based on the first distance matrix and the second distance matrix.

[0108] In an embodiment of the present disclosure, the edit distance calculation algorithm is used to calculate the first distance matrix and the second distance matrix. The first distance matrix is used to represent the edit distance from the first text string to the sampling string; the second distance matrix is used to represent the edit distance from the second text string to the sampling string.

[0109] In one embodiment of the present disclosure, the text similarity measurement module 94 includes: a first distance representation vector extraction unit configured to perform feature extraction on the first distance matrix to obtain a first distance representation vector, where the first distance representation vector includes a column representation vector, a diagonal representation vector, and a row representation vector corresponding to the first distance matrix; a second distance representation vector extraction unit configured to perform feature extraction on the second distance matrix to obtain a second distance representation vector, where the second distance representation vector includes a column representation vector, a diagonal representation vector, and a row representation vector corresponding to the second distance matrix; and a text similarity calculation unit configured to calculate the vector similarity between the first distance representation vector and the second distance representation vector as the similarity between the first string and the second text string.

[0110] In one embodiment of the present disclosure, the feature extraction includes: extracting the last row sequence in the distance matrix as the row representation vector; extracting the last column sequence in the distance matrix as the column representation vector; and extracting the diagonal sequence in the distance matrix as the diagonal representation vector.

[0111] In one embodiment of the present disclosure, the sampling string determination module 92 is specifically configured to randomly downsample the joint probability distribution based on a preset sampling ratio to obtain a sampling string, where the preset sampling ratio is inversely proportional to the length of the text string.

[0112] In one embodiment of the present disclosure, the Euclidean distance or cosine similarity is used to determine the vector similarity.

[0113] In one embodiment of the present disclosure, the apparatus further includes: a first text string determination module configured to obtain the request response data transmitted by the first network interface as the first text string; a second text string determination module configured to obtain the request response data transmitted by the second network interface as the second text string; and a security risk similarity determination module configured to determine whether there is a similar security risk between the first network interface and the second network interface according to the similarity between the first text string and the second text string.

[0114] In one embodiment of the present disclosure, the first text string determination module is further configured to obtain the test result obtained from white box security testing as the first text string; the second text string determination module is further configured to obtain the uniform resource locator system (URL) interface information as the second text string; and a correspondence determination module configured to determine the URL interface information corresponding to the test result according to the similarity between the first text string and the second text string.

[0115] The text similarity measurement device provided by an embodiment of the present disclosure can execute the steps performed in the text similarity measurement method provided by the method embodiment of the present disclosure. The execution steps and beneficial effects are not described herein again.

[0116] Figure 10 It is a schematic structural diagram of an electronic device in an embodiment of the present disclosure. Specifically, refer to Figure 10 , which shows a schematic structural diagram of the electronic device 1000 suitable for implementing the present disclosure. The electronic device 1000 in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable terminal devices, etc., and fixed terminals such as digital TVs, desktop computers, smart home devices, etc. Figure 10 The illustrated electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0117] As Figure 10 shown, the electronic device 1000 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1001, which can execute various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1008 into the random access memory (RAM) 1003 to implement the text similarity measurement method of the embodiments as described in the present disclosure. In the RAM 1003, various programs and data required for the operation of the terminal device 1000 are also stored. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.

[0118] Generally, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the terminal device 1000 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 10 the illustrated terminal device 1000 has various devices, it should be understood that it is not required to implement or include all the illustrated devices. Instead, more or fewer devices may be implemented or included.

[0119] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a non-transitory computer-readable medium. The computer program contains program code for performing the methods shown in the flowcharts, thereby implementing the text similarity measurement method as described above. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1009, or installed from the storage device 1008, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.

[0120] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0121] In some embodiments, the client and the server can communicate using any known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any known or future-developed network to be tested.

[0122] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0123] The above computer-readable medium carries one or more programs, which, when executed by the terminal device, cause the terminal device to: obtain a first text string and a second text string; construct a joint probability distribution of the two based on the first text string and the second text string, and sample the joint probability distribution to obtain a sampled string; calculate the distance from the first text string to the sampled string to obtain a first distance matrix, and calculate the distance from the second text string to the sampled string to obtain a second distance matrix; determine the similarity between the first text string and the second text string based on the first distance matrix and the second distance matrix.

[0124] Optionally, when the above one or more programs are executed by the terminal device, the terminal device can also perform the other steps described in the above embodiments.

[0125] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0127] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0128] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, the types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0129] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0130] According to one or more embodiments of the present disclosure, the present disclosure provides a method for measuring text similarity, including: obtaining a first text string and a second text string; constructing a joint probability distribution of the two based on the first text string and the second text string, and sampling the joint probability distribution to obtain a sampled string; calculating the distance from the first text string to the sampled string to obtain a first distance matrix, and calculating the distance from the second text string to the sampled string to obtain a second distance matrix; determining the similarity between the first text string and the second text string based on the first distance matrix and the second distance matrix.

[0131] According to one or more embodiments of the present disclosure, the present disclosure provides a device for measuring text similarity, including: a text string obtaining module, configured to obtain a first text string and a second text string; a sampled string determining module, configured to construct a joint probability distribution of the two based on the first text string and the second text string, and sample the joint probability distribution to obtain a sampled string; a distance matrix calculating module, configured to calculate the distance from the first text string to the sampled string to obtain a first distance matrix, and calculate the distance from the second text string to the sampled string to obtain a second distance matrix; a text similarity measuring module, configured to determine the similarity between the first text string and the second text string based on the first distance matrix and the second distance matrix.

[0132] According to one or more embodiments of the present disclosure, the present disclosure provides an electronic device, including:

[0133] One or more processors;

[0134] A memory for storing one or more programs;

[0135] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for measuring text similarity as provided in any one of the present disclosure.

[0136] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for measuring text similarity as provided in any one of the present disclosure.

[0137] The embodiments of the present disclosure further provide a computer program product, which includes a computer program or instruction, and when the computer program or instruction is executed by a processor, it implements the method for measuring text similarity as described above.

[0138] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0139] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0140] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms for implementing the claims.

Claims

1. A method for measuring text similarity, characterized in that, the method includes: Obtain a first text string and a second text string; Based on the first text string and the second text string, construct a joint probability distribution of the two, and sample the joint probability distribution to obtain a sampled string; Calculate the distance from the first text string to the sampled string to obtain a first distance matrix, and calculate the distance from the second text string to the sampled string to obtain a second distance matrix; Extract features from the first distance matrix to obtain a first distance representation vector, where the first distance representation vector includes a column representation vector, a diagonal representation vector, and a row representation vector corresponding to the first distance matrix; Extract features from the second distance matrix to obtain a second distance representation vector, where the second distance representation vector includes a column representation vector, a diagonal representation vector, and a row representation vector corresponding to the second distance matrix; Calculate the vector similarity between the first distance representation vector and the second distance representation vector as the similarity between the first text string and the second text string.

2. The method according to claim 1, characterized in that, Use an edit distance calculation algorithm to calculate the first distance matrix and the second distance matrix. The first distance matrix is used to represent the edit distance from the first text string to the sampled string; the second distance matrix is used to represent the edit distance from the second text string to the sampled string.

3. The method according to claim 1, characterized in that, The feature extraction includes: Extract the last row sequence in the distance matrix as the row representation vector; Extract the last column sequence in the distance matrix as the column representation vector; Extract the diagonal sequence in the distance matrix as the diagonal representation vector.

4. The method according to any one of claims 1 to 3, characterized in that, The sampling the joint probability distribution to obtain a sampled string includes: Randomly downsample the joint probability distribution based on a preset sampling ratio to obtain a sampled string, and the preset sampling ratio is inversely proportional to the length of the text string.

5. The method according to claim 1, characterized in that, Use Euclidean distance or cosine similarity to determine the vector similarity.

6. The method according to claim 1, characterized in that, The method further includes: Obtain the request response data transmitted by the first network interface as the first text string; Obtain the request response data transmitted by the second network interface as the second text string; Determine whether there are similar security risks between the first network interface and the second network interface according to the similarity between the first text string and the second text string.

7. The method according to claim 1, characterized in that, The method further includes: Obtain the test result obtained from white box security testing as the first text string; Obtain the Uniform Resource Locator (URL) interface information as the second text string; Determine the URL interface information corresponding to the test result according to the similarity between the first text string and the second text string.

8. A text similarity measurement device, characterized in that, the device includes: a text string acquisition module for acquiring a first text string and a second text string; a sampled string determination module for constructing a joint probability distribution of the first text string and the second text string and sampling the joint probability distribution to obtain a sampled string; a distance matrix calculation module for calculating the distance from the first text string to the sampled string to obtain a first distance matrix, and calculating the distance from the second text string to the sampled string to obtain a second distance matrix; a text similarity measurement module for extracting features from the first distance matrix to obtain a first distance representation vector, the first distance representation vector including a column representation vector, a diagonal representation vector, and a row representation vector corresponding to the first distance matrix; extracting features from the second distance matrix to obtain a second distance representation vector, the second distance representation vector including a column representation vector, a diagonal representation vector, and a row representation vector corresponding to the second distance matrix; calculating the vector similarity between the first distance representation vector and the second distance representation vector as the similarity between the first text string and the second text string.

9. An electronic device, characterized in that, the electronic device includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, the program, when executed by a processor, implements the method according to any one of claims 1-7.

11. A computer program product, the computer program product includes a computer program or instruction, and the computer program or instruction, when executed by a processor, implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Text topic mining method and device, computer equipment and storage medium

    CN112883154A

  • Text data processing method and device

    CN112926341A

Cited By

  • Text similarity measurement method and apparatus, device, storage medium, and program product

    EP4607376A1