Text similarity detection method and device, electronic device and storage medium
Through text splitting, sentence screening and word screening, the accuracy problem of neural network models when detecting text similarity in long-term texts is solved, and similarity detection at sentence and word levels is realized, which improves the accuracy of text similarity detection and product recommendation efficiency.
Patent Information
- Application Number
- CN202310720391.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-16
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2043-06-16
AI Technical Summary
Existing neural network models are insufficiently accurate when detecting text similarity with longer texts, making it difficult to effectively reflect the similarity between texts.
By performing text splitting, sentence filtering, word segmentation and word filtering on the text, the target sentence vector and keyword vector are constructed, and the similarity scores at the sentence and word level are combined to improve the accuracy of text similarity detection.
The diversity detection of text similarity at the sentence and word levels is realized, the accuracy of text similarity detection is improved, and the efficiency of product recommendations and the transaction rate of financial products are improved.
Smart Images

Figure CN116719999B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of financial technology, and in particular to a text similarity detection method and device, an electronic device, and a storage medium. Background Art
[0002] With the development of science and technology, acquiring data and using data modeling to solve corresponding problems has become a very common technical means. For example, e-commerce platforms collect data such as users' product browsing history and build product recommendation models based on this data to recommend products to users.
[0003] To improve product recommendation efficiency, e-commerce platforms leverage the similarity between different users' product browsing histories. This allows them to recommend products recommended to one user to another user whose product browsing history is even more similar, thus achieving cross-product recommendation. This process often relies heavily on the similarity of textual data, such as users' product browsing histories.
[0004] Currently, text similarity detection often relies on neural network models, but the accuracy of neural network models is often limited by the length of the text. For example, most neural network models can achieve good detection results for short texts, but often cannot accurately detect the similarity between longer texts. Therefore, how to improve the accuracy of text similarity detection has become a pressing issue. Summary of the Invention
[0005] The main purpose of the embodiments of the present application is to provide a text similarity detection method and device, an electronic device and a storage medium, aiming to improve the accuracy of text similarity detection.
[0006] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application proposes a text similarity detection method, the method comprising:
[0007] Get the first text and the second text;
[0008] Performing text splitting on the first text to obtain a plurality of first original text sentences, and performing text splitting on the second text to obtain a plurality of second original text sentences;
[0009] The first original text sentence is screened based on the preset first sentence weight data to obtain a first candidate text sentence, and the second original text sentence is screened based on the preset second sentence weight data to obtain a second candidate text sentence;
[0010] Segmenting the first candidate text sentence to obtain a plurality of first text words, and segmenting the second candidate text sentence to obtain a plurality of second text words;
[0011] Filtering the first text words based on the first sentence weight data and the preset first word weight data to obtain first candidate words, and filtering the second text words based on the second sentence weight data and the preset second word weight data to obtain second candidate words;
[0012] Constructing a first target sentence vector and a first keyword vector for the first text based on the first candidate word, and constructing a second target sentence vector and a second keyword vector for the second text based on the second candidate word;
[0013] Based on the first target sentence vector, the first keyword vector, the second target sentence vector and the second keyword vector, the similarity between the first text and the second text is detected to obtain text detection data; wherein, the text detection data is used to represent the similarity of the text content between the first text and the second text.
[0014] In some embodiments, the filtering of the first original text sentence based on the preset first sentence weight data to obtain the first candidate text sentence, and the filtering of the second original text sentence based on the preset second sentence weight data to obtain the second candidate text sentence, includes:
[0015] calculating first sentence weight data of each first original text sentence based on a preset algorithm, and selecting a predetermined number of first original text sentences having the largest first sentence weight data as the first candidate text sentences;
[0016] The second sentence weight data of each second original text sentence is calculated based on the preset algorithm, and a predetermined number of second original text sentences having the largest second sentence weight data are selected as the second candidate text sentences.
[0017] In some embodiments, the filtering of the first text words based on the first sentence weight data and the preset first word weight data to obtain the first candidate word, and filtering of the second text words based on the second sentence weight data and the preset second word weight data to obtain the second candidate word, includes:
[0018] For each of the first text words, performing weighted calculation on the first sentence weight data of the first candidate sentence in which the first text word is located and the first word weight data of the first text word to obtain first target weight data of the first text word;
[0019] screening the first text words based on the first target weight data to obtain the first candidate words;
[0020] For each second text word, performing weighted calculation on the second sentence weight data of the second candidate sentence in which the second text word is located and the second word weight data of the second text word to obtain second target weight data of the second text word;
[0021] The second text words are screened based on the second target weight data to obtain the second candidate words.
[0022] In some embodiments, constructing a first target sentence vector and a first keyword vector for the first text based on the first candidate word, and constructing a second target sentence vector and a second keyword vector for the second text based on the second candidate word, includes:
[0023] Concatenating all of the first candidate words to obtain a first target sentence, and concatenating all of the second candidate words to obtain a second target sentence;
[0024] Performing embedding processing on the first target sentence to obtain the first target sentence vector, and performing embedding processing on the second target sentence to obtain the second target sentence vector;
[0025] Merging the first target sentence and the second target sentence according to a preset delimiter to obtain a merged text sentence, wherein the sentence content before the delimiter in the merged text sentence comes from the first target sentence, and the sentence content after the delimiter comes from the second target sentence;
[0026] Perform keyword extraction on the merged text sentence to obtain a first keyword group and a second keyword group;
[0027] The first keyword group is embedded to obtain the first keyword vector, and the second keyword group is embedded to obtain the second keyword vector.
[0028] In some embodiments, extracting keywords from the merged text sentences to obtain the first keyword group and the second keyword group includes:
[0029] Performing importance scoring on the text phrases of the merged text sentences based on a preset algorithm to obtain importance scoring data for each text phrase;
[0030] Screening the text phrases before the separator in the merged text sentence based on the importance score data, and extracting the text phrase with the highest importance score data as the first keyword group;
[0031] The text phrases after the separator in the merged text sentence are screened based on the importance score data, and the text phrase with the highest importance score data is extracted as the second keyword group.
[0032] In some embodiments, detecting the similarity between the first text and the second text based on the first target sentence vector, the first keyword vector, the second target sentence vector, and the second keyword vector to obtain text detection data includes:
[0033] Performing a similarity score on the first target sentence vector and the second target sentence vector to obtain first score data;
[0034] Performing a similarity score on the first keyword vector and the second keyword vector to obtain second score data;
[0035] The text detection data is obtained based on the first scoring data and the second scoring data.
[0036] In some embodiments, obtaining the text detection data based on the first scoring data and the second scoring data includes:
[0037] Performing a weighted sum operation on the first scoring data and the second scoring data to obtain target scoring data;
[0038] If the target score data is greater than a preset threshold, determining the text detection data as the text contents of the first text and the second text are similar;
[0039] If the target score data is not greater than a preset threshold, the text detection data is determined to indicate that the text contents of the first text and the second text are not similar.
[0040] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a text similarity detection device, comprising:
[0041] A text acquisition module, used to acquire a first text and a second text;
[0042] a text splitting module, configured to perform text splitting on the first text to obtain a plurality of first original text sentences, and perform text splitting on the second text to obtain a plurality of second original text sentences;
[0043] a sentence screening module, configured to screen the first original text sentence based on preset first sentence weight data to obtain a first candidate text sentence, and to screen the second original text sentence based on preset second sentence weight data to obtain a second candidate text sentence;
[0044] A word segmentation module, configured to segment the first candidate text sentence to obtain a plurality of first text words, and segment the second candidate text sentence to obtain a plurality of second text words;
[0045] a word screening module, configured to screen the first text words based on the first sentence weight data and preset first word weight data to obtain first candidate words, and to screen the second text words based on the second sentence weight data and preset second word weight data to obtain second candidate words;
[0046] a sentence and phrase construction module, configured to construct a first target sentence vector and a first keyword vector for the first text based on the first candidate word, and to construct a second target sentence vector and a second keyword vector for the second text based on the second candidate word;
[0047] A similarity detection module is used to detect the similarity between the first text and the second text based on the first target sentence vector, the first keyword vector, the second target sentence vector and the second keyword vector to obtain text detection data; wherein the text detection data is used to represent the degree of similarity between the text content of the first text and the second text.
[0048] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect is implemented.
[0049] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.
[0050] The text similarity detection method, text similarity detection device, electronic device and storage medium proposed in the present application obtain a first text and a second text; perform text segmentation on the first text and the second text to obtain a plurality of first original text sentences and a plurality of second original text sentences; screen the first original text sentences based on preset first sentence weight data to obtain a first candidate text sentence, and screen the second original text sentences based on preset second sentence weight data to obtain a second candidate text sentence; perform word segmentation on the first candidate text sentence and the second candidate sentence to obtain a plurality of first text words and a plurality of second text words; screen the first text words based on the first sentence weight data and the preset first word weight data to obtain a first candidate word, and screen the second text words based on the second sentence weight data and the preset second word weight data to obtain a second candidate word; and construct a first target sentence vector for the first text based on the first candidate word. and the first keyword vector, and construct a second target sentence vector and a second keyword vector for the second text based on the second candidate word; finally, the similarity between the first text and the second text is detected based on the first target sentence vector, the first keyword vector, the second target sentence vector and the second keyword vector to obtain text detection data. This method can detect the similarity between the first text and the second text at the sentence level and the word level at the same time, which can improve the diversity of similarity detection and the accuracy of text similarity detection, and then recommend products according to the similarity between text data such as product browsing records and query records of different objects. The degree of similarity of interests and preferences of different objects can be determined according to the similarity, and the product recommended to a certain object can be recommended again to another object with relatively similar interests and preferences, which can effectively improve the efficiency and accuracy of product recommendations, and is also conducive to improving the transaction rate of financial products such as insurance products and wealth management products. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a flowchart of a text similarity detection method provided in an embodiment of the present application;
[0052] Figure 2 yes Figure 1 Flowchart of step S103 in FIG.
[0053] Figure 3 yes Figure 1 Flowchart of step S105 in FIG.
[0054] Figure 4 yes Figure 1 Flowchart of step S106 in FIG.
[0055] Figure 5 yes Figure 4 Flowchart of step S404 in FIG.
[0056] Figure 6 yes Figure 1 Flowchart of step S107 in FIG.
[0057] Figure 7 yes Figure 6 Flowchart of step S603 in FIG.
[0058] Figure 8 Schematic diagram of the structure of the text similarity detection device provided in an embodiment of the present application;
[0059] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0060] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0061] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0063] First, let’s analyze some of the terms used in this application:
[0064] Artificial intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.
[0065] Natural language processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese and English). A branch of artificial intelligence, NLP is an interdisciplinary field between computer science and linguistics, often referred to as computational linguistics. Natural language processing encompasses grammatical analysis, semantic analysis, and discourse comprehension. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It encompasses data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistics research related to language computing.
[0066] Information Extraction (NER): A text processing technology that extracts specified types of entity, relationship, event, and other factual information from natural language text and forms structured data output. Information extraction is a technology that extracts specific information from text data. Text data is composed of some specific units, such as sentences, paragraphs, and chapters. Text information is composed of some small specific units, such as characters, words, phrases, sentences, paragraphs, or a combination of these specific units. Extracting noun phrases, names, place names, etc. from text data is all text information extraction. Of course, the information extracted by text information extraction technology can be of various types.
[0067] With the development of science and technology, acquiring data and using data modeling to solve corresponding problems has become a very common technical means. For example, e-commerce platforms collect data such as users' product browsing history and build product recommendation models based on this data to recommend products to users.
[0068] To improve product recommendation efficiency, e-commerce platforms leverage the similarity between different users' product browsing histories. This allows them to recommend products recommended to one user to another user whose product browsing history is even more similar, thus achieving cross-product recommendation. This process often relies heavily on the similarity of textual data, such as users' product browsing histories.
[0069] For example, in the insurance recommendation process, the insurance consultation texts of object A and object B are obtained, and the text similarity of the insurance consultation texts of object A and object B is compared. If the text similarity of the insurance consultation texts of object A and object B is higher than the preset threshold, it indicates that the needs of object A and object B for insurance products are relatively similar. Therefore, the insurance product recommended to object A can be directly recommended to object B to improve the recommendation efficiency and accuracy of insurance products.
[0070] Currently, text similarity detection often relies on neural network models, but the accuracy of neural network models is often limited by the length of the text. For example, most neural network models can achieve good detection results for short texts, but often cannot accurately detect the similarity between longer texts. Therefore, how to improve the accuracy of text similarity detection has become a pressing issue.
[0071] Based on this, embodiments of the present application provide a text similarity detection method, a text similarity detection device, an electronic device, and a storage medium, aiming to improve the accuracy of text similarity detection.
[0072] The text similarity detection method and device, electronic device, and storage medium provided in the embodiments of the present application are specifically described through the following embodiments. First, the text similarity detection method in the embodiments of the present application is described.
[0073] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0074] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0075] The text similarity detection method provided in the embodiment of the present application relates to the field of artificial intelligence technology. The text similarity detection method provided in the embodiment of the present application can be applied in a terminal, can be applied in a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the text similarity detection method, etc., but is not limited to the above forms.
[0076] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0077] Figure 1 This is an optional flowchart of the text similarity detection method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S107.
[0078] Step S101, obtaining a first text and a second text;
[0079] Step S102: performing text splitting on the first text to obtain a plurality of first original text sentences, and performing text splitting on the second text to obtain a plurality of second original text sentences;
[0080] Step S103, screening the first original text sentence based on the preset first sentence weight data to obtain a first candidate text sentence, and screening the second original text sentence based on the preset second sentence weight data to obtain a second candidate text sentence;
[0081] Step S104: segmenting the first candidate text sentence to obtain a plurality of first text words, and segmenting the second candidate text sentence to obtain a plurality of second text words;
[0082] Step S105, screening the first text words based on the first sentence weight data and the preset first word weight data to obtain first candidate words, and screening the second text words based on the second sentence weight data and the preset second word weight data to obtain second candidate words;
[0083] Step S106, constructing a first target sentence vector and a first keyword vector for the first text based on the first candidate word, and constructing a second target sentence vector and a second keyword vector for the second text based on the second candidate word;
[0084] Step S107 : detecting the similarity between the first text and the second text based on the first target sentence vector, the first keyword vector, the second target sentence vector, and the second keyword vector to obtain text detection data.
[0085] The steps S101 to S107 shown in the embodiment of the present application are as follows: obtaining a first text and a second text; performing text splitting on the first text and the second text to obtain a plurality of first original text sentences and a plurality of second original text sentences; screening the first original text sentences based on the preset first sentence weight data to obtain a first candidate text sentence, and screening the second original text sentences based on the preset second sentence weight data to obtain a second candidate text sentence; performing word segmentation on the first candidate text sentence and the second candidate sentence to obtain a plurality of first text words and a plurality of second text words; screening the first text words based on the first sentence weight data and the preset first word weight data to obtain a first candidate word, and screening the first text words based on the preset second sentence weight data to obtain a first candidate word. The second text words are screened based on the second sentence weight data and the preset second word weight data to obtain second candidate words; a first target sentence vector and a first keyword vector are constructed for the first text based on the first candidate words, and a second target sentence vector and a second keyword vector are constructed for the second text based on the second candidate words; finally, the similarity between the first text and the second text is detected based on the first target sentence vector, the first keyword vector, the second target sentence vector and the second keyword vector to obtain text detection data. This method can detect the similarity between the first text and the second text at both the sentence level and the word level, which can improve the diversity of similarity detection and the accuracy of text similarity detection.
[0086] In step S101 of some embodiments, the first and second texts can be obtained by programming a web crawler, setting up a data source, and then crawling data in a targeted manner. Alternatively, the first and second texts can be downloaded from a network platform, extracted from a pre-set text database, or obtained through other means, without limitation. The first and second texts include text data from various business fields, such as economic and financial texts, daily life texts, medical disease texts, sports event texts, commentary text data, question text data, and the like, without limitation.
[0087] In the financial field, the first text and the second text can be product consulting texts for insurance products and financial products, consulting texts for handling certain financial services, or other text data containing financial business requirements, without restriction.
[0088] In step S102 of some embodiments, when the first text is subjected to text splitting to obtain a plurality of first original text sentences, and the second text is subjected to text splitting to obtain a plurality of second original text sentences, text sentence splitting can be performed according to a preset regular expression, that is, the first text and the second text are converted into the form of a string, each string is matched with a preset regular expression, the sentence position of the first text and the second text is determined according to the matching situation, and the first text and the second text are subjected to text splitting at the sentence position. For example, when the string meets the format requirements of the regular expression, the end position of the string is determined to be the sentence position. In addition, text splitting can also be performed according to a preset text splitting rule, specifically, first the quotation mark portion of the first text and the second text is replaced with other fixed characters, and then the text sentence is performed according to the position of the period, question mark, exclamation mark, and ellipsis in the first text and the second text. After the sentence is completed, the quotation mark portion replaced with the fixed character is restored to its content, thereby completing text splitting and obtaining a plurality of first original text sentences and a plurality of second original text sentences. This method can more conveniently split the first text and the second text into sentences, convert the text-level similarity detection into sentence-level similarity detection, and realize the fine-grained text detection, which is conducive to improving the detection accuracy.
[0089] See also Figure 2 In some embodiments, step S103 may include but is not limited to steps S201 to S202:
[0090] Step S201, calculating the first sentence weight data of each first original text sentence based on a preset algorithm, and selecting a predetermined number of first original text sentences with the largest first sentence weight data as first candidate text sentences;
[0091] Step S202 : calculating the second sentence weight data of each second original text sentence based on a preset algorithm, and selecting a predetermined number of second original text sentences with the largest second sentence weight data as second candidate text sentences.
[0092] In step S201 of some embodiments, the preset algorithm may be a TextRank algorithm or other algorithms, without limitation. Taking the TextRank algorithm as an example, all original text sentences of the first text are formed into a sentence network, each original text sentence is a sentence node, and the links between sentence nodes represent the semantic relationship between sentences. The specific process of calculating the first sentence weight data of each first original text sentence in the sentence network using the TextRank algorithm can be expressed as shown in formula (1):
[0093]
[0094] Among them, WS(Vi ) represents the weight data of the first sentence of the first original text sentence i; WS(V j ) represents the weight data of the first sentence of the first original text sentence j, which can be obtained in the previous iteration process; d represents the damping coefficient, which is generally 0.85; In(V i ) is the set of all other first original text sentences adjacent to the first original text sentence i; Out(V j ) is the set of all other first original text sentences adjacent to the first original text sentence j; W ji is the similarity between the first original text sentence i and the first original text sentence j, W jk is the similarity between the first original text sentence j and the first original text sentence k, wherein the similarity between the first original text sentences can be calculated based on the cosine similarity algorithm.
[0095] It should be noted that, in the embodiment of the present application, it is assumed that all the first original text sentences are adjacent.
[0096] After calculating the first sentence weight data of each first original text sentence, since the first sentence weight data of each first original text sentence can intuitively reflect the importance of the first original text sentence in the first text, the larger the first sentence weight data is, the more important the first original text sentence data is. Therefore, the first original text sentences whose first sentence weight data are ranked in the front of a predetermined number are screened out as the first candidate text sentences. The predetermined number can be set according to actual business needs and is not limited.
[0097] In step S202 of some embodiments, the second sentence weight data of each second original text sentence is calculated based on a preset algorithm, and a predetermined number of second original text sentences with the largest second sentence weight data are selected as second candidate text sentences. The implementation process is similar to the above-mentioned step S201 and will not be repeated here.
[0098] Through the above steps S201 to S202, the TextRank algorithm can be used to calculate the sentence weight of each text sentence in the first text and the second text, determine the more important text sentences in the first text and the second text, and select several text sentences with higher importance as the first candidate text sentences and the second candidate text sentences, so that representative text sentences can be used to perform text similarity detection, which can improve the rationality and accuracy of text similarity detection.
[0099] In step S104 of some embodiments, when the first candidate text sentence is segmented to obtain multiple first text words, and the second candidate text sentence is segmented to obtain multiple second text words, the first candidate text sentence and the second candidate text sentence can be segmented based on the preset Jieba segmenter, and the sentences can be segmented according to the part-of-speech categories of the text words in the first candidate text sentence to obtain multiple first text words, and the sentences can be segmented according to the part-of-speech categories of the text words in the second candidate text sentence to obtain multiple second text words, wherein the part-of-speech categories include modifiers, nouns, verbs, adjectives, adverbs, etc. This method can more conveniently split the first candidate text sentence and the second candidate text sentence into text words, further convert the sentence-level similarity detection into the word-level similarity detection, and can achieve fine-grained text detection, so that text similarity detection can be performed at both the sentence level and the word level, which is conducive to improving detection accuracy.
[0100] See also Figure 3 In some embodiments, step S105 may include but is not limited to steps S301 to S304:
[0101] Step S301: For each first text word, weighted calculation is performed on the first sentence weight data of the first candidate sentence in which the first text word is located and the first word weight data of the first text word to obtain first target weight data of the first text word;
[0102] Step S302: screening the first text words based on the first target weight data to obtain a first candidate word;
[0103] Step S303: for each second text word, performing weighted calculation on the second sentence weight data of the second candidate sentence in which the second text word is located and the second word weight data of the second text word to obtain second target weight data of the second text word;
[0104] Step S304: Screen the second text words based on the second target weight data to obtain second candidate words.
[0105] In step S301 of some embodiments, for each first text word, the above-mentioned TextRank algorithm is used to calculate the word weight of the first text word to obtain the first word weight data of each first text word. The calculation process is similar to the above-mentioned step S201, except that step S201 uses the first original text sentence as a sentence network, the first original text sentence as a node, the semantic connection between the two first original text sentences as an edge, and the similarity between the first original text sentences as an edge weight; while step S301 uses words as a word network, the first text word as a node, the semantic connection between the two first original text sentences as an edge, and the similarity between the first original text sentences as an edge weight. Its implementation process is similar to the implementation details of step S201 and will not be repeated here. Here, for each first text word, the first sentence weight data of the first candidate sentence in which the first text word is located is extracted, the first sentence weight data is multiplied by the first word weight data, and the first target weight data of the first text word is obtained.
[0106] For example, the first sentence weight data of the first original text sentence i is 0.35, the first sentence weight data of the first original text sentence j is 0.3, the first word weight data of the first text word A and the first text word B in the first original text sentence i are 0.44 and 0.2 respectively, and the first word weight data of the first text word C and the first text word D in the first original text sentence j are 0.12 and 0.33 respectively. Then the first target weight data of the first text word A is 0.35*0.44=0.154, the first target weight data of the first text word B is 0.35*0.2=0.07, the first target weight data of the first text word C is 0.3*0.12=0.036, and the first target weight data of the first text word D is 0.3*0.33=0.099.
[0107] In step S302 of some embodiments, since the first target weight data can reflect the importance of each first text word to the first text, the larger the first target weight data is, the higher the importance of the first text word is. Therefore, according to the first target data, the first text word in the top m positions is selected as the first candidate word.
[0108] For example, the first target weight data of the aforementioned first text word A, first text word B, first text word C, and first text word D are 0.154, 0.07, 0.036, and 0.099, respectively. Therefore, the first text words with the first target weight data in the top 2 are selected as the first candidate words, and the first candidate words are the first text word A and the first text word D.
[0109] In step S303 and step S304 of some embodiments, the specific calculation process of the second target weight data of the second text words and the word screening process are basically the same as those in step S301 and step S302, and will not be repeated here.
[0110] Through the above steps S301 to S304, the sentence weight data of the candidate sentences in which the first text word and the second text word are located and their own word weight data can be conveniently combined to comprehensively determine the final weight data of each first text word and the second text word, which can improve the accuracy of judging the importance of the first text word and the second text word in the text, and is conducive to improving the accuracy of word screening, so that the screened first candidate words and second candidate words can be more representative and better reflect the real semantic information of the text, thereby improving the accuracy of text similarity detection.
[0111] See also Figure 4 In some embodiments, step S106 may include but is not limited to steps S401 to S405:
[0112] Step S401: concatenate all first candidate words to obtain a first target sentence, and concatenate all second candidate words to obtain a second target sentence;
[0113] Step S402: embedding the first target sentence to obtain a first target sentence vector, and embedding the second target sentence to obtain a second target sentence vector;
[0114] Step S403: Merging the first target sentence and the second target sentence according to a preset delimiter to obtain a merged text sentence, wherein the sentence content before the delimiter in the merged text sentence is derived from the first target sentence, and the sentence content after the delimiter is derived from the second target sentence;
[0115] Step S404: extract keywords from the merged text sentences to obtain a first keyword group and a second keyword group;
[0116] Step S405 : Embedding the first keyword group to obtain a first keyword vector, and embedding the second keyword group to obtain a second keyword vector.
[0117] In step S401 of some embodiments, all first candidate words are concatenated according to the order in which each first candidate word appears in the first text to obtain a first target sentence. Since each first candidate word is a word of high importance in the first text, the concatenated first target sentence can more accurately reflect the key content of the first text. Similarly, according to the order in which each second candidate word appears in the second text, all second candidate words are concatenated to obtain a second target sentence. Since each second candidate word is a word of high importance in the second text, the concatenated second target sentence can more accurately reflect the key content of the second text.
[0118] In step S402 of some embodiments, the first target sentence can be embedded using a preset Bert model, and the first target sentence can be mapped from the text data space to the vector space to obtain a first target sentence vector; similarly, the second target sentence can be embedded using the Bert model, and the second target sentence can be mapped from the text data space to the vector space to obtain a second target sentence vector.
[0119] In step S403 of some embodiments, the preset delimiters include CLS characters and SEP characters, and the preset delimiters are used to merge the first target sentence and the second target sentence to obtain a merged text sentence, wherein the CLS character is set at the beginning of the entire merged text sentence, the first SEP character is set at the end of the first target sentence, and the second SEP character is set at the end of the second target sentence, and the first SEP character is used as a key delimiter for distinguishing the front part and the back part of the merged text sentence, that is, the sentence content between the CLS character and the first SEP character comes from the first target sentence, and the sentence content between the first SEP character and the second SEP character comes from the second target sentence.
[0120] In step S404 of some embodiments, when performing keyword extraction on the merged text sentence, a preset algorithm can be used to calculate the importance score of each word in the merged text sentence, and the degree of importance of the word in the entire merged text sentence is determined by the importance score. A word with the highest importance score is selected from the part of the text sentence before the separator as the first keyword, and a first keyword group is determined based on the adjacent words of the first keyword. The first keyword is composed of the first keyword and its adjacent words; a word with the highest importance score is selected from the part of the text sentence after the separator as the second keyword, and a second keyword group is determined based on the adjacent words of the second keyword. The second keyword is composed of the second keyword and its adjacent words. The first keyword group is a phrase that can be used to represent the text content of the first text, and the second keyword group is a phrase that can be used to represent the text content of the second text.
[0121] In step S405 of some embodiments, the first keyword group can be embedded using the Bert model, and the first keyword group can be mapped to the word vector space to obtain a first keyword vector; similarly, the second keyword group can be embedded using the Bert model, and the second keyword group can be mapped to the word vector space to obtain a second keyword vector.
[0122] Through the above steps S401 to S405, the first target sentence and the first keyword group used to represent the text content of the first text, as well as the second target sentence and the second keyword group used to represent the text content of the second text can be determined more conveniently, so that the similarity between the first text and the second text can be detected at the sentence level and the word level at the same time, which can improve the detection diversity and detection accuracy.
[0123] See also Figure 5 In some embodiments, step S404 may include but is not limited to steps S501 to S503:
[0124] Step S501: performing importance scoring on the text phrases of the merged text sentences based on a preset algorithm to obtain importance score data for each text phrase;
[0125] Step S502: screening the text phrases before the separator in the merged text sentence based on the importance score data, and extracting the text phrase with the highest importance score data as the first keyword phrase;
[0126] Step S503 : screening the text phrases after the separator in the merged text sentence based on the importance score data, and extracting the text phrase with the highest importance score data as the second keyword phrase.
[0127] In step S501 of some embodiments, the preset algorithm may be the above-mentioned TextRank algorithm. The process of using the TextRank algorithm to perform importance scoring on the text phrases of the merged text sentences is similar to the implementation process of the above-mentioned step S301. The difference is that step S301 performs weight calculation on the words in the first original text sentence, while step S501 performs weight calculation on the text phrases of the merged text sentences. The result of the weight calculation is used as the importance scoring data for each text phrase, which will not be repeated here.
[0128] In step S502 of some embodiments, since the importance score data can reflect the importance of each text phrase to the merged text sentence, the larger the importance score data, the higher the importance of the text phrase. Therefore, based on the importance score data, the text phrase with the highest importance score data is extracted as the first keyword phrase among the text phrases before the separator in the merged text sentence (that is, the sentence content between the CLS character and the first SEP character in the merged text sentence).
[0129] In step S503 of some embodiments, based on the importance score data, the text phrase with the highest importance score data is extracted from the text phrases after the separator in the merged text sentence (i.e., the sentence content between the first SEP character and the second SEP character in the merged text sentence) as the second keyword phrase.
[0130] It should be noted that, in a financial scenario, the first keyword group and the second keyword group may include common words such as finance, transaction, insurance, bank, online banking, interest, credit, mortgage, stock, investment, fund, asset portfolio, pension, etc.
[0131] Through the above steps S501 to S503, the accuracy of word screening can be improved, so that the screened first keyword group and second keyword group can be more representative and can better reflect the key content information of the first text and the second text, thereby improving the accuracy of text similarity detection.
[0132] See also Figure 6 In some embodiments, step S107 includes but is not limited to steps S601 to S603:
[0133] Step S601, performing similarity scoring on the first target sentence vector and the second target sentence vector to obtain first scoring data;
[0134] Step S602: Perform similarity scoring on the first keyword vector and the second keyword vector to obtain second scoring data;
[0135] Step S603: Obtain text detection data based on the first scoring data and the second scoring data.
[0136] In step S601 of some embodiments, when the first target sentence vector and the second target sentence vector are scored for similarity to obtain the first scoring data, the similarity scoring process may be to calculate the Euclidean distance, Manhattan distance, or cosine similarity value between the first target sentence vector and the second target sentence vector, without limitation. For example, the first target sentence vector and the second target sentence vector are scored for similarity using the cosine similarity algorithm, and the calculated cosine similarity value is used as the first scoring data. The calculation process can be expressed as shown in formula (2):
[0137]
[0138] Among them, sim(u, v) is the first scoring data, the first target sentence vector is u, the second target sentence vector is v, u T is the result of transposing the second target sentence vector.
[0139] In step S602 of some embodiments, when performing similarity scoring on the first keyword vector and the second keyword vector to obtain second scoring data, the similarity scoring process can be calculating the Euclidean distance, Manhattan distance, or cosine similarity value between the first keyword vector and the second keyword vector, etc., without limitation.
[0140] In step S603 of some embodiments, since the first scoring data is the similarity score between the first text and the second text at the sentence level, and the second scoring data is the similarity score between the first text and the second text at the word level, scoring weights can be set separately for different feature levels, the first scoring data and the second scoring data are weighted and calculated, and the calculation result is compared with a preset threshold. Based on the comparison, the text detection data is determined as the text contents of the first text and the second text are similar or the text contents of the first text and the second text are dissimilar.
[0141] Through the above steps S601 to S603, the text similarity between the first text and the second text can be detected based on the two granularities of word level and sentence level, and different weights can be set for different granularities, so that the weight ratios of the word-level similarity score and the sentence-level similarity score can be flexibly adjusted according to the different degrees of influence on text similarity detection in specific application scenarios, thereby improving the accuracy and flexibility of text similarity detection.
[0142] See also Figure 7 In some embodiments, step S603 may include but is not limited to steps S701 to S703:
[0143] Step S701, performing a weighted sum operation on the first scoring data and the second scoring data to obtain target scoring data;
[0144] Step S702: If the target score data is greater than a preset threshold, the text detection data is determined to indicate that the text contents of the first text and the second text are similar;
[0145] Step S703: If the target score data is not greater than a preset threshold, the text detection data is determined to indicate that the text contents of the first text and the second text are not similar.
[0146] In step S701 of some embodiments, for texts of different lengths, the word-level scoring and sentence-level scoring have different degrees of influence on text similarity. For example, for texts of shorter lengths, both word granularity and sentence granularity can represent the basic text content of the text. Therefore, there is no significant difference in the degree of influence of word-level and sentence-level similarity scoring on text similarity detection. However, for texts of longer lengths, word granularity often cannot represent the actual text content of the text. Therefore, sentence-level similarity scoring will have a greater impact on text similarity detection than word-level similarity scoring. Based on this, for different text lengths, different weight values can be assigned to scoring data of different granularities, and then a weighted calculation is performed on the weighted first scoring data and the second scoring data, that is, the first scoring data and the second scoring data are multiplied by their corresponding weight values to obtain a product result, and the two product results are summed to obtain the target scoring data. For example, when the length of both the first and second texts is no greater than a preset text length threshold, the weight values of the first and second scoring data are both set to 0.5, so the target scoring data = 0.5 * first scoring data + 0.5 * second scoring data; when one of the first and second texts has a length greater than the preset text length threshold, the weight value of the first scoring data is set to 0.7, and the weight value of the second scoring data is set to 0.3, so the target scoring data = 0.7 * first scoring data + 0.3 * second scoring data. This method can assign different weights to sentence-level similarity scoring data and word-level similarity scoring data based on different situations, which is beneficial to improving the calculation accuracy of the target scoring data and the accuracy of text similarity detection.
[0147] Furthermore, the target scoring data is compared with a preset threshold, and whether the text content similarity between the first text and the second text meets the requirement is determined based on the size of the target scoring data and the preset threshold.
[0148] In step S702 of some embodiments, if the target score data is greater than a preset threshold, it indicates that the degree of similarity between the text content of the first text and the second text meets the requirements, and the first text and the second text can represent the same semantic content information. Therefore, the text detection data is determined as the text content of the first text and the second text is similar.
[0149] In step S703 of some embodiments, if the target score data is less than or equal to a preset threshold, it indicates that the degree of similarity between the text content of the first text and the second text does not meet the requirements, and the semantic content information represented by the first text and the second text is different. Therefore, the text detection data is determined as the text content of the first text and the second text is not similar.
[0150] Through the above steps S701 to S703, it is possible to determine whether the similarity between the text contents of the first text and the second text meets the requirements by means of threshold comparison, thereby making it more convenient to determine whether the first text and the second text are similar texts or dissimilar texts, thereby improving the accuracy of text similarity detection and the rationality of judgment.
[0151] The text similarity detection method of the embodiment of the present application obtains a first text and a second text; performs text splitting on the first text and the second text to obtain a plurality of first original text sentences and a plurality of second original text sentences; screens the first original text sentences based on preset first sentence weight data to obtain a first candidate text sentence, and screens the second original text sentences based on preset second sentence weight data to obtain a second candidate text sentence; performs word segmentation on the first candidate text sentence and the second candidate sentence to obtain a plurality of first text words and a plurality of second text words; screens the first text words based on the first sentence weight data and the preset first word weight data to obtain a first candidate word, and screens the second text words based on the second sentence weight data and the preset second word weight data to obtain a second candidate word; constructs a first target sentence vector and a first keyword vector for the first text based on the first candidate word, and Based on the second candidate words, a second target sentence vector and a second keyword vector are constructed for the second text; finally, based on the first target sentence vector, the first keyword vector, the second target sentence vector and the second keyword vector, the similarity between the first text and the second text is detected to obtain text detection data. This method can detect the similarity between the first text and the second text at the sentence level and the word level at the same time, which can improve the diversity of similarity detection and the accuracy of text similarity detection, and then make product recommendations based on the similarity between text data such as product browsing records and query records of different objects. The degree of similarity of interests and preferences of different objects can be determined based on the similarity, and the product recommended to a certain object can be recommended again to another object with relatively similar interests and preferences. This can effectively improve the efficiency and accuracy of product recommendations, and is also conducive to improving the transaction rate of financial products such as insurance products and wealth management products.
[0152] See also Figure 8 The present application also provides a text similarity detection device that can implement the above-mentioned text similarity detection method. The device includes:
[0153] A text acquisition module 801 is used to acquire a first text and a second text;
[0154] A text splitting module 802 is configured to perform text splitting on the first text to obtain a plurality of first original text sentences, and to perform text splitting on the second text to obtain a plurality of second original text sentences;
[0155] The sentence screening module 803 is configured to screen the first original text sentence based on the preset first sentence weight data to obtain a first candidate text sentence, and screen the second original text sentence based on the preset second sentence weight data to obtain a second candidate text sentence;
[0156] A word segmentation module 804 is configured to segment the first candidate text sentence to obtain a plurality of first text words, and segment the second candidate text sentence to obtain a plurality of second text words;
[0157] A word screening module 805 is configured to screen the first text words based on the first sentence weight data and the preset first word weight data to obtain first candidate words, and to screen the second text words based on the second sentence weight data and the preset second word weight data to obtain second candidate words;
[0158] a sentence and phrase construction module 806 for constructing a first target sentence vector and a first keyword vector for the first text based on the first candidate word, and constructing a second target sentence vector and a second keyword vector for the second text based on the second candidate word;
[0159] The similarity detection module 807 is used to detect the similarity between the first text and the second text based on the first target sentence vector, the first keyword vector, the second target sentence vector, and the second keyword vector to obtain text detection data; wherein the text detection data is used to represent the degree of similarity between the text content of the first text and the second text.
[0160] The specific implementation of the text similarity detection device is substantially the same as the specific embodiment of the above-mentioned text similarity detection method, and will not be described in detail here.
[0161] The present application also provides an electronic device comprising: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for enabling communication between the processor and the memory. When the program is executed by the processor, the above-described text similarity detection method is implemented. The electronic device can be any intelligent terminal, including a tablet computer and an in-vehicle computer.
[0162] See also Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0163] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0164] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the text similarity detection method of the embodiments of this application.
[0165] Input / output interface 903, used to implement information input and output;
[0166] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0167] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );
[0168] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0169] An embodiment of the present application further provides a computer-readable storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the above-mentioned text similarity detection method.
[0170] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0171] The text similarity detection method, text similarity detection device, electronic device and computer-readable storage medium provided by the embodiments of the present application are as follows: obtaining a first text and a second text; performing text splitting on the first text and the second text to obtain a plurality of first original text sentences and a plurality of second original text sentences; screening the first original text sentences based on preset first sentence weight data to obtain a first candidate text sentence, and screening the second original text sentences based on preset second sentence weight data to obtain a second candidate text sentence; performing word segmentation on the first candidate text sentence and the second candidate sentence to obtain a plurality of first text words and a plurality of second text words; screening the first text words based on the first sentence weight data and the preset first word weight data to obtain a first candidate word, and screening the second text words based on the second sentence weight data and the preset second word weight data to obtain a second candidate word; constructing a first keyword for the first text based on the first candidate word. The first target sentence vector and the first keyword vector are marked, and a second target sentence vector and a second keyword vector are constructed for the second text based on the second candidate word; finally, the similarity between the first text and the second text is detected based on the first target sentence vector, the first keyword vector, the second target sentence vector and the second keyword vector to obtain text detection data. This method can detect the similarity between the first text and the second text at the sentence level and the word level at the same time, which can improve the diversity of similarity detection and the accuracy of text similarity detection, and then recommend products according to the similarity between text data such as product browsing records and query records of different objects. The similarity between the interests and preferences of different objects can be determined based on the similarity, and the products recommended to a certain object can be recommended again to another object with similar interests and preferences. This can effectively improve the efficiency and accuracy of product recommendations, and is also conducive to improving the transaction rate of financial products such as insurance products and wealth management products.
[0172] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0173] It will be understood by those skilled in the art that Figure 1-7 The technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or a combination of certain steps, or different steps.
[0174] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0175] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0176] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0177] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0178] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0179] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0180] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0181] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0182] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A text similarity detection method, characterized in that: The method comprises: Get the first text and the second text; Performing text splitting on the first text to obtain a plurality of first original text sentences, and performing text splitting on the second text to obtain a plurality of second original text sentences; The first original text sentence is screened based on the preset first sentence weight data to obtain a first candidate text sentence, and the second original text sentence is screened based on the preset second sentence weight data to obtain a second candidate text sentence; Segmenting the first candidate text sentence to obtain a plurality of first text words, and segmenting the second candidate text sentence to obtain a plurality of second text words; Filtering the first text words based on the first sentence weight data and the preset first word weight data to obtain first candidate words, and filtering the second text words based on the second sentence weight data and the preset second word weight data to obtain second candidate words; Constructing a first target sentence vector and a first keyword vector for the first text based on the first candidate word, and constructing a second target sentence vector and a second keyword vector for the second text based on the second candidate word; Based on the first target sentence vector, the first keyword vector, the second target sentence vector and the second keyword vector, the similarity between the first text and the second text is detected to obtain text detection data; wherein, the text detection data is used to represent the similarity of the text content between the first text and the second text.
2. The text similarity detection method according to claim 1, characterized in that: The filtering of the first original text sentence based on the preset first sentence weight data to obtain a first candidate text sentence, and filtering of the second original text sentence based on the preset second sentence weight data to obtain a second candidate text sentence, includes: calculating first sentence weight data of each first original text sentence based on a preset algorithm, and selecting a predetermined number of first original text sentences having the largest first sentence weight data as the first candidate text sentences; The second sentence weight data of each second original text sentence is calculated based on the preset algorithm, and a predetermined number of second original text sentences having the largest second sentence weight data are selected as the second candidate text sentences.
3. The text similarity detection method according to claim 1, characterized in that: The step of screening the first text words based on the first sentence weight data and the preset first word weight data to obtain a first candidate word, and screening the second text words based on the second sentence weight data and the preset second word weight data to obtain a second candidate word includes: For each of the first text words, performing weighted calculation on the first sentence weight data of the first candidate sentence in which the first text word is located and the first word weight data of the first text word to obtain first target weight data of the first text word; screening the first text words based on the first target weight data to obtain the first candidate words; For each second text word, performing weighted calculation on the second sentence weight data of the second candidate sentence in which the second text word is located and the second word weight data of the second text word to obtain second target weight data of the second text word; The second text words are screened based on the second target weight data to obtain the second candidate words.
4. The text similarity detection method according to claim 1, characterized in that: The step of constructing a first target sentence vector and a first keyword vector for the first text based on the first candidate word, and constructing a second target sentence vector and a second keyword vector for the second text based on the second candidate word, includes: Concatenating all of the first candidate words to obtain a first target sentence, and concatenating all of the second candidate words to obtain a second target sentence; Performing embedding processing on the first target sentence to obtain the first target sentence vector, and performing embedding processing on the second target sentence to obtain the second target sentence vector; Merging the first target sentence and the second target sentence according to a preset delimiter to obtain a merged text sentence, wherein the sentence content before the delimiter in the merged text sentence comes from the first target sentence, and the sentence content after the delimiter comes from the second target sentence; Perform keyword extraction on the merged text sentence to obtain a first keyword group and a second keyword group; The first keyword group is embedded to obtain the first keyword vector, and the second keyword group is embedded to obtain the second keyword vector.
5. The text similarity detection method according to claim 4, characterized in that: The keyword extraction of the merged text sentence to obtain the first keyword group and the second keyword group includes: Performing importance scoring on the text phrases of the merged text sentences based on a preset algorithm to obtain importance scoring data for each text phrase; Screening the text phrases before the separator in the merged text sentence based on the importance score data, and extracting the text phrase with the highest importance score data as the first keyword group; The text phrases after the separator in the merged text sentence are screened based on the importance score data, and the text phrase with the highest importance score data is extracted as the second keyword group.
6. The text similarity detection method according to claim 1, characterized in that: The detecting the similarity between the first text and the second text based on the first target sentence vector, the first keyword vector, the second target sentence vector, and the second keyword vector to obtain text detection data includes: Performing a similarity score on the first target sentence vector and the second target sentence vector to obtain first score data; Performing a similarity score on the first keyword vector and the second keyword vector to obtain second score data; The text detection data is obtained based on the first scoring data and the second scoring data.
7. The text similarity detection method according to claim 6, characterized in that: The obtaining of the text detection data based on the first scoring data and the second scoring data includes: Performing a weighted sum operation on the first scoring data and the second scoring data to obtain target scoring data; If the target score data is greater than a preset threshold, determining the text detection data as the text contents of the first text and the second text are similar; If the target score data is not greater than a preset threshold, the text detection data is determined to indicate that the text contents of the first text and the second text are not similar.
8. A text similarity detection device, characterized in that: The device comprises: A text acquisition module, used to acquire a first text and a second text; a text splitting module, configured to perform text splitting on the first text to obtain a plurality of first original text sentences, and perform text splitting on the second text to obtain a plurality of second original text sentences; a sentence screening module, configured to screen the first original text sentence based on preset first sentence weight data to obtain a first candidate text sentence, and to screen the second original text sentence based on preset second sentence weight data to obtain a second candidate text sentence; A word segmentation module, configured to segment the first candidate text sentence to obtain a plurality of first text words, and segment the second candidate text sentence to obtain a plurality of second text words; a word screening module, configured to screen the first text words based on the first sentence weight data and preset first word weight data to obtain first candidate words, and to screen the second text words based on the second sentence weight data and preset second word weight data to obtain second candidate words; a sentence and phrase construction module, configured to construct a first target sentence vector and a first keyword vector for the first text based on the first candidate word, and to construct a second target sentence vector and a second keyword vector for the second text based on the second candidate word; A similarity detection module is used to detect the similarity between the first text and the second text based on the first target sentence vector, the first keyword vector, the second target sentence vector and the second keyword vector to obtain text detection data; wherein the text detection data is used to represent the degree of similarity between the text content of the first text and the second text.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the text similarity detection method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the text similarity detection method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Text similarity computing method and device
CN107644010A
Text similarity analysis method and apparatus, electronic device, and computer readable storage medium
CN108804421A