Short text retrieval method and device based on ordered compactness

By calculating the word position order matching and word frequency inverse text frequency model of input text and database text, the problem of word matching error in short text retrieval is solved, achieving more accurate search results and better user experience.

CN120256608APending Publication Date: 2025-07-04GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410014556.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the existing short text search methods, there are errors in searching based on the degree of word matching, which makes it difficult for the retrieved short text to meet user needs, especially when the word position span is too large, the semantic correlation is small.

Method used

By calculating the sequential matching degree of the word position of the input text and the database text, combining the word frequency inverse text frequency model, the orderly compactness score is calculated to ensure that the order of the word position of the search target and the input text meet the correlation requirements and reduce short text retrieval errors.

Benefits of technology

It improves the accuracy of short text retrieval, makes the search results more in line with user needs, improves the search experience, reduces errors, and eliminates the need for a large amount of labeled data and manual writing rules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256608A_ABST
    Figure CN120256608A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a short text retrieval method and device based on ordered compactness. According to the technical scheme provided by the embodiment of the invention, the to-be-retrieved input text is obtained, and the word matching degree score of the input text and the corresponding database short text is determined; calculating an ordered compactness score of the input text and the corresponding database short text based on the word position sequence information, wherein the ordered compactness score is used for representing the word position sequence matching degree of the input text and the database short text; and determining a retrieval target from the database short text according to the word matching degree score and the ordered compactness score, and outputting the retrieval target. It can be seen that on the basis of word matching, by calculating the word position sequence matching degree of the input text and the database text, it is ensured that the word position sequence of the retrieval target and the word position sequence of the input text meet the correlation degree requirement, the short text retrieval error is reduced, and the more accurate retrieval target is output; and a short text retrieval output result better meets user requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular, to a short text retrieval method and device based on ordered compactness. Background Art

[0002] Short text retrieval refers to the process of retrieving short texts that meet the user's information needs from a large set of data with irregular structures and short text lengths. Currently, when performing short text search, the text to be retrieved is usually input into a pre-trained text matching model, and the model calculates the word matching degree between the text to be retrieved and each short text in the database, and then outputs the associated short text according to the word matching degree.

[0003] However, for multiple words that match between the text to be retrieved and the associated short text, the positions where they appear in the short text may span too large, resulting in that although the matched short text matches the text to be retrieved in terms of words, the actual semantics may have a small degree of association. Determining the associated short text simply based on the word matching degree has retrieval errors, making it difficult for the retrieved short text to meet the user's needs. Summary of the Invention

[0004] The embodiments of the present application provide a short text retrieval method and device based on ordered compactness, which can improve the short text retrieval accuracy, make the retrieved short text more in line with the user's needs, and solve the error problem of retrieving short texts according to the word matching degree.

[0005] In a first aspect, the embodiments of the present application provide a short text retrieval method based on ordered compactness, including:

[0006] Obtain the input text to be retrieved, and determine the word matching degree score between the input text and the short text in the corresponding database;

[0007] Calculate the ordered compactness score between the input text and the corresponding short text in the database based on the word position order information, and the ordered compactness score is used to represent the word position order matching degree between the input text and the short text in the database;

[0008] Determine and output the retrieval target from the short texts in the database according to the word matching degree score and the ordered compactness score.

[0009] It can be seen that based on word matching, the present application calculates the word position order matching degree between the input text and the database text to ensure that the word position order of the retrieval target and the input text meets the association degree requirement, reduces the short text retrieval error, outputs a more accurate retrieval target, makes the short text retrieval output result more in line with the user's needs, and improves the short text retrieval experience.

[0010] Further, calculating the ordered compactness score of the input text and the corresponding short database text based on the word position sequence information, including:

[0011] Comparing each word of the input text with the short database text based on the word position sequence information to determine the position vector of each word of the input text relative to the database text;

[0012] Determining the order score of the input text relative to the database text according to each position vector, and calculating the ordered compactness score of the input text relative to the database text based on the order score and each position vector.

[0013] Further, determining the order score of the input text relative to the database text according to each position vector, including:

[0014] Determining the value of the order score of the input text relative to the database text according to whether there is a position vector with a specified value in each position vector, where the specified value is used to identify that there is no corresponding word for the current word of the input text in the database text, or the position order of the corresponding word is before the marked position of the previous word of the input text.

[0015] Further, determining the term frequency-inverse document frequency vector of the words co-existing in the input text and the corresponding short database text based on the term frequency-inverse document frequency model, including:

[0016] Determining the term frequency of each word in the input text that appears in the corresponding database text based on the term frequency-inverse document frequency model;

[0017] Calculating the inverse document frequency of the corresponding word in the input text according to the total number of short database texts and the number of short texts containing the corresponding word;

[0018] Calculating the term frequency-inverse document frequency vector of the words co-existing in the input text and the corresponding short database text according to the term frequency and inverse document frequency of each word.

[0019] Comparing the input text with the database text in the way of word order and word position matching to obtain a more accurate word position order matching result, thereby improving the short text retrieval accuracy, and the short text retrieval output result better meets the user's needs.

[0020] Further, determining the word matching degree score of the input text and the corresponding short database text, including:

[0021] Determining the term frequency-inverse document frequency vector of the words co-existing in the input text and the corresponding short database text based on the term frequency-inverse document frequency model;

[0022] Calculate the cosine similarity between the input text and the short text in the corresponding database based on the term frequency-inverse document frequency vector, and use the cosine similarity to represent the word matching degree score between the input text and the short text in the corresponding database.

[0023] Furthermore, determine the term frequency-inverse document frequency vector of the words that coexist in the input text and the short text in the corresponding database based on the term frequency-inverse document frequency model, including:

[0024] Determine the term frequency of each word in the input text that appears in the corresponding database text based on the term frequency-inverse document frequency model;

[0025] Calculate the inverse document frequency of the corresponding word in the input text according to the total number of short texts in the database and the number of short texts containing the corresponding word;

[0026] Calculate the term frequency-inverse document frequency vector of the words that coexist in the input text and the short text in the corresponding database according to the term frequency and inverse document frequency of each word.

[0027] Compare the input text with the database text by means of term frequency and inverse document frequency detection to obtain accurate word results, so that the short text retrieval meets the basic accuracy requirements and improves the relevance between the short text output result and the input text.

[0028] Furthermore, determine the target text output from the short texts in the database according to the word matching degree score and the ordered compactness score, including:

[0029] Determine a set number of candidate texts from the short texts in the database in descending order according to the word matching degree score;

[0030] Sort the candidate texts based on the word matching degree score and the ordered compactness score of each candidate text, and use the sorting result of the candidate texts as the retrieval target output.

[0031] By selecting candidate texts according to the word matching degree and then comprehensively sorting the candidate texts according to the word matching degree score and the ordered compactness score, the short text retrieval output result can intuitively reflect the relevance with the input text, which is convenient for users to select the required short text according to the sorting and improves the user's short text retrieval experience.

[0032] In a second aspect, an embodiment of the present application provides a short text retrieval device based on ordered compactness, including:

[0033] An input module, configured to obtain the input text to be retrieved and determine the word matching degree score between the input text and the short text in the corresponding database;

[0034] A calculation module, configured to calculate an ordered compactness score of the input text and the corresponding short text in the database based on the word position sequence information, where the ordered compactness score is used to characterize the matching degree of the word position sequence between the input text and the short text in the database;

[0035] An output module, configured to determine and output a retrieval target from the short texts in the database according to the word matching degree score and the ordered compactness score.

[0036] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0037] A memory and one or more processors;

[0038] The memory is configured to store one or more programs;

[0039] When the one or more programs are executed by the one or more processors, the one or more processors implement the short text retrieval method based on ordered compactness as described in the first aspect.

[0040] In a fourth aspect, an embodiment of the present application provides a storage medium containing computer-executable instructions, where the computer-executable instructions are used to execute the short text retrieval method based on ordered compactness as described in the first aspect when executed by a computer processor. Description of the Drawings

[0041] Figure 1 is a flowchart of a short text retrieval method based on ordered compactness provided in Embodiment 1 of the present application;

[0042] Figure 2 is a flowchart of calculating the word matching degree score in Embodiment 1 of the present application;

[0043] Figure 3 is a flowchart of calculating the ordered compactness score in Embodiment 1 of the present application;

[0044] Figure 4 is a flowchart of outputting the retrieval target in Embodiment 1 of the present application;

[0045] Figure 5 is a schematic structural diagram of a short text retrieval device based on ordered compactness provided in Embodiment 2 of the present application;

[0046] Figure 6 is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present application. Detailed Embodiments

[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following further describes specific embodiments of this application in detail with reference to the accompanying drawings. It can be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. Additionally, it should be noted that, for ease of description, only parts related to this application rather than all content are shown in the accompanying drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0048] The short text retrieval method based on ordered compactness provided by this application aims to improve the accuracy of short text retrieval, making the retrieved short text more in line with user needs.

[0049] Short text retrieval refers to the process of retrieving short texts that meet the user's information needs from a large collection of data with an irregular structure and short text lengths. Common application scenarios include: search association words in a retrieval system, retrieval-based chatbots, question matching in a question-and-answer system, retrieval of text entities such as product names or company names, and work order retrieval in a customer service management system, etc.

[0050] Currently, the technical solutions for short text retrieval are as follows: (1) The character-based hard matching scheme, which obtains the matching degree between the user input text and each short text in the database by manually writing character-level hard matching rules. Although this scheme is easy to implement, it is time-consuming and laborious because a large number of matching rules need to be manually written, and it does not utilize other information of short texts, such as word frequency information, semantic information, etc.; (2) The short text matching scheme based on machine learning, which trains a text matching model based on a machine learning algorithm through labeled data. The trained matching model can judge the matching degree between the user input text and each short text in the database. However, for multiple words that match between the text to be retrieved and the associated short text, the positions where they appear in the short text may span too large, resulting in the retrieved short text having a small semantic association degree although it matches the text to be retrieved in terms of words. Simply determining the associated short text based on the word matching degree has retrieval errors, making it difficult for the retrieved short text to meet user needs. Moreover, this scheme requires a large amount of labeled data to train the model and has high requirements for the quality of data annotation.

[0051] Based on this, an ordered-compactness-based short text retrieval method according to an embodiment of the present application is provided to solve the error problem of retrieving short texts according to the word matching degree. On the basis of word matching, by calculating the matching degree of the word position order between the input text and the database text, it is ensured that the retrieval target and the word position order of the input text meet the relevance requirements, reducing the short text retrieval error, outputting a more accurate retrieval target, making the short text retrieval output result more in line with user needs, and improving the short text retrieval experience.

[0052] Embodiment 1:

[0053] Figure 1 A flowchart of an ordered-compactness-based short text retrieval method provided by Embodiment 1 of the present application is given. The ordered-compactness-based short text retrieval method provided in this embodiment can be executed by an ordered-compactness-based short text retrieval device. The ordered-compactness-based short text retrieval device can be implemented in software and / or hardware. The ordered-compactness-based short text retrieval device can be composed of two or more physical entities, or can be composed of one physical entity. Generally speaking, the ordered-compactness-based short text retrieval device can be a processing device such as a computer, a mobile phone, a tablet, a text retrieval server, etc.

[0054] The following takes this short text retrieval device as the main body for executing the ordered-compactness-based short text retrieval method for description. Refer to Figure 1 , the ordered-compactness-based short text retrieval method specifically includes:

[0055] S110. Obtain the input text to be retrieved, and determine the word matching degree score between the input text and the corresponding database short text.

[0056] When performing short text retrieval in the embodiment of the present application, the short text retrieval result is determined by comprehensively considering the word matching degree score between the input text and the corresponding database short text. Among them, the word matching degree score is used to represent the matching degree between the words of the input text and the words of the database short text. It can be understood that. The more the number of words in the database short text that are the same as those in the input text, the higher the degree of association between the two in terms of words.

[0057] Exemplarily, when performing short text retrieval, for the text of the short text "construction equipment" input by the user, it is defined as the input text. By comparing the words of this input text with the short texts stored in the database, the corresponding word matching degree score can be determined.

[0058] Among them, the calculation of the word matching degree can be realized by a pre-trained text matching model based on a machine learning algorithm. Based on the word matching degree calculation ability obtained from model training, the text matching model compares the input text with the short texts in the pre-constructed database, and then outputs the corresponding word matching degree score.

[0059] Optionally, referring to Figure 2 , determining the word matching degree score between the input text and the corresponding short text in the database includes:

[0060] S1101. Determining the term frequency-inverse document frequency vector of the words that coexist in the input text and the corresponding short text in the database based on the term frequency-inverse document frequency model;

[0061] S1102. Calculating the cosine similarity between the input text and the corresponding short text in the database based on the term frequency-inverse document frequency vector, and using the cosine similarity to represent the word matching degree score between the input text and the corresponding short text in the database.

[0062] Different from the way of text matching by traditional machine learning models, this application also calculates the word matching degree through the TFIDF model. Among them, TFIDF (term frequency–inverse document frequency) is a commonly used weighting technique for information retrieval and data mining. The main idea of TF-IDF is that if a certain word or phrase has a high frequency TF in an article and rarely appears in other articles, it is considered that this word or phrase has good category discrimination ability and is suitable for classification. TFIDF is actually TF * IDF, where TF (Term Frequency) is the frequency of a term in document d, and IDF (Inverse Document Frequency) is the inverse document frequency. The main idea of IDF is that if the number of documents containing term t is less, that is, n is smaller, and the IDF is larger, it means that term t has good category discrimination ability. If the number of documents containing term t in a certain category of documents C is m, and the total number of documents containing t in other categories is k, obviously the total number of documents containing t, n = m + k. When m is large, n is also large, and the IDF value obtained according to the IDF formula will be small, indicating that the category discrimination ability of term t is not strong. However, in fact, if a term appears frequently in the documents of a certain category, it means that this term can well represent the characteristics of the text of this category. Such terms should be given higher weights and selected as the feature words of this category of text to distinguish them from the documents of other categories. This is the shortcoming of IDF. In a given document, the term frequency (TF) refers to the frequency of a given word in the document. This number is a normalization of the term count to prevent it from biasing towards long documents.

[0063] Based on the above term frequency–inverse document frequency model, by comparing the input text with the corresponding database text one by one to determine the term frequency–inverse document frequency vector, and then the corresponding cosine similarity can be calculated according to this term frequency–inverse document frequency vector, and the cosine similarity is used to represent the word matching degree score between the input text and the corresponding database short text.

[0064] Among them, determining the term frequency–inverse document frequency vector of the words that coexist in the input text and the corresponding database short text based on the term frequency–inverse document frequency model includes:

[0065] Determining the term frequency of each word in the input text that appears in the corresponding database text based on the term frequency–inverse document frequency model;

[0066] Calculating the inverse document frequency of the corresponding words in the input text according to the total number of database short texts and the number of short texts containing the corresponding words;

[0067] Calculate the TF-IDF vector of the words that coexist in the input text and the short text in the corresponding database according to the word frequency and inverse document frequency of each word.

[0068] Exemplarily, when calculating the word matching degree score, assume that the user input text is Q. After segmenting Q by the Chinese word segmentation tool, we can get Q = [q1,..., q N , where N is the number of words in the user input text. Assume that the i-th short text in the database is D i , then after Chinese word segmentation, we can get M i is the number of words in the i-th short text. It should be noted that there are many ways to split words based on short texts. Each word after segmentation can be a single character or composed of multiple characters. The embodiments of this application do not make a fixed limit on the specific word segmentation method and will not elaborate here one by one.

[0069] Before that, build a TF-IDF model based on the short texts in the database. For the i-th short text D i in the database, first remove the duplicates of all its characters, that is, get the word set i of D M′ i is the length of the word set of the i-th short text. Then count the frequency of the j-th word i appearing in the short text D

[0070]

[0071] T1 represents the number of times the j-th word appears in the short text D i . Then count the inverse document frequency of the j-th word

[0072]

[0073] where α represents the total number of short texts in the database, and β represents the number of short texts containing plus 1.

[0074] Finally, the TF-IDF vector of the i-th short text D i can be obtained:

[0075]

[0076] For the user input text Q, first remove the duplicates of all its characters, that is, get the word set {q′1,..., q′ N′}, where N′ is the length of the set of words in the user input text. Then, count the frequency (TF j in the user input text Q) of the j-th word q′ j :

[0077]

[0078] Let T2 denote the number of times the j-th word q′ j appears in the user input text Q. Then, count the inverse document frequency (IDF j of the j-th word q′ j ):

[0079]

[0080] where λ represents the number of short texts containing q′ j plus 1.

[0081] Finally, the TFIDF vector (term frequency-inverse document frequency vector) of the user input text Q can be obtained:

[0082] TFIDF(Q) = [TF1 × IDF1,..., TF N′ × IDF N′

[0083] Furthermore, based on this TFIDF vector, calculate the cosine similarity between the TFIDF vectors of the user input text Q and the i-th short text D i :

[0084]

[0085] represents matrix multiplication calculation based on the TFIDF(Q) and TFIDF(D i for the word pairs that coexist in Q and D i ). The larger the value of cos(Q, D i ), the higher the matching degree between the user input text Q and the i-th short text D i .

[0086] ​Thus, by calculating the cosine similarity between the input text and the corresponding database text, the cosine similarity can be used as the word matching degree score between the two. It should be noted that in the calculation of cosine similarity based on TFIDF, if the TF values of the relevant words in the candidate short text are equal to those of the relevant words in the user input text, the more the number of irrelevant words, the smaller the cosine similarity and the lower the matching degree; the smaller the TFIDF value of the irrelevant words, that is, the more common the irrelevant words in the candidate short text, the greater the cosine similarity and the higher the matching degree. Among them, the irrelevant words are the words that exist in the candidate short text but do not exist in the user input text, and the relevant words are the words that exist in both the candidate short text and the user input text.

[0087] Thus, by comparing the input text with the database text through the method of term frequency and inverse document frequency detection, accurate word results can be obtained, enabling the short text retrieval to meet the basic accuracy requirements and improving the relevance between the short text output result and the input text.

[0088] S120. Calculate the ordered compactness score of the input text and the corresponding database short text based on the word position order information. The ordered compactness score is used to represent the word position order matching degree between the input text and the database short text.

[0089] On the other hand, the embodiment of the present application also calculates the ordered compactness score in combination with the word position order matching degree between the input text and the database text to comprehensively determine the short text retrieval result. It should be noted that assuming that a short sentence in the user input text sequentially includes four words A, B, C, and D, in the database short text, a corresponding short sentence also sequentially includes four words A, B, C, and D, then the word position order matching degree between the two short sentences is relatively high. On the contrary, if in the database short text, the order of the four words A, B, C, and D is reversed, or they are distributed in different sentences, then the word position order matching degree between the two is relatively low. Therefore, the embodiment of the present application further performs the calculation of word position order matching on the basis of word matching to obtain a database short text with a higher degree of relevance.

[0090] Among them, the ordered compactness score can be calculated by comparing each short sentence one by one, and according to the word position order matching situation in each short sentence, calculating the word position order matching degree score of each short sentence between the input text and the database text. If the number of words with matching position order in the two short sentences compared between the input text and the database text is more, the word position order matching degree score is higher. The word position order matching degree score can be determined according to the proportion of the number of words with matching position order to the total number of words in the short sentence. Then, the average value of the word position order matching degree scores of all short sentences is obtained as the ordered compactness score.

[0091] Optionally, refer to Figure 3, calculating the ordered compactness score of the input text and the corresponding database short text based on the word position sequence information, including:

[0092] S1201. Compare each word of the input text with the database short text based on the word position sequence information to determine the position vector of each word of the input text relative to the database text;

[0093] S1202. Determine the order score of the input text relative to the database text according to each position vector, and calculate the ordered compactness score of the input text relative to the database text based on the order score and each position vector.

[0094] Since the TFIDF model only utilizes the word frequency information of the database short text and the user input text, without considering the order of the words in the user input text in the database short text and the positional compactness degree between the words. Therefore, the ordered compactness score is introduced in the embodiments of the present application for comprehensive short text retrieval.

[0095] When calculating the ordered compactness score, the embodiments of the present application respectively compare the word positions and orders of the input text and the database text, and determine the position deviation of each word of the input text relative to the corresponding word of the database short text one by one, which is defined as the position vector. And according to this part of the position vector representing the word position deviation, determine the sorting situation of the order of each word of the input text relative to the corresponding word order of the database short text, which is defined as the order score. Based on this order score and the position vector, the corresponding ordered compactness score can be calculated.

[0096] Among them, determining the order score of the input text relative to the database text according to each position vector includes:

[0097] Determine the value of the order score of the input text relative to the database text according to whether there is a position vector with a specified value in each position vector. The specified value is used to identify that the current word of the input text does not have a corresponding word in the database text, or the position order of the corresponding word is before the marked position of the previous word of the input text.

[0098] Based on the user input text Q, gradually mark the position P where each word of Q appears in the corresponding database short text j is the position of the jth word q j in the database short text. If the jth word q j of the user input text Q is in the database short text before the position of the (j - 1)th word q j-1 , then P j = -1. If the jth word q j of the user input text Q does not exist in the candidate short text, then P j= -1. If the j-th word q of the user input text Q j has been marked by other identical words at its position in the candidate short text, continue to search backward for an unmarked position. In other cases, if the current word has a corresponding word in the database short text, use the position serial number of the corresponding word as the value of P j . Finally, the position vector P of the user input text Q in the candidate short text can be obtained as P = [P1,..., P N . For example, assume the input text Q = ['a', 'c', 'b', 'c', 'a'], and the database short text D = ['a', 'b', 'c', 'c'], then P = [1, 3, -1, 4, -1].

[0099] The formula for calculating the sequential score is as follows:

[0100]

[0101] The sequential score isOrder indicates whether each word of the user input text Q is in order or exists in the candidate short text. If there is an element value of -1 in the position vector P, it means that there is a situation of disordered order or the absence of a certain word in the user input text Q in the candidate short text.

[0102] Furthermore, based on the term frequency-inverse document frequency model, determine the term frequency-inverse document frequency vector of the words co-existing in the input text and the corresponding database short text, including:

[0103] Based on the term frequency-inverse document frequency model, determine the term frequency of each word in the input text that appears in the corresponding database text;

[0104] According to the total number of database short texts and the number of short texts containing the corresponding word, calculate the inverse document frequency of the corresponding word in the input text;

[0105] According to the term frequency and inverse document frequency of each word, calculate the term frequency-inverse document frequency vector of the words co-existing in the input text and the corresponding database short text.

[0106] Among them, the formula for calculating the ordered compactness score is as follows:

[0107]

[0108] The ordered compactness score represents the degree of compactness among each word of the user input text Q in the candidate short text. This score is calculated by computing the position differences between adjacent words, then taking the reciprocal of the mean value after taking the absolute value of all differences. Therefore, the smaller the compactness score, the more compact the user input text Q is in the candidate short text. If the user input text has only one word, the ordered compactness score S = 0. If the number of words in the user input text is greater than 1, the ordered compactness score S is the product of the sequential score isOrder and the compactness score.

[0109] By comparing the input text with the database text in the way of word order and word position matching, a more accurate word position order matching result can be obtained, thereby improving the short text retrieval accuracy and making the short text retrieval output result more in line with the user's needs.

[0110] S130. Determine and output the retrieval target from the database short texts according to the word matching degree score and the ordered compactness score.

[0111] Finally, based on the determined word matching degree score and ordered compactness score, the corresponding short text retrieval result output can be determined from the database short texts according to the actual retrieval requirements. Define this short text retrieval result as the retrieval target. It should be noted that according to the actual retrieval requirements, the retrieval target can be one database short text or multiple database short texts. In the case of outputting multiple database short texts, the word matching degree score and the ordered compactness score can also be used to sort them according to the corresponding sorting rules. According to the actual needs, the retrieval target can have various output forms, which are not fixedly restricted here.

[0112] For example, select a1 database short texts from the database in descending order of the word matching degree score, and select a2 database short texts from the database short texts in descending order of the ordered compactness score, and then select the overlapping database short texts from a1 and a2 as the retrieval target output. When outputting the retrieval target, it can be sorted in descending order according to the ordered compactness score, the word matching degree score, or the weighted score of both.

[0113] It can be seen that based on word matching, this application calculates the word position order matching degree between the input text and the database text to ensure that the word position order of the retrieval target meets the relevance requirements with the input text, reduces the short text retrieval error, outputs a more accurate retrieval target, makes the short text retrieval output result more in line with the user's needs, and improves the short text retrieval experience.

[0114] Optionally, determining and outputting the target text from the database short texts according to the word matching degree score and the ordered compactness score includes:

[0115] Determine a set number of candidate texts from the database short texts in descending order of the word matching degree scores;

[0116] Sort the candidate texts based on the word matching degree scores and the ordered compactness scores of each candidate text, and use the sorting result of the candidate texts as the retrieval target for output.

[0117] Refer to Figure 4 , when determining the retrieval target based on the ordered compactness score and the word matching degree score, first select K (set number) candidate texts from the database in descending order of the word matching degree scores. Then calculate the ordered compactness scores of the K candidate short texts retrieved roughly by the TF IDF model, and then add the ordered compactness scores of each candidate short text to the original word matching degree scores to obtain new sorting scores. Finally, re - sort the K candidate texts in descending order based on the new sorting scores, and output the retrieval target accordingly.

[0118] In this way, by selecting candidate texts according to the word matching degree, and then comprehensively sorting the candidate texts according to the word matching degree scores and the ordered compactness scores, the short - text retrieval output result can intuitively reflect the degree of association with the input text, facilitating the user to select the required short text according to the sorting, and improving the user's short - text retrieval experience.

[0119] Moreover, compared with the character - based hard - matching scheme, the embodiments of the present application do not require manual writing of character - level hard - matching rules. Only need to pre - train a TF IDF model based on the short texts in the database, and then the word matching degree scores between the user - input text and each short text in the database can be calculated; compared with the short - text matching scheme of machine learning, the embodiments of the present application do not require a large amount of labeled data, because the construction and calculation of the TF IDF model and the ordered compactness both belong to unsupervised algorithms and do not require the participation of high - quality labeled data, only the short - text data to be matched needs to be provided; compared with the traditional TF IDF retrieval algorithm that only considers word - frequency information, the embodiments of the present application introduce the ordered compactness score, adding sequential structure and compactness structure features from the perspective of the text structure between the user - input text and the short text. By re - sorting the TF IDF rough - recall results through the ordered compactness score, short - text information more in line with the user input can be recalled.

[0120] Embodiment 2:

[0121] Based on the above - mentioned embodiment, Figure 5 is a schematic structural diagram of a short - text retrieval device based on ordered compactness provided by the second embodiment of the present application. Refer to Figure 5 , the short - text retrieval device based on ordered compactness provided in this embodiment specifically includes: an input module 21, a calculation module 22, and an output module 23.

[0122] Among them, the input module 21 is used to obtain the input text to be retrieved and determine the word matching degree score between the input text and the short text in the corresponding database;

[0123] The calculation module 22 is used to calculate the ordered compactness score between the input text and the corresponding short text in the database based on the word position order information, and the ordered compactness score is used to characterize the matching degree of the word position order between the input text and the short text in the database;

[0124] The output module 23 is used to determine and output the retrieval target from the short text in the database according to the word matching degree score and the ordered compactness score.

[0125] It can be seen that based on word matching, by calculating the matching degree of the word position order between the input text and the database text, this application ensures that the word position order of the retrieval target and the input text meets the relevance requirements, reduces the short text retrieval error, outputs a more accurate retrieval target, makes the short text retrieval output result more in line with the user's needs, and improves the short text retrieval experience.

[0126] Specifically, calculating the ordered compactness score between the input text and the corresponding short text in the database based on the word position order information includes:

[0127] Comparing each word of the input text with the short text in the database based on the word position order information to determine the position vector of each word of the input text relative to the database text;

[0128] Determine the order score of the input text relative to the database text according to each position vector, and calculate the ordered compactness score of the input text relative to the database text based on the order score and each position vector.

[0129] Further, determining the order score of the input text relative to the database text according to each position vector includes:

[0130] Determine the value of the order score of the input text relative to the database text according to whether there is a position vector with a specified value in each position vector. The specified value is used to indicate that there is no corresponding word for the current word of the input text in the database text, or the position order of the corresponding word is before the marked position of the previous word of the input text.

[0131] Specifically, determining the inverse document frequency vector of the words co-existing in the input text and the corresponding short text in the database based on the term frequency-inverse document frequency model includes:

[0132] Determine the term frequency of each word in the input text that appears in the corresponding database text based on the term frequency-inverse document frequency model;

[0133] Calculate the inverse document frequency of the corresponding words in the input text based on the total number of short texts in the database and the number of short texts containing the corresponding words.

[0134] Calculate the term frequency-inverse document frequency vector of the words that coexist in the input text and the corresponding short texts in the database based on the term frequency and inverse document frequency of each word.

[0135] Compare the input text with the database text by matching the word order and word positions to obtain a more accurate matching result of the word position order, thereby improving the short text retrieval accuracy and making the short text retrieval output result more in line with the user's needs.

[0136] Specifically, determine the word matching degree score between the input text and the corresponding short texts in the database, including:

[0137] Based on the term frequency-inverse document frequency model, determine the term frequency-inverse document frequency vector of the words that coexist in the input text and the corresponding short texts in the database.

[0138] Calculate the cosine similarity between the input text and the corresponding short texts in the database based on the term frequency-inverse document frequency vector, and use the cosine similarity to represent the word matching degree score between the input text and the corresponding short texts in the database.

[0139] Specifically, based on the term frequency-inverse document frequency model, determine the term frequency-inverse document frequency vector of the words that coexist in the input text and the corresponding short texts in the database, including:

[0140] Based on the term frequency-inverse document frequency model, determine the term frequency of each word in the input text that appears in the corresponding database text.

[0141] Calculate the inverse document frequency of the corresponding words in the input text based on the total number of short texts in the database and the number of short texts containing the corresponding words.

[0142] Calculate the term frequency-inverse document frequency vector of the words that coexist in the input text and the corresponding short texts in the database based on the term frequency and inverse document frequency of each word.

[0143] Compare the input text with the database text by means of term frequency and inverse document frequency detection to obtain accurate word results, so that the short text retrieval meets the basic accuracy requirements and improves the relevance between the short text output result and the input text.

[0144] Specifically, determine the target text output from the short texts in the database according to the word matching degree score and the ordered compactness score, including:

[0145] Determine a set number of candidate texts from the short texts in the database in descending order of the word matching degree score.

[0146] Sort the candidate texts based on the word matching degree scores and ordered compactness scores of each candidate text, and use the sorting result of the candidate texts as the retrieval target output.

[0147] Select candidate texts according to the word matching degree, and then comprehensively sort the candidate texts according to the word matching degree scores and ordered compactness scores, so that the short text retrieval output result can intuitively reflect the degree of association with the input text, facilitating the user to select the required short text according to the sorting and improving the user's short text retrieval experience.

[0148] The short text retrieval device based on ordered compactness provided in the second embodiment of the present application can be used to execute the short text retrieval method based on ordered compactness provided in the first embodiment, and has corresponding functions and beneficial effects.

[0149] Embodiment Three:

[0150] The third embodiment of the present application provides an electronic device. Referring to Figure 6 , the electronic device includes: a processor 31, a memory 32, a communication module 33, an input device 34, and an output device 35. The number of processors in the electronic device can be one or more, and the number of memories in the electronic device can be one or more. The processor, memory, communication module, input device, and output device of the electronic device can be connected through a bus or other means.

[0151] As a computer-readable storage medium, the memory can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the short text retrieval method based on ordered compactness described in any embodiment of the present application (for example, the input module, calculation module, and output module in the short text retrieval device based on ordered compactness). The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the device. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory can further include a memory remotely provided relative to the processor, and these remote memories can be connected to the device through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise internal network, a local area network, a mobile communication network, and combinations thereof.

[0152] The communication module is used for data transmission.

[0153] The processor executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory, that is, implements the above-mentioned short text retrieval method based on ordered compactness.

[0154] The input device can be used to receive input digital or character information and generate key signal inputs related to the user settings and function control of the device. The output device can include display devices such as a display screen.

[0155] The electronic device provided above can be used to execute the short text retrieval method based on ordered compactness provided in the first embodiment above, and has corresponding functions and beneficial effects.

[0156] Embodiment 4:

[0157] The embodiment of the present application also provides a storage medium containing computer-executable instructions. When the computer-executable instructions are executed by a computer processor, they are used to execute a short text retrieval method based on ordered compactness. The short text retrieval method based on ordered compactness includes: obtaining the input text to be retrieved, and determining the word matching degree score between the input text and the short text in the corresponding database; calculating the ordered compactness score between the input text and the corresponding database short text based on the word position sequence information, and the ordered compactness score is used to represent the word position sequence matching degree between the input text and the database short text; determining the retrieval target from the database short text according to the word matching degree score and the ordered compactness score and outputting

[0158] Storage medium - any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media such as CD-ROM, floppy disk or magnetic tape devices; computer system memory or random access memory such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory such as flash memory, magnetic media (such as hard disk or optical storage); registers or other similar types of memory elements, etc. The storage medium can also include other types of memory or combinations thereof. Additionally, the storage medium can be located in the first computer system in which the program is executed, or can be located in a different second computer system, and the second computer system is connected to the first computer system through a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term "storage medium" can include two or more storage media residing in different locations (such as in different computer systems connected through a network). The storage medium can store program instructions (such as specifically implemented as a computer program) executable by one or more processors.

[0159] Of course, for a storage medium containing computer-executable instructions provided by the embodiment of the present application, the computer-executable instructions are not limited to the short text retrieval method based on ordered compactness as described above, and can also execute related operations in the short text retrieval method based on ordered compactness provided in any embodiment of the present application.

[0160] The short text retrieval device, storage medium, and electronic device provided in the above embodiments can execute the short text retrieval method based on ordered compactness provided in any embodiment of the present application. For technical details not described in detail in the above embodiments, reference can be made to the short text retrieval method based on ordered compactness provided in any embodiment of the present application.

[0161] The above is only the preferred embodiment of the present application and the technical principles applied. The present application is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions that can be made by those skilled in the art will not depart from the protection scope of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments. Without departing from the concept of the present application, it may also include more other equivalent embodiments, and the scope of the present application is determined by the scope of the claims.

Claims

1. A short text retrieval method based on ordered compactness, characterized in that, Including: Obtain the input text to be retrieved, and determine the word matching degree score between the input text and the short text in the corresponding database; Calculate the ordered compactness score between the input text and the corresponding short text in the database based on the word position order information, and the ordered compactness score is used to characterize the matching degree of the word position order between the input text and the short text in the database; Determine the retrieval target from the short text in the database according to the word matching degree score and the ordered compactness score, and output it.

2. The short text retrieval method based on ordered compactness according to claim 1, wherein The calculating the ordered compactness score between the input text and the corresponding short text in the database based on the word position order information includes: Based on the word position order information, compare each word of the input text with the short text in the database, and determine the position vector of each word of the input text relative to the database text; Determine the order score of the input text relative to the database text according to each position vector, and calculate the ordered compactness score of the input text relative to the database text based on the order score and each position vector.

3. The short text retrieval method based on ordered compactness according to claim 2, characterized in that, The determining the order score of the input text relative to the database text according to each position vector includes: Determine the value of the order score of the input text relative to the database text according to whether there is a position vector with a specified value in each position vector, and the specified value is used to indicate that there is no corresponding word for the current word of the input text in the database text, or the position order of the corresponding word is before the marked position of the previous word of the input text.

4. The short text retrieval method based on ordered compactness according to claim 2, wherein The calculating the ordered compactness score of the input text relative to the database text based on the order score and each position vector includes: Determine the vector difference between the two adjacent position vectors, and calculate the ordered compactness score of the input text relative to the database text according to the number of position vectors, the vector difference and the order score.

5. The short text retrieval method based on ordered compactness according to claim 1, characterized in that The determining the word matching degree score between the input text and the corresponding short text in the database includes: Determine the term frequency-inverse document frequency vector of the words co-existing in the input text and the corresponding short text in the database based on the term frequency-inverse document frequency model; Calculate the cosine similarity between the input text and the corresponding short text in the database based on the term frequency-inverse document frequency vector, and use the cosine similarity to characterize the word matching degree score between the input text and the corresponding short text in the database.

6. The short text retrieval method based on ordered compactness according to claim 5, wherein The determining the term frequency-inverse document frequency vector of the words co-existing in the input text and the corresponding short text in the database based on the term frequency-inverse document frequency model includes: Determine the term frequency of each word in the input text that appears in the corresponding database text based on the term frequency-inverse document frequency model; Calculate the inverse document frequency of the corresponding word in the input text according to the total number of short texts in the database and the number of short texts containing the corresponding word; Calculate the term frequency-inverse document frequency vector of the words co-existing in the input text and the corresponding short text in the database according to the term frequency and the inverse document frequency of each word.

7. The short text retrieval method based on ordered compactness according to claim 1, characterized in that Determining and outputting a target text from the short database texts according to the word matching degree score and the ordered compactness score includes: Determining a set number of candidate texts from the short database texts in descending order of the word matching degree score; Sorting the candidate texts based on the word matching degree score and the ordered compactness score of each candidate text, and using the sorting result of the candidate texts as the retrieval target output.

8. A short text retrieval device based on ordered compactness, characterized in that, Including: An input module, configured to obtain an input text to be retrieved and determine the word matching degree score between the input text and the corresponding short database text; A calculation module, configured to calculate the ordered compactness score between the input text and the corresponding short database text based on the word position order information, where the ordered compactness score is used to characterize the matching degree of the word position order between the input text and the short database text; An output module, configured to determine and output a retrieval target from the short database texts according to the word matching degree score and the ordered compactness score.

9. An electronic device, characterized in that, Including: A memory and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the short text retrieval method based on ordered compactness according to any one of claims 1-7.

10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to execute the short text retrieval method based on ordered compactness according to any one of claims 1-7.