Retrieval method and device based on term set synonymous expansion and responsiveness keyword groups

Through the search method based on term set synonym expansion and responsive keyword groups, the problem of large uncertainty in model training and low search accuracy in existing text search technology is solved, and efficient and accurate text search results are achieved.

CN119938915APending Publication Date: 2025-05-06CHENGDU HANLAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510024383.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing text retrieval technology has the problem of high uncertainty in model training and low retrieval accuracy, especially in asymmetric search tasks, which are difficult to provide high-quality results.

Method used

The search method based on the synonym expansion and responsive keyword group of term sets is adopted. By obtaining the original query sentence, synonym expansion dictionary and responsive keyword dictionary, keyword extraction, synonym expansion and responsive keyword expansion are carried out to narrow the scope of search results and improve the search accuracy.

Benefits of technology

It effectively improves the accuracy and accuracy of text retrieval, reduces the uncertainty and maintenance costs of model training, and is suitable for asymmetric retrieval tasks in the field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938915A_ABST
    Figure CN119938915A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to a term set synonymous expansion and responsiveness keyword group-based retrieval method and device. According to the method, a technical term set, a synonymous expansion dictionary and long-term effective auxiliary retrieval data of responsive keywords are used, and reusable expert knowledge in the field is efficiently fixed in a low-dimension form. According to the key keyword extraction and synonymous expansion based on the professional term set, on one hand, retrieval results irrelevant to professional themes are eliminated, the retrieval accuracy is improved, and on the other hand, the use of the synonymous expansion avoids possible content omission. And the responsiveness keywords establish a bridge between the question and the target query result, part of asymmetric retrieval questions are converted into symmetric retrieval questions, more retrieval clues are provided while the retrieval difficulty is reduced, the preliminary retrieval result is further corrected, and finally a small number of retrieval results with high accuracy are obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a retrieval method and device based on synonymous expansion of a term set and responsive keyword groups. Background Art

[0002] In text retrieval, retrieval tasks can be divided into two categories: symmetric and asymmetric. Symmetric retrieval is used to search for sentences of the same or similar length in the corpus. The order of the two sentences can be interchanged, for example, to find two sentences with similar meanings. Asymmetric retrieval uses a short query statement to search for a longer paragraph to answer the question. For asymmetric retrieval, the positions of the query sequence and the queried sequence cannot be interchanged. Symmetric retrieval tasks can be understood as finding synonyms or synonymous expressions. The two paragraphs used for comparison are relatively close in length and form. In asymmetric retrieval, the short query statement - usually a question asked by a person - is always the initiator of the retrieval, and the target query text is usually longer than the query statement. In terms of wording, the query statement is usually a question, and the target statement is usually a statement.

[0003] In certain professional fields, there are a large number of internal professional texts, and the various sub-fields involved in these texts have strong professional attributes. The professional scope involved in different business positions in each sub-field is small, but the knowledge depth is deep and the volume is large. A large number of professional texts will naturally lead to frequent query and search operations, and give rise to a strong demand for high retrieval accuracy.

[0004] In actual text queries, asymmetric retrieval methods, that is, the "question-answer" method, are more often used. For queries in the public domain, due to the huge volume of text, the text form and writing style vary greatly, so it is difficult to use a single solution to obtain a high accuracy rate for queries in the development field. Therefore, query tools such as search engines need to return a large number of results and require secondary manual screening and click jumps to obtain sufficient query coverage and hit rates. Text content in different fields usually has some commonalities that do not shift due to the diversification of professional content. First, the text in a specific field is basically unified in writing style, such as format, wording, and idioms; second, a field usually has a set of discourse systems that conform to its internal logic, such as limited professional terms and structured expressions that conform to specific businesses. These commonalities provide a common handle for improving the accuracy of text retrieval in a field.

[0005] In actual search applications, symmetric search usually has a higher accuracy rate than asymmetric search, firstly because the query parties matched by symmetric search have similar number of query keywords, and secondly, they usually use the same tone (statement, question, etc.). However, for text queries within a field, asymmetric search is a more common form, because the queryer usually does not know the answer before searching, and therefore cannot provide enough keywords to be included in the answer.

[0006] Currently, the commonly used retrieval and search technologies mainly include various retrieval algorithms such as TF-IDF related algorithms, various semantic search models based on machine learning such as sentence-bert, and generative language models such as the GPT series of models. The most commonly used application forms are search engines such as Google and Baidu, and conversational chatbots such as Chat-GPT.

[0007] Search engine applications usually combine retrieval algorithms (such as TF-IDF) and language processing models. Their advantage lies in the breadth of queries, but their disadvantage is that the range of results is too wide. The number of results returned by ordinary queries is generally in the tens of thousands or even millions, which cannot provide a small number of high-quality results for specific fields, especially for asymmetric retrieval tasks. In the symmetric retrieval of search engines, there is a special complete string matching method, which uses a specified syntax to specify the string that needs to be fully matched. For example, in Google's search service, if the query string is wrapped in double quotes, only the result web pages that completely contain the string will be returned. If a sentence in the target query text is known in advance, this method can greatly narrow the scope of the query results, but this situation is not common in actual queries. First, it is almost impossible for the questioner to know a complete sentence contained in the answer, and second, if there is any deviation in the complete match clue provided, it is likely to return 0 results. Therefore, the queryer can usually only ask questions for asymmetric retrieval, and in many cases the queryer cannot even be sure whether the question has an answer. In addition, since search engines need to process massive amounts of content in different fields, they need to introduce a lot of complex technologies and their maintenance and operation costs are also huge. For queries within a field, the search engine is too large, which will cause a lot of waste of resources and cannot provide sufficient accuracy.

[0008] Generative language model, chat tools based on generative language model such as OpenAI's Chat-GPT, many people mistakenly regard it as another search engine because they do not understand its essence, and because it only gives one answer to each question, many inquirers have a great dependence on it. In essence, what the generative language model does is not search but "generate". It only counts the probability of different words appearing after a word sequence (sentence) in a large amount of text, and based on this probability, it continuously extends the word sequence to obtain a "smooth answer". But in fact, the generative language model does not ensure that the answer given is consistent with facts or logic. Therefore, for searches within the field, the inability to guarantee the authenticity of the answer violates the basic requirements of text retrieval within the field and is not applicable.

[0009] Other language models: Some search engines and other search services use language models to increase search accuracy, but there are several problems with using these models: (1) The construction of model training sets requires a lot of manpower, financial resources and time investment. However, the text content terms in some professional fields are commercial secrets or other content that cannot be made public. Therefore, it is impossible to use outsourcing to construct training sets. Even if it can be outsourced, the quality of the constructed training sets cannot be guaranteed because the external personnel do not have professional knowledge in the field. (2) A sufficient number and quality of training sets can only increase the probability of training a good model. The final model training effect is affected by many variables. (3) As the amount of text increases, the effect of the original model is likely to decline, because the text features fitted by the model contain information dimensions related to the specific content. When the text content changes, some of the original effective features may no longer be effective.

[0010] Therefore, in general, the cost of initial model training and subsequent maintenance is high, the uncertainty is large, and there is a risk of failure in the long run, resulting in a low retrieval accuracy. Summary of the invention

[0011] In view of this, the purpose of the present invention is to provide a retrieval method and device based on synonymous expansion of term sets and responsive keyword groups, so as to at least solve the problems of large model training uncertainty and low retrieval accuracy in existing text retrieval.

[0012] The present invention solves the above technical problems by the following technical means:

[0013] In a first aspect, an embodiment of the present invention provides a retrieval method based on synonymous expansion of a term set and responsive keyword groups, comprising the following steps:

[0014] Obtaining an original query statement to be searched, and obtaining a synonym expansion dictionary and a responsive key dictionary associated with the original query statement domain;

[0015] Extract keywords from the original query sentence based on the professional term set to obtain a second word sequence Qi containing at least one keyword;

[0016] Use a search algorithm to score and sort the text segments containing any keyword in the second word sequence to obtain a first text segment sequence Si;

[0017] Based on the query task data F, the first F text segments with the highest scores are selected from the first text sequence as the second text sequence Sf;

[0018] Based on the synonym expansion dictionary and the responsive key dictionary, perform keyword synonym expansion and responsive keyword expansion on the second word sequence to obtain the third word sequence Qix;

[0019] Remove the text segments in the first text segment sequence that do not contain any keyword in the third word sequence, and use the remaining text segments as the third text segment sequence;

[0020] According to the number of text segments in the third text segment sequence, the text segment most relevant to the original query statement is retrieved.

[0021] In some implementations, extracting keywords from the original query statement based on the professional term set to obtain a second word sequence Qi containing at least one keyword includes:

[0022] Perform a word segmentation operation on the original query statement with reference to the professional term set to obtain a first word sequence containing at least two word segments;

[0023] Associatively matching each word segment in the first word sequence with the candidate text segments to obtain the number of text segments corresponding to each word segment in the first word sequence;

[0024] The stop words in the first word sequence are removed to retain the keywords, and the second word sequence Qi is obtained.

[0025] In some implementations, the use of a search algorithm to score and sort text segments containing any keyword in the second word sequence to obtain a first text segment sequence Si includes:

[0026] Among all the candidate text segments, collect and summarize the text segments containing any keyword in the second word sequence;

[0027] Using the search algorithm to score each text segment containing any of the keywords in the second word sequence;

[0028] The related text segments are sorted according to the scores obtained to obtain the first text segment sequence Si.

[0029] In some implementations, the method of selecting the first F text segments with the highest scores from the first text segment sequence based on the query task data F as the second text segment sequence Sf includes:

[0030] Setting the final result data F based on the query task requirement, wherein the query task is retrieval accuracy or retrieval breadth;

[0031] In the first text segment sequence, the first F text segments with the highest scores are retained as the second text segment sequence.

[0032] In some embodiments, the responsive keywords include general responsive keywords and domain-limited responsive keywords. General responsive keywords refer to corresponding relationships that are completely independent of any professional knowledge content; domain-limited responsive keywords refer to responses to the same type of questions in a specific professional field that explicitly or implicitly follow a fixed word structure.

[0033] In some implementations, if the number of text segments in the third text segment sequence is 0, the text segment most relevant to the original query statement is the text segment in the second text segment sequence;

[0034] If the number of text segments in the third text segment sequence is less than or equal to F, the text segment most relevant to the original query statement is the text segment in the third text segment sequence;

[0035] If the number of text fragments in the third text sequence is greater than F, the matching conditions of the responsive keywords are converted into weight values ​​for auxiliary retrieval screening to obtain the text fragments most relevant to the original query statement.

[0036] In some implementations, converting the matching conditions of the responsive keywords into weight values ​​for auxiliary search and screening to obtain the text fragments most relevant to the original query statement includes:

[0037] Get the number x of responsive keywords in the third word sequence;

[0038] Using a retrieval algorithm to score each text segment in the third text segment sequence, to obtain a retrieval score for each text segment in the third text segment sequence;

[0039] Multiply the search score of each text segment in the third text segment sequence by a weight value to obtain a final score corresponding to each text segment, wherein the weight value is related to x;

[0040] The text segments in the third text segment sequence are sorted according to the final scores, and the first F text segments with the highest scores are selected, that is, the text segments most relevant to the original query statement.

[0041] In some embodiments, the weight value is a base y raised to the power of x, and the base 1.5>y>1.0.

[0042] In the second aspect, the present invention also provides a retrieval device based on synonymous expansion of a term set and responsive keyword groups, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the retrieval method described in the first aspect above when executing the computer program.

[0043] In a third aspect, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the retrieval method described in the first aspect are implemented.

[0044] The retrieval method based on synonymous expansion of term sets and responsive keyword groups of the present invention is a complete and universal text retrieval method in the field based on professional term sets, synonymous expansion dictionaries and responsive key dictionaries. The present invention uses three kinds of auxiliary retrieval data, namely professional term sets, synonymous expansion dictionaries and responsive keywords, which are simple in form, low in construction and maintenance costs, and long-term effective, to efficiently fix the reusable expert knowledge in the field in a low-dimensional form. Compared with traditional retrieval algorithms, the thinking and response methods in the field fixed in these three auxiliary data effectively improve the accuracy of retrieval, and their low-dimensional characteristics provide sufficient versatility. For retrieval algorithms that focus on language models, the processing method provided by the present invention eliminates a large amount of repeated construction and updating of training sets, reduces the risk cost of bearing the uncertainty of model training effects, and greatly reduces the complexity of applications, increasing the comprehensibility of the implementation process.

[0045] The core keyword extraction and synonym expansion based on the professional terminology set used in the retrieval method based on synonym expansion of the terminology set and responsive keyword group of the present invention, on the one hand, excludes the retrieval results irrelevant to the professional subject, greatly narrows the result range, and increases the retrieval accuracy, and on the other hand, the use of synonym expansion avoids the possible omission of content. The responsive keywords build a bridge between the problem and the target query result, converting some asymmetric retrieval problems into symmetric retrieval problems, reducing the difficulty of retrieval while providing more retrieval clues, providing auxiliary retrieval information of another dimension other than the core content for the retrieval task, further correcting the preliminary retrieval results, and finally obtaining a small number of retrieval results with high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a flowchart of the search method based on synonymous expansion of term set and responsive keyword group;

[0047] Figure 2It is a schematic diagram of the influence of the search method of the present invention on the scope of search results;

[0048] Figure 3 It is a retrieval flow chart of the retrieval method based on synonymous expansion of term sets and responsive keyword groups. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0050] The terms "first" and "second" in the specification and claims herein are used to distinguish different objects rather than to describe a specific order of objects. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" refers to two or more, for example, multiple processing units refer to two or more processing units, etc., multiple elements refer to two or more elements, etc.

[0051] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0052] ChatGPT (Chat Generative Pre-trained Transformer), yes OpenAI A developed chat Sky Robot Program. ChatGPT is AI Technology-driven Natural Language Processing A tool that generates responses based on patterns and statistics seen during the pre-training phase, and can also generate responses based on the conversation Context ChatGPT is an intelligent tool that improves the human brain's ability to collect, organize, calculate, and analyze various information. It is a tool system that provides rich and accurate plans, diagrams, and other materials or conditions for the human brain's "concept construction".

[0053] Retrieval algorithm refers to the process and method of finding specific information in a data set. According to different application scenarios and requirements, retrieval algorithms can be divided into many types, the most common of which are linear search, binary search, interpolation search, Fibonacci search, block search, hash search and backtracking search. Linear search is the simplest search method. It starts from the first element and compares it in sequence until the target element is found or all elements are traversed. It is applicable to both unordered and ordered data sets. Binary search is applicable to ordered data sets. It can quickly locate the target element by continuously narrowing the search range. The premise is that the data has been sorted. Interpolation search is to dynamically adjust the search interval in an ordered data set according to the distribution of keywords. It is applicable to ordered data sets with uniform distribution. Fibonacci search uses the Fibonacci sequence for search. It is applicable to ordered data sets and can reduce the number of comparisons. Block search is to divide the data set into several blocks, first determine the block where the data is located, and then search within the block. Hash search is to map keywords to a position in the table through a hash function for search. It is suitable for fast search but needs to resolve conflicts. Backtracking search finds a solution that meets the conditions by step-by-step trial and backtracking, and is suitable for complex combinatorial optimization problems.

[0054] The present invention performs word segmentation on the query statement (question) in asymmetric retrieval, extracts core keywords based on a professional terminology set, expands core keywords based on a synonym dictionary, and finally expands the original one-dimensional word sequence of the query statement into a two-dimensional word sequence based on a responsive keyword dictionary summarized based on the terminology conventions and characteristics in the field. The new two-dimensional word sequence is then used for asymmetric retrieval tasks.

[0055] For details, please refer to Figure 1 , the retrieval method based on synonymous expansion of term set and responsive keyword group includes the following steps:

[0056] Step 100, obtaining the original query statement to be searched, and obtaining the synonym expansion dictionary and responsive key dictionary associated with the original query statement domain.

[0057] Step 200: extract keywords from the original query statement based on the professional term set to obtain a second word sequence Qi containing at least one keyword.

[0058] Step 300: Use a search algorithm to score and sort the text segments containing any keyword in the second word sequence to obtain a first text segment sequence Si.

[0059] Step 400 : selecting the first F text segments with the highest scores from the first text sequence based on the query task data F as the second text sequence Sf.

[0060] Step 500 , based on the synonym expansion dictionary and the responsive key dictionary, perform keyword synonym expansion and responsive keyword expansion on the second word sequence to obtain a third word sequence Qix.

[0061] Step 600: remove text segments in the first text segment sequence that do not contain any keyword in the third word sequence, and the remaining text segments are used as the third text segment sequence.

[0062] Step 700: retrieve the text segment most relevant to the original query statement according to the number of text segments in the third text segment sequence.

[0063] Through the above technical solution, the core keyword extraction and synonym expansion based on the professional terminology set are used to exclude search results that are not related to professional topics, greatly narrow the result range, and increase the search accuracy. On the other hand, the use of synonym expansion avoids possible content omissions. Responsive keywords build a bridge between the problem and the target query results, converting some asymmetric search problems into symmetric search problems, reducing the difficulty of search while providing more search clues, providing auxiliary search information of another dimension for the search task in addition to the core content, further correcting the preliminary search results, and finally obtaining a small number of search results with high accuracy.

[0064] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods. Figure 1-3 :

[0065] In step 100, the original query sentence to be retrieved is generally a sentence expressed in natural language, that is, a sentence expressed in the language people use in daily life, such as "What are the functions of the ABC system?".

[0066] The difference between the synonym expansion dictionary in this application and the traditional synonym dictionary is that the synonym expansion dictionary used in the embodiments of this application is based on a professional terminology set and does not consider the content in the traditional general synonym dictionary. The synonym expansion dictionary required for use in this application includes the following contents: (1) Chinese synonyms corresponding to each other, for example, the meaning of "ABC system" is the same as "ABC subsystem" and "ABC module"; (2) the corresponding relationship between English (abbreviation) and Chinese, for example, "ABC system" may be abbreviated as "ABCS" in its field.

[0067] Essentially, the use of a synonym expansion dictionary is to pay for the cost of inconsistent use of terms in texts within a field. In theory, if only one Chinese noun and English abbreviation are used for the same term in a text within a field, then the synonym expansion step can be omitted, that is, only a professional term set and possible Chinese-English comparison tables are required. The construction of a synonym expansion dictionary requires certain expert knowledge, but like a professional term set, once a high-quality and high-coverage synonym dictionary is built, the maintenance cost is also low, and it is a task that requires a one-time investment in the early stage and enjoys long-term benefits.

[0068] In the responsive key dictionary, there are general responsive keywords and domain-specific responsive keywords. The so-called responsive keywords refer to the words that are usually included in the corresponding answers to a type of question or question word. For example, when we ask "what causes XXX", the words that may be included in the answer to the word "cause" are "because", "cause", "due to", etc. When we ask "what is the function of XXX", the answer to the word "function" is likely to include "for", "purpose", "complete", etc. When the question is "what are XXX", the answer to the word "which" is likely to include "include", "include", "constitute", etc. This type of correspondence that does not rely on any professional knowledge content belongs to the "general" responsive keyword relationship. Its advantage lies in its universality. That is, no matter what field the text is in, in theory, as long as the text is written in a standardized manner and does not pursue simplicity too much, this correspondence must exist universally. The disadvantage is that the writing of the text does not require the inclusion of this correspondence. Different text writers have their own text styles, and this part of the factor is uncontrollable.

[0069] For certain specific fields, especially those with a long history, responses to the same type of questions usually follow a fixed set of word structures, either explicitly or implicitly. This is especially true for fields with a high degree of standardization, which are called field-specific response keywords. For example, in the fields of medicine and law, the answers to a certain type of question, regardless of the topic of the question, need to follow a relatively fixed logical framework, or there is a long-standing response process within the profession. For example, in the medical field, after a patient gives a preliminary description of his or her condition to a doctor, the doctor's response usually includes:

[0070] “Medical history”: analysis of other diseases the patient has had in the past or the impact of underlying diseases;

[0071] “Symptoms”: Provide analysis of symptoms for patients;

[0072] “Cause”: Analyze possible causes;

[0073] “Recommendations”: medication or other suggestions;

[0074] "Attention": Regarding medication and other subsequent precautions;

[0075] For a query statement "What are the treatment suggestions for XXX disease", for the term "treatment suggestions", it can be mapped to the above-mentioned responsive keywords.

[0076] Domain-specific responsive keyword dictionaries also require the participation of domain experts in construction, but they are also one-time upfront investments with long-term benefits. Essentially, it abstracts the common logical framework within a domain into responsive keywords grouped by unit. This logical framework itself is not related to the specific work content within the domain. It is a summary of the general thinking mode within the domain and a refinement of the key dimensional information of various problems within the domain.

[0077] In step 200, keyword extraction is performed on the original query statement based on the professional term set to obtain a second word sequence Qi containing at least one keyword, including:

[0078] Step 210, perform word segmentation on the original query statement with reference to the professional term set to obtain a first word sequence containing at least two word segments.

[0079] Step 220, perform associated matching between each word segment in the first word sequence and the alternative text fragments to obtain the number of text fragments corresponding to each word segment in the first word sequence;

[0080] Step 230, remove the stop words in the first word sequence and retain the keywords to obtain the second word sequence Qi.

[0081] Although the professional term set is simple in form, it can significantly narrow the scope of search results and improve the hit rate during the retrieval process. For example, there is a query question: "What are the functions of the ABC system?". The first step in the query is word segmentation. For the word segmentation engine, the query statement will probably be segmented into: "Excuse me, ABC, system, of, functions, have, which". Stop words refer to high-frequency words that have no substantial contribution to semantic parsing, such as "of", "le", "is". Removing stop words can improve the retrieval efficiency. For example, after removing stop words from "What are the functions of the ABC system", we get "ABC system functions".

[0082] If there are 100,000 text fragments in a certain query system, when performing result association, the number of matching text fragments that each word can increase may be as follows:

[0083] "Excuse me": +0;

[0084] "ABC": +200;

[0085] "System": +500;

[0086] "of": +20000;

[0087] "function": +800;

[0088] "have": +300;

[0089] "which": +0.

[0090] After removing the stop words "excuse me", "of", and "which" with the least and most number of text fragments, the second word sequence Qi: "ABC, system, function, have".

[0091] In step 300, use the retrieval algorithm to score and sort the text fragments containing any keyword in the second word sequence to obtain the first text passage sequence Si, including:

[0092] Step 310, among all alternative text fragments, collect and summarize the text fragments containing any keyword in the second word sequence.

[0093] Step 320, use the retrieval algorithm to score each text fragment containing any keyword in the second word sequence respectively.

[0094] Step 330, sort the relevant text fragments according to the scores obtained to obtain the first text passage sequence Si.

[0095] For example, among all alternative text fragments, collect and summarize the text fragments containing any keyword in "ABC, system, function, have". The total number of associated text fragments that can be obtained is 1800. These document fragments are scored by a retrieval algorithm such as TF-IDF and arranged in descending order of scores, and the top n results with the highest scores are selected. Suppose the number of correct text fragments is 2. The probability that these n results exactly contain these 2 correct results is not very high, and the probability that they contain these 2 correct results and they are ranked靠前 is further reduced. The reasons for this are as follows:

[0096] (1) The basic number of text fragments is large, and the retrieval algorithm has a very low basic probability of hitting 2 correct results from 1800 text fragments;

[0097] (2) The limitations of the retrieval algorithm itself. Since the design of the retrieval algorithm is more general, and the way of using information such as word frequency and inverse document frequency in the algorithm considers the usual writing style of people. When the number of repetitions of keywords in some texts is large, the scoring of the algorithm will be interfered;

[0098] (3) Interference caused by inconsistent terminology. For example, in this field, the meaning of "ABC system" is the same as "ABC subsystem" and "ABC module", and if the word "ABC system" is not used in the two correct results, the correct result is likely to be missed;

[0099] (4) Asymmetry between question keywords and response keywords. In this example, the question is about the “function” of the ABC system. However, if the description in the correct result is “the ABC subsystem is mainly used for…”, then the search algorithm does not match the word “function”, so the score of the correct result will be lowered.

[0100] If, as described in the present application, professional terminology sets are used to assist word segmentation and matching, there are 200 original 1800 text segments that contain "ABC", and 500 that contain "system" but not "ABC". Assuming that among the 200 text segments that contain "ABC":

[0101] There are 10 that include the “ABC system”;

[0102] 5 contain “ABC subsystems”;

[0103] 2 contain "ABC modules".

[0104] One of the correct results uses the word "ABC system" and happens to contain the word "function". After excluding all the preliminary text fragments that do not contain the word "ABC system", the result range is narrowed to 10 document fragments, and the probability of the correct result with the highest score will be greatly increased.

[0105] The use of professional terminology is actually a form of information noise reduction, and its effectiveness is based on the following premises: (1) there are limited professional terms in the field, and (2) questions in the field are likely to be related to professional concepts. Excluding all text fragments that are not related to professional topics can reduce a large number of text fragments that do not contain the question topic but whose features are very consistent with the search algorithm and therefore have high scores, greatly narrowing the basic range of hits.

[0106] In step 400, the first F text segments with the highest scores are selected from the first text segment sequence based on the query task data F as the second text segment sequence Sf, including:

[0107] Step 410, final result data F is set based on the query task requirements, and a desired number of final results F is preset, such as 2 or 10. The specific value needs to refer to whether the query task focuses more on query accuracy or query breadth.

[0108] Step 420: In the first text segment sequence Si, retain the first F text segments with the highest scores as the second text segment sequence Sf.

[0109] In step 500, after completing the keyword extraction based on the professional terminology set, it is necessary to perform synonym expansion and responsive keyword expansion on the keywords.

[0110] The premise of the application of synonym expansion is to have a synonym expansion dictionary. The difference between the synonym expansion dictionary and the traditional synonym dictionary is that the synonym dictionary used in the embodiment of the present application is based on a professional terminology set and does not consider the content in the traditional general synonym dictionary. For example, according to the content of the synonym expansion dictionary, a keyword "ABC system" will be expanded to:

[0111] [“ABC system”, “ABC subsystem”, “ABC module”, “ABCS”]

[0112] The function of keyword synonym expansion based on professional terminology is to expand the search scope as needed based on the original search scope and the information provided by the synonym dictionary to reduce the omission of correct results. Using the keyword group after synonym expansion for text screening can collect all 17 text fragments mentioned in the above example, and neither of the two correct answers will be missed.

[0113] In the query processing steps described above, keyword extraction based on a professional terminology set is a kind of information noise reduction processing, keyword synonym expansion can be regarded as information enhancement, and responsive keywords belong to the conversion of some asymmetric attributes in the original query task into symmetric attributes, so that the parts of the original query that cannot be effective only using the retrieval algorithm can also become available to the retrieval algorithm after conversion. Keyword extraction and synonym expansion based on a professional terminology set narrow the search scope by determining the subject of the question, while responsive keywords further narrow the search scope by determining the common dimension or key dimension pointing to the subject of the question. The original query statement after expansion changes from a one-dimensional word sequence to two-dimensional data containing multiple phrases. Taking the question "What are the functions of the ABC system?" as an example, keyword synonym expansion and responsive keyword expansion are performed. The example is as follows:

[0114] (1) Keyword extraction based on professional terminology

[0115] That is, refer to a professional term set to guide the word segmentation of the query statement:

[0116] Sequence 0: [What are the specific functions of the ABC system?]

[0117] (2) Remove stop words from sequence 0 and obtain a word sequence:

[0118] Sequence Qi: [ABC system, functions, what are they]

[0119] (3) Perform keyword synonym expansion based on professional terminology

[0120] After the synonym expansion dictionary is expanded, we will get:

[0121] Sequence 2: [[ABC system, ABC subsystem, ABC module, ABCS], [used for, purpose, completed], what are]

[0122] (4) Perform responsive keyword expansion

[0123] There are two ways to use responsive keywords: one is expansion and the other is substitution. For example, for words that will also appear in the response, such as "function", direct expansion is used, and for words that will not appear in the response, such as "which", substitution is used. After the expansion of responsive keywords, two-dimensional data can be obtained:

[0124] Sequence Qix: [[ABC system, ABC subsystem, ABC module, ABCS], [function, used for, purpose, completed], [include, include, constitute]]

[0125] In step 600, any keyword in the third word sequence is searched in the first text segment sequence, text segments that do not contain any keyword in the third word sequence are excluded, and text segments that contain any keyword in the third word sequence are retained as the third text segment sequence Son_topic.

[0126] In step 700, if the number of text segments in the third text segment sequence is 0, the text segment most relevant to the original query statement is the text segment in the second text segment sequence Sf;

[0127] If the number of text segments in the third text segment sequence is less than or equal to F, the text segment most relevant to the original query statement is the text segment in the third text segment sequence Son_topic;

[0128] If the number of text fragments in the third text sequence is greater than F, the matching conditions of the responsive keywords are converted into weight values ​​for auxiliary retrieval screening, that is, responsive keyword correction is performed to obtain the text fragment most relevant to the original query statement.

[0129] Specifically, perform the following operations:

[0130] Obtain x responsive keywords in the third word sequence Qix; use a retrieval algorithm to score each text segment in the third text segment sequence to obtain the retrieval scores of each text segment in the third text segment sequence; multiply the retrieval scores of each text segment Son_topic in the third text segment sequence by a weight value to obtain the final score corresponding to each text segment. The weight value is related to x and is the xth power of the base y, where the base y should be between 1.0 < y ≤ 1.5. The lower limit value 1.0 of y ensures that the weight value amplifies (rather than weakens) the retrieval score, and the upper limit value 1.5 of y avoids unreasonable bias effects caused by the increase of x. Sort the text segments in the third text segment sequence according to the final score size, and select the top F text segments with the highest scores, that is, the text segments most relevant to the original query statement.

[0131] For example, the score of a text segment containing x responsive keywords can be multiplied by 1.2 to the power of x. For example, if a text segment contains both "function" and "include" at the same time, and the original retrieval algorithm scores it as 0.78, then this score can be multiplied by 1.2 squared: 0.78 x 1.44 = 1.12. After performing the score weighting operation on all text segments in Son_topic, re-sort them based on the final score to obtain sequence R, and then obtain the text segment sequence Rf with the highest scores according to the preset result quantity and return it.

[0132] The above-mentioned retrieval method based on term set synonym expansion and responsive keyword groups is a complete and general text retrieval method in the field based on a professional term set, a synonym expansion dictionary, and a responsive keyword dictionary. This application uses three forms of professional term sets, synonym expansion dictionaries, and responsive keywords, which are simple in form, low in construction and maintenance costs, and long-term effective auxiliary retrieval data, and efficiently fixes the reusable expert knowledge in the field in a low-dimensional form. Please refer to Figure 2 , compared with the traditional retrieval algorithm, the thinking and response methods in the field fixed in these three types of auxiliary data effectively improve the retrieval accuracy, and their low-dimensional characteristics provide sufficient generality. For retrieval algorithms that focus on language models, the processing method provided by this application eliminates a large amount of repetitive construction and update of training sets, reduces the risk cost of bearing the uncertainty of model training effects, and greatly reduces the complexity of the application and increases the understandability of the implementation process. In practical applications, the retrieval method of this application can be easily integrated with other retrieval algorithms to effectively improve the retrieval accuracy rate. In addition, although the application of the retrieval method of this application focuses on solving asymmetric retrieval tasks, it is actually also fully applicable to symmetric retrieval tasks.

[0133] Another embodiment of the present application provides a retrieval device based on synonymous expansion of a term set and responsive keyword groups, including: a processor, a memory, and a computer program in the memory that can be run on the processor, such as a retrieval method program based on synonymous expansion of a term set and responsive keyword groups. When the processor executes the computer program, the steps in the above-mentioned retrieval method embodiments based on synonymous expansion of a term set and responsive keyword groups are implemented, such as Figure 1 steps.

[0134] Exemplarily, the above-mentioned computer program can be divided into one or more modules / units, one or more modules / units are stored in a memory and executed by a processor to complete the present invention. One or more modules / units can be a series of computer program instruction segments that can complete specific functions. The instruction segments are used to describe the execution process of the computer program in a retrieval device based on synonymous expansion of a term set and responsive keyword groups. For example, the computer program can be divided into an acquisition module, a keyword extraction module, a scoring module, a first text segment screening module, an expansion module, a second text segment screening module, and a result acquisition module. The specific functions of each module are as follows:

[0135] The acquisition module is used to acquire the original query statement to be retrieved, and acquire the synonym expansion dictionary and responsive key dictionary associated with the original query statement domain.

[0136] The keyword extraction module is used to extract keywords from the original query sentence based on the professional term set to obtain a second word sequence Qi containing at least one keyword.

[0137] The scoring module is used to use a retrieval algorithm to score and sort the text segments containing any keyword in the second word sequence to obtain a first text segment sequence Si.

[0138] The first text segment screening module is used to select the first F text segments with the highest scores from the first text segment sequence based on the query task data F as the second text segment sequence Sf.

[0139] The expansion module is used to perform keyword synonym expansion and responsive keyword expansion on the second word sequence based on the synonym expansion dictionary and the responsive key dictionary to obtain the third word sequence Qix.

[0140] The second text segment screening module is used to remove the text segments in the first text segment sequence that do not contain any keyword in the third word sequence, and the remaining text segments are used as the third text segment sequence.

[0141] The result acquisition module is used to retrieve the text segment most relevant to the original query statement according to the number of text segments in the third text segment sequence.

[0142] The retrieval device based on the synonymous expansion of the term set and the responsive keyword group may include, but is not limited to, a processor and a memory, for example, it may also include an output device, a network access device, a bus, etc. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), off-the-shelf programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the retrieval device based on the synonymous expansion of the term set and the responsive keyword group, and various interfaces and lines are used to connect the various parts of the retrieval device based on the synonymous expansion of the term set and the responsive keyword group. The memory may be used to store computer programs and / or modules, and the processor realizes various functions of the retrieval device based on the synonymous expansion of the term set and the responsive keyword group by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory.

[0143] If the module / unit integrated in the retrieval device based on the synonymous expansion of the term set and the responsive keyword group is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of each embodiment of the above-mentioned retrieval method based on the synonymous expansion of the term set and the responsive keyword group.

[0144] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. Computer readable media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0145] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention is described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions of the present invention, which should be included in the scope of the claims of the present invention. The techniques, shapes, and structural parts not described in detail in the present invention are all known technologies.

Claims

1. A search method based on synonymous expansion of term sets and responsive keyword groups, characterized in that: The following steps are involved: Obtaining an original query statement to be searched, and obtaining a synonym expansion dictionary and a responsive key dictionary associated with the original query statement domain; Extract keywords from the original query sentence based on the professional term set to obtain a second word sequence Qi containing at least one keyword; Use a search algorithm to score and sort the text segments containing any keyword in the second word sequence to obtain a first text segment sequence Si; Based on the query task data F, the first F text segments with the highest scores are selected from the first text sequence as the second text sequence Sf; Based on the synonym expansion dictionary and the responsive key dictionary, perform keyword synonym expansion and responsive keyword expansion on the second word sequence to obtain the third word sequence Qix; Remove the text segments in the first text segment sequence that do not contain any keyword in the third word sequence, and use the remaining text segments as the third text segment sequence; According to the number of text segments in the third text segment sequence, the text segment most relevant to the original query statement is retrieved.

2. The retrieval method based on synonymous expansion of term set and responsive keyword group according to claim 1 is characterized in that: The step of extracting keywords from the original query statement based on the professional term set to obtain a second word sequence Qi containing at least one keyword includes: Perform a word segmentation operation on the original query statement with reference to the professional term set to obtain a first word sequence containing at least two word segments; Associatively matching each word segment in the first word sequence with the candidate text segments to obtain the number of text segments corresponding to each word segment in the first word sequence; The stop words in the first word sequence are removed to retain the keywords, and the second word sequence Qi is obtained.

3. The retrieval method based on synonymous expansion of term set and responsive keyword group according to claim 1 is characterized in that: The retrieval algorithm is used to score and sort the text segments containing any keyword in the second word sequence to obtain a first text segment sequence Si, including: Among all the candidate text segments, collect and summarize the text segments containing any keyword in the second word sequence; Using the search algorithm to score each text segment containing any of the keywords in the second word sequence; The related text segments are sorted according to the scores obtained to obtain the first text segment sequence Si.

4. The retrieval method based on synonymous expansion of term set and responsive keyword group according to claim 3 is characterized in that: The method of selecting the first F text segments with the highest scores from the first text segment sequence based on the query task data F as the second text segment sequence Sf includes: Setting the final result data F based on the query task requirement, wherein the query task is retrieval accuracy or retrieval breadth; In the first text segment sequence, the first F text segments with the highest scores are retained as the second text segment sequence.

5. The retrieval method based on synonymous expansion of term set and responsive keyword group according to claim 4 is characterized in that: The responsive keywords include general responsive keywords and domain-limited responsive keywords. General responsive keywords refer to corresponding relationships that are completely independent of any professional knowledge content; domain-limited responsive keywords refer to responses to the same type of questions in a specific professional field that explicitly or implicitly follow a fixed word structure.

6. The retrieval method based on synonymous expansion of term set and responsive keyword group according to claim 5 is characterized in that: If the number of text segments in the third text segment sequence is 0, the text segment most relevant to the original query statement is the text segment in the second text segment sequence; If the number of text segments in the third text segment sequence is less than or equal to F, the text segment most relevant to the original query statement is the text segment in the third text segment sequence; If the number of text fragments in the third text sequence is greater than F, the matching conditions of the responsive keywords are converted into weight values ​​for auxiliary retrieval screening to obtain the text fragments most relevant to the original query statement.

7. The retrieval method based on synonymous expansion of term set and responsive keyword group according to claim 6 is characterized in that: The matching of the responsive keywords is converted into weight values ​​for auxiliary search and screening to obtain the text fragments most relevant to the original query statement, including: Get the number x of responsive keywords in the third word sequence; Using a retrieval algorithm to score each text segment in the third text segment sequence, to obtain a retrieval score for each text segment in the third text segment sequence; Multiply the search score of each text segment in the third text segment sequence by a weight value to obtain a final score corresponding to each text segment, wherein the weight value is related to x; The text segments in the third text segment sequence are sorted according to the final scores, and the first F text segments with the highest scores are selected, that is, the text segments most relevant to the original query statement.

8. The retrieval method based on synonymous expansion of term set and responsive keyword group according to claim 7 is characterized in that: The weight value is the xth power of the base y, and the base y satisfies 1.5>y>1.

0.

9. A retrieval device based on synonymous expansion of a term set and responsive keyword groups, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the retrieval method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the search method according to any one of claims 1 to 8 are implemented.