An LLM text processing method

Through word segmentation and vocabulary database analysis, combined with initial and precise inspection methods to screen LLM answer information, the LLM search accuracy problem was solved and more accurate answer output was achieved.

CN119884355BActive Publication Date: 2025-07-22SHANGHAI JILIXUN INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510370860.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-22
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

During the process of comparing the problem text and knowledge base search, the knowledge base transfers the entire paragraph to affect the search accuracy, resulting in inaccurate search results.

Method used

The input problem information is analyzed through the word segmentation database, and the problem vocabulary is selected, and the query vocabulary is defined using synonyms, mutual exclusion words, and abbreviation vocabulary. Combined with the initial and fine inspection methods, the benchmark answer information is matched with priority, and the fine answer information is finally output.

Benefits of technology

It improves the accuracy and efficiency of LLM retrieval, reduces the upload of non-related texts, and ensures that the answer text is highly correlated with the input question information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884355B_ABST
    Figure CN119884355B_ABST
Patent Text Reader

Abstract

The present invention relates to an LLM text processing method, belonging to the technical field of text processing. The method includes: obtaining input question information; inputting the input question information into a word segmentation database to output question words; inputting the question words into a preset vocabulary database to output synonymous words, mutually exclusive words, and abbreviated words, and defining the question words, synonymous words, mutually exclusive words, and abbreviated words as query words; outputting final question information according to the query words; using a preliminary retrieval method according to the final question information to find out the reference answer information and the corresponding path; inputting the reference answer information and the path into a priority database to match the priority; determining the preliminary retrieval answer information from the matched priority; using a refined retrieval method according to the preliminary retrieval answer information to find out the refined retrieval answer information; taking the refined retrieval answer information as the final answer text and uploading and displaying it. This application has the effect of improving the accuracy of LLM retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text processing, and more particularly to an LLM text processing method. Background Art

[0002] LLM refers to a large language model, and the technology of LLM is widely applied to automatic reply, text classification, sentiment analysis, machine translation, intelligent question answering, information extraction, and summary generation, etc.

[0003] LLM includes retrieval-augmented generation technology. Currently, the augmented generation technology generally first collects materials on the Internet and stores them to form a knowledge base, then retrieves and compares the input question text with the knowledge base to obtain alternative texts, sorts the alternative texts according to the comparison between the question text and the alternative texts, so that the fragment most relevant to the question text is placed at the beginning of the alternative text, and the sorted alternative text is sent to the LLM for answering.

[0004] In the process of retrieving and comparing the question text with the knowledge base, the knowledge base will send the entire paragraph related to the question text, thus affecting the accuracy of LLM retrieval. Summary of the Invention

[0005] In order to improve the accuracy of LLM retrieval, the present invention provides an LLM text processing method.

[0006] In a first aspect, the present invention provides an LLM text processing method, adopting the following technical solution:

[0007] An LLM text processing method includes:

[0008] Obtain the input question information in a preset recognition area;

[0009] Input the input question information into a preset word segmentation database to output question words;

[0010] Input the question words into a preset vocabulary database to output synonymous words, mutually exclusive words, and abbreviated words, and define the question words, synonymous words, mutually exclusive words, and abbreviated words as query words;

[0011] Output the final question information according to the query words;

[0012] Find out the reference answer information and the corresponding path according to the final question information through a preset preliminary inspection method;

[0013] Input the reference answer information and the path into a preset priority database to match the priority;

[0014] Determine the benchmark response information with the highest priority from the matched priorities, and define the benchmark response information with the highest priority as the preliminary inspection response information;

[0015] Search for the refined inspection response information according to the preliminary inspection response information through a preset refined inspection method;

[0016] Use the refined inspection response information as the final response text and upload it for display.

[0017] By adopting the above technical solution, analyze the vocabulary in the input question information and the vocabulary database to obtain the final question information, obtain the benchmark response information and the corresponding path through the preliminary inspection method, then distinguish the priorities of the benchmark response information and the path through the priority database to obtain the preliminary inspection response information, search for the refined inspection response information through the refined inspection method, and finally use the refined inspection response information as the final response text for upload and display. Through the preliminary inspection method and the refined inspection method, the final response text that meets the input question information can be further screened to improve the accuracy of LLM retrieval.

[0018] Optionally, the preliminary inspection method includes:

[0019] Define the response information in the preset text set that contains all the query vocabulary as the first level;

[0020] Use the response information corresponding to the first level as the benchmark response information.

[0021] Optionally, the preliminary inspection method further includes:

[0022] Define the last level in the path as the last path, the vocabulary found by the query vocabulary in the last path as known vocabulary, and the vocabulary not found by the query vocabulary in the last path as unknown vocabulary;

[0023] Define the response information on the last path that contains at least two known vocabularies and at least one unknown vocabulary in the preset text set as the first level.

[0024] Optionally, the preliminary inspection method further includes:

[0025] Define the vocabulary found by the query vocabulary in the preset text set as text-known vocabulary, and the vocabulary not found by the query vocabulary in the preset text set as text-unknown vocabulary;

[0026] Define the response information in the path that contains at least one text-known vocabulary and at most two unknown vocabularies in the preset text set as the second level;

[0027] Use the response information corresponding to the second level as the benchmark response information.

[0028] Optionally, the preliminary inspection method further includes:

[0029] Count the known words on the last path in the first level to filter out the response information corresponding to the most known words and define it as the first level, and define the response information corresponding to the remaining known words as the third level;

[0030] Use the response information corresponding to the third level as the reference response information.

[0031] Optionally, the preliminary inspection method further includes:

[0032] Define the response information that contains two known words on the last path and does not contain unknown words in the preset text set as the third level.

[0033] By adopting the above technical solution, the response information related to the input question information is divided into the first level, the second level, and the third level by asking about the existence of words in the text set, so as to reduce the retrieval range of the LLM, improve the retrieval efficiency of the LLM, and then output the response information with the highest priority, so as to directly display the response related to the input question information.

[0034] Optionally, the preliminary inspection method further includes:

[0035] Obtain the lexical distance value of adjacent query words in the preset text set;

[0036] Determine whether the lexical distance value falls within the preset reference position distance values;

[0037] If the lexical distance value does not fall within the preset reference position distance values, then determine whether the query word falls within the preset table features;

[0038] If the query word does not fall within the preset table features, then eliminate it;

[0039] If the lexical distance value falls within the preset reference position distance values or the query word falls within the preset table features, then mark the text between the query words whose lexical distance values fall within the reference position distance values in the preset text set;

[0040] When the number of query words exceeds 1, define the query words as a query word group;

[0041] If there are repeated query word groups in the text set, then mark the text between adjacent repeated query word groups in the text set.

[0042] By adopting the above technical solution, by understanding the lexical distance values of adjacent query words in the text set, and screening the texts in the text set according to the falling situation of the lexical distance value and the reference position distance value, the length of the text fragment finally transmitted to the LLM can be dynamically adjusted according to the lexical distance value, which can be large or small, avoiding the defect that key knowledge points are discarded due to fixed-length texts or the text is too long to interfere with the answer accuracy of the LLM. Thus, texts related to the input question information can be further screened to improve the accuracy of LLM retrieval.

[0043] Optionally, the preliminary inspection method further includes:

[0044] Correcting the final question information according to the query word with a preset missing quantity.

[0045] By adopting the above technical solution, retrieving according to the quantities corresponding to different query words, so as to continue to screen within the semantic range corresponding to the input question information, thereby improving the accuracy of LLM retrieval.

[0046] Optionally, the refined inspection method includes:

[0047] Defining the text marked in the text set as the marked text;

[0048] Inputting the path and the marked text into a preset scoring database to calculate the target score;

[0049] Arranging the target scores in ascending order, and selecting the smallest target score as the selected score according to the arranged target scores;

[0050] Calculating the product between the selected score and a preset reference coefficient, and defining the product value as the reference score;

[0051] Selecting the paths and marked texts corresponding to the target scores less than the reference score from the target scores as the refined inspection answer information.

[0052] By adopting the above technical solution, calculating the target score through a preset scoring database for the path and the marked text, and using the paths and marked texts corresponding to the target scores less than the reference score as the refined inspection answer information, and further screening through the scores corresponding to the paths and marked texts, so as to improve the accuracy of LLM retrieval.

[0053] Optionally, the calculation method of the scoring database:

[0054] Score 综合 = (Score 向量 / log 16 [1 / 6(n - 1) 3 + (n - 1) 2+5 / 6 (n - 1)+1])×1 / [(2 - index×0.1)×(10e log2(m+1) +1)],

[0055] Score 综合 is the target score, index is the ranking or score corresponding to the input question information in the preset title database, Score 向量 is the score obtained by comparing the vectors of the input question information and the preset sliced text database. n is the number of text words of different lengths intercepted from the input question information, and m is the ranking corresponding to the marked text after sorting by the preset rearrangement model.

[0056] In summary, the present application includes at least one of the following beneficial technical effects:

[0057] 1. By analyzing the vocabulary in the input question information and the vocabulary database to obtain the final question information, and using the preliminary inspection method to obtain the reference answer information and the corresponding path, then using the priority database to distinguish the priorities of the reference answer information and the path to obtain the preliminary inspection answer information, and using the refined inspection method to find the refined inspection answer information. Finally, the refined inspection answer information is uploaded and displayed as the final answer text. Through the preliminary inspection method and the refined inspection method, the final answer text that meets the input question information can be further screened to improve the accuracy of LLM retrieval;

[0058] 2. By understanding the lexical distance values of adjacent query words in the text set, and screening the texts in the text set according to the falling situation of the lexical distance values and the reference position distance values, the texts related to the input question information can be further screened to improve the accuracy of LLM retrieval;

[0059] 3. By calculating the target score for the path and the marked text in the preset scoring database, and using the path and the marked text corresponding to the target score less than the reference score as the refined inspection answer information, and further screening through the scores corresponding to the path and the marked text to improve the accuracy of LLM retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is a flowchart of a method for processing LLM texts according to an embodiment of the present invention;

[0061] Figure 2 is the method flow of the preliminary inspection method according to an embodiment of the present invention Figure 1 ;

[0062] Figure 3 is the method flow of the preliminary inspection method according to an embodiment of the present invention Figure 2 ;

[0063] Figure 4It is the method flow of the preliminary inspection method of the embodiment of the present invention Figure 3 ;

[0064] Figure 5 It is the method flow of the preliminary inspection method of the embodiment of the present invention Figure 4 ;

[0065] Figure 6 It is the method flow of the preliminary inspection method of the embodiment of the present invention Figure 5 ;

[0066] Figure 7 It is the method flow chart of the detailed inspection method of the embodiment of the present invention. Specific implementation manners

[0067] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0068] An LLM text processing method understands through the vocabulary in the input question information, and screens the text materials retrieved by the LLM through the preliminary inspection method and the detailed inspection method to obtain the final answer information and upload and output it. Thus, while the LLM uploads the materials related to the input question information, it can reduce the length of the uploaded answer information and the irrelevant text materials, and further improve the accuracy of the LLM retrieval.

[0069] Refer to Figure 1 This application embodiment discloses an LLM text processing method, including the following steps:

[0070] Step S100: Obtain the input question information in the preset recognition area.

[0071] The input question information refers to the text input by the user to the LLM for question asking, and the text in the preset input module in the system is retrieved as the question text. The input module refers to the module in the LLM for receiving the text input by the user, that is, the text input box. The input module is preset by those skilled in the art and will not be elaborated here.

[0072] Step S101: Input the input question information into the preset word segmentation database to output question vocabulary.

[0073] The word segmentation database refers to a database that stores various word vocabularies and is used to split text to form individual words. The word segmentation database is preset by those skilled in the art and will not be elaborated here.

[0074] The question vocabulary refers to the vocabulary in the input question information, and the vocabulary obtained by splitting and segmenting the input question information into the word segmentation database is used as the question vocabulary.

[0075] Step S102: Input the problem vocabulary into a preset vocabulary database to output synonymous vocabulary, mutually exclusive vocabulary, and abbreviated vocabulary, and define the problem vocabulary, synonymous vocabulary, mutually exclusive vocabulary, and abbreviated vocabulary as query vocabulary.

[0076] The vocabulary database includes a synonym database, a mutually exclusive word library, an abbreviated word library, etc. The synonym database stores different words corresponding to the same semantics. The mutually exclusive word library stores a set of words that cannot simultaneously satisfy the text semantics in the vocabulary set where the text appears. The abbreviated word library stores the abbreviated vocabulary represented by the professional vocabulary in the corresponding field of the text. The vocabulary database is preset by those skilled in the art and will not be elaborated here. Synonymous vocabulary refers to the vocabulary corresponding to the same semantics as the problem vocabulary. Mutually exclusive vocabulary refers to the vocabulary that appears in the text with the problem vocabulary and cannot satisfy the text semantics. Abbreviated vocabulary refers to the vocabulary obtained by abbreviating the professional vocabulary in the corresponding field of the problem vocabulary. By inputting the problem vocabulary into the vocabulary database for query and comparison, synonymous vocabulary, mutually exclusive vocabulary, and abbreviated vocabulary are output.

[0077] Query vocabulary refers to the vocabulary corresponding to the system's query to the LLM. By defining the problem vocabulary, synonymous vocabulary, mutually exclusive vocabulary, and abbreviated vocabulary as query vocabulary.

[0078] Step S103: Output the final question information according to the query vocabulary.

[0079] The final question information refers to the text information corresponding to the final query to the LLM. By using the query vocabulary to reorganize the input text information to obtain the final question information.

[0080] Step S104: Use a preset preliminary inspection method to find the reference answer information and the corresponding path according to the final question information.

[0081] The preliminary inspection method is a method for initially retrieving and querying the final question information. The specific operation steps refer to Step S200 to Step S800. The reference answer information refers to the reference text information related to the final question information. The path refers to the query path corresponding to the materials storing the information related to the final question information. The path can be the address description of a text file in a computer and the location of resources on a web server, etc. By using the preliminary inspection method to query the final question information, the reference answer information and the corresponding path are obtained.

[0082] Step S105: Input the reference answer information and the path into a preset priority database to match the priority.

[0083] The priority database refers to a database used to classify benchmark answer information and paths. The priority database is set in advance by those skilled in the art and will not be elaborated here. In this embodiment, the priority database classifies benchmark answer information and paths into three levels, grades the benchmark answer information and paths through a preliminary inspection method, and inputs each level into the priority database to obtain priorities.

[0084] Step S106: Determine the benchmark answer information with the highest priority from the matched priorities, and define the benchmark answer information with the highest priority as the preliminary inspection answer information.

[0085] The preliminary inspection answer information refers to the benchmark answer information obtained after screening through a preliminary method and the priority database. By determining the benchmark answer information with the highest priority from the matched priorities and defining the benchmark answer information with the highest priority as the preliminary inspection answer information.

[0086] Step S107: Search for refined inspection answer information according to the preliminary inspection answer information through a preset refined inspection method.

[0087] The refined inspection method refers to a method for finely retrieving the preliminary inspection answer information. The specific operation steps refer to Step S900 to Step S904. The refined inspection answer information refers to the answer information obtained after screening the preliminary inspection answer information through the refined inspection method. The preliminary inspection answer information obtained by screening the preliminary inspection answer information through the refined inspection method is used as the refined inspection answer information.

[0088] Step S108: Use the refined inspection answer information as the final answer text and upload it for display.

[0089] By using the refined inspection answer information as the final answer text and uploading it to the LLM for display to the user, the final answer text that meets the input question information is screened out through the preliminary inspection method and the refined inspection method to improve the accuracy of LLM retrieval.

[0090] Refer to Figure 2 , the preliminary inspection method includes the following steps:

[0091] Step S200: Define the answer information containing all the query words in the preset text set as the first level.

[0092] The text set refers to the set of text materials stored in each path. The text set is queried and set in advance by those skilled in the art and will not be elaborated here. The first level refers to the answer information with the highest relevance degree marked as the first to the input question information in the text set. This answer information includes the path and the text materials corresponding to the path query.

[0093] When the preset text set contains the response information for all the query terms, it means that 1 query term can be found on the complete path, and all other query terms can be found in the text, or all the query terms can be found in the text. Therefore, this text set is defined as the first level.

[0094] Step S201: Use the response information corresponding to the first level as the reference response information.

[0095] By using the response information in the text set of the first level as the reference response information, the texts related to the input question information are screened, improving the accuracy of LLM retrieval.

[0096] Refer to Figure 3 , the initial inspection method further includes the following steps:

[0097] Step S300: Define the last level in the path as the last path. The terms found for the query terms in the last path are known terms, and the terms not found for the query terms in the last path are unknown terms.

[0098] The last path refers to the shortest path where text materials can be queried in the text set. By defining the last level in the path as the last path. Known terms refer to the query terms that can be queried in the last path of the text set, and unknown terms refer to the query terms that cannot be queried in the last path of the text set. The query terms found in the last path of the text set are used as known terms, and the query terms that cannot be queried in the last path of the text set are used as unknown terms.

[0099] Step S301: Define the response information that contains at least two known terms on the last path and at least one unknown term in the preset text set as the first level.

[0100] When there are at least two known terms on the last path and at least one unknown term in the preset text set, it means that 2 or more query terms are found at the last level of the path and at least one other query term is found in the text. Therefore, this text set is defined as the first level.

[0101] Refer to Figure 4 , the initial inspection method further includes the following steps:

[0102] Step S400: The terms found for the query terms in the preset text set are text known terms, and the terms not found for the query terms in the preset text set are text unknown terms.

[0103] The known vocabulary in the text refers to the query vocabulary that can be found in the text set, and the unknown vocabulary in the text refers to the query vocabulary that cannot be found in the text set. The query vocabulary found in the text set is used as the known vocabulary in the text, and the query vocabulary that cannot be found in the text set is used as the unknown vocabulary in the text.

[0104] Step S401: Define the response information whose path contains at least one known vocabulary in the text and whose preset text set contains at most two unknown vocabularies as the second level.

[0105] When the path contains at least one known vocabulary in the text and the preset text set contains at most two unknown vocabularies, it indicates that one or two vocabularies are missing in the query vocabulary found in the text compared to all the query vocabularies, or one query vocabulary is found on the complete path, and one or two vocabularies are missing in the query vocabulary found in the text compared to the remaining query vocabularies. Therefore, this text set is defined as the second level.

[0106] Step S402: Use the response information corresponding to the second level as the reference response information.

[0107] By using the response information in the text set of the second level as the reference response information, continue to screen the text related to the input question information to improve the accuracy of LLM retrieval.

[0108] Refer to Figure 5 , the initial inspection method further includes the following steps:

[0109] Step S500: Count the known vocabularies on the last path in the first level to screen out the response information corresponding to the most known vocabularies and define it as the first level, and define the response information corresponding to the remaining known vocabularies as the third level.

[0110] When there are multiple known vocabularies in the last path of the first level, sort the quantities corresponding to the known vocabularies on the last path in the first level, screen out the response information corresponding to the most numerous known vocabularies in the text set and define it as the first level, and define the remaining response information as the third level.

[0111] Step S501: Use the response information corresponding to the third level as the reference response information.

[0112] By using the response information in the text set of the third level as the reference response information, continue to screen the text related to the input question information to improve the accuracy of LLM retrieval.

[0113] The initial inspection method further includes the following steps:

[0114] Step S600: Define the response information whose last path contains two known vocabularies and whose preset text set does not contain unknown vocabularies as the third level.

[0115] When there are two known words on the last path and the response information that does not contain unknown words in the preset text set is defined as the third level, it means that only 2 keywords are found at the last level of the path, but no other keyword is found in the text. Therefore, the response information in this text set is defined as the third level.

[0116] Refer to Figure 6 , the preliminary inspection method further includes the following steps:

[0117] Step S700: Obtain the lexical distance value between adjacent query words in the preset text set.

[0118] The lexical distance value refers to the distance between adjacent query words in the text set. By retrieving the number of characters between adjacent query words in the text set and using the corresponding numerical value of the number of characters as the lexical distance value. In this embodiment, the lexical distance value is only valid in the text, and the words on the path do not participate in the distance value. If there are adjacent query words located in the table, special processing can be performed on the lexical distance value, so there is no restriction on the query words located in the table.

[0119] Step S701: Determine whether the lexical distance value falls between the preset reference position distance values.

[0120] The reference position distance value refers to the maximum allowable distance value of the text corresponding to the text related to the input question information between adjacent query words. The reference position distance value is set in advance by those skilled in the art and will not be elaborated here. In this embodiment, the reference position distance value is 2000. By determining whether the lexical distance value falls between the reference position distance values, it is determined whether the text where the lexical distance value is located is related to the input question information.

[0121] Step S7011: If the lexical distance value does not fall between the preset reference position distance values, determine whether the query word falls into the preset table features.

[0122] The table feature refers to the arrangement and shape of the table corresponding to the text in the text. The table feature is set in advance by those skilled in the art and will not be elaborated here. When the lexical distance value does not fall between the reference position distance values, it indicates that the text where the lexical distance value is located has a high probability of being unrelated to the input question information. Therefore, it is determined whether the query word falls into the table feature to determine whether the text corresponding to the table is related to the input question information.

[0123] Step S70111: If the query word does not fall into the preset table features, perform elimination.

[0124] When the query terms do not fall within the table features, it indicates that the text corresponding to the table in the text set is not relevant to the input question information. Therefore, the text corresponding to the lexical distance values that do not fall within the benchmark position distance values is excluded.

[0125] Step S7012: If the lexical distance value falls within the preset benchmark position distance values or the query term falls within the preset table features, then mark the text between the query terms whose lexical distance values in the preset text set fall within the benchmark position distance values.

[0126] When the lexical distance value falls within the benchmark position distance values, or the query term falls within the table features, it indicates that the text corresponding to the lexical distance values falling within the benchmark position distance values and the table are relevant to the input question information. Therefore, mark the text between the query terms whose lexical distance values in the text set fall within the benchmark position distance values.

[0127] Step S702: When the query term exceeds 1, define the query term as a query term group.

[0128] A query term group refers to a group of query terms for combined search. When the query term exceeds 1, it indicates that there are multiple query terms for search in the text. Therefore, define the query term as a query term group.

[0129] Step S7021: If there are repeated query term groups in the text set, then mark the text between adjacent repeated query term groups in the text set.

[0130] When there are no repeated query term groups in the text set, it indicates that there is only one piece of content in the text set that is relevant to the input question information. Therefore, continue to mark the text between the query terms whose lexical distance values in the text set fall within the benchmark position distance values. When there are repeated query term groups in the text set, it indicates that there are multiple pieces of content in the text set that are relevant to the input question information. Therefore, mark the text between adjacent repeated query term groups in the text set.

[0131] The preliminary inspection method further includes the following steps:

[0132] Step S800: Modify the final question information according to the query term with a preset missing quantity.

[0133] The missing quantity refers to the number of query terms reduced in the input question for query retrieval. The missing quantity is set in advance by those skilled in the art and will not be elaborated here. In this embodiment, the missing quantity is 1 or 2. Since the situation of retrieving after reducing the query terms corresponding to the missing quantity has sufficiently covered the semantic range of the question, by searching for the query terms with the reduced missing quantity, the corresponding path can still be retrieved quite accurately.

[0134] By re - executing step S103 on the query vocabulary for reducing the missing quantity, new text information is obtained as the corrected final question information.

[0135] Refer to Figure 7 , the refined inspection method includes the following steps:

[0136] Step S900: Define the text marked in the text set as the marked text.

[0137] The marked text refers to the text marked in the text set, and the text marked by steps S7012 and S7021 is used as the marked text.

[0138] Step S901: Input the path and the marked text into a preset scoring database to calculate the target score.

[0139] The target score is the comprehensive score corresponding to the marked text. The path and the marked text are input into a preset scoring database to calculate the target score.

[0140] Calculation method of the scoring database:

[0141] Score 综合 = (Score 向量 / log 16 [1 / 6(n - 1) 3 + (n - 1) 2 + 5 / 6(n - 1)+1])×1 / [(2 - index×0.1)×(10e log2(m+1) +1)],

[0142] Score 综合 is the target score, index is the ranking or score corresponding to the input question information in the preset title database, Score 向量 is the score obtained by performing vector comparison between the input question information and the preset sliced - text database, n is the number of text words of different lengths intercepted in the input question information, and m is the ranking corresponding to the marked text after being sorted by the preset rearrangement model.

[0143] The title database stores a knowledge base created separately for all titles. The title database is a database set by humans and will not be elaborated here. By inputting the question information into the title database, the top 10 titles are obtained. When the path corresponding to the input question information exists in the top 10 titles, the index is the ranking in the top 10 titles. When the path corresponding to the input question information does not exist in the top 10 titles, then the index is 1.

[0144] The sliced text database stores a text set and paths after slicing, and is a database for vectorized storage. The sliced text database is a database set by humans and will not be elaborated here.

[0145] The rearrangement model is a model used to rank marked text. The rearrangement model is a model set by humans and will not be elaborated here.

[0146] Step S902: Arrange the respective target scores in ascending order, and select the smallest target score as the selected score according to the arranged target scores.

[0147] The selected score is the target score corresponding to the text related to the input question information selected by arranging the respective target scores in ascending order and selecting the smallest target score as the selected score.

[0148] Step S903: Calculate the product between the selected score and a preset benchmark coefficient, and define the product value as the benchmark score.

[0149] The benchmark coefficient is a coefficient used to correct the selected score. The benchmark coefficient is set in advance by those skilled in the art and will not be elaborated here. In this embodiment, the benchmark coefficient can be 1.3. The benchmark score is the score corresponding to the text related to the input question information further selected after correction by calculating the product between the selected score and the preset benchmark coefficient and defining the product value as the benchmark score.

[0150] Step S904: Select the paths and marked texts corresponding to the target scores less than the benchmark score from the target scores as the refined retrieval answer information.

[0151] By selecting the paths and marked texts corresponding to the target scores less than the benchmark score from the target scores as the refined retrieval answer information, the text related to the input question information can be output, thereby improving the accuracy of LLM retrieval.

[0152] Furthermore, the method after determining the query vocabulary includes the following steps:

[0153] Step S1021: Determine whether the input question information is completely input according to the query vocabulary and a preset vocabulary continuous database.

[0154] By querying the completeness of the input question information corresponding to the query terms in the preset continuous vocabulary database, and determining whether the question text is completely input based on the excess situation between the completeness and the preset benchmark completeness, it is possible to determine whether it is necessary to supplement the question text. The benchmark completeness refers to the maximum completeness required to supplement the question text, which is preset by those skilled in the art and will not be elaborated here. The continuous vocabulary database stores the completeness of the question text corresponding to different question terms. The continuous vocabulary database is a database set by humans and will not be elaborated here.

[0155] Step S1022: When the input question information is not completely input, obtain the question word distance values corresponding to each query term in the input question information.

[0156] The question word distance value refers to the position corresponding to each word in the input question information. When the input question information is not completely input, it means that the completeness does not exceed the benchmark completeness and the question text needs to be supplemented. Therefore, by retrieving the number of words between each word in the input question information and the word at the beginning of the sentence, and using the corresponding value of this number of words as the question word distance value.

[0157] Step S1023: Arrange the query terms corresponding to each question word distance value in ascending order, and select the query term corresponding to the smallest question word distance value as the priority query term.

[0158] The priority query term refers to the query term that is preferentially read and written in the input question information. By arranging the query terms corresponding to each question word distance value in ascending order, and selecting the query term corresponding to the smallest question word distance value as the priority query term.

[0159] Step S1024: Determine the priority problem area according to the priority query term and the vocabulary database.

[0160] The priority problem area refers to the technical area where the priority query term is used in the text. The priority problem area is matched from the vocabulary database through the priority query term. The vocabulary database also stores the technical areas corresponding to different vocabulary, which will not be elaborated here.

[0161] Step S1025: After determining the priority problem area, based on the question word distance value, use the next query term as the new priority query term, and update the priority problem area according to the new priority query term and the vocabulary database until the priority query term is the last query term.

[0162] After determining the priority problem area, to know the general area direction of the input problem information, so based on the problem text distance value, continue to use the next query term as the new priority query term, and match a new area from the vocabulary database through the new priority query term, and combine the new area with the priority problem area to form a new priority problem area, and repeat the above steps until the priority query term is the last query term.

[0163] Step S1026: Match domain terms from the vocabulary database according to the updated priority problem area.

[0164] Domain terms refer to the terms related to the updated priority problem area, and match domain terms from the vocabulary database through the updated priority problem area.

[0165] Step S1027: Recombine the input problem information according to the domain terms.

[0166] By adding the domain terms after the input problem information to form new input problem information, the content of the input problem information can be supplemented.

[0167] Step S1028: Update the final problem information according to the recombined input problem information.

[0168] Execute steps S101 to S103 again through the recombined input problem information to obtain new final problem information, so that it is not easy for incomplete input problem information to affect the LLM to retrieve irrelevant texts, thereby improving the accuracy of LLM retrieval.

[0169] Furthermore, the method after determining the query term further includes the following steps:

[0170] Step S10221: When the input problem information is completed and a preset trigger information is output, obtain the pause time for outputting the trigger information.

[0171] The trigger information refers to the information output for asking questions to the LLM. The trigger information is set in advance by those skilled in the art and will not be elaborated here. The pause time refers to the time paused between the completion of the input problem information and the output of the trigger information. When the input problem information is completed and a preset trigger information is output, it means that questions need to be asked to the LLM. Mark the time point when the input problem information is completed by the system, and mark the time point when the system outputs the trigger information, and calculate the time parameter value between the two marked time points as the pause time.

[0172] Step S10222: Determine whether the input problem information is updated during the pause time.

[0173] By determining whether the input question information is updated during the pause time, it is determined whether the user actively adds text or accidentally touches to input other text.

[0174] Step S10223: If not updated, continue to determine the final question information.

[0175] When the input question information during the pause time is not updated, it means that the user does not actively add text or accidentally touch to input other text, so continue to execute step S103.

[0176] Step S10224: If updated, determine the added text based on the input question information before and after the update, and update the pause time according to the added text and the trigger information.

[0177] The added text refers to the text added to the input question information during the pause time. When the input question information during the pause time is updated, it means that the user actively adds text or accidentally touches to input other text. Therefore, the input question information before and after the update is compared for overlap, and the non-overlapping text is used as the added text.

[0178] Step S10225: Match the added vocabulary from the vocabulary database according to the added text and the query vocabulary.

[0179] The added vocabulary refers to the vocabulary in the input question information after adding the added text. By combining the added text with the query vocabulary and matching the combined text from the vocabulary database for the added vocabulary.

[0180] Step S10226: Determine the query vocabulary field according to the query vocabulary, and determine the added vocabulary field according to the added vocabulary.

[0181] The query vocabulary field refers to the technical field in which the query vocabulary is used in the text, and the added vocabulary field refers to the technical field in which the added vocabulary is used in the text. The query vocabulary field is matched from the vocabulary database through the query vocabulary, and the added vocabulary field is matched from the vocabulary database through the added vocabulary.

[0182] Step S10227: When the query vocabulary field is consistent with the added vocabulary field, update the final question information according to the updated input question information.

[0183] When the query vocabulary field is consistent with the added vocabulary field, it means that the added text will not affect the retrieval of the input question information by the LLM. Therefore, steps S101 to S103 are re-executed with the updated input question information.

[0184] Step S10228: When the query vocabulary field is inconsistent with the added vocabulary field, continue to determine the final question information.

[0185] When the vocabulary field of inquiry is inconsistent with the vocabulary field to be increased, it indicates that adding text will affect the retrieval of the input question information by the LLM. Therefore, continue to execute steps S101 to S103 with the input question information before the update.

[0186] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. An LLM text processing method, characterized in that, including: Obtaining input question information in a preset recognition area; Inputting the input question information into a preset word segmentation database to output question words; Inputting the question words into a preset vocabulary database to output synonymous words, mutually exclusive words, and abbreviated words, and defining the question words, synonymous words, mutually exclusive words, and abbreviated words as query words; Outputting final question information according to the query words; Searching for reference answer information and corresponding paths according to the final question information through a preset preliminary inspection method; Inputting the reference answer information and paths into a preset priority database to match priorities; Determining the reference answer information with the highest priority from the matched priorities, and defining the reference answer information with the highest priority as the preliminary inspection answer information; Searching for refined inspection answer information according to the preliminary inspection answer information through a preset refined inspection method; Taking the refined inspection answer information as the final answer text and uploading and displaying it; The preliminary inspection method includes: Defining the answer information in a preset text set that contains all query words as the first level; Taking the answer information corresponding to the first level as the reference answer information; The preliminary inspection method also includes: Obtaining the lexical distance value of adjacent query words in a preset text set; Determining whether the lexical distance value falls within the preset reference position distance values; If the lexical distance value does not fall within the preset reference position distance values, determining whether the query words fall within a preset table feature; If the query words do not fall within the preset table feature, eliminating them; If the lexical distance value falls within the preset reference position distance values or the query words fall within the preset table feature, marking the text between the query words in the preset text set whose lexical distance values fall within the reference position distance values; When the number of query words exceeds 1, defining the query words as a query word group; When there are repeated query word groups in the text set, marking the text between adjacent repeated query word groups in the text set; The refined inspection method includes: Defining the marked text in the text set as the marked text; Inputting the path and the marked text into a preset scoring database to calculate the target score; Sorting the target scores in ascending order, and selecting the smallest target score as the selected score according to the sorted target scores; Calculating the product between the selected score and a preset reference coefficient, and defining the product value as the reference score; Selecting the paths and marked texts corresponding to the target scores less than the reference score from the target scores as the refined inspection answer information; Calculation method of the scoring database: Score comprehensive = (Score vector / log16[1 / 6(n - 1)3 + (n - 1)2 + 5 / 6(n - 1) + 1]) × 1 / [(2 - index × 0.1) × (10elog2(m + 1) + 1)], Score is the target score, index is the ranking or score corresponding to the input question information in the preset title database, the Score vector is the score obtained by comparing the input question information with the preset sliced text database, n is the number of text words of different lengths intercepted from the input question information, and m is the ranking corresponding to the marked text after sorting by the preset rearrangement model.

2. The method for processing LLM text according to claim 1, wherein, The preliminary inspection method further includes: Define the last level in the path as the last path, the words found for the query word in the last path as known words, and the words not found for the query word in the last path as unknown words; Define the response information on the last path that contains at least two known words and at least one unknown word in the preset text set as the first level.

3. A method for processing LLM text according to claim 1, characterized in that The preliminary inspection method further includes: The words found for the query word in the preset text set are text known words, and the words not found for the query word in the preset text set are text unknown words; Define the response information in the path that contains at least one text known word and at most two unknown words in the preset text set as the second level; Use the response information corresponding to the second level as the reference response information.

4. A method for processing LLM text according to claim 2, characterized in that, The preliminary inspection method further includes: Count the known words on the last path in the first level to screen out the response information corresponding to the most known words and define it as the first level, and define the response information corresponding to the remaining known words as the third level; Use the response information corresponding to the third level as the reference response information.

5. An LLM text processing method according to claim 4, characterized in that, The preliminary inspection method further includes: Define the response information on the last path that contains two known words and does not contain unknown words in the preset text set as the third level.

6. A method for processing LLM text according to claim 1, characterized in that, The preliminary inspection method further includes: Modify the final question information according to the query word with the preset missing quantity.

Citation Information

Patent Citations

  • Intelligent customer service question and answer method based on large language model technology

    CN118364084A

  • Visual data processing method and device based on large language model, equipment and medium

    CN119248921A

Cited By

  • LLM output stability control method and system based on reinforcement learning

    CN120804310A

  • Reinforcement learning-based llm output stability control method and system

    CN120804310B