A method and system for retrieving large-capacity text content

By generating a keyword search table containing offsets, the problem that traditional text search is not applicable to super-large text files is solved, and fast and efficient retrieval is achieved.

CN114218373BActive Publication Date: 2025-05-27CHINA TELECOM DIGITAL INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111555700.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-17
Publication Date
2025-05-27
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

The traditional text retrieval method is not applicable to super large text files, and the workload of traversal and comparison is large and resource consumption is large.

Method used

By extracting keywords containing offsets in text information, a search table with keywords is generated. When a search request is received, the search table is matched with the keywords in the search entry, the corresponding offset is found, the target information is determined and displayed.

Benefits of technology

There is no need to traverse text information, and the search speed is extremely fast, suitable for super large text files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114218373B_ABST
    Figure CN114218373B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data retrieval, and specifically discloses a large-capacity text content retrieval method and system. The method includes intercepting the stored text information according to a preset interval word length to obtain a text to be retrieved; extracting keywords of the text to be retrieved, and inserting the label of the text to be retrieved into the keywords; counting the keywords containing labels to obtain a query table sorted based on the labels; wherein, the query table includes a keyword item and a corresponding frequency item; wherein, the keyword further includes an offset relative to the head byte of the text to be retrieved. By extracting keywords with offsets of the text information, the present invention generates a retrieval table with keywords as the content. When a retrieval request containing a retrieval term is received, according to the keyword matching in the retrieval term, the corresponding offset is found, the target information is determined and displayed, without traversing the text information, and the retrieval speed is extremely fast.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data retrieval, and specifically to a method and system for retrieving large-capacity text content. Background Art

[0002] The traditional retrieval method is to traverse the text, and continuously compare during the traversal process to obtain the retrieved content. However, for extremely large text files, the workload of traversing once is extremely large, and the traditional retrieval method is obviously not applicable. Therefore, a retrieval method dedicated to extremely large text files is needed. Summary of the Invention

[0003] An object of the present invention is to provide a method and system for retrieving large-capacity text content to solve the problems raised in the above background art.

[0004] To achieve the above object, the present invention provides the following technical solutions:

[0005] A method for retrieving large-capacity text content includes:

[0006] Determine a reading node according to a preset interval word length;

[0007] Intercept the stored text information according to the reading node to obtain a text to be inspected with labels;

[0008] Extract the keywords of the text to be inspected, and insert the label of the text to be inspected into the keywords;

[0009] Count the keywords with labels to obtain a query table sorted based on labels; wherein, the query table includes keyword items and corresponding frequency items;

[0010] When receiving a retrieval request containing a retrieval term, extract the keywords in the retrieval term, traverse the query table based on the keywords in the retrieval term, and determine and display the target information;

[0011] Among them, the keyword further includes an offset relative to the byte at the head of the text to be inspected.

[0012] As a further limitation of the technical solution of the present invention: the step of determining a reading node according to a preset interval word length includes:

[0013] Randomly determine an interval word length, and randomly intercept a text to be inspected from a preset reference text based on the interval word length;

[0014] Extract keywords from the text to be inspected to obtain the unit extraction time;

[0015] Calculate the total extraction time according to the unit extraction time, and determine the preset interval word length according to the total extraction time.

[0016] As a further limitation of the technical solution of the present invention: the step of intercepting the stored text information according to the reading node to obtain the text to be inspected containing labels includes:

[0017] Sort the reading nodes, and generate labels in a mapping relationship with the reading nodes according to the sorting result;

[0018] Read the text information, and intercept the text information with the reading node as an end point to obtain the text to be inspected;

[0019] Obtain the label of the reading node at the head of the text to be inspected, and insert the label into the text to be inspected.

[0020] As a further limitation of the technical solution of the present invention: the step of extracting the keywords of the text to be inspected includes:

[0021] Traverse the text to be inspected, locate the whitespace characters, and convert the text to be inspected into a multi-segment text array based on the whitespace characters;

[0022] Obtain the array lengths of the multi-segment text arrays in sequence, and compare the array lengths with a preset length threshold;

[0023] When the array length is less than the length threshold, extract the content in the corresponding text array as the keyword;

[0024] When the array length is greater than the length threshold, perform content recognition on the corresponding text array and extract keywords.

[0025] As a further limitation of the technical solution of the present invention: the step of, when the array length is greater than the length threshold, performing content recognition on the corresponding text array and extracting keywords includes:

[0026] Input the text array into a trained part-of-speech analysis model to obtain a preprocessed text with part-of-speech tags;

[0027] Remove the function words in the preprocessed text to obtain a preliminary screening text;

[0028] Traverse the modifiers in the preliminary screening text, obtain the common usage degree of the modifiers based on the thesaurus, and mark the modifiers with a common usage degree less than a preset common usage degree threshold;

[0029] Read the adjacent main words according to the marked modifiers as the keywords;

[0030] Among them, the modifiers include adjectives and adverbs, and the main words include nouns and verbs.

[0031] As a further limitation of the technical solution of the present invention: the step of statistically analyzing the keywords with labels to obtain a query table sorted based on labels includes:

[0032] Read the keywords with labels, classify the keywords according to the labels, and obtain a sub-word library named after the labels;

[0033] Traverse the sub-word library to generate a sub-query table with labels, where the sub-query table includes keywords and their repetition times;

[0034] Connect the sub-query tables according to the label order to generate a query table.

[0035] As a further limitation of the technical solution of the present invention: the method further includes:

[0036] Receive the feedback information of the user, and obtain the expected word length according to the feedback information;

[0037] Correct the preset interval word length according to the expected word length.

[0038] The technical solution of the present invention also provides a large-capacity text content retrieval system, and the system includes:

[0039] A node determination module for determining a read node according to a preset interval word length;

[0040] An interception module for intercepting the stored text information according to the read node to obtain a text to be inspected with labels;

[0041] A label insertion module for extracting the keywords of the text to be inspected and inserting the labels of the text to be inspected into the keywords;

[0042] A query table generation module for statistically analyzing the keywords with labels to obtain a query table sorted based on labels; wherein, the query table includes a keyword item and a corresponding frequency item;

[0043] A retrieval module for, when receiving a retrieval request containing a retrieval term, extracting the keywords in the retrieval term, traversing the query table based on the keywords in the retrieval term, determining the target information and displaying it;

[0044] Wherein, the keyword further includes an offset relative to the byte at the head of the text to be inspected.

[0045] As a further limitation of the technical solution of the present invention: the label insertion module includes:

[0046] A conversion unit for traversing the text to be inspected, locating the whitespace characters, and converting the text to be inspected into a multi-segment text array based on the whitespace characters;

[0047] A comparison unit for sequentially obtaining the array lengths of the multi-segment text arrays and comparing the array lengths with a preset length threshold;

[0048] An extraction unit for extracting the content in the corresponding text array as keywords when the array length is less than the length threshold;

[0049] A content recognition unit for performing content recognition on the corresponding text array and extracting keywords when the array length is greater than the length threshold.

[0050] As a further limitation of the technical solution of the present invention: the content recognition unit includes:

[0051] A part-of-speech analysis subunit for inputting the text array into a trained part-of-speech analysis model to obtain a preprocessed text with part-of-speech tags;

[0052] A removal subunit for removing function words in the preprocessed text to obtain a preliminary screening text;

[0053] A marking subunit for traversing the modifiers in the preliminary screening text, obtaining the common usage degree of the modifiers based on a thesaurus, and marking the modifiers with a common usage degree less than a preset common usage degree threshold;

[0054] A reading subunit for reading adjacent main words according to the marked modifiers as keywords;

[0055] Wherein, the modifiers include adjectives and adverbs, and the main words include nouns and verbs.

[0056] Compared with the prior art, the beneficial effect of the present invention is that the traditional retrieval method traverses the text. However, for extremely large text files, the workload of traversing once is extremely large and the resource consumption is relatively high. The present invention extracts keywords with offsets from text information to generate a retrieval table with keywords as the content. When a retrieval request containing a retrieval term is received, it matches the keywords in the retrieval term, finds the corresponding offset, determines the target information and displays it, without traversing the text information, and the retrieval speed is extremely fast. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention.

[0058] Figure 1 Shows a flowchart of a large-capacity text content retrieval method;

[0059] Figure 2Shows the first sub - process block diagram of the large - capacity text content retrieval method;

[0060] Figure 3 Shows the second sub - process block diagram of the large - capacity text content retrieval method;

[0061] Figure 4 Shows the third sub - process block diagram of the large - capacity text content retrieval method;

[0062] Figure 5 Shows the fourth sub - process block diagram of the large - capacity text content retrieval method;

[0063] Figure 6 Shows the fifth sub - process block diagram of the large - capacity text content retrieval method;

[0064] Figure 7 Shows the composition structure block diagram of the large - capacity text content retrieval system;

[0065] Figure 8 Shows the composition structure block diagram of the label insertion module in the large - capacity text content retrieval system;

[0066] Figure 9 Shows the composition structure block diagram of the content recognition unit in the label insertion module. Detailed implementation mode

[0067] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0068] Embodiment 1

[0069] Figure 1 Shows the flow block diagram of the large - capacity text content retrieval method. In the embodiment of the present invention, a large - capacity text content retrieval method is provided, including steps S100 to S500:

[0070] Step S100: Determine the reading node according to a preset interval word length;

[0071] Step S200: Intercept the stored text information according to the reading node to obtain the text to be inspected containing labels;

[0072] The purpose of steps S100 to S200 is to "split" the text to be inspected to obtain the text to be inspected. It can also be understood that a large file is split into multiple small files. As can be imagined, operating on small files is much easier than on large files. However, in the process of splitting the stored text information into the text to be inspected, it is also necessary to be able to synthesize the text information based on the text to be inspected. This means that each text to be inspected needs to be marked, that is, the label mentioned above. With the help of the label, regardless of the storage method of the text to be inspected, the text information can be synthesized based on the text to be inspected.

[0073] Step S300: Extract the keywords of the text to be inspected and insert the label of the text to be inspected into the keywords.

[0074] For the keywords extracted from a certain text to be inspected, the source needs to be clarified.

[0075] Step S400: Count the keywords containing labels to obtain a query table sorted based on labels; wherein, the query table includes keyword items and corresponding frequency items.

[0076] The query table generated in step S400 is a combination of multiple sub-tables. In each sub-table, there are keyword items and their corresponding repetition frequency items. The purpose of the repetition frequency item is to sort the keywords. When a retrieval request is received, we should start the retrieval from the keywords with more repetitions.

[0077] Step S500: When a retrieval request containing a retrieval term is received, extract the keywords in the retrieval term, traverse the query table based on the keywords in the retrieval term, determine the target information and display it.

[0078] The retrieval term may be a sentence, and the retrieval principle of the technical solution of the present invention is the comparison between words. Therefore, it is necessary to extract the keywords in the retrieval term.

[0079] It is worth mentioning that among the keywords mentioned above, there is also an offset relative to the head byte of the text to be inspected. The offset is an address, which may not be displayed. For example, in the final display process, only the specific characters are displayed, not the address. However, from the perspective of the computer, the keyword defined by the technical solution of the present invention is a data type including the address. For example, if the word "novel" is located at a

[30] in the computer language, then 30 can be used as the offset. If the head element of a certain text to be inspected is a

[20] , then the offset can be 10.

[0080] Figure 2 Shows the first sub-flow block diagram of the large-capacity text content retrieval method. The step of determining the reading node according to the preset interval word length includes steps S101 to S103:

[0081] Step S101: Randomly determine the interval word length, and randomly extract a text to be inspected from a preset reference text based on the interval word length;

[0082] Step S102: Extract keywords from the text to be inspected to obtain the unit extraction time;

[0083] Step S103: Calculate the total extraction time according to the unit extraction time, and determine the preset interval word length according to the total extraction time.

[0084] The purpose of Steps S101 to S103 is to generate a rough interval word length. The concept of the interval word length is very simple, which is the length of the text to be inspected. Its significance lies in that the user hopes to intercept a large text into small texts of a certain length.

[0085] Furthermore, a text to be inspected requires a certain amount of time to go through the various processes in the technical solution of the present invention. Then, the number of texts to be inspected is calculated according to the length of the large text and the interval word length, and then the total extraction time can be calculated; obviously, the total extraction time is related to the specific division method of the large text, that is, the above-mentioned interval word length; the optimal interval word length for different hardware devices is also different, but we do not need to determine the optimal interval word length, we only need to randomly determine several interval word lengths to obtain a better interval word length.

[0086] Figure 3 The second sub-process block diagram of the large-capacity text content retrieval method is shown. The step of intercepting the stored text information according to the reading node to obtain the text to be inspected with labels includes Steps S201 to S203:

[0087] Step S201: Sort the reading nodes, and generate labels in a mapping relationship with the reading nodes according to the sorting result;

[0088] Step S202: Read the text information, and intercept the text information with the reading node as the end point to obtain the text to be inspected;

[0089] Step S203: Obtain the label of the reading node at the head of the text to be inspected, and insert the label into the text to be inspected.

[0090] Steps S201 to S203 provide specific steps to obtain the text to be inspected with labels. It is worth mentioning that the label of the reading node at the head of the text to be inspected is used as the label of the text to be inspected because in the specific program design process, there is generally a flag at the tail element of the text file, and it will be very troublesome to use it as the label of the text to be inspected.

[0091] Figure 4The third sub-flow chart of the method for retrieving large-capacity text content is shown, wherein the step of extracting keywords of the text to be inspected includes steps S301 to S304:

[0092] Step S301: traverse the text to be checked, locate blank characters, and convert the text to be checked into a multi-segment text array based on the blank characters;

[0093] Step S302: sequentially obtaining the array lengths of the plurality of text arrays, and comparing the array lengths with a preset length threshold;

[0094] Step S303: when the array length is less than the length threshold, extract the content in the corresponding text array as a keyword;

[0095] Step S304: when the array length is greater than the length threshold, content recognition is performed on the corresponding text array to extract keywords.

[0096] In the above content, the content of the text to be inspected is classified based on blank characters. This is because the content between blank characters, if it is short, then it is the title or keyword part of the text, which can naturally be used as a keyword; if it is long, then it needs to be identified.

[0097] Figure 5 The fourth sub-flow chart of the method for retrieving large-capacity text content is shown. When the length of the array is greater than the length threshold, the steps of performing content recognition on the corresponding text array and extracting keywords include:

[0098] Step S3041: input the text array into the trained part-of-speech analysis model to obtain pre-processed text with part-of-speech tags;

[0099] Step S3042: removing function words from the preprocessed text to obtain a preliminary screening text;

[0100] Step S3043: traversing the modifiers of the primary screening text, obtaining the commonness of the modifiers based on the vocabulary, and marking the modifiers whose commonness is less than a preset commonness threshold;

[0101] Step S3044: Read the adjacent main words according to the marked modifier words as keywords;

[0102] The modifiers include adjectives and adverbs, and the main words include nouns and verbs.

[0103] First, input the text array into the trained part-of-speech analysis model, which is quite common in some typing software and its code is also open-source; through the part-of-speech analysis model, preprocessed text with part-of-speech tags can be obtained; in a text, the possibility of function words being keywords is almost zero, so function words need to be removed; then, there are a large number of nouns and verbs. As to which ones can be used as keywords and which ones cannot, it needs to be judged according to their modifiers. If a word has more or more important modifiers, then it can be used as a keyword; for example: "a magnificent counter" and "many counters". In the above two descriptions, although the subject is the counter, relatively speaking, the counter in the former is more important in its overall text.

[0104] Figure 6 The fifth sub-flow block diagram of the large-capacity text content retrieval method is shown. The steps of statistically counting the keywords with labels and obtaining the query table sorted based on the labels include steps S401 to S403:

[0105] Step S401: Read the keywords with labels and classify the keywords according to the labels to obtain a sub-word library named after the labels;

[0106] Step S402: Traverse the sub-word library to generate a sub-query table with labels. The sub-query table includes keywords and their repetition times;

[0107] Step S403: Connect the sub-query tables according to the label order to generate a query table.

[0108] Steps S401 to S403 are the process of generating a query table. Its essence is a connection process, using the label as the mark of each data, and there is a mapping relationship between various data types with the same label.

[0109] It is worth mentioning that in a preferred embodiment of the technical solution of the present invention, the method further includes:

[0110] Receiving the feedback information of the user and obtaining the expected word length according to the feedback information;

[0111] Correcting the preset interval word length according to the expected word length.

[0112] The above content is a supplementary solution to the technical solution of the present invention. By correcting the interval word length by the user, the interval word length is the size of the content observed by the user. For example, when the user inputs a word, if the retrieval is successful, the size of the content displayed by the system is the size of the interval word length. If it is too large, although the exact offset is known, the reading process is still difficult. Therefore, according to the feedback information of the user, a more appropriate interval word length is determined, and then the used interval word length can be corrected regularly according to the expected word length.

[0113] Example 2

[0114] Figure 7 The composition structure block diagram of a large-capacity text content retrieval system is shown. In an embodiment of the present invention, a large-capacity text content retrieval system is further provided. The system 10 includes:

[0115] A node determination module 11, configured to determine a reading node according to a preset interval word length;

[0116] An interception module 12, configured to intercept the stored text information according to the reading node to obtain a text to be inspected containing labels;

[0117] A label insertion module 13, configured to extract keywords of the text to be inspected and insert the label of the text to be inspected into the keywords;

[0118] A query table generation module 14, configured to count the keywords containing labels to obtain a query table sorted based on labels; wherein, the query table includes a keyword item and a corresponding frequency item;

[0119] A retrieval module 15, configured to, when receiving a retrieval request containing a retrieval term, extract the keyword in the retrieval term, traverse the query table based on the keyword in the retrieval term, determine the target information and display it;

[0120] Wherein, the keyword further includes an offset relative to the head byte of the text to be inspected.

[0121] Figure 8 The composition structure block diagram of the label insertion module in the large-capacity text content retrieval system is shown. The label insertion module 13 includes:

[0122] A conversion unit 131, configured to traverse the text to be inspected, locate whitespace characters, and convert the text to be inspected into a multi-segment text array based on the whitespace characters;

[0123] A comparison unit 132, configured to sequentially obtain the array lengths of the multi-segment text arrays and compare the array lengths with a preset length threshold;

[0124] An extraction unit 133, configured to, when the array length is less than the length threshold, extract the content in the corresponding text array as a keyword;

[0125] A content recognition unit 134, configured to, when the array length is greater than the length threshold, perform content recognition on the corresponding text array and extract keywords.

[0126] Figure 9 The composition structure block diagram of the content recognition unit in the label insertion module is shown. The content recognition unit 134 includes:

[0127] A part-of-speech analysis subunit 1341, configured to input the text array into a trained part-of-speech analysis model to obtain a preprocessed text with part-of-speech tags.

[0128] An elimination subunit 1342, configured to eliminate function words in the preprocessed text to obtain a preliminary screening text.

[0129] A marking subunit 1343, configured to traverse the modifiers in the preliminary screening text, obtain the common usage degree of the modifiers based on a thesaurus, and mark the modifiers whose common usage degree is less than a preset common usage degree threshold.

[0130] A reading subunit 1344, configured to read adjacent main words according to the marked modifiers as keywords.

[0131] Wherein, the modifiers include adjectives and adverbs, and the main words include nouns and verbs.

[0132] All functions that the above large-capacity text content retrieval method can achieve are completed by a computer device. The computer device includes one or more processors and one or more memories. At least one program code is stored in the one or more memories, and the program code is loaded and executed by the one or more processors to implement the functions of the large-capacity text content retrieval method.

[0133] The processor fetches instructions one by one from the memory, analyzes the instructions, and then completes corresponding operations according to the requirements of the instructions, generating a series of control commands to make each part of the computer act automatically, continuously and coordinately, becoming an organic whole, realizing the input of the program, the input of data, and the operation and output of results. All arithmetic operations or logical operations generated in this process are completed by the arithmetic unit; the memory includes a read-only memory (ROM), and the read-only memory is used to store computer programs, and a protection device is provided outside the memory.

[0134] Exemplarily, a computer program can be divided into one or more modules. One or more modules are stored in the memory and executed by the processor to complete the present invention. One or more modules can be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.

[0135] Those skilled in the art can understand that the above description of the service device is only an example and does not constitute a limitation on the terminal device. It may include more or fewer components than the above description, or combine some components, or different components. For example, it may include input / output devices, network access devices, buses, etc.

[0136] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The above-mentioned processor is the control center of the above-mentioned terminal device, and connects various parts of the entire user terminal through various interfaces and lines.

[0137] The above-mentioned memory can be used to store computer programs and / or modules. The above-mentioned processor realizes various functions of the above-mentioned terminal device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as information collection template display function, product information release function, etc.); the data storage area can store data created according to the use of the berth status display system (such as product information collection templates corresponding to different product types, product information to be released by different product providers, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, memory, plug-in hard disks, Smart Media Cards (SMCs), Secure Digital (SD) cards, Flash Cards, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0138] If the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the modules / units in the above-mentioned embodiment system of the present invention, it can also be completed by instructing the relevant hardware through a computer program. The above computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can realize the functions of the above-mentioned various system embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0139] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.

[0140] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for retrieving large-capacity text content, characterized in that, the method includes: determining a reading node according to a preset interval word length; intercepting the stored text information according to the reading node to obtain a text to be inspected with labels; extracting keywords of the text to be inspected and inserting the labels of the text to be inspected into the keywords; counting the keywords with labels to obtain a query table sorted based on the labels; wherein, the query table includes keyword items and corresponding frequency items; when receiving a retrieval request containing a retrieval term, extracting the keywords in the retrieval term, traversing the query table based on the keywords in the retrieval term, determining the target information and displaying it; the keyword further includes an offset relative to the byte at the head of the text to be inspected; wherein, the step of intercepting the stored text information according to the reading node to obtain a text to be inspected with labels includes: sorting the reading nodes and generating labels in a mapping relationship with the reading nodes according to the sorting result; reading the text information, intercepting the text information with the reading node as an endpoint to obtain the text to be inspected; acquiring the label of the reading node at the head of the text to be inspected and inserting the label into the text to be inspected.

2. The method for retrieving large-capacity text content according to claim 1, characterized in that, the step of determining a reading node according to a preset interval word length includes: randomly determining an interval word length, and randomly intercepting a text to be inspected from a preset reference text based on the interval word length; extracting keywords from the text to be inspected to obtain a unit extraction time; calculating a total extraction time according to the unit extraction time and determining the preset interval word length according to the total extraction time.

3. The method for retrieving large-capacity text content according to claim 1, characterized in that, the step of extracting keywords of the text to be inspected includes: traversing the text to be inspected, locating whitespace characters, and converting the text to be inspected into a multi-segment text array based on the whitespace characters; sequentially obtaining the array lengths of the multi-segment text arrays and comparing the array lengths with a preset length threshold; when the array length is less than the length threshold, extracting the content in the corresponding text array as a keyword; when the array length is greater than the length threshold, performing content recognition on the corresponding text array and extracting keywords.

4. The method for retrieving large-capacity text content according to claim 3, characterized in that, the step of, when the array length is greater than the length threshold, performing content recognition on the corresponding text array and extracting keywords includes: inputting the text array into a trained part-of-speech analysis model to obtain a preprocessed text with part-of-speech tags; removing function words from the preprocessed text to obtain a preliminary screening text; traversing the modifiers in the preliminary screening text, obtaining the common usage degree of the modifiers based on a thesaurus, and marking the modifiers with a common usage degree less than a preset common usage degree threshold; reading adjacent main words according to the marked modifiers as keywords; wherein, the modifiers include adjectives and adverbs, and the main words include nouns and verbs.

5. The method for retrieving large-capacity text content according to claim 1, It is characterized in that the step of counting the keywords with labels and obtaining a query table sorted based on the labels includes: reading the keywords with labels, classifying the keywords according to the labels, and obtaining a sub-word library named after the labels; traversing the sub-word library to generate a sub-query table with labels, where the sub-query table includes the keywords and their repetition times; connecting the sub-query tables according to the label order to generate a query table.

6. The large-capacity text content retrieval method according to claim 5, It is characterized in that the method further includes: receiving the feedback information of the user, and obtaining the expected word length according to the feedback information; correcting the preset interval word length according to the expected word length.

7. A large-capacity text content retrieval system, It is characterized in that the system includes: a node determination module for determining a reading node according to a preset interval word length; a truncation module for truncating the stored text information according to the reading node to obtain a text to be inspected with labels; a label insertion module for extracting the keywords of the text to be inspected and inserting the labels of the text to be inspected into the keywords; a query table generation module for counting the keywords with labels to obtain a query table sorted based on the labels; wherein, the query table includes a keyword item and a corresponding frequency item; a retrieval module for, when receiving a retrieval request containing a retrieval term, extracting the keywords in the retrieval term, traversing the query table based on the keywords in the retrieval term, determining the target information and displaying it; the keyword further includes an offset relative to the byte at the head of the text to be inspected; wherein, the functions of the truncation module specifically include: sorting the reading nodes, and generating labels in a mapping relationship with the reading nodes according to the sorting result; reading the text information, and truncating the text information with the reading nodes as endpoints to obtain the text to be inspected; obtaining the label of the reading node at the head of the text to be inspected, and inserting the label into the text to be inspected.

8. The large-capacity text content retrieval system according to claim 7, It is characterized in that the label insertion module includes: a conversion unit for traversing the text to be inspected, locating the whitespace characters, and converting the text to be inspected into a multi-segment text array based on the whitespace characters; a comparison unit for sequentially obtaining the array lengths of the multi-segment text arrays, and comparing the array lengths with a preset length threshold; an extraction unit for, when the array length is less than the length threshold, extracting the content in the corresponding text array as the keyword; a content recognition unit for, when the array length is greater than the length threshold, performing content recognition on the corresponding text array and extracting the keyword.

9. The large-capacity text content retrieval system according to claim 8, It is characterized in that the content recognition unit includes: a part-of-speech analysis sub-unit for inputting the text array into a trained part-of-speech analysis model to obtain a preprocessed text with part-of-speech tags; a deletion sub-unit for deleting the function words in the preprocessed text to obtain a preliminary screening text; A marking subunit, configured to traverse the modifiers of the pre-screened text, obtain the commonness of the modifiers based on a thesaurus, and mark the modifiers whose commonness is less than a preset commonness threshold; A reading subunit, configured to read adjacent main words according to the marked modifiers as keywords; Wherein, the modifiers include adjectives and adverbs, and the main words include nouns and verbs.

Citation Information

Patent Citations

  • Efficient inverted index storage method for full-text retrieval

    CN111639151A

  • Methods and systems for providing potential search queries that may be targeted by one or more keywords

    US20150012560A1