Text search method and device, electronic equipment and storage medium

By generating text tag information as search results, the problem of large storage space required for displaying large amounts of information is solved, enabling text search on personal devices and improving user experience.

CN114416920BActive Publication Date: 2026-01-23BEIJING PIXEL SOFTWARE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111622131.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2026-01-23
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

Existing search engines display large amounts of text content that takes up a lot of storage space when faced with large amounts of information, and are not suitable for individuals or scenarios without large devices.

Method used

By generating text tag information as search results, storage space usage is reduced. Tag information, including categories and summaries, is generated using a preset scoring mechanism and classifier, and the text identifier and tag information are displayed.

Benefits of technology

Enables text search without the need for large devices, saving storage space and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114416920B_ABST
    Figure CN114416920B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a text search method and device, electronic equipment and storage medium, and relate to the technical field of retrieval. First, according to a plurality of word data, an index library and preset search information, all texts containing the preset search information are filtered to obtain at least one first text; then, the at least one first text is sorted to obtain at least one second text; and according to the category and content of the second text, label information corresponding to each second text is generated; finally, the identification and label information corresponding to each second text are taken as search results and displayed. According to the category and content of the filtered texts, the corresponding label information is generated, and the texts and the label information are taken as the search results for display, thereby saving storage space, and text search can be realized without large equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of retrieval technology, and more specifically, to a text search method, apparatus, electronic device, and storage medium. Background Technology

[0002] Currently, all industries need to process a large amount of information, and many projects also require the storage, classification and retrieval of information. Search engines are particularly important when dealing with large amounts of information.

[0003] Due to the sheer volume of information, search results are numerous, often numbering in the tens of thousands. Displaying these results involves not only a large amount of text but also significant storage space requirements. Therefore, existing search engines often place high demands on equipment, making them suitable only for large enterprises or projects. For individuals or those without access to large-scale equipment, more miniaturized search technologies are needed. Summary of the Invention

[0004] The objectives of this invention include, for example, providing a text search method, apparatus, electronic device, and storage medium that can generate corresponding tag information based on the category and content of the filtered text, and display the text and tag information as search results, thereby saving storage space and improving user experience.

[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows:

[0006] In a first aspect, embodiments of the present invention provide a text search method applied to an electronic device, wherein the electronic device pre-stores multiple texts, each text having a corresponding identifier; the method includes:

[0007] Based on multiple word data, an index library, and preset search information, at least one first text is obtained, wherein the word data is obtained by dividing the multiple texts, the index library is used to represent the mapping relationship between each word data and at least one text corresponding to each word data, and each first text includes the preset search information;

[0008] According to a preset scoring mechanism, the at least one first text is sorted to obtain at least one second text;

[0009] Generate tag information for each of the second texts, wherein the tag information is used to characterize the category and content of the second text;

[0010] The identifier corresponding to each of the second texts and the tag information corresponding to each of the second texts are used as search results and displayed.

[0011] In one possible implementation, prior to the step of obtaining at least one first text based on multiple word data, an index, and preset search information, the method further includes:

[0012] Using a preset word segmentation tool, the multiple texts are divided to obtain multiple word data, wherein each word data has at least one corresponding text.

[0013] The index library is established for the multiple word data.

[0014] In one possible implementation, the step of sorting the at least one first text according to a preset scoring mechanism to obtain at least one second text includes:

[0015] According to the preset scoring mechanism, each of the first texts is scored to obtain a score corresponding to each of the first texts, wherein the score represents the frequency of the preset search information appearing in the first texts;

[0016] The at least one first text is sorted according to the scores from largest to smallest to obtain the at least one second text.

[0017] In one possible implementation, the step of generating tag information corresponding to each of the second texts includes:

[0018] According to a preset classifier, the at least one second text is classified to generate classification information corresponding to each second text, wherein the classification information is used to characterize the category of the second text;

[0019] Generate summary information for each of the second texts to obtain the tag information for each of the second texts, wherein the tag information includes the classification information and the summary information, and the summary information is used to characterize the content of the second text.

[0020] In one possible implementation, the step of generating summary information corresponding to each second text according to a preset algorithm includes:

[0021] Take any one of the at least one second texts as the target second text;

[0022] According to the text sorting algorithm, the target second text is extracted to obtain several feature sentences, wherein the several feature sentences are used to characterize the core content of the second text;

[0023] The aforementioned feature statements are used as the summary information corresponding to the target second text;

[0024] Traverse the at least one second text to obtain summary information corresponding to each second text.

[0025] Secondly, embodiments of the present invention also provide a text search device applied to an electronic device, wherein the electronic device pre-stores multiple texts, and the text search device includes:

[0026] The acquisition module is used to acquire at least one first text based on multiple word data, an index library, and preset search information. The word data is obtained by dividing the multiple texts. The index library is used to represent the mapping relationship between each word data and at least one text corresponding to each word data. Each first text includes the preset search information.

[0027] The sorting module is used to sort the at least one first text according to a preset scoring mechanism to obtain at least one second text;

[0028] A generation module is used to generate tag information for each of the second texts, wherein the tag information is used to characterize the category and content of the second text;

[0029] The display module is used to obtain search results based on the at least one second text and the tag information corresponding to each second text.

[0030] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:

[0031] One or more processors;

[0032] A memory for storing one or more programs that, when executed by one or more processors, enable the one or more processors to implement the text search method described above.

[0033] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described text search method.

[0034] Compared to existing technologies, the text search method, apparatus, electronic device, and storage medium provided in this invention firstly filter out all texts containing the preset search information based on multiple word data, an index, and preset search information to obtain at least one first text; then, sort the at least one first text to obtain at least one second text; and generate tag information corresponding to each second text based on its category and content; finally, display the identifier and tag information corresponding to each second text as the search result. This method generates corresponding tag information based on the category and content of the filtered text and displays the text and tag information as the search result, thereby saving storage space and enabling text search without large-scale equipment. Attached Figure Description

[0035] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 A block diagram of an electronic device provided in an embodiment of the present invention.

[0037] Figure 2 This is a flowchart illustrating a text search method provided in an embodiment of the present invention.

[0038] Figure 3 This is a schematic diagram of the index library provided in an embodiment of the present invention.

[0039] Figure 4 An example image of a search results page provided in an embodiment of the present invention.

[0040] Figure 5 An example image of the details page of search results provided in an embodiment of the present invention.

[0041] Figure 6 This is another flowchart illustrating the text search method provided in an embodiment of the present invention.

[0042] Figure 7 for Figure 2 The flowchart of step S120 in the text search method is shown.

[0043] Figure 8 for Figure 2 The flowchart of step S130 in the text search method is shown.

[0044] Figure 9 for Figure 8 The flowchart of step S1302 in the text search method is shown.

[0045] Figure 10 A block diagram illustrating a text search device provided in an embodiment of the present invention.

[0046] Icons: 100 - Electronic device; 101 - Memory; 102 - Processor; 103 - Bus; 200 - Text search device; 201 - Acquisition module; 202 - Sorting module; 203 - Generation module; 204 - Display module. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0048] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0049] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0050] It should be noted that, where there is no conflict, the features in the embodiments of the present invention can be combined with each other.

[0051] Currently, all industries need to process a large amount of information, and many projects also require the storage, classification and retrieval of information. Search engines are particularly important when dealing with large amounts of information.

[0052] In traditional technology, search engines are often used for searching web pages on the internet. Given the vast number of web pages and the enormous amount of stored information, filtering out the information a user wants from this massive amount of data generally involves the following process:

[0053] First, web page information is discovered and collected from the Internet and stored in a local database. Then, the information is extracted and organized to build an index. Next, the search engine quickly retrieves documents from the index based on the user's input query keywords, evaluates the relevance of the documents to the query, sorts the results to be output, and returns the query results to the user.

[0054] Because web pages contain a huge amount of information and are updated in real time, the number of documents retrieved is often in the tens of thousands. Traditional search methods sort the search results and then display the sorted results directly on the page. The displayed content includes the document title and the first paragraph or the first few paragraphs of the document, which requires a lot of storage space. Therefore, traditional search methods are often based on large-scale equipment and are only suitable for large enterprises or projects. For individuals or those who do not have large-scale equipment, more miniaturized search technology is needed.

[0055] To address this issue, this embodiment provides a text search method that generates corresponding tag information based on the content of the filtered text, and displays the filtered text and tag information as search results. This results in search results occupying less storage space, thus enabling text search without large-scale equipment.

[0056] The following is a detailed introduction.

[0057] Please refer to Figure 1 , Figure 1 The diagram shows a block illustration of an electronic device 100 provided in this embodiment. The electronic device 100 may be, but is not limited to, a mobile phone, tablet computer, laptop computer, server, or other electronic device with processing capabilities. The electronic device 100 includes a memory 101, a processor 102, and a bus 103. The memory 101 and the processor 102 are connected via the bus 103.

[0058] The memory 101 is used to store programs, such as the text search device 200. It should be noted that the text search device 200 in this embodiment is a modified version of a traditional search engine.

[0059] The text search device 200 includes at least one software function module that can be stored in the memory 101 in the form of software or firmware. After receiving an execution instruction, the processor 102 executes the program to implement the text search method in this embodiment.

[0060] The memory 101 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0061] The processor 102 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the text search method in this embodiment can be completed by the integrated logic circuitry in the processor 102 or by software instructions.

[0062] The processor 102 mentioned above can be a general-purpose processor, including a central processing unit (CPU), a microcontroller unit (MCU), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), an embedded ARM chip, etc.

[0063] exist Figure 1 Based on the electronic device 100 shown, the text search method provided in this embodiment will be introduced. Please refer to... Figure 2 , Figure 2 A flowchart illustrating the text search method provided in this embodiment is shown. This method is applied to an electronic device 100, which pre-stores multiple texts, each with a corresponding identifier. The method includes the following steps:

[0064] S110, based on multiple word data, an index library, and preset search information, obtain at least one first text, wherein the word data is obtained by dividing multiple texts, the index library is used to represent the mapping relationship between each word data and at least one text corresponding to each word data, and each first text includes preset search information.

[0065] In this embodiment, the text may be article data stored in the local database of the electronic device 100, or it may be article data obtained from the network.

[0066] The text can be regarded as a set of words composed of several word data. Among them, word data refers to the smallest word element that makes up the text.

[0067] For example, for an English text, an English word is a word data, such as "we"; for a Chinese text, a word segmenter needs to be used to perform word segmentation to obtain word data. The obtained word data may consist of a single Chinese character, such as "我", or may consist of multiple Chinese characters, for example, "我们".

[0068] The index library is an index database established by the inverted index method based on word data and the text corresponding to the word data, including a dictionary composed of word data and an inverted file. The inverted file is used to store inverted lists. An inverted list is a correspondence table between word data and at least one text corresponding to the word data. The dictionary is used to store the word data and the position of the inverted list corresponding to the word data in the inverted file.

[0069] It should be noted that a word data may appear in multiple texts at the same time. Therefore, the correspondence stored in the above inverted list is often that one word data corresponds to multiple texts. And here, the text is generally represented by an identifier, and the identifier can be the name of the text, the number of the text, etc.

[0070] For example, Figure 3 is a schematic diagram of the index library, including a dictionary and an inverted file. Each word data in the dictionary has a corresponding inverted list in the inverted file. According to the inverted list, at least one text corresponding to the word data can be obtained.

[0071] Next, taking the preset search information as word data 1 as an example, the process of obtaining at least one first text will be described:

[0072] First, determine word data 1 in the dictionary according to the preset search information. According to the position of the inverted list corresponding to word data 1 in the inverted file, obtain inverted list 1 in the inverted file. Inverted list 1 stores all texts corresponding to word data 1. Obtain all texts in inverted list 1 as at least one first text.

[0073] The preset search information refers to the search information input by the user through the interaction interface of the electronic device 100. It can be a single word or a combination of multiple words. When it is a combination of multiple words, a separator needs to be used to separate the multiple words. <000's0170>

[0074] The first text refers to all texts that contain the preset search information. Taking the preset search information as a single word as an example, the process of obtaining the first text will be specifically described:

[0075] First, the target word data that matches the preset search information is queried in the dictionary in the index library; then, the inverted list of the target word data is obtained from the inverted file according to the storage location of the inverted list of the target word data; finally, the corresponding text is obtained as the first text according to the identifier in the inverted list.

[0076] S120, according to the preset scoring mechanism, sort at least one first text to obtain at least one second text.

[0077] In this embodiment, the second text is obtained by sorting the first text. It should be noted that the second text and the first text are not essentially different; the distinction is merely for ease of understanding and has no special meaning.

[0078] In the second text, the text that appears earlier in the list is more relevant to the preset search information.

[0079] S130, Generate tag information for each second text, wherein the tag information is used to characterize the category and content of the second text.

[0080] S140, the identifier and tag information corresponding to each second text are used as search results and displayed.

[0081] In this embodiment, the search results are displayed on the webpage in the order in which the second text is arranged. The displayed content includes: the identifier of the second text and the tag information of the second text. The identifier of the second text can be the title or number of the second text. When the user clicks on the identifier of the second text, he / she can obtain the full content of the second text.

[0082] The above web interface is implemented based on the Flask web application framework.

[0083] For example, if you enter the preset search term "universe" in the search box of the search interface, the search results will be as follows: Figure 4 As shown, the current page displays two search results. The order in which the two search results are displayed is determined by the score of the second text. That is, in the first search result "Cosmic Calendar", the word data "Cosmic" appears more frequently than in the second search result.

[0084] The first search result includes: the identifier of the second text "Cosmic Calendar" and the tag information of the second text. The tag information of the second text includes categories and summaries, where the summary is the three most central sentences extracted from the full text of "Cosmic Calendar".

[0085] Clicking the "Cosmic Calendar" icon in the first search result will redirect you to a webpage displaying detailed text content. The webpage interface looks like this: Figure 5as shown

[0086] Compared with the prior art, the text search method provided in this embodiment generates corresponding tag information according to the categories and contents of the filtered texts, and displays the identifiers corresponding to the filtered texts and the tag information as search results, so that the storage space occupied by the search results is smaller. Therefore, it is possible to implement text search without large-scale equipment.

[0087] Based on Figure 2 please refer to Figure 6 Before step S110, the following steps are further included:

[0088] S108. Use a preset word segmentation tool to divide multiple texts to obtain multiple word data, where each word data has at least one corresponding text.

[0089] In this embodiment, the preset word segmentation tool can be the jieba Chinese word segmentation tool. During the word segmentation process, some words without actual meaning are often filtered out. For example, when segmenting the sentence "The weather today is very sunny", the word data "today", "weather", and "sunny" are obtained.

[0090] S109. Establish an index library for multiple word data values.

[0091] Compared with the prior art, this embodiment uses an inverted index method to establish an index library for each word data value, which can quickly determine the text that matches the preset search information from a large amount of word data values, improving the search speed.

[0092] The following is a detailed introduction to step S120. Based on Figure 2 please refer to Figure 7 Step S120 may include the following detailed steps:

[0093] S1201. Score each first text according to a preset scoring mechanism to obtain the score corresponding to each first text, where the score represents the frequency of occurrence of the preset search information in the first text.

[0094] In this embodiment, the preset scoring mechanism can be the TF-IDF (term frequency–inverse document frequency) scoring function. TF-IDF is a commonly used weighting technique for information retrieval and data mining. It is a statistical method used to evaluate the importance of a word for a document in a document set or a corpus.

[0095] For example, if the word "we" appears 10 times in the first text and 5 times in the second text, then the score of the first text is greater than the score of the second text.

[0096] S1202, sort at least one first text according to the score from largest to smallest to obtain at least one second text.

[0097] In this embodiment, among the at least one second text obtained, the second text that appears earlier in the sorting indicates that the preset search information appears more frequently in the second text, so that users can more easily obtain the text they want, thus improving the user experience.

[0098] The following is a detailed description of step S130. Figure 2 Based on this, please refer to Figure 8 Step S130 may include the following detailed steps:

[0099] S1301, According to a preset classifier, at least one second text is classified to generate classification information corresponding to each second text, wherein the classification information is used to characterize the category of the second text.

[0100] In this embodiment, the preset classifier can be a Naive Bayes classifier. Before using the Naive Bayes classifier to classify at least one second text, the Naive Bayes classifier needs to be trained. The Naive Bayes classifier is trained by dividing all the text stored in the electronic device 100 into training and test sets in a 7:3 ratio using the train_test_split (split training and test sets) function.

[0101] The trained classifier is used to classify all second texts. The classification criteria can be the domain to which the content of the second text belongs, and the generated classification information can be the name of the domain to which the content of the second text belongs, such as "military", "economics", "literature", etc.

[0102] S1302, Generate summary information for each second text and obtain tag information for each second text. The tag information includes classification information and summary information. The summary information is used to characterize the content of the second text.

[0103] In this embodiment, the summary information is information extracted from the content of the second text, which is usually a short sentence that can reflect the central content of the second text.

[0104] Compared to existing technologies, the text search method in this embodiment, after obtaining at least one second text, also generates corresponding tag information based on the content of each second text. Users can use the category information and summary information to determine whether it is content they are interested in, and then decide whether to further obtain the full text, thereby improving the user experience.

[0105] The following is a detailed description of step S1302. Figure 8 Based on this, please refer to Figure 9 Step S1302 may include the following detailed steps:

[0106] S13021, take any one of the at least one second texts as the target second text.

[0107] S13022, Based on the text sorting algorithm, extract the target second text to obtain several feature sentences, among which several feature sentences are used to characterize the core content of the second text.

[0108] In this embodiment, the text ranking algorithm refers to the extractive, unsupervised text summarization method TextRank, which can extract the keywords and keyword groups of a given text and use an extractive automatic summarization method to extract the key sentences of the text.

[0109] Generally, the extracted characteristic sentences are the three most central sentences in the article.

[0110] S13023, several feature statements are used as summary information corresponding to the target second text.

[0111] S13024, traverse at least one second text to obtain each summary information corresponding to each second text.

[0112] Compared to existing technologies, the text search method provided in this embodiment extracts several key feature sentences from the text as summary information. Users can understand the general content of the text by browsing the summary information, saving users time and improving the user experience.

[0113] Compared with the prior art, this embodiment has the following beneficial effects:

[0114] First, the method provided in this embodiment generates corresponding tag information based on the content of the filtered text, and displays the text and tag information as search results, thereby saving storage space and enabling text search without large equipment.

[0115] Then, using a preset classifier, classification information for the second text is generated. Based on the text sorting algorithm, summary information is generated. Users can use the classification information and summary information to determine whether the content is of interest to them and then decide whether to obtain the full text, thereby saving users time and improving the user experience.

[0116] Please refer to Figure 10 , Figure 10 A block diagram of the text search device 200 provided in this embodiment is shown. The text search device 200 is applied to the electronic device 100 and includes: an acquisition module 201, a sorting module 202, a generation module 203, and a display module 204.

[0117] The acquisition module 201 is used to acquire at least one first text based on multiple word data, an index library and preset search information. The word data is obtained by dividing multiple texts, the index library is used to represent the mapping relationship between each word data and at least one text corresponding to each word data, and each first text includes preset search information.

[0118] The sorting module 202 is used to sort at least one first text according to a preset scoring mechanism to obtain at least one second text.

[0119] The generation module 203 is used to generate tag information for each second text, wherein the tag information is used to characterize the category and content of the second text.

[0120] The display module 204 is used to display the identifier and tag information corresponding to each second text as search results.

[0121] Optionally, module 201 is also used for:

[0122] Using a preset word segmentation tool, multiple texts are divided into multiple word data, where each word data has at least one corresponding text.

[0123] Build an index for multiple word data.

[0124] Optional, sorting module 202, specifically used for:

[0125] According to the preset scoring mechanism, each first text is scored to obtain the score corresponding to each first text. The score represents the frequency of the preset search information in the first text.

[0126] Sort at least one first text according to the score from largest to smallest to obtain at least one second text.

[0127] Optionally, module 203 is generated, specifically for:

[0128] According to a preset classifier, at least one second text is classified, and classification information corresponding to each second text is generated, wherein the classification information is used to characterize the category of the second text;

[0129] Generate summary information for each second text and obtain tag information for each second text. The tag information includes classification information and summary information. The summary information is used to characterize the content of the second text.

[0130] Optionally, module 203 is generated, specifically for:

[0131] Take any one of the at least two second texts as the target second text;

[0132] Based on the text sorting algorithm, the target second text is extracted to obtain several feature sentences, among which several feature sentences are used to characterize the core content of the second text;

[0133] Several feature statements are used as the summary information corresponding to the target second text;

[0134] Iterate through at least one second text to obtain the summary information for each second text.

[0135] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the text search device 200 described above is as follows: Refer to the corresponding processes in the foregoing method embodiments; further details will not be repeated here.

[0136] This embodiment also provides a computer-readable storage medium storing a computer program thereon, which, when executed by the processor 102, implements the text search method disclosed in the above embodiment.

[0137] In summary, the text search method, apparatus, electronic device, and storage medium provided by this invention firstly filter out all texts containing the preset search information based on multiple word data, an index, and preset search information to obtain at least one first text; then, sort the at least one first text to obtain at least one second text; and generate tag information corresponding to each second text based on its category and content; finally, display the identifier and tag information corresponding to each second text as the search result. This method generates corresponding tag information based on the category and content of the filtered text and displays the text and tag information as the search result, thereby saving storage space and enabling text search without large-scale equipment.

[0138] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A text search method, characterized in that, Applied to an electronic device, the electronic device pre-stores multiple texts, each of which has a corresponding identifier; the method includes: Based on multiple word data, an index library, and preset search information, at least one first text is obtained, wherein the word data is obtained by dividing the multiple texts, the index library is used to represent the mapping relationship between each word data and at least one text corresponding to each word data, and each first text includes the preset search information; According to a preset scoring mechanism, the at least one first text is sorted to obtain at least one second text, including: According to the preset scoring mechanism, each of the first texts is scored to obtain a score corresponding to each of the first texts, wherein the score represents the frequency of the preset search information appearing in the first texts; The at least one first text is sorted according to the scores from largest to smallest to obtain the at least one second text; Generate tag information for each of the second texts, wherein the tag information is used to characterize the category and content of the second text, and generating tag information for each of the second texts includes: According to a preset classifier, the at least one second text is classified to generate classification information corresponding to each second text, wherein the classification information is used to characterize the category of the second text; Generate summary information for each second text to obtain tag information for each second text, wherein the tag information includes the classification information and the summary information, and the summary information is used to characterize the content of the second text; generating summary information for each second text includes: Take any one of the at least one second texts as the target second text; According to the text sorting algorithm, the target second text is extracted to obtain several feature sentences, wherein the several feature sentences are used to characterize the core content of the second text; The aforementioned feature statements are used as the summary information corresponding to the target second text; Traverse the at least one second text to obtain the summary information corresponding to each second text; The identifier corresponding to each of the second texts and the tag information corresponding to each of the second texts are used as search results and displayed.

2. The method according to claim 1, characterized in that, Before the step of obtaining at least one first text based on multiple word data, an index, and preset search information, the method further includes: Using a preset word segmentation tool, the multiple texts are divided to obtain multiple word data, wherein each word data has at least one corresponding text. The index library is established for the multiple word data.

3. A text search device, characterized in that, This is applied to an electronic device, which pre-stores multiple texts, each of which has a corresponding identifier; The text search device includes: The acquisition module is used to acquire at least one first text based on multiple word data, an index library, and preset search information. The word data is obtained by dividing the multiple texts. The index library is used to represent the mapping relationship between each word data and at least one text corresponding to each word data. Each first text includes the preset search information. The sorting module is used to sort the at least one first text according to a preset scoring mechanism to obtain at least one second text; the sorting module is also used to: According to the preset scoring mechanism, each of the first texts is scored to obtain a score corresponding to each of the first texts, wherein the score represents the frequency of the preset search information appearing in the first texts; The at least one first text is sorted according to the scores from largest to smallest to obtain the at least one second text; A generation module is used to generate tag information corresponding to each of the second texts, wherein the tag information is used to characterize the category and content of the second text, and the generation module is further used to: According to a preset classifier, the at least one second text is classified to generate classification information corresponding to each second text, wherein the classification information is used to characterize the category of the second text; Generate summary information for each of the second texts to obtain the tag information for each of the second texts, wherein the tag information includes the classification information and the summary information, and the summary information is used to characterize the content of the second text; the generation module is further configured to: Take any one of the at least one second texts as the target second text; According to the text sorting algorithm, the target second text is extracted to obtain several feature sentences, wherein the several feature sentences are used to characterize the core content of the second text; The aforementioned feature statements are used as the summary information corresponding to the target second text; Traverse the at least one second text to obtain the summary information corresponding to each second text; The display module is used to display the identifier corresponding to each of the second texts and the tag information corresponding to each of the second texts as search results.

4. The apparatus according to claim 3, characterized in that, The acquisition module is also used for: Using a preset word segmentation tool, the multiple texts are divided to obtain multiple word data, wherein each word data has at least one corresponding text. The index library is established for the multiple word data.

5. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the text search method as described in any one of claims 1-2.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the text search method as described in any one of claims 1-2.

Citation Information

Patent Citations

  • Information display method and device and information search method and device

    CN111859195A

  • Text abstract extraction method and device

    CN113342968A