Text processing method and device, storage medium, and electronic device
By determining the first text in the text library and performing word segmentation processing, using the index word set to find the relevant text and calculating the correlation score, the problem of low text processing efficiency in the prior art is solved, and the effect of quickly establishing text attribute diagrams is achieved.
Patent Information
- Application Number
- CN202210128335.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-11
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-02-11
AI Technical Summary
When the prior art determines the correlation between articles in online reading, the method is relatively simple and the calculation content is complicated, resulting in low text processing efficiency and inconvenient operation.
By determining the first text from the text library, word segmentation processing is performed to determine the query word, search for the relevant text using the index word set, calculate the correlation score, and determine the correlation text based on the score to create the attribute graph.
It improves the efficiency of text processing, simplifies operations, and can quickly determine the relevant text of the text, thereby facilitating the establishment of attribute diagrams between the text and the related text.
Smart Images

Figure CN114492371B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of computer technology, and more specifically, relate to a text processing method and device, a storage medium, and an electronic device. Background Art
[0002] This section is intended to provide a background or context to the embodiments of the disclosure that are recited in the claims, and no description herein is admitted to be prior art by inclusion in this section.
[0003] With the development of Internet technology, online reading has gradually become a common reading method for more and more users. When reading an article online, users often pay attention to articles related to the article. Therefore, in order to increase the speed and breadth of article dissemination, it is necessary to determine the correlation between different articles and establish an attribute graph between the same article and different articles. In related technologies, the method of establishing an article attribute graph is often relatively simple, and the calculation content is relatively complicated, which makes the efficiency of text processing extremely low and inconvenient to operate. Summary of the invention
[0004] In order to overcome the problems existing in the related art, the present disclosure provides a text processing method and device, a storage medium, and an electronic device.
[0005] According to one aspect of the present disclosure, a text processing method is provided, the method comprising:
[0006] determining a first text from the text library;
[0007] Performing word segmentation processing on the first text to determine a query word corresponding to the first text;
[0008] Determine the candidate text corresponding to the first text by searching the query word in the index word set; the index word set is determined based on the plurality of second texts in the text library;
[0009] Calculating a correlation score between the first text and the candidate text;
[0010] According to the relevance score, an associated text corresponding to the first text is determined from the candidate texts, so as to establish an attribute graph of the first text based on the associated text.
[0011] Optionally, before searching the index word set by using the query word to determine the candidate text corresponding to the first text, the method further includes:
[0012] Determine the second text according to other texts in the text library except the first text;
[0013] Performing word segmentation processing on the second text to generate keywords corresponding to the second text;
[0014] Using a preset full-text search algorithm, an index word corresponding to the second text is established based on the keyword.
[0015] Optionally, the step of establishing an index word corresponding to the second text based on the keyword includes:
[0016] Determining a mapping relationship between the keyword and the second text;
[0017] Generate an inverted index word corresponding to the second text according to the mapping relationship.
[0018] Optionally, the inverted index words include the author, text category, tag words, and title of the second text.
[0019] Optionally, searching the index word set by the query word to determine the candidate text corresponding to the first text includes:
[0020] In the index word set, the inverted index word that is the same as the query word is used as the target index word;
[0021] According to the mapping relationship between the second text and the inverted index term, the second text corresponding to the target index term is used as the candidate text corresponding to the first text.
[0022] Optionally, calculating the correlation score between the first text and the candidate text includes:
[0023] The first text and the candidate text are processed and calculated using a preset text correlation algorithm to determine a correlation score between the first text and the candidate text.
[0024] Optionally, performing word segmentation processing on the first text to determine a query word corresponding to the first text includes:
[0025] Performing word segmentation processing on the first text to obtain a word segmentation result of the first text;
[0026] The word segmentation results are filtered, and the filtered word segmentation results are used as query words corresponding to the first text.
[0027] Optionally, determining the first text from the text library includes:
[0028] For each text in the text library, determining the amount of interaction between the user and the text;
[0029] The text whose interaction amount is less than a preset interaction threshold is used as the first text.
[0030] According to one aspect of the present disclosure, a text processing device is provided, the device comprising:
[0031] A first determining module, used for determining a first text from a text library;
[0032] A first word segmentation module, used to perform word segmentation processing on the first text to determine a query word corresponding to the first text;
[0033] A search module, configured to search an index word set using the query word to determine an alternative text corresponding to the first text; the index word set is determined based on a plurality of second texts in the text library;
[0034] A calculation module, used for calculating a correlation score between the first text and the candidate text;
[0035] The second determination module is used to determine the associated text corresponding to the first text from the candidate texts according to the relevance score, so as to establish an attribute graph of the first text based on the associated text.
[0036] Optionally, the device further comprises:
[0037] A selection module, used for determining the second text according to other texts in the text library except the first text;
[0038] A second word segmentation module, used to perform word segmentation processing on the second text to generate keywords corresponding to the second text;
[0039] An establishing module is used to establish index words corresponding to the second text based on the keywords by using a preset full-text search algorithm.
[0040] Optionally, the establishment module is further used to:
[0041] Determine a mapping relationship between the keyword and the second text;
[0042] Generate an inverted index word corresponding to the second text according to the mapping relationship.
[0043] Optionally, the inverted index words include the author, text category, tag words, and title of the second text.
[0044] Optionally, the search module is further used to:
[0045] In the index word set, the inverted index word that is the same as the query word is used as the target index word;
[0046] According to the mapping relationship between the second text and the inverted index term, the second text corresponding to the target index term is used as the candidate text corresponding to the first text.
[0047] Optionally, the computing module is further used for:
[0048] The first text and the candidate text are processed and calculated using a preset text correlation algorithm to determine a correlation score between the first text and the candidate text.
[0049] Optionally, the first word segmentation module is further used to:
[0050] Performing word segmentation processing on the first text to obtain a word segmentation result of the first text;
[0051] The word segmentation results are filtered, and the filtered word segmentation results are used as query words corresponding to the first text.
[0052] Optionally, the first determining module includes:
[0053] For each text in the text library, determining the amount of interaction between the user and the text;
[0054] The text whose interaction amount is less than a preset interaction threshold is used as the first text.
[0055] According to one aspect of the present disclosure, a storage medium is provided, on which a computer program is stored, and the computer program is the above-mentioned text processing method when executed by a processor.
[0056] According to one aspect of the present disclosure, there is provided an electronic device, including:
[0057] Processor; and
[0058] A memory, configured to store executable instructions of the processor;
[0059] The processor is configured to execute any one of the above-mentioned text processing methods by executing the executable instructions.
[0060] In summary, the text processing method provided by the embodiment of the present disclosure can first determine the first text from the text library, then perform word segmentation processing on the first text, determine the query word corresponding to the first text, search in the index word set through the query word, and determine the alternative text corresponding to the first text. The index word set is determined based on multiple second texts in the text library, and the correlation score between the first text and the alternative text is calculated. Finally, based on the correlation score, the associated text corresponding to the first text is determined from the alternative text, so as to establish an attribute graph of the first text based on the associated text. In this way, by determining the query word and index word of the text, the related text of the text can be quickly determined, so that it is easy to establish an attribute graph between the text and the related text, which simplifies the operation content and improves the efficiency of text processing to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, in which:
[0062] Figure 1 is a flowchart of a text processing method provided by an embodiment of the present disclosure;
[0063] Figure 2 is a flowchart of a method for determining a first text provided by an embodiment of the present disclosure;
[0064] Figure 3 is a flowchart of a method for establishing a second text index term provided by an embodiment of the present disclosure;
[0065] Figure 4 is a flowchart of a method for determining a first text query term provided by an embodiment of the present disclosure;
[0066] Figure 5 is a flowchart of a method for determining an alternative text corresponding to a first text provided by an embodiment of the present disclosure;
[0067] Figure 6 is a schematic diagram of a text processing flow provided by an embodiment of the present disclosure;
[0068] Figure 7 is a block diagram of a text processing device provided by an embodiment of the present disclosure;
[0069] Figure 8 is a schematic diagram of a storage medium provided by an embodiment of the present disclosure; and
[0070] Fig. 9 It is a block diagram of an electronic device provided by an embodiment of the present disclosure.
[0071] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION
[0072] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0073] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, device, equipment, method or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software. The data involved in the present disclosure can be data authorized by the user or fully authorized by all parties.
[0074] In this document, any number of elements in the drawings is used for illustration rather than limitation, and any naming is used only for distinction and does not have any limiting meaning.
[0075] The principle and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure.
[0076] Figure 1 is a flowchart of a text processing method provided by an embodiment of the present disclosure, such as Figure 1 As shown, the method may include:
[0077] Step S101: determine a first text from a text library.
[0078] In the embodiment of the present disclosure, a text library may be pre-established and store multiple texts, wherein the texts in the text library may be news uploaded to a platform, or may be blogs published by individuals, or may be public papers, etc., and the present disclosure does not limit this. Determining the first text from the text library may be selecting a text as the first text based on the attribute characteristics of the text stored in the text library, and the attribute characteristics may be the amount of interaction between the text and the user, for example, the number of user clicks on the text, the number of citations of the text, or the user reading time of the text, etc. The attribute characteristics may also be the characteristics of the text itself, for example, the total number of characters in the text, the field to which the text belongs, the author of the text, etc. Among them, the first text may be an article containing only plain text content, or an article containing information such as pictures, videos, links, etc., and the present disclosure does not limit this.
[0079] Step S102: performing word segmentation processing on the first text to determine a query word corresponding to the first text.
[0080] In the disclosed embodiment, the plain text content in the first text may be determined first, and then the plain text content may be segmented using a preset segmentation method to obtain a segmentation result, and finally, the query word corresponding to the first text may be determined based on the segmentation result. The preset segmentation method may be a machine learning method based on statistics, such as natural language processing (NLP), or a dictionary-based rule matching method, etc. Specifically, the dictionary-based rule matching method may be used as an example, and the plain text string in the first text may be obtained first, and then the plain text string may be searched and matched with a pre-stored dictionary string. If the plain text string has a dictionary string hit, the hit dictionary string may be used as a segmentation result of the plain text string. Finally, the plain text string may be traversed to determine all hit dictionary strings, so that all hit dictionary strings may be used as all segmentation results of the plain text string, that is, all hit dictionary strings may be the segmentation results corresponding to the first text.
[0081] In the embodiment of the present disclosure, the query word corresponding to the first text is determined according to the segmentation result of the first text. The segmentation result obtained from the first text can be directly used as the query word corresponding to the first text, or the segmentation result obtained from the first text can be firstly screened, and the segmentation result that meets the preset conditions after screening can be used as the query word corresponding to the first text, wherein the query word can be a segmentation result for querying articles with similar content to the first text, and the preset condition can be to use the segmentation result that expresses actual meaning as the query word. For example, the segmentation results of "bullet screen", "platform", "recommendation", "personalization", etc. can be determined as segmentation results with actual expression meaning, and then the segmentation results of this type with actual expression meaning can be used as query words, and the segmentation results of "we", "they", "what", "right", etc. can be determined as segmentation results without actual expression meaning, and then the segmentation results of this type without actual expression meaning can be screened out and not used as query words.
[0082] Step S103: searching an index word set by using the query word to determine a candidate text corresponding to the first text; the index word set is determined based on a plurality of second texts in the text library.
[0083] In the disclosed embodiment, it is possible to search for index terms that match the query terms in the index term set, for example, it is possible to search for index terms that are the same as the query terms, or it is possible to search for index terms that are similar in meaning to the query terms, and then determine the text indicated by the matching index term based on the correspondence between the index term and the text, and determine the indicated text as the candidate text corresponding to the index term, thereby obtaining the candidate text corresponding to the first text. It should be noted that there may be multiple query terms corresponding to the first text, that is, it is possible to search for multiple query terms in the index term set to find index terms that match each query term. Furthermore, any index term in the index term set may correspond to one text, or may correspond to multiple texts, and then the text indicated by the index term is determined as the candidate text, and one text indicated by the index term may be determined as the candidate text corresponding to the first text, or multiple texts indicated by the index term may be determined as the candidate text corresponding to the first text. Therefore, by searching the index word set with multiple query words to determine the alternative texts corresponding to the first text, it can be obtained that the number of alternative texts is multiple, and the minimum number of alternative texts can be the same as the number of query words for the first text. Therefore, the more query words for the first text, the more alternative texts corresponding to the first text can be obtained.
[0084] In the disclosed embodiment, the index word set may be obtained by combining the index words of multiple second texts, and the index words of the second text may be the segmentation results for searching the second text. Specifically, the second text may be segmented first, and then the index words of the second text may be determined based on the segmentation results of the second text. The segmentation results of the second text may be obtained by segmenting using a preset segmentation method, or the author, text category, tag word, title, etc. of the second text may be used as the segmentation results of the second text. Among them, the second text may be a text used to determine whether there is a correlation with the first text. Specifically, the second text may be other texts in the text library except the first text, or may be texts randomly downloaded from the Internet.
[0085] It should be noted that in order to ensure that the correlation score between the selected alternative text and the first text is high, in one implementation of the present disclosure, a text that is correlated with the first text can be selected as the second text. For example, a text belonging to the same or similar technical field as the first text can be selected as the second text. For example, text 1 belongs to the field of communication technology, text 2 belongs to the field of software application, and text 3 belongs to the field of humanities. Among them, text 1 is the first text. Since text 1 and text 2 belong to similar fields, text 2 can be determined as the second text. However, since the fields to which text 3 belongs are quite different from those of text 1, text 3 cannot be determined as the second text.
[0086] Step S104: Calculate the correlation score between the first text and the candidate text.
[0087] In an embodiment of the present disclosure, the first text and the alternative text may be compared, and the correlation between the two contents may be scored based on the content expressed in the first text and the content expressed in the alternative text, and the scoring result may be used as the correlation score between the first text and the alternative text. For example, the technical fields to which the first text and the alternative text belong may be compared. If the overlap of the fields is higher, a higher score may be assigned to the correlation between the first text and the alternative text; if the fields are quite different, a lower score may be assigned to the correlation between the first text and the alternative text; or the repetition between the keywords of the first text and the keywords of the alternative text may be compared. If the repetition of the keywords of the two texts is higher, a higher score may be assigned to the correlation between the first text and the alternative text; if the repetition of the keywords of the two texts is lower, a lower score may be assigned to the correlation between the first text and the alternative text; or the publication addresses of the first text and the alternative text may be compared. If the publication addresses of the first text and the alternative text are closer, a higher score may be assigned to the correlation between the first text and the alternative text; and if the difference between the publication addresses of the first text and the alternative text is greater, a lower score may be assigned to the correlation between the first text and the alternative text.
[0088] Step S105: determining, according to the relevance score, an associated text corresponding to the first text from the candidate texts, so as to establish an attribute graph of the first text based on the associated text.
[0089] In an embodiment of the present disclosure, the alternative texts may be sorted according to their relevance scores to the first text, and the top N alternative texts may be selected as associated texts corresponding to the first text, wherein N may be a preset value, for example, N may be set to 100, so that the top 100 alternative texts ranked from high to low in relevance scores may be selected as associated texts corresponding to the first text.
[0090] In the disclosed embodiment, the property graph of the first text is established by taking the first text as the main node and the associated text corresponding to the first text as the neighbor node. The correlation score between each associated text and the first text can be used as the correlation between the neighbor node corresponding to the associated text and the main node, so that the property graph of the first text can be generated based on the main node, the neighbor node, and the correlation between each neighbor node and the main node. In this way, through the property graph of the first text, the text that has a correlation with the first text can be quickly obtained, and the correlation between the text and the first text can be determined, so that the processing steps of the text can be reduced and the processing efficiency of the text can be improved.
[0091] In summary, the text processing method provided by the embodiment of the present disclosure can first determine the first text from the text library, then perform word segmentation processing on the first text, determine the query word corresponding to the first text, search in the index word set through the query word, and determine the alternative text corresponding to the first text. The index word set is determined based on multiple second texts in the text library, and the correlation score between the first text and the alternative text is calculated. Finally, based on the correlation score, the associated text corresponding to the first text is determined from the alternative text, so as to establish an attribute graph of the first text based on the associated text. In this way, by determining the query word and index word of the text, the related text of the text can be quickly determined, so that it is easy to establish an attribute graph between the text and the related text, which simplifies the operation content and improves the efficiency of text processing to a certain extent.
[0092] Optionally, in the embodiment of the present disclosure, the operation of determining the first text from the text library is as follows: Figure 2 As shown, it may specifically include:
[0093] Step S1011: for each text in the text library, determine the amount of interaction between the user and the text.
[0094] In the embodiments of the present disclosure, for each text stored in the text library, the number of clicks received by the text by the users may be counted, and the number of clicks may be used as the amount of user interaction with the text. Alternatively, the number of paid readings received by the text by the users may be counted, and the number of paid readings may be used as the amount of user interaction with the text. Alternatively, the reading time of the users for the text may be counted, and the corresponding interaction amount may be determined according to the length of the reading time, and this may be used as the amount of user interaction with the text.
[0095] Step S1012: The text whose interaction amount is less than a preset interaction threshold is used as the first text.
[0096] In the disclosed embodiment, the preset interaction threshold may be an interaction threshold pre-set according to actual conditions. For example, the preset interaction threshold may be 1000 user clicks, and the text with an interaction amount less than the preset interaction threshold of 1000 may be used as the first text. In order to solve the problem of low text interaction and low promotion effectiveness, the disclosed embodiment may select a text with a low interaction amount as the first text so as to establish an attribute graph for the first text. The first text may be recommended based on the attribute graph later. Since people tend to pay attention to texts with related relationships, recommending texts with a low interaction amount under related texts can increase users' attention to texts with a low interaction amount, and to a certain extent, can increase the recommendation success rate of texts with a low interaction amount.
[0097] Optionally, in the embodiment of the present disclosure, before the operation of searching the index word set by the query word to determine the candidate text corresponding to the first text, Figure 3 As shown, it may also include:
[0098] Step S21: Determine the second text according to other texts in the text library except the first text.
[0099] In the embodiment of the present disclosure, since the selected first text is a text in the text library whose interaction amount is less than a preset interaction threshold, and other texts in the text library except the first text may have texts with an interaction amount less than the preset interaction threshold, and may also have texts with an interaction amount greater than the preset interaction threshold, when determining the second text, one implementation method may be to use the text in the text library except the first text and with an interaction amount greater than the preset interaction threshold as the second text, that is, the second text may only be the text with an interaction amount greater than the preset interaction threshold. In this way, when determining the alternative text from the second text later, the determined alternative text can also be the text with an interaction amount greater than the preset interaction threshold. When establishing the attribute graph of the first text using the alternative text, the correlation between the first text and the text with a higher interaction amount can be determined. When recommending the first text based on the attribute graph later, the first text can be recommended under the text with a higher interaction amount. This can improve the interaction amount of the first text to a certain extent, thereby improving the effectiveness of the text recommendation.
[0100] It should be noted that when determining the second text, another implementation method may be to directly use the text other than the first text in the text library as the second text, that is, the second text may be a text with an interaction amount greater than a preset interaction threshold, or a text with an interaction amount less than a preset interaction threshold. In this way, by increasing the number of second texts, when determining alternative texts from the second texts later, the number of alternative texts can also be increased, and when establishing the attribute graph of the first text using the alternative texts, the correlation between the first text and multiple different texts can be determined. When the first text is subsequently recommended based on the attribute graph, the number of recommendations and the recommendation range of the first text can be increased, thereby increasing the amount of interaction between the first text and the user to a certain extent.
[0101] Step S22: performing word segmentation processing on the second text to generate keywords corresponding to the second text.
[0102] In the embodiment of the present disclosure, the second text may be first segmented using a preset segmentation method, and then the segmentation results obtained may be filtered to remove non-keywords such as punctuation marks and stop words, and the segmentation results obtained by filtering may be used as keywords corresponding to the second text. The preset segmentation method may be specifically as described above, and will not be described in detail here.
[0103] Step S23: using a preset full-text search algorithm, establish index terms corresponding to the second text based on the keywords.
[0104] In the embodiment of the present disclosure, the preset full-text search algorithm can be a sequential scanning method or an index scanning method. For example, the preset full-text search algorithm can be Elasticsearch (distributed full-text search), WHOOSH (full-text search), or SOLR (enterprise-level search application server), which is not limited to the embodiment of the present disclosure. Using the preset full-text search algorithm, an index word corresponding to the second text is established based on a keyword. After obtaining the keyword corresponding to each second text through word segmentation processing, multiple keywords can be obtained for multiple second texts, and the multiple keywords are scanned using the preset full-text search algorithm. When the keyword appears only once, a corresponding second text can be obtained based on the keyword search, and the keyword can be used as the index word corresponding to the second text. When the number of times the keyword appears is X, and X is a positive integer greater than 1, the corresponding X second texts can be obtained based on the keyword search, and the keyword can be used as the index word corresponding to the X second texts.
[0105] Optionally, in the embodiment of the present disclosure, the operation of establishing the index term corresponding to the second text based on the keyword may specifically include:
[0106] Determine a mapping relationship between the keyword and the second text; and generate an inverted index term corresponding to the second text according to the mapping relationship.
[0107] For example, the keywords in the second text 1 may include [communication, 5G, micro base station], and the mapping relationship between the keywords and the second text can be expressed as second text 1-[communication, 5G, micro base station], the keywords in the second text 2 may include [software, communication, micro base station], and the mapping relationship between the keywords and the second text can be expressed as second text 2-[software, communication, micro base station], the keywords in the second text 3 may include [communication, 5G, micro base station, partition], and the mapping relationship between the keywords and the second text can be expressed as second text 3-[communication, 5G, micro base station, partition], the keywords in the second text 4 may include [communication, 5G, configuration ], the mapping relationship between the keyword and the second text can be expressed as second text 4-[communication, 5G, configuration]. The inverted index terms corresponding to the second text are generated according to the above mapping relationship, and the inverted index terms are "communication", "5G", "micro base station", "software", "partition", and "configuration". Among them, the inverted index term "communication" can retrieve the corresponding second text 1, second text 2, second text 3, and second text 4; the inverted index term "5G" can retrieve the corresponding second text 1, second text 3, and second text 4; the inverted index term "micro base station" can retrieve the corresponding second text 1, second text 2, and second text 3, and so on.
[0108] It should be noted that in one implementation, the author, text category, tag word, and title in the second text can also be used as keywords corresponding to the second text, and the mapping relationship between the keyword and the second text can be determined, and the inverted index words corresponding to the second text can be generated according to the mapping relationship, so that the author, text category, tag word, and title can be used as the inverted index words corresponding to the second text. For example, the mapping relationship can be second text 1-[author 1, text category 1, tag word 1, title 1, communication, 5G, micro base station], and the inverted index words obtained are "author 1", "text category 1", "tag word 1", "title 1", "communication", "5G", "micro base station", and the above inverted index words can be retrieved as corresponding to the second text 1 respectively.
[0109] Optionally, in the embodiment of the present disclosure, the operation of performing word segmentation on the first text to determine the query word corresponding to the first text is as follows: Figure 4 As shown, it may specifically include:
[0110] Step S1021: perform word segmentation processing on the first text to obtain a word segmentation result of the first text.
[0111] In an embodiment of the present disclosure, a preset word segmentation method can be used to perform word segmentation processing on the first text to obtain a word segmentation result of the first text. Specifically, a character string contained in the first text can be obtained, and the character string can be segmented using a preset word segmentation method, and the obtained result can be used as the word segmentation result of the first text.
[0112] Step S1022: Filter the word segmentation results, and use the filtered word segmentation results as query words corresponding to the first text.
[0113] In the disclosed embodiment, the obtained word segmentation results can be filtered to delete punctuation marks, conjunctions, modal particles, etc., and the filtered word segmentation results can be used as the query words corresponding to the first text. For example, if the word segmentation results are "first, finance, stocks, is it", the word segmentation results can be filtered first, and "first" and "is it" can be deleted, and the query words corresponding to the first text can be obtained as "finance, stocks".
[0114] Optionally, in the embodiment of the present disclosure, the operation of searching the index word set by the query word to determine the candidate text corresponding to the first text is as follows: Figure 5 As shown, it may specifically include:
[0115] Step S1031: In the index word set, the inverted index word that is the same as the query word is used as the target index word.
[0116] For example, the query words may be "finance, stock", and the inverted index words in the index word set may include "finance", "stock", "futures", and "funds". The inverted index words that are the same as the query words are "finance" and "stock", so the target index words can be determined to be "finance" and "stock".
[0117] Step S1032: according to the mapping relationship between the second text and the inverted index term, use the second text corresponding to the target index term as the candidate text corresponding to the first text.
[0118] For example, the inverted index terms that are the same as the query terms are "finance, stock", that is, the target index terms are "finance" and "stock", and the mapping relationship between the second text and the inverted index terms can be finance-text 012, stock-text 035. It can be determined that the second text corresponding to the target index term "finance" is text 012, and the second text corresponding to the target index term "stock" is text 035, that is, the alternative texts corresponding to the first text are text 012 and text 035.
[0119] Optionally, in the embodiment of the present disclosure, the step of calculating the correlation score between the first text and the candidate text includes:
[0120] The first text and the candidate text are processed and calculated using a preset text correlation algorithm to determine a correlation score between the first text and the candidate text.
[0121] In the disclosed embodiment, the text relevance algorithm may be a semantic matching algorithm or a non-semantic matching algorithm. For example, the text relevance algorithm may be a TF-IDF (term frequency–inverse document frequency) or a BM25 algorithm. Specifically, the first text may be subjected to morpheme analysis to generate morpheme information, and then, in the candidate text, each morpheme information is searched to obtain search results for each morpheme information, and the relevance score between each morpheme information and the corresponding search result is calculated. Finally, the relevance scores between all morpheme information and the corresponding search results are weighted and summed, and the sum result is used as the relevance score between the first text and the candidate text.
[0122] In the disclosed embodiment, compared with the related text processing method, when establishing the attribute graph of the text, the entire content of the text is directly loaded into the memory to establish the attribute graph of the text, resulting in a large consumption of memory during text processing and a large limit on the scale of the attribute graph. In the disclosed embodiment, by determining the query terms and index terms of the text, the associated text of the text can be quickly determined, so that it is easy to establish the attribute graph between the text and the associated text, simplifying the operation content and improving the efficiency of text processing to a certain extent.
[0123] In the disclosed embodiment, since the selected first text is a text with a small amount of interaction with the user, by establishing an attribute graph of the first text, the associated text corresponding to the first text in the attribute graph is often a text with a large amount of interaction with the user. In actual applications, for example, when recommending the first text to the user, the first text can be pushed at the associated text with a large amount of interaction, thereby improving the interaction of the first text to a certain extent and achieving personalized recommendation of the first text.
[0124] For example, Figure 6 is a schematic diagram of a text processing flow provided by an embodiment of the present disclosure, such as Figure 6 As shown, 11, perform word segmentation processing on each text in the text library, 12, determine the second text from the text library, and establish an inverted index term for the second text, 13, determine the first text from the text library, and generate a query term for the first text, 14, search for an alternative text corresponding to the first text in the inverted index term, 15, calculate the correlation score between the first text and the alternative text, 16, determine the associated text corresponding to the first text according to the correlation score.
[0125] It should be noted that the text processing method provided in the embodiment of the present disclosure can be executed by a text processing device, or by a control module in the text processing device for executing the loading text processing method. The text processing method provided in the embodiment of the present disclosure is described by taking the text processing device executing the loading text processing method as an example. Figure 7 A text processing device according to an exemplary embodiment of the present disclosure is described.
[0126] Figure 7 is a block diagram of a text processing device provided by an embodiment of the present disclosure, such as Figure 7 As shown, the text processing device 50 may include:
[0127] A first determining module 501 is used to determine a first text from a text library;
[0128] A first word segmentation module 502, configured to perform word segmentation processing on the first text to determine a query word corresponding to the first text;
[0129] A search module 503, configured to search an index word set using the query word to determine an alternative text corresponding to the first text; the index word set is determined based on a plurality of second texts in the text library;
[0130] A calculation module 504, configured to calculate a correlation score between the first text and the candidate text;
[0131] The second determination module 505 is used to determine the associated text corresponding to the first text from the candidate texts according to the relevance score, so as to establish an attribute graph of the first text based on the associated text.
[0132] In summary, the text processing device provided by the embodiment of the present disclosure can first determine the first text from the text library, then perform word segmentation processing on the first text, determine the query word corresponding to the first text, search in the index word set through the query word, and determine the alternative text corresponding to the first text. The index word set is determined based on multiple second texts in the text library, and the correlation score between the first text and the alternative text is calculated. Finally, based on the correlation score, the associated text corresponding to the first text is determined from the alternative text, so as to establish an attribute graph of the first text based on the associated text. In this way, by determining the query word and index word of the text, the related text of the text can be quickly determined, so that it is easy to establish an attribute graph between the text and the related text, which simplifies the operation content and improves the efficiency of text processing to a certain extent.
[0133] Optionally, the device 50 further includes:
[0134] A selection module, used for determining the second text according to other texts in the text library except the first text;
[0135] A second word segmentation module, used to perform word segmentation processing on the second text to generate keywords corresponding to the second text;
[0136] An establishing module is used to establish index words corresponding to the second text based on the keywords by using a preset full-text search algorithm.
[0137] Optionally, the establishment module is further used to:
[0138] Determine a mapping relationship between the keyword and the second text;
[0139] Generate an inverted index word corresponding to the second text according to the mapping relationship.
[0140] Optionally, the inverted index words include the author, text category, tag words, and title of the second text.
[0141] Optionally, the search module 503 is further used to:
[0142] In the index word set, the inverted index word that is the same as the query word is used as the target index word;
[0143] According to the mapping relationship between the second text and the inverted index term, the second text corresponding to the target index term is used as the candidate text corresponding to the first text.
[0144] Optionally, the calculation module 504 is further used for:
[0145] The first text and the candidate text are processed and calculated using a preset text correlation algorithm to determine a correlation score between the first text and the candidate text.
[0146] Optionally, the first word segmentation module 502 is further used for:
[0147] Performing word segmentation processing on the first text to obtain a word segmentation result of the first text;
[0148] The word segmentation results are filtered, and the filtered word segmentation results are used as query words corresponding to the first text.
[0149] Optionally, the first determining module 501 includes:
[0150] For each text in the text library, determining the amount of interaction between the user and the text;
[0151] The text whose interaction amount is less than a preset interaction threshold is used as the first text.
[0152] After introducing the text processing method and device according to the exemplary embodiment of the present disclosure, Figure 8 A storage medium according to an exemplary embodiment of the present disclosure is described.
[0153] refer to Figure 8 As shown, a storage medium 600 for implementing the above method according to an embodiment of the present disclosure is described, which can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.
[0154] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0155] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0156] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0157] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user computing device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0158] After introducing the storage medium of the exemplary embodiment of the present disclosure, next, reference is made to Fig. 9 An electronic device according to an exemplary embodiment of the present disclosure is described.
[0159] Fig. 9 The electronic device 800 shown in the figure is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0160] like Fig. 9 As shown, the electronic device 800 is in the form of a general computing device. The components of the electronic device 800 may include, but are not limited to: the at least one processing unit 810, the at least one storage unit 820, a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810), and a display unit 840.
[0161] The storage unit stores a program code, and the program code can be executed by the processing unit 810, so that the processing unit 810 executes the steps described in the "Exemplary Method" section of the specification according to various exemplary embodiments of the present disclosure. For example, the processing unit 810 can execute step S101, determining a first text from a text library; step S102, performing word segmentation processing on the first text to determine the query word corresponding to the first text; step S103, searching in an index word set through the query word to determine the candidate text corresponding to the first text; the index word set is determined based on multiple second texts in the text library; step S104, calculating the correlation score between the first text and the candidate text; step S105, determining the associated text corresponding to the first text from the candidate text according to the correlation score, so as to establish an attribute graph of the first text based on the associated text.
[0162] The storage unit 820 may include a volatile storage unit, such as a random access storage unit (RAM) 8201 and / or a cache storage unit 8202 , and may further include a read-only storage unit (ROM) 8203 .
[0163] The storage unit 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, such program modules 8205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0164] The bus 830 may include a data bus, an address bus, and a control bus.
[0165] The electronic device 800 can also communicate with one or more external devices 70 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), and this communication can be performed through an input / output (I / O) interface 850. The electronic device 800 also includes a display unit 840, which is connected to the input / output (I / O) interface 850 for display. In addition, the electronic device 800 can also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs) and / or public networks, such as the Internet) through a network adapter 860. As shown in the figure, the network adapter 860 communicates with other modules of the electronic device 800 through a bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0166] It should be noted that although several modules or sub-modules of the audio playback device and the audio sharing device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided to be embodied by multiple units / modules.
[0167] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present disclosure.
[0168] The embodiments of the present disclosure are described above in conjunction with the accompanying drawings, but the present disclosure is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present disclosure, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present disclosure and the claims, all of which are within the protection of the present disclosure.
Claims
1. A text processing method, characterized in that: The method comprises: For each text in the text library, determining the amount of interaction between the user and the text; taking the text whose interaction amount is less than a preset interaction threshold as the first text; Performing word segmentation processing on the first text to determine a query word corresponding to the first text; Determine a second text based on other texts in the text library except the first text; perform word segmentation processing on the second text to generate keywords corresponding to the second text; determine a mapping relationship between the keywords and the second text using a preset full-text search algorithm; and generate an inverted index word corresponding to the second text based on the mapping relationship; In the index word set, the inverted index word identical to the query word is used as a target index word; according to the mapping relationship between the second text and the inverted index word, the second text corresponding to the target index word is used as an alternative text corresponding to the first text; the index word set is determined according to the plurality of second texts in the text library; Calculating a correlation score between the first text and the candidate text; According to the relevance score, an associated text corresponding to the first text is determined from the candidate texts, so as to establish an attribute graph of the first text based on the associated text.
2. The method according to claim 1, characterized in that The inverted index words include the author, text category, tag words, and title of the second text.
3. The method according to claim 1, characterized in that The calculating the correlation score between the first text and the candidate text includes: The first text and the candidate text are processed and calculated using a preset text correlation algorithm to determine a correlation score between the first text and the candidate text.
4. The method according to claim 1, characterized in that The performing word segmentation processing on the first text to determine the query word corresponding to the first text includes: Performing word segmentation processing on the first text to obtain a word segmentation result of the first text; The word segmentation results are filtered, and the filtered word segmentation results are used as query words corresponding to the first text.
5. A text processing device, characterized in that: The device comprises: A first determination module is used to determine the amount of interaction between the user and each text in the text library; and to take the text whose interaction amount is less than a preset interaction threshold as the first text; A first word segmentation module, used to perform word segmentation processing on the first text to determine a query word corresponding to the first text; A selection module is used to determine a second text based on other texts in the text library except the first text; a second word segmentation module is used to perform word segmentation processing on the second text to generate keywords corresponding to the second text; an establishment module is used to establish index words corresponding to the second text based on the keywords using a preset full-text search algorithm; the establishment module is also used to: determine a mapping relationship between the keywords and the second text; and generate an inverted index word corresponding to the second text according to the mapping relationship; A search module, configured to search an index word set by using the query word to determine an alternative text corresponding to the first text; the index word set is determined based on a plurality of second texts in the text library; the search module is further configured to: in the index word set, use an inverted index word identical to the query word as a target index word; and based on a mapping relationship between the second text and the inverted index word, use the second text corresponding to the target index word as an alternative text corresponding to the first text; A calculation module, used for calculating a correlation score between the first text and the candidate text; The second determination module is used to determine the associated text corresponding to the first text from the candidate texts according to the relevance score, so as to establish an attribute graph of the first text based on the associated text.
6. The device according to claim 5, characterized in that The inverted index words include the author, text category, tag words, and title of the second text.
7. The device according to claim 5, characterized in that The computing module is further used for: The first text and the candidate text are processed and calculated using a preset text correlation algorithm to determine a correlation score between the first text and the candidate text.
8. The device according to claim 5, characterized in that The first word segmentation module is further used for: Performing word segmentation processing on the first text to obtain a word segmentation result of the first text; The word segmentation results are filtered, and the filtered word segmentation results are used as query words corresponding to the first text.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the text processing method according to any one of claims 1 to 4 is implemented.
10. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the text processing method according to any one of claims 1 to 4 by executing the executable instructions.
Citation Information
Patent Citations
Document retrieval method and device based on block index structure, medium and equipment
CN112199461A
Semantic-based approximate text search method and device, computer equipment and medium
CN113434636A