Processing method and electronic equipment

Through word segmentation processing and word combination and text features generation, the problem of poor machine learning's cross-domain text information extraction effect is solved, and efficient and cross-domain text core information extraction and search accuracy are improved.

CN120257985APending Publication Date: 2025-07-04LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510398655.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When existing machine learning methods have fewer training data, it is difficult to effectively extract cross-domain text information, resulting in poor text information extraction effect.

Method used

Ordered word elements are generated through word segmentation processing, phrases are constructed and phrases with the same content are merged, phrases are filtered and spliced ​​according to predetermined rules, text features are generated, core information of the text to be processed, and text features are optimized using phrase weights and similarity.

Benefits of technology

It realizes efficient extraction of text core information in different fields, reduces the pressure of vector storage, and improves the accuracy of search results and system response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257985A_ABST
    Figure CN120257985A_ABST
Patent Text Reader

Abstract

The invention provides a processing method and electronic equipment, and the method comprises the steps: obtaining a to-be-processed text, carrying out the word segmentation of the to-be-processed text, and obtaining a plurality of ordered lexical elements; the lexical elements comprise semantic words, and the semantic words are vocabularies representing text semantics in the to-be-processed text; for any semantic word of the plurality of semantic words, sequentially connecting the next semantic word from any semantic word until the text content of the next semantic word is the same as that of any semantic word, and obtaining a plurality of first word groups; combining any two first phrases with the same text content of the final lexical elements in the plurality of first phrases to obtain a plurality of second phrases; screening at least part of the second phrases from the plurality of second phrases according to a predetermined rule, and splicing the at least part of the second phrases to obtain a plurality of text features; and generating core information of the to-be-processed text according to the plurality of text features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and more particularly, to a method for processing text and an electronic device. Background Art

[0002] Extracting key text information from text has applications in various scenarios. For example, in a search scenario, it can reduce the volume of stored indexes and improve the accuracy of search results.

[0003] Currently, the way to extract text information is mainly based on machine learning. Since machine learning requires a large amount of data for training and manual marking of the data. When the training data is scarce, machine learning may have the disadvantage of poor adaptability to data from different fields, resulting in poor extraction effect of text information. Summary of the Invention

[0004] In view of this, the present disclosure provides a processing method and an electronic device.

[0005] A first aspect of the present disclosure provides a processing method, including: obtaining a text to be processed, performing word segmentation processing on the text to be processed to obtain a plurality of ordered word elements; a word element includes a semantic word, and a semantic word is a vocabulary in the text to be processed that represents the semantics of the text; for any semantic word among the plurality of semantic words, connecting the next semantic word in sequence from any semantic word until the text content of the next semantic word is the same as that of any semantic word, to obtain a plurality of first word groups; merging any two first word groups with the same text content of the word element at the end of the sorting among the plurality of first word groups to obtain a plurality of second word groups; screening at least some of the second word groups from the plurality of second word groups according to a predetermined rule, and splicing at least some of the second word groups to obtain a plurality of text features; generating core information of the text to be processed according to the plurality of text features.

[0006] According to an embodiment of the present disclosure, the method further includes: for any semantic word among the plurality of semantic words, connecting the next semantic word in sequence from any semantic word until the number of semantic words between the next semantic word and any semantic word reaches a first preset number, stopping the connection and obtaining a first word group.

[0007] According to an embodiment of the present disclosure, splicing at least some of the second word groups includes: splicing at least some of the second word groups according to a second preset number; before generating core information of the text to be processed according to the text features, the method further includes: for each of the at least some of the second word groups, determining the weight of the second word group; according to the group weights of the second preset number of second word groups that are spliced into each text feature, determining the weight of each text feature; screening some of the text features from the plurality of text features according to the weight of each text feature and a preset weight threshold.

[0008] According to an embodiment of the present disclosure, for each second phrase in at least a part of the second phrases, determining the weight of the second phrase includes: determining the weight of each first phrase; determining the weight of the second phrase according to the weights of the respective first phrases that are combined into the second phrase; wherein, the weight of the second phrase is inversely proportional to the number of semantic words, and the number of semantic words is the total number of semantic words in the respective first phrases that are combined into the second phrase.

[0009] According to an embodiment of the present disclosure, the weight of the second phrase is further determined by the following method: from at least a part of the second phrases, determining any two second phrases with the same text content of the last-ranked token; calculating the similarity between any two second phrases; in the case where the similarity is greater than a preset similarity threshold, respectively increasing the phrase weights of any two second phrases according to the phrase similarity.

[0010] According to an embodiment of the present disclosure, a token further includes a stop word, and a stop word is a vocabulary in the text to be processed that does not represent the text semantics; connecting from any semantic word to the next semantic word in sequence until the next semantic word has the same text content as any semantic word includes: connecting from any token to the next token in sequence and skipping any stop word until the next semantic word has the same text content as any semantic word.

[0011] According to an embodiment of the present disclosure, determining the weight of each first phrase includes: determining the weight of each token among a plurality of tokens; respectively obtaining the weights of the first and last semantic words in the first phrase according to the weights of each token, and the word frequencies of the first and last semantic words among the plurality of tokens; determining the number of stop words skipped in the first phrase; obtaining the weight of each first phrase according to the weights, word frequencies of the first and last semantic words in the first phrase, and the number of stop words skipped in the first phrase.

[0012] According to an embodiment of the present disclosure, the text information includes statement information; generating the text information of the text to be processed according to the text features further includes: determining at least one semantic word that constitutes each text feature; obtaining the position information of at least one semantic word in the text to be processed; determining the corresponding statement information in the text to be processed according to the position information.

[0013] According to an embodiment of the present disclosure, the core information further includes paragraph information; generating the core information of the text to be processed according to the text features further includes: determining the weights of the respective second phrases in at least a part of the second phrases; obtaining at least one second phrase corresponding to each semantic word in any text statement; in any text statement, obtaining the weight of any text statement according to the sum of the weights of at least one second phrase corresponding to each semantic word therein; screening a plurality of text statements according to the weight of the text statement to obtain paragraph information.

[0014] A second aspect of the present disclosure provides an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the following method: obtaining a text to be processed, performing word segmentation on the text to be processed to obtain a plurality of ordered word tokens; the word tokens include semantic words; for any semantic word among the plurality of semantic words, connecting the next semantic word in sequence from the any semantic word until the text content of the next semantic word is the same as that of the any semantic word, to obtain a plurality of first word groups; merging any two first word groups with the same text content of the word token at the end of the sorting among the plurality of first word groups, to obtain a plurality of second word groups; screening at least some of the second word groups from the plurality of second word groups according to a predetermined rule, and splicing the at least some of the second word groups to obtain a plurality of text features; generating core information of the text to be processed according to the plurality of text features. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above and other objects, features and advantages of the present disclosure will become more apparent. In the drawings: Figure 1 Schematically shows a flowchart of a processing method according to an embodiment of the present disclosure; Figure 2 Schematically shows a schematic diagram of an ordered word set according to an embodiment of the present disclosure; Figure 3 Schematically shows one of the schematic diagrams of word group merging according to an embodiment of the present disclosure; Figure 4 Schematically shows another schematic diagram of word group merging according to an embodiment of the present disclosure; Figure 5 Schematically shows a flowchart of screening text features according to an embodiment of the present disclosure; Figure 6 Schematically shows a flowchart of determining the weight of the first word group according to an embodiment of the present disclosure; Figure 7 Schematically shows a block diagram of a processing device according to an embodiment of the present disclosure; Figure 8 Schematically shows a block diagram of an electronic device suitable for implementing the above-described electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.

[0017] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0018] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0019] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, having only B, having only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0020] The processing method and electronic device provided by the embodiments of the present disclosure can be applied to the vectorization application scenario to reduce the vector storage pressure. Among them, the vectorization application scenario refers to that after converting the words in the text to be processed into vectors, the computer can better understand the semantic information of the text, so as to classify the text. For example, classifying news articles into categories such as politics, economy, sports, etc., or classifying user comments into emotional categories such as positive, negative, and neutral. It can also be applied to the Retrieval-augmented Generation (RAG) scenario, which can effectively reduce Token consumption and significantly improve the system response speed and efficiency.

[0021] Figure 1 A flowchart of a processing method according to an embodiment of the present disclosure is schematically shown.

[0022] Specifically, as Figure 1 shown, the processing method includes operations S101 to S104.

[0023] In operation S101, the text to be processed is obtained, and the text to be processed is segmented to obtain a plurality of ordered word tokens; the word tokens include semantic words, and the semantic words are the words in the text to be processed that represent the text semantics.

[0024] In an embodiment of the present disclosure, the text to be processed can be a long text string, which contains a plurality of words and punctuation marks. By segmenting the text to be processed, the text to be processed is segmented into a plurality of sequentially arranged word tokens; among them, the types of word tokens include semantic words, and any one semantic word has a specific text semantics in the text to be processed. The types of word tokens also include stop words, and the stop words do not represent specific text semantics in the text to be processed, or the degree of representing text semantics is weak.

[0025] For n sequentially arranged word tokens, an ordered word set can be obtained, defined as , where represents the th word token in the ordered word set; for the th semantic word, it is defined as ; for the j th stop word, it is defined as .

[0026] Figure 2 FIG. schematically shows a schematic diagram of an ordered word set according to an embodiment of the present disclosure.

[0027] As Figure 2 shown, the text to be processed "Hello, please come here" is segmented to obtain an ordered word set, and each word token in the ordered word set has corresponding text content and serial number.

[0028] In an embodiment of the present disclosure, before segmenting the text to be processed, the text to be processed is first subjected to stemming and part-of-speech restoration.

[0029] In operation S102, for any one of the plurality of semantic words, starting from any one semantic word, the next semantic word is connected in sequence until the text content of the next semantic word is the same as that of any one semantic word, and a plurality of first phrases are obtained.

[0030] In an embodiment of the present disclosure, after segmenting the text to be processed, a plurality of ordered word tokens are obtained; starting from any one of the semantic words, the next semantic word is connected in sequence until the same semantics of the text content is encountered, and the connection is stopped, and all the semantic words encountered during the connection process are used as the first phrase.

[0031] Exemplarily, as Figure 2As shown, starting from the first semantic word, connect the second semantic word, the third semantic word, and the fifth semantic word in sequence; the text content of the first semantic word is the same as that of the fifth semantic word, stop connecting, and output the above four semantic words as a first phrase group, denoted as C1,5.

[0032] In operation S103, any two first phrase groups with the same text content of the word elements at the end of the sorting among the multiple first phrase groups are merged to obtain multiple second phrase groups.

[0033] Figure 3 Schematically shows one of the schematic diagrams of word group merging according to an embodiment of the present disclosure.

[0034] In the embodiments of the present disclosure, as Figure 3 shown, multiple first phrase groups are constructed into a word connection diagram as shown in Figure 3 . For example, the text contents of the word elements at the end of the sorting of the three first phrase groups C(14,15), C(1,7), and C(40,43) are the same; for example, the text contents of the word elements at the end of the sorting of the two first phrase groups C(36,37) and C(20,22) are the same. Since the text contents of the word elements at the end of the sorting of the multiple first phrase groups to be merged are the same, and the text contents of the word elements at the beginning of the sorting are also the same, therefore, the word elements at the end of the sorting and the word elements at the beginning of the sorting of the multiple first phrase groups to be merged can be used to represent the merged second phrase group.

[0035] Figure 4 Schematically shows another schematic diagram of word group merging according to an embodiment of the present disclosure.

[0036] As Figure 4 shown, the three first phrase groups C(14,15), C(1,7), and C(40,43) are merged, and the resulting second phrase group is represented as . The two first phrase groups C(36,37) and C(20,22) are merged, and the resulting second phrase group is represented as . Among them, S(N) represents a word element with the text content of N, and N belongs to a, b, c......

[0037] In operation S104, at least some second phrase groups are screened from the multiple second phrase groups according to a predetermined rule, and at least some second phrase groups are spliced to obtain multiple text features; the core information of the text to be processed is generated according to the multiple text features.

[0038] In the embodiments of the present disclosure, the length of the text feature can be preset in advance; according to the preset length of the text feature, it is determined that each text feature is spliced by a predetermined number of second phrase groups.

[0039] Exemplarily, the length of the text feature is configured to be 3 tokens. At this time, each text feature is formed by concatenating 2 second word groups. As Figure 4 shown, and collectively contain three tokens: S(a), S(b), and S(d). Connecting S(a), S(b), and S(d) in sequence gives the text feature S(a)S(b)S(d); similarly, connecting S(a), S(b), and S(f) in sequence gives the text feature S(a)S(b)S(f). Traverse Figure 4 all possible concatenation methods in

[0040] to obtain multiple text features. It should be noted that in the embodiments of the present disclosure, during the process of concatenating at least some of the second word groups, the concatenation needs to be performed in the direction of the second word group. For any second word group , its direction points from S(x) to S(y). Therefore, the concatenation direction also needs to point from S(x) to S(y).

[0041] By performing word segmentation on the text to be processed to obtain multiple tokens, representative text features can be extracted, and then a text representing the core information of the text to be processed can be generated through the text features. Its implementation process does not depend on training data in a specific domain and has better cross-domain application capabilities, and can effectively extract the core information of the text in various fields.

[0042] In the embodiments of the present disclosure, the method further includes: for any semantic word in multiple semantic words, connecting the next semantic word in sequence from any semantic word until the number of semantic words between the next semantic word and any semantic word reaches a first preset number, and then stop connecting and obtain a first word group.

[0043] In operation S102, when connecting the next semantic word in sequence, there may be a situation where a semantic word with the same text content as the first semantic word never appears. Therefore, during the process of obtaining the first word group, when the number of connected semantic words reaches the first preset number, but a semantic word with the same text content as the first semantic word has not been encountered, stop connecting and obtain the first word group.

[0044] By setting the number of connected semantic words, it is possible to avoid the situation where the number of connected semantic words is too long, and to avoid the failure to obtain the first word group caused by the fact that a semantic word with the same text content as the first semantic word has not been connected all the time.

[0045] In the embodiments of the present disclosure, concatenating at least some of the second word groups includes: concatenating at least some of the second word groups according to a second preset number.

[0046] Figure 5A flowchart for screening text features according to an embodiment of the present disclosure is schematically shown.

[0047] As Figure 5 shown, before generating the core information of the text to be processed according to the text features, operations S501 to S503 are further included.

[0048] Operation S501: For each of at least some of the second phrases, determine the weight of the second phrase.

[0049] Operation S502: According to the phrase weights of the second preset number of second phrases that make up each text feature, determine the weight of each text feature.

[0050] Operation S503: According to the weight of each text feature and a preset weight threshold, screen out some text features from multiple text features.

[0051] Since the number of obtained text features is large, it is necessary to determine the weight of each text feature according to the phrase weights of the second preset number of second phrases that make up each text feature. Among them, the phrase weight of the second group of words is positively correlated with the weight of the text feature. After obtaining the weight of each text feature, according to the preset weight threshold, screen out the text features whose weights of the text features are greater than the preset weight threshold for subsequent generation of the core information of the text to be processed.

[0052] In an embodiment of the present disclosure, in operation S501, for each of at least some of the second phrases, determining the weight of the second phrase includes: determining the weight of each first phrase; according to the weights of the respective first phrases that are combined into the second phrase, determining the weight of the second phrase; wherein, the weight of the second phrase is inversely proportional to the number of semantic words, and the number of semantic words is the total number of semantic words in the respective first phrases that are combined into the second phrase.

[0053] Exemplarily, a calculation formula for the weight of a second phrase is given as follows: ;

[0054] Wherein, is the sum of the weights of the respective first phrases that are combined into the second phrase; is the number of phrases of the respective first phrases that are combined into the second phrase; is the total number of semantic fields of the respective first phrases that are combined into the second phrase.

[0055] In an embodiment of the present disclosure, for the process of determining the weight of the first phrase, reference may be made to the specific content of subsequent embodiments.

[0056] By considering the weights of each first phrase merged into the second phrase and the influence of the number of semantic words on the weights, the weight of the second phrase can be evaluated more accurately.

[0057] In an embodiment of the present disclosure, the weight of the second phrase is further determined in the following manner: from at least part of the second phrases, determine any two second phrases with the same text content of the word element at the end of the sorting; calculate the similarity between any two second phrases; in the case where the similarity is greater than a preset similarity threshold, respectively increase the phrase weights of any two second phrases according to the phrase similarity.

[0058] Specifically, from at least part of the second phrases, determine any two second phrases with the same text content of the word element at the end of the sorting, see Figure 4 in and two second phrases. Calculate the similarity between the two second phrases . In the case where the similarity is greater than the similarity threshold, according to the phrase similarity, to increase and the phrase weights of the two second phrases.

[0059] Exemplarily, the similarity between two second phrases , is determined in the following manner: Since and the text content of the word elements at the end of the sorting of the two second phrases is the same, by calculating and the correlation degree of the word elements at the beginning of the sorting of the two second phrases, to determine the similarity between the two second phrases .

[0060] First, obtain and the word elements S(a) and S(c) at the beginning of the sorting of the two second phrases. Subsequently, convert the word elements S(a) and S(c) into vector forms respectively, and then is determined by the following formula: ;

[0061] In after determining that it is greater than the preset similarity threshold, according to the phrase similarity, respectively increase and the phrase weights of, where the higher the phrase similarity, the greater the degree of increase in the phrase weights of the two second phrases.

[0062] Exemplarily, the increased and phrase weights are determined by the following formula: ; ;

[0063] Among them, is the basic weight of the second phrase . is the basic weight of the second phrase . is the increased weight of the second phrase . is the increased weight of the second phrase .

[0064] In the embodiments of the present disclosure, the word units further include stop words, and stop words are words in the text to be processed that do not represent the semantics of the text; connecting the next semantic word in sequence from any semantic word until the text content of the next semantic word is the same as that of any semantic word, including: connecting the next word unit in sequence from any word unit and skipping any stop word until the text content of the next semantic word is the same as that of any semantic word.

[0065] Specifically, when obtaining the first phrase composed of semantic words, since the ordered word set composed of multiple word units also includes stop words, when obtaining the first phrase by connecting semantic words in sequence, it is necessary to avoid stop words during the connection process.

[0066] By skipping stop words, the first phrase can be constructed more efficiently, while retaining key information and improving the accuracy and efficiency of phrase extraction.

[0067] Figure 6 Schematically shows a flowchart of determining the weight of the first phrase according to an embodiment of the present disclosure.

[0068] As Figure 6 shown, determining the weight of each first phrase includes: operations S601 to S604.

[0069] Operation S601, determining the weight of each word unit among multiple word units.

[0070] Specifically, for 's weight; is determined by the following formula: ;

[0071] Among them, exp is the exponential function, origin is defined as the starting parameter, and offset is defined as the offset parameter; scale and decay are preset parameters for controlling the decreasing rate of the weight of the word unit as i changes.

[0072] Exemplarily, in this solution, offset=origin×0.2, origin=t× 0.4, decay = 0.3, scale= t×0.8 . Among them, t is the total number of lemmas in the ordered lexicon.

[0073] Operation S602: According to the weights of each lemma, respectively obtain the weights of the semantic words ranked first and last in the first phrase, and the word frequencies of the semantic words ranked first and last among multiple lemmas.

[0074] After obtaining the weight calculation formula for each lemma, according to the serial numbers of the semantic words ranked first and last in the first phrase, obtain the weights of the semantic words ranked first and last in the first phrase. At the same time, determine the word frequencies of the semantic words ranked first and last among multiple lemmas.

[0075] Exemplarily, for any lemma the word frequency is determined by the following formula: ;

[0076] Among them, is the number of lemmas in the entire ordered lexicon that have the same text content as ; is the number of semantic words in the entire ordered lexicon.

[0077] Operation S603: Determine the number of stop words skipped in the first phrase.

[0078] Exemplarily, during the process of obtaining the first phrase, the number of stop words skipped is configured as k.

[0079] Operation S604: According to the weights, word frequencies of the semantic words ranked first and last in the first phrase, and the number of stop words skipped in the first phrase, obtain the weight of each first phrase.

[0080] Specifically, the weight W(i, j) of any first phrase C(i, j) is determined by the following formula: ;

[0081] Among them, m is the first preset number, that is, the maximum number of semantic words that can be connected during the process of obtaining the first phrase; is the weight coefficient of the stop word, and the weight coefficients of each stop word in the text to be processed are configured with the same value.

[0082] By comprehensively considering the weights, word frequencies of the lemmas, and the number of stop words skipped, the weight of the first phrase can be evaluated more accurately.

[0083] Next, a method for generating the core information of the text to be processed according to the text features in the embodiments of the present disclosure will be described.

[0084] In the embodiments of the present disclosure, the text information includes phrase information; to generate the text information of the text to be processed according to the text features, traverse the feature set { L 1 ,L 2 …L n} in descending order of the text feature weights, and connect the respective second phrases that make up the text feature weights L n head to tail in the subscript order to form phrase information, and each phrase information represents an important feature information point of the original text to be processed.

[0085] In the embodiments of the present disclosure, the text information includes sentence information; to generate the text information of the text to be processed according to the text features, it further includes: determining at least one semantic word that makes up each text feature; obtaining the position information of at least one semantic word in the text to be processed. According to the position information, determine the corresponding sentence information in the text to be processed.

[0086] Specifically, according to the weight of the second phrase, select a second phrase in the text feature, and obtain the subscript information of the second phrase, that is, the serial numbers of the semantic word at the first position in the sorting and the semantic word at the last position in the sorting; locate the specific positions in the text to be processed through the two serial numbers represented by the subscript information. Expand forward from the semantic word at the first position in the sorting and expand backward from the semantic word at the last position in the sorting until the end of the sentence. The sentence obtained during the expansion process is used as the sentence information of the text to be processed.

[0087] In the embodiments of the present disclosure, the core information further includes paragraph information; to generate the core information of the text to be processed according to the text features, it further includes: determining the weights of the respective second phrases in at least some of the second phrases; obtaining at least one second phrase corresponding to each semantic word in any text sentence; in any text sentence, obtain the weight of any text sentence according to the sum of the weights of at least one second phrase corresponding to each semantic word therein; screen multiple text sentences according to the weight of the text sentence to obtain paragraph information.

[0088] Specifically, according to the weight of the second phrase, select a second phrase in the text feature, and obtain the subscript information of the semantic word at the first position in the sorting and the semantic word at the last position in the sorting in the second phrase to obtain the coordinate information of the first and last words of the second phrase, and superimpose the weight of the second phrase on the corresponding subscripts in the text to be processed. Initially, the subscript weights of the text to be processed are all 0. For the above-mentioned second phrase , which is formed by merging C(14, 15), C(1, 7), and C(40, 43), with a total of 6 punctuation marks. Assume the second phrase has a weight of 10. The starting and ending word subscripts of the second phrase share the weight of this second phrase. Then, for the subscript set {1, 7, 14, 15, 40, 43} corresponding to C(14, 15), C(1, 7), and C(40, 43), the coordinate weights of each token are all 10. The weights of different second phrases are mapped to the subscripts of the text to be processed and superimposed. The higher the subscript weight of a text block, the higher the information density and the higher the relevance. On the contrary, the lower the subscript weight of a text block, the lower the information density and the lower the relevance.

[0089] Exemplarily, taking the text to be processed as the dimension, by calculating the average weight of the subscript weights of each token in the statement corresponding to the second phrase to determine the weight of the text statement corresponding to the second phrase The calculation formula is as follows: W ;

[0090] where is the weight of the token with subscript i in the statement, Count ( s ) is the total number of tokens in this text statement, Count ( S stop ) is the number of stop words in this text statement.

[0091] It should be noted that the text statements in the above embodiments can be configured as statements generated during the process of obtaining statement information in the text information, or can also be any statement in a certain paragraph of the text to be processed.

[0092] Subsequently, the text to be processed is segmented according to the paragraphs in the text. Within each original paragraph, statements with weights higher than a certain threshold are extracted in the order of the text to be processed to form new paragraphs.

[0093] By extracting phrase information, statement information, and paragraph information from the text to be processed, while removing redundant or excessive text content, the information density of the core information of the text to be processed is improved.

[0094] Based on the above processing method, the present disclosure also provides a processing device. The following will be combined with Figure 7 to describe this device in detail.

[0095] Figure 7 Schematically shows a block diagram of a processing device according to an embodiment of the present disclosure.

[0096] As Figure 7As shown, the processing device 700 of this embodiment includes a word segmentation processing module 710, a phrase acquisition module 720, a phrase merging module 730, and an information acquisition module 740.

[0097] The word segmentation processing module 710 is configured to obtain a text to be processed, perform word segmentation processing on the text to be processed, and obtain a plurality of ordered word elements; the word elements include semantic words, and the semantic words are the words in the text to be processed that represent the text semantics.

[0098] The phrase acquisition module 720 is configured to, for any one of the plurality of semantic words, sequentially connect the next semantic word from any one of the semantic words until the text content of the next semantic word is the same as that of any one of the semantic words, and obtain a plurality of first phrases; The phrase merging module 730 is configured to merge any two first phrases with the same text content of the word elements at the end of the sorting among the plurality of first phrases to obtain a plurality of second phrases; The information acquisition module 740 is configured to screen at least some of the second phrases from the plurality of second phrases according to a predetermined rule, splice at least some of the second phrases, and obtain a plurality of text features; generate core information of the text to be processed according to the plurality of text features.

[0099] According to the embodiments of the present disclosure, any plurality of modules, sub-modules, units, and sub-units, or at least some functions of any plurality of them can be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present disclosure can be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present disclosure can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging the circuit in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present disclosure can be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.

[0100] For example, any combination of the word segmentation processing module 710, phrase acquisition module 720, phrase merging module 730, and information acquisition module 740 can be combined and implemented in one module / unit / sub-unit, or any one of the modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least some functions of one or more of these modules / units / sub-units can be combined with at least some functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to an embodiment of the present disclosure, at least one of the word segmentation processing module 710, phrase acquisition module 720, phrase merging module 730, and information acquisition module 740 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), programmable logic array (PLA), system on chip, system on substrate, system on package, application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, at least one of the word segmentation processing module 710, phrase acquisition module 720, phrase merging module 730, and information acquisition module 740 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0101] It should be noted that the processing device part in the embodiments of the present disclosure corresponds to the processing method part in the embodiments of the present disclosure, and their specific implementation details are the same, so they will not be elaborated here.

[0102] Figure 8 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present disclosure is schematically shown. Figure 8 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0103] As Figure 8 shown, the electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The processor 801 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 801 can also include on-board memory configured for caching purposes. The processor 801 can include a single processing unit or multiple processing units configured to perform different actions of the method flow according to an embodiment of the present disclosure.

[0104] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to the embodiments of the present disclosure by executing programs in the ROM 802 and / or the RAM 803. It should be noted that the programs may also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing programs stored in the one or more memories.

[0105] According to an embodiment of the present disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage portion 808 as needed.

[0106] According to an embodiment of the present disclosure, the method flow according to the embodiments of the present disclosure may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program code configured to execute the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via the communication portion 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, the above-described functions defined in the system according to the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described system, device, apparatus, module, unit, etc. may be implemented by computer program modules.

[0107] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the methods according to the embodiments of the present disclosure are implemented.

[0108] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0109] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the above-described ROM 802 and / or RAM 803 and / or ROM 802 and RAM 803.

[0110] Embodiments of the present disclosure further include a computer program product, which includes a computer program. The computer program includes program code configured to execute the method provided by the embodiments of the present disclosure. When the computer program product runs on an electronic device, the program code is configured to cause the electronic device to implement the remote sensing image detection method based on a deep neural network provided by the embodiments of the present disclosure.

[0111] When the computer program is executed by the processor 801, the above functions defined in the system / apparatus of the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.

[0112] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part 809, and / or installed from the removable medium 811. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0113] According to embodiments of the present disclosure, program code configured to execute the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions configured to implement the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system configured to perform the specified functions or operations, or by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0115] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. A processing method, comprising: Obtaining the text to be processed, performing word segmentation on the text to be processed, and obtaining a plurality of ordered word tokens; The word tokens include semantic words, and the semantic words are the words in the text to be processed that represent the text semantics; For any one of the plurality of semantic words, connecting the next semantic word in sequence from the any one of the semantic words until the text content of the next semantic word is the same as that of the any one of the semantic words, and obtaining a plurality of first phrases; Merging any two first phrases with the same text content of the word token at the end of the sorting among the plurality of first phrases, and obtaining a plurality of second phrases; Screening at least some of the second phrases from the plurality of second phrases according to a predetermined rule, and splicing the at least some of the second phrases to obtain a plurality of text features; Generating the core information of the text to be processed according to the plurality of text features.

2. The method according to claim 1, the method further comprising: For any one of the plurality of semantic words, connecting the next semantic word in sequence from the any one of the semantic words until the number of semantic words between the next semantic word and the any one of the semantic words reaches a first preset number, stopping the connection and obtaining a first phrase.

3. According to the method described in claim 1, splicing the at least part of the second phrase includes: Splicing the at least some of the second phrases according to a second preset number; Before generating the core information of the text to be processed according to the text features, further comprising: Determining the weight of each of the second phrases among the at least some of the second phrases; Determining the weight of each of the text features according to the phrase weights of the second preset number of second phrases that are spliced into each of the text features; Screening some of the text features from the plurality of text features according to the weight of each of the text features and a preset weight threshold.

4. The method according to claim 3, for each of the second phrases among the at least some of the second phrases, determining the weight of the second phrase, comprising: Determining the weight of each first phrase; Determining the weight of the second phrase according to the weights of the first phrases that are merged into the second phrase; wherein, the weight of the second phrase is inversely proportional to the number of semantic words, and the number of semantic words is the total number of semantic words in the first phrases that are merged into the second phrase.

5. The method according to claim 3 or 4, the weight of the second phrase is further determined by the following method: Determining any two second phrases with the same text content of the word token at the end of the sorting from the at least some of the second phrases; Calculating the similarity between the any two second phrases; When the similarity is greater than a preset similarity threshold, respectively increasing the phrase weights of the any two second phrases according to the phrase similarity.

6. The method according to claim 4, the word token further includes stop words, and the stop words are the words in the text to be processed that do not represent the text semantics; connecting the next word token in sequence from the any one of the semantic words until the text content of the next semantic word is the same as that of the any one of the semantic words, including: Connecting the next word token in sequence from the any one word token, and skipping any one of the stop words until the text content of the next semantic word is the same as that of the any one of the semantic words.

7. The method according to claim 6, wherein determining the weight of each first phrase includes: Determining the weight of each token among the multiple tokens; According to the weight of each token, respectively obtaining the weights of the semantic words at the first and last positions in the first phrase, and the word frequencies of the semantic words at the first and last positions among the multiple tokens; Determining the number of stop words skipped in obtaining the first phrase; According to the weights of the semantic words at the first and last positions in the first phrase, the word frequencies, and the number of stop words skipped in the first phrase, obtaining the weight of each first phrase.

8. The method according to claim 1 or 3, wherein the text information includes statement information; Generating the text information of the text to be processed according to the text features further includes: Determining at least one semantic word constituting each text feature; Obtaining the position information of the at least one semantic word in the text to be processed; According to the position information, determining the corresponding statement information in the text to be processed.

9. The method according to claim 8, wherein the core information further includes paragraph information; generating the core information of the text to be processed according to the text features further includes: Determining the weights of the respective second phrases in at least some of the second phrases; Obtaining at least one second phrase corresponding to each semantic word in any one of the text statements; In any one of the text statements, obtaining the weight of any one of the text statements according to the sum of the weights of at least one second phrase corresponding to each semantic word therein; According to the weights of the text statements, screening the multiple text statements to obtain the paragraph information.

10. An electronic device, comprising at least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the following method: Obtaining a text to be processed, performing word segmentation processing on the text to be processed to obtain an ordered plurality of tokens; the tokens include semantic words; For any one of the multiple semantic words, sequentially connecting the next semantic word from the any one of the semantic words until the text content of the next semantic word is the same as that of the any one of the semantic words, to obtain a plurality of first phrases; Merging any two first phrases with the same text content of the tokens at the last position in the plurality of first phrases to obtain a plurality of second phrases; Screening at least some of the second phrases from the plurality of second phrases according to a predetermined rule, and splicing the at least some of the second phrases to obtain a plurality of text features; generating the core information of the text to be processed according to the plurality of text features.