Text processing method and device, electronic equipment, storage medium and program product
By constructing a network of association relationships between characters, words, and sentences, the problem of difficulty in capturing semantics caused by word segmentation errors in Chinese text sentiment analysis is solved, achieving more accurate text processing and sentiment analysis.
Patent Information
- Application Number
- CN202510713643.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-09
AI Technical Summary
Existing Chinese text sentiment analysis methods are prone to ambiguity or incorrect segmentation because the word segmentation results are limited by the word segmentation algorithm and dictionary, making it difficult for the model to accurately capture the semantic information of the text.
By generating a network of association relationships between characters, words, and sentences, a fourth network is established to fuse character-level, word-level, and sentence-level features to capture global semantics and contextual information.
The quality of text processing is improved, and it can more systematically capture the semantic and contextual information in the text, generating more accurate sentiment analysis results.
Smart Images

Figure CN120611713A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a text processing method, device, electronic device, storage medium, and program product. Background Art
[0002] In the field of natural language processing, Chinese text analysis technology has been widely used in multiple scenarios such as public opinion monitoring, user feedback analysis, and market trend forecasting.
[0003] In related technologies, sentiment analysis methods for Chinese texts mostly rely on word segmentation techniques, dividing text into word units and then constructing word vectors for sentiment classification. Because Chinese word segmentation results are limited by word segmentation algorithms and dictionaries, they are prone to ambiguity or incorrect segmentation, resulting in a loose model structure and difficulty in accurately capturing the text's semantic information.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention
[0005] The present disclosure provides a text processing method, device, electronic device, storage medium and program product, which at least to some extent overcome the problem of the related art in that it is difficult to accurately capture text semantic information.
[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.
[0007] According to one aspect of the present disclosure, a text processing method is provided, comprising: obtaining a text to be processed; determining a plurality of characters, a plurality of words, and a plurality of sentences in the text to be processed; generating a first network based on the relationship between the characters; generating a second network based on the relationship between the words; generating a third network based on the relationship between the sentences; establishing an association relationship among the first network, the second network, and the third network to obtain a fourth network; and generating a processed text based on the fourth network.
[0008] In a possible embodiment, establishing an association relationship among the first network, the second network, and the third network to obtain the fourth network includes: determining a first association relationship between the first network and the second network based on a relationship between characters and words; determining a second association relationship between the second network and the third network based on a relationship between words and sentences; determining a third association relationship between the first network and the third network based on a relationship between characters and sentences; and determining the fourth network based on the first association relationship, the second association relationship, and the third association relationship.
[0009] In one possible embodiment, generating a first network based on the relationship between each character includes: calculating the weight of each character in the corresponding word; generating a character node corresponding to each character, wherein the character node stores the character and the weight of the character; and establishing edges of each character node based on the adjacent relationship between each character to obtain the first network.
[0010] In one possible embodiment, the weight of a character in its own word is determined by the position of the character in the own word and the total number of characters in the own word.
[0011] In a possible embodiment, the weight of a character in the word to which it belongs is positively correlated with a first distance, wherein the first distance is the distance between the character and the first character in the word to which it belongs.
[0012] In one possible embodiment, generating a second network based on the relationship between each word includes: generating word nodes corresponding to each word; establishing a first edge between each word node based on the co-occurrence relationship between each word; determining the semantic relationship between each word; and establishing a second edge between each word node based on the semantic relationship between each word to obtain a second network.
[0013] In one possible embodiment, the semantic relationship between words includes the semantic similarity between the words; the text to be processed includes multiple texts; determining the semantic relationship between each word includes: counting the number of texts to be processed, the frequency of occurrence of each word in each text to be processed, and the number of different words appearing in the multiple texts to be processed; constructing a first matrix based on the number of texts to be processed, the frequency of occurrence of each word in each text to be processed, and the number of different words appearing in the multiple texts to be processed; determining the weight of each word; using the weight of each word to perform dimensionality reduction processing on the first matrix; and using the first matrix after dimensionality reduction to calculate the semantic similarity between each word.
[0014] In a possible embodiment, determining the weight of each word includes: calculating the weight of each word based on the amount of text to be processed, the frequency of occurrence of each word in each text to be processed, and a word weight calculation function.
[0015] In a possible embodiment, generating a third network based on the relationship between each sentence includes: generating sentence nodes and domain nodes corresponding to each sentence; establishing third edges between each sentence node based on the contextual relationship and semantic relationship between each sentence; determining the similarity between each sentence and each domain; and establishing fourth edges between each sentence node and the domain node based on the similarity between each sentence and each domain, to obtain the third network.
[0016] In one possible embodiment, determining the similarity between each of the sentences and each of the fields includes: determining the sentence vector of each of the sentences and the field vector of each field, wherein the sentence vector and the field vector are represented in the form of a two-dimensional matrix, the first row of the two-dimensional matrix is the term feature, and the other rows of the two-dimensional matrix except the first row are the feature attributes corresponding to each of the terms; setting an influence coefficient for each of the feature attributes in the sentence vector and the field vector; for each feature attribute, calculating the mean of each feature dimension in the sentence vector and the field vector; modifying the feature attribute value in the sentence vector using the mean of the sentence vector, and then modifying the feature attribute in the field vector using the mean of the field vector; and calculating the similarity between the modified sentence vector and the modified field vector using the cosine similarity algorithm.
[0017] In a possible embodiment, the method further includes: extracting entity information from the text to be processed and identifying the attribute relationship between each entity information; determining the corresponding node of each entity information in the fourth network; and establishing edges between the nodes corresponding to each entity information based on the attribute relationship between each entity information.
[0018] In a possible embodiment, the multiple characters of the text to be processed include a first character and a second character, the first character is a character extracted from the text to be processed, and the second character includes a character expanded from the first character; the multiple words of the text to be processed include a first word and a second word, the first word is a word extracted from the text to be processed, and the second word includes a word expanded from the first word; the multiple sentences of the text to be processed include a first sentence and a second sentence, the first sentence is a sentence extracted from the text to be processed, and the second sentence includes a sentence expanded from the first sentence based on the topic type.
[0019] According to another aspect of the present disclosure, a text processing device is also provided, characterized in that it includes: a text acquisition module for acquiring a text to be processed; a determination module for determining multiple characters, multiple words and multiple sentences in the text to be processed; a first network generation module for generating a first network based on the relationship between each character; a second network generation module for generating a second network based on the relationship between each word; a third network generation module for generating a third network based on the relationship between each sentence; a fourth network establishment module for establishing an association relationship between the first network, the second network and the third network to obtain a fourth network; and a text generation module for generating a processed text based on the fourth network.
[0020] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any one of the above-mentioned text processing methods by executing the executable instructions.
[0021] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any one of the above-mentioned text processing methods is implemented.
[0022] According to another aspect of the present disclosure, a computer program product is provided, including: a computer program or instructions, which implements any one of the above text processing methods when executed by a processor.
[0023] The text processing method provided in the embodiments of the present disclosure obtains a text to be processed; then determines multiple characters, multiple words, and multiple sentences in the text to be processed; then generates a first network based on the relationships between the characters; then generates a second network based on the relationships between the words; and then generates a third network based on the relationships between the sentences; then establishes an association relationship between the first network, the second network, and the third network to obtain a fourth network; and finally, generates processed text based on the fourth network. In this embodiment, character-level, word-level, and sentence-level features are integrated to obtain a fourth network that includes global feature information. This method can more systematically capture the semantic and contextual information in the text, thereby improving the quality of the generated text.
[0024] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0026] Figure 1 A schematic diagram of an exemplary application system architecture to which the text processing method in the embodiments of the present disclosure can be applied is shown;
[0027] Figure 2 A flow chart of a text processing method according to an embodiment of the present disclosure is shown;
[0028] Figure 3 A flowchart of another text processing method according to an embodiment of the present disclosure is shown;
[0029] Figure 4 A flow chart of a first network establishment method according to an embodiment of the present disclosure is shown;
[0030] Figure 5 A flow chart of a method for establishing a second network in an embodiment of the present disclosure is shown;
[0031] Figure 6 A flow chart of a method for calculating word similarity in an embodiment of the present disclosure is shown;
[0032] Figure 7 A flowchart of a third network establishment method according to an embodiment of the present disclosure is shown;
[0033] Figure 8 A flow chart of a method for calculating sentence and domain similarity in an embodiment of the present disclosure is shown;
[0034] Figure 9 A flow chart showing a method for establishing a knowledge graph according to an embodiment of the present disclosure is shown;
[0035] Figure 10 A flowchart of text processing according to an embodiment of the present disclosure is shown;
[0036] Figure 11 A schematic diagram of a text processing device according to an embodiment of the present disclosure is shown;
[0037] Figure 12 A structural block diagram of an electronic device in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0038] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0039] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0040] The specific implementation of the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.
[0041] Figure 1FIG. 1 shows an exemplary application system architecture diagram to which the text processing method in the embodiment of the present disclosure can be applied. Figure 1 As shown, the system architecture may include a terminal device 101 , a network 102 and a server 103 .
[0042] The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103 , and can be a wired network or a wireless network.
[0043] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of a virtual private network). In some embodiments, technologies and / or formats including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged over the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPSec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.
[0044] The terminal device 101 can be various electronic devices, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, wearable devices, augmented reality devices, virtual reality devices, etc.
[0045] Optionally, the client of the application installed in different terminal devices 101 is the same, or the client of the same type of application based on different operating systems. Based on the different terminal platforms, the specific form of the client of the application can also be different, for example, the application client can be a mobile phone client, a PC client, etc.
[0046] The server 103 may be a server that provides various services, such as a background management server that provides support for the devices operated by the user using the terminal device 101. The background management server may analyze and process the received request and other data, and feed back the processing results to the terminal device.
[0047] Optionally, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0048] Those skilled in the art will know that Figure 1 The number of terminal devices, networks, and servers in the embodiment is merely illustrative, and any number of terminal devices, networks, and servers may be provided based on actual needs. This embodiment of the present disclosure does not limit this.
[0049] Under the above system architecture, an embodiment of the present disclosure provides a text processing method, which can be executed by any electronic device with computing and processing capabilities.
[0050] In some embodiments, the text processing method provided in the embodiments of the present disclosure can be executed by a terminal device of the above-mentioned system architecture; in other embodiments, the text processing method provided in the embodiments of the present disclosure can be executed by a server in the above-mentioned system architecture; in other embodiments, the text processing method provided in the embodiments of the present disclosure can be implemented by the terminal device and server in the above-mentioned system architecture through interaction.
[0051] Figure 2 A flow chart of a text processing method according to an embodiment of the present disclosure is shown as follows: Figure 2 As shown, the text processing method provided in the embodiment of the present disclosure includes the following steps:
[0052] S202: Obtain the text to be processed.
[0053] Unprocessed text refers to raw text data collected from various channels, providing the raw material for subsequent character, word, and sentence-level analysis and network modeling. Unprocessed text can be a single text, such as a sentence, paragraph, or article. It can also be a text collection, i.e., a data set consisting of multiple texts. Furthermore, the unprocessed text can also be in the form of books, news, social media posts, Weibo, web pages, and other texts containing rich information.
[0054] The ways to obtain the text to be processed include the following:
[0055] 1) Directly receive text through the user interface. For example: receive the text manually entered by the user in the text input box.
[0056] 2) Use a programming language to read a local file. For example: use the open() function in Python to read a local file.
[0057] 3) Obtain the web page content through an HTTP (Hypertext Transfer Protocol) request, and extract the text in combination with an HTML parsing library.
[0058] 4) Obtain structured text data through an API (Application Programming Interface).
[0059] 5) Use a speech recognition tool to convert audio into text.
[0060] 5) Extract the text in the picture through OCR (Optical Character Recognition) technology.
[0061] S204. Determine multiple characters, multiple words, and multiple sentences in the text to be processed.
[0062] A character refers to the smallest basic unit in the text, mainly referring to single letters, Chinese characters, numbers, punctuation marks, etc. A word is a language unit composed of one or more characters and having independent meaning. For example: "car" in Chinese, "apple" in English, etc. A sentence is a language unit formed by combining words according to grammatical rules and expressing a complete meaning, usually ending with a full stop, a question mark, an exclamation mark, etc. For example: The sentence is "The battery life is too short".
[0063] In a possible implementation, after receiving the text to be processed, perform data preprocessing to remove redundant data, incorrect characters, etc.
[0064] Essentially, the text is a sequence of characters. You can directly traverse the text in order to obtain all characters. Exemplarily, use the "for char in text" function to traverse character by character to obtain multiple characters in the text to be processed. The text to be processed is "The battery life is too short", and the extracted characters are "电", "池", "续", "航", "太", "短".
[0065] In one possible implementation, the text to be processed is segmented into multiple words through word segmentation. For example, a predefined dictionary is used for string matching, using matching algorithms such as forward maximum match and reverse maximum match. Another example is training a word segmentation model using a machine learning model, learning the probability of word occurrence and contextual features based on a corpus, and using the trained word segmentation model to extract words from the text to be processed. For example, using forward maximum match word segmentation for "battery life is too short" yields the following words: "battery," "life," "too," and "short."
[0066] In one possible implementation, a natural language processing library is used to perform part-of-speech tagging on the text, identifying nouns within the text. The tagged words are then filtered to exclude nouns and retained. For example, regular expressions or filtering functions can be used to retain only words marked as nouns. Word frequency statistics can also be used to retain high-frequency nouns, improving relevance while reducing noise.
[0067] Further, the vocabulary is expanded, for example, by adding related nouns. A vocabulary is used to search for synonyms for nouns and add related nouns. Semantic models are used to find words similar to the original nouns based on semantic similarity. Pre-trained models such as BERT are used to adapt to specific domains to extract related nouns. By analyzing large amounts of text, the contextual relevance of nouns is identified and new nouns are generated. For example, a model can be trained to generate "fruit" and "healthy food" based on the context of "apple."
[0068] In one possible implementation, the text is split into independent sentences based on punctuation marks, sentence structure, or semantic boundaries. For example, sentences are split based on punctuation marks such as periods, question marks, and exclamation marks. Alternatively, complete sentences are identified in combination with grammatical rules. Grammatical rules include subject-predicate-object structures, etc. Another example is: a classification model is used to determine whether a position in the text is a sentence boundary, and the text to be processed is split into multiple sentences based on the identified sentence boundaries. For example, the text to be processed is "The battery life is too short. The charging speed is also slow. Hope it can be improved!", and the identified sentences are "The battery life is too short", "The charging speed is also slow", and "Hope it can be improved".
[0069] Identify and extract character sequences, word lists, and sentence sets from the text to be processed, providing structured data for subsequent character relationship analysis, word relationship analysis, and sentence relationship analysis.
[0070] S206: Generate a first network based on the relationship between the characters.
[0071] Characters are the most basic elements of text. The first network is a character-level network graph constructed primarily through the proximity relationships between characters. Character-level networks are typically sequential. In the first network, each character is a node, and the edges between adjacent characters represent the order in which the characters appear in the text. At this level, details such as spelling and character merging in the text are identified and processed.
[0072] For example, in the first network, each character, such as "electricity," "cell," "continue," and "aircraft," corresponds to a character node. Determine the order of adjacent characters. For example, "electricity" and "cell" are adjacent, with an edge of "sequential"; "continue" and "aircraft" are adjacent, with an edge of "sequential." Generate the character nodes "electricity" and "cell" from "battery," with the edge between them being "sequential-adjacent." Generate the character nodes "continue" and "aircraft" from "continue," with the edge between them being "sequential-adjacent."
[0073] S208: Generate a second network based on the relationship between the words.
[0074] A word is a basic unit composed of characters. The relationships between words are often analyzed through methods such as lexical semantics or word frequency. The second network is a word-level network graph structure. In the second network, each word is a node. The relationships between nodes can be co-occurrence relationships (for example, words that frequently appear together in a text), semantic associations (for example, synonyms and antonyms), or dependencies (for example, the subject-verb-object relationship in a grammatical structure).
[0075] The network graph at this level can be optimized through methods such as word embedding, so that the second network not only includes the frequency of word occurrence, but also captures their semantic associations.
[0076] In one possible implementation, the second network is a lexical graph structure. Specifically, keywords are extracted from the text processed at the sentence level, and their co-occurrence frequencies within the text are calculated. If two words co-occur in a certain number of sentences, an edge is established between them in the second network, and the edge type can be represented by the co-occurrence frequency.
[0077] In one possible implementation, the second network is a sentence similarity graph. Specifically, sentences are converted into vectors using TF-IDF, Word2Vec, or BERT. The similarity between sentences is calculated using cosine similarity or other similarity metrics. A similarity threshold is set, and only sentence pairs exceeding this threshold are established. The edge type can be represented by similarity.
[0078] In one possible implementation, the processed sentences or words are converted into feature vectors. Hierarchical clustering is used to cluster similar sentences or words into clusters. The sentences in each cluster are treated as nodes, and sentences in the same cluster are connected by edges.
[0079] In one possible implementation, the second network is a semantic network. Specifically, words and their relationships (synonyms, antonyms, hyponyms, etc.) are extracted from the text, each word is used as a node, and the edge type between the nodes is a semantic relationship.
[0080] In one possible implementation, the second network is a knowledge graph. Specifically, named entity recognition (NER) technology is used to extract entities from text, such as names, places, and organizations. Relationships between entities, such as "belongs to" and "is located in," are also extracted. Entities are represented as nodes, and edges between nodes are represented as relationships between entities.
[0081] For example, the segmented words include "battery," "battery life," and "charging speed." The co-occurrence relationship indicates that "battery" and "battery life" co-occur in multiple sentences, and the edge type is co-occurrence frequency. Semantic association: "battery life is too short" and "not durable" are synonyms, and the edge type is "synonymous." "Battery" and "battery life" are connected by a co-occurrence edge, and "too short" and "not durable" are connected by a synonymous edge.
[0082] S210: Generate a third network based on the relationship between the sentences.
[0083] The third network is a sentence-level network graph with a more complex structure. It needs to consider the logical relationships between sentences in the text being processed. For example, sentences may have the following relationships: contextual association: connecting different sentences through context; causal relationship: sentence A may be the cause of sentence B; grammatical dependency: the grammatical relationship between each word in a sentence and other words.
[0084] In the sentence-level tertiary network, each sentence can be considered a node, and the relationships between sentence nodes can be defined by contextual dependencies, semantic relationships, and discourse structure. For example, the segmented sentences are "Battery life is too short," "Charging speed is also slow," and "Hope for improvement." Both "Battery life is too short" and "Charging speed is also slow" discuss "battery performance," and the edge type is "topically similar." The logical relationship "Hope for improvement" summarizes the previous sentence, and the edge type is "Cause-Suggestion."
[0085] In one possible implementation, list topics of interest, such as "environmental protection," "technological innovation," and "social issues." Obtain sentences from news articles, social media, books, and other sources. Ensure that there are sufficient samples for each topic. Use a word segmentation tool to segment the sentences. Apply a stop word list to filter out common meaningless words. Reduce the words to their basic form for a unified representation. Use TF-IDF to obtain a feature vector for each sentence, which represents the importance of the key words in the sentence. You can choose to use a pre-trained model such as Word2Vec or GloVe to convert words into vectors. You can also use a deep learning model such as BERT to generate a vector representation of the sentence. Select an appropriate classification algorithm, divide the labeled sentence samples into training and test sets, and train the model. Perform the same preprocessing and feature extraction on new sentences. Use the trained model to classify the topics. Use the validation set to evaluate the model performance and calculate metrics such as precision, recall, and F1-score.
[0086] S212: Establish an association relationship among the first network, the second network, and the third network to obtain a fourth network.
[0087] The first network includes character nodes composed of characters, the second network includes word nodes composed of phrases, and the third network includes sentence nodes composed of sentences. Since the characters, phrases, and sentences are all extracted from the same text to be processed, edges are established between the character nodes of the first network, the word nodes of the second network, and the sentence nodes of the third network based on the relationships among the characters, phrases, and sentences, thereby obtaining a fourth network. The fourth network includes the first, second, and third networks.
[0088] From characters to words, and then to sentences, a hierarchical structure of increasing abstraction is formed. The character level provides basic building blocks, the word level organizes them into meaningful units through semantic and structural relationships, and the sentence level organizes them into complete expressions through grammar and discourse structure.
[0089] Sentences are composed of words arranged according to grammatical rules. The grammatical relationships between words are the basic structural units of a sentence. These relationships form the logical framework of a sentence, ensuring the integrity of its meaning. For example, the "subject-verb-object" structure clearly defines the initiator of an action, the action itself, and the object of the action.
[0090] Semantics and context are the "connecting threads" between sentences, forming higher-level paragraphs or entire texts through semantic associations and contextual information. Semantic associations include causality, progression, and transitions, while contextual information includes reference and thematic consistency.
[0091] Make multiple sentences form a coherent semantic whole rather than isolated language units, such as connecting sentences with logical words such as "therefore" and "however", or maintaining the consistency of the theme by repeating key words.
[0092] These network graphs, from the bottom-level character network to the upper-level sentence network, each have their own unique structure and function. Through layer-by-layer combination, a comprehensive network with multiple dimensions of syntax, semantics, and context is formed.
[0093] S214: Generate processed text based on the fourth network.
[0094] The fourth network is used to generate processed text. This expanded text collection, which includes the original reviews and the generated new sentences, is used to train customer service robots or product improvement models.
[0095] The text processing method provided in the embodiments of the present disclosure obtains a text to be processed; then determines multiple characters, multiple words, and multiple sentences in the text to be processed; then generates a first network based on the relationships between the characters; then generates a second network based on the relationships between the words; and then generates a third network based on the relationships between the sentences; then establishes an association relationship between the first network, the second network, and the third network to obtain a fourth network; and finally, generates processed text based on the fourth network. In this embodiment, character-level, word-level, and sentence-level features are integrated to obtain a fourth network that includes global feature information. This method can more systematically capture the semantic and contextual information in the text, thereby improving the quality of the generated text.
[0096] In a possible implementation, the above text processing method in this embodiment is further optimized, such as Figure 3 As shown, the optimized text processing method includes the following steps.
[0097] S302: Obtain the text to be processed.
[0098] S304: Determine multiple characters, multiple words, and multiple sentences in the text to be processed. The multiple characters of the text to be processed include a first character and a second character, where the first character is a character extracted from the text to be processed, and the second character includes a character expanded from the first character; the multiple words of the text to be processed include a first word and a second word, where the first word is a word extracted from the text to be processed, and the second word includes a word expanded from the first word; the multiple sentences of the text to be processed include a first sentence and a second sentence, where the first sentence is a sentence extracted from the text to be processed, and the second sentence includes a sentence expanded from the first sentence based on a topic type.
[0099] The first character refers to the native character directly extracted from the text to be processed, such as each Chinese character, letter, punctuation mark, etc. in the text to be processed. The second character is a new character generated by expanding the first character through a rule or algorithm, and is not a character directly contained in the original text.
[0100] The first word is a native word directly extracted from the text to be processed, such as the result of word segmentation. The second word is a new word generated by expanding the first word through semantic, grammatical and other rules.
[0101] The first sentence is a native sentence directly extracted from the text to be processed, such as the result of punctuation segmentation. The second sentence is a new sentence generated by expanding the first sentence based on the topic type, such as synonymous sentences or topic-related sentences. The topic type is the core topic or field covered in the text, which guides the direction of sentence expansion, such as science and technology, education, and healthcare.
[0102] In one possible implementation, a second character related to a first character is generated through rules or algorithms, including: completing character attributes, such as uppercase and lowercase conversion, traditional and simplified Chinese conversion; generating character-associated symbols, such as mathematical symbols and special characters; and language conversion, such as Chinese pinyin and English phonetic symbols.
[0103] In one possible implementation, a second word that is semantically related, grammatically derived, or domain-related is generated based on the first word, including: synonym / antonym expansion, such as "beautiful" expanding to "pretty"; hyponym expansion, such as "dog" expanding to "animal"; part-of-speech conversion, such as verb expansion to noun, "running" expanding to "running sports"; domain term expansion, such as "artificial intelligence" expanding to "machine learning" and "deep learning".
[0104] In one possible implementation, a second sentence is generated based on the topic type of the first sentence, which is semantically related, has a different sentence structure, or supplements information. This includes: generating synonymous sentences, such as changing "he ate the apple" to "the apple was eaten by him"; extending the topic, such as expanding "artificial intelligence is developing rapidly" to "machine learning is the core technology of artificial intelligence"; and expanding details, such as changing a short sentence into a long sentence and supplementing adverbials and attributives.
[0105] Through the hierarchical expansion of characters, words, and sentences, structured generation from original text (first element) to derived elements (second element) is achieved, providing richer node materials for the subsequent construction of multi-level networks (such as character relationship networks and word association networks), and enhancing the depth and flexibility of text processing.
[0106] S306 : Generate a first network based on the relationship between the characters, generate a second network based on the relationship between the words, and generate a third network based on the relationship between the sentences.
[0107] S308: Determine a first association relationship between the first network and the second network according to the relationship between the characters and the words.
[0108] The first network includes a plurality of character nodes, the second network includes a plurality of word nodes, and the first association relationship refers to an association relationship between the character nodes and the word nodes.
[0109] The relationship between characters and words includes a composition relationship, meaning that a word is composed of characters. For example, "lithium," "electricity," and "cell" form "lithium battery." The character node "electricity" and the word node "battery" have a "composition-left" relationship.
[0110] Determining the relationships between character nodes and word nodes links the underlying structural characters of a language with higher-level semantic words, forming a multi-granular knowledge network. Cross-layer connections capture the compositional patterns and semantic transmission paths of a language, providing richer structural information for subsequent text processing.
[0111] S310: Determine a second association relationship between the second network and the third network according to the relationship between the words and the sentences.
[0112] The second network includes multiple word nodes, the third network includes multiple sentence nodes, and the second association relationship refers to the association relationship between the word nodes and the sentence nodes.
[0113] The relationship between words and sentences includes inclusion. Words are the grammatical components of sentences, such as subject, predicate, and object. Word nodes are "endurance" and sentence nodes are "insufficient endurance." The relationship between word nodes and sentence nodes is "subject-core word."
[0114] The relationship between words and sentences includes logical associations. The grammatical relationship between words determines sentence structure. For example, "leads to" connects two clauses in a causal relationship. The relationship between the word node "leads to" and the sentence node "...leads to the user's need..." is a "causal-predicate trigger" relationship.
[0115] S312: Determine an association relationship between the first network and the third network according to a relationship between the characters and the sentences.
[0116] The first network includes a plurality of character nodes, the third network includes a plurality of sentence nodes, and the third association relationship refers to an association relationship between the character nodes and the sentence nodes.
[0117] The relationships between words and sentences include: semantic triggering relationships. Characters carry key semantics in sentences, such as the negative word "not" and the degree word "frequency". For the character "not" and the sentence "The battery life is insufficient", the association relationship between the character node and the sentence node is "negation - semantic core".
[0118] S314. Determine the fourth network according to the first association relationship, the second association relationship, and the third association relationship.
[0119] According to the first association relationship, establish an edge between the character node of the first network and the word node of the second network, and the edge type is the first association relationship. For example: establish an edge between the character node "electric" and the word node "battery", and the edge type is "composition - left part".
[0120] According to the second association relationship, establish an edge between the word node of the second network and the sentence node of the third network, and the edge type is the second association relationship. For example: establish an edge between the word node "battery life" and the sentence node "The battery life is insufficient", and the edge type is "subject - core word".
[0121] According to the third association relationship, establish an edge between the character node of the first network and the sentence node of the third network, and the edge type is the third association relationship. For example: establish an edge between the word node "cause" and the sentence node "... causes the user to need...", and the edge type is "causality - predicate trigger".
[0122] After establishing the edges between the character node and the word node, the word node and the sentence node, and the character node and the sentence node, the obtained network is the fourth network. In other words, the fourth network includes the first network, the second network, the third network, the edges between the character nodes of the first network and the word nodes of the second network, the edges between the word nodes of the second network and the sentence nodes of the third network, and the edges between the character nodes of the first network and the sentence nodes of the third network.
[0123] S316. Generate the processed text based on the fourth network.
[0124] These network graphs range from the underlying character network to the upper - level sentence network, and each level has its unique structure and function. Through layer - by - layer combination, a comprehensive network with multiple dimensions of syntax, semantics, and context is formed.
[0125] Based on the above - mentioned embodiments, this embodiment optimizes the construction method of the first network. As Figure 4 shown, the construction method of the first network provided in this embodiment mainly includes the following steps S402 - S406.
[0126] S402. Calculate the weight of each character in the word it belongs to; the weight of a character in the word it belongs to is determined by the position of the character in the word and the total number of characters in the word. The weight of a character in the word it belongs to is positively correlated with the first distance, where the first distance is the distance between the character and the first character in the word it belongs to.
[0127] The weight of a character in the word it belongs to is used to reflect the importance degree of the character in the word and its contribution to the semantics of the word.
[0128] The weights of each character in the word it belongs to are related to the position, rarity, and semantic information. The first and last characters usually have higher weights. For example, in "apple", the weight of "苹" is greater than that of the middle characters. Rare characters have higher weights. For example, in "饕餮", the weight of "饕" is higher than that of the common character "的". Characters carrying more semantic information have higher weights. For example, in "computer", the weight of "计" is higher than that of "机".
[0129] In this embodiment, a character weight calculation rule is provided. The weight of a character in a word is jointly determined by its position in the word and the length of the word. Characters that are more forward or backward in the word have a stronger decisive effect on the semantics of the word, so higher weights can be assigned. The shorter the word, the higher the weight of a single character. For example, the weight of "书" in "书" is 1, and the weight of "书" in "书籍" may be .6.
[0130] Introduce position-sensitive semantic analysis in text processing to improve the model's ability to understand language structure.
[0131] The first distance refers to the position interval between a certain character and the first character in the word it is in. For example: In "我爱语文", the position of the first character of "文" is 3. The positive correlation between the weight and the first distance means that the greater the first distance, the higher the weight value. That is, the more backward the character position, the higher the weight value.
[0132] In the composition of Chinese words, the feature of "center of gravity moving backward" is very obvious. For a word with a specific and specific concept, its theme center is often in the latter part of the word. Literally, the more backward the character, the greater the role it plays in expressing the theme concept. So here the character weight is determined by the position of the character in the word.
[0133] Calculating the weight of each character in the word it belongs to can be calculated by formula (1).
[0134]
[0135] Here, x represents the specific position of a character within a word, reflecting its order within the word. i is a position index from 1 to n, used to accumulate the sum of character positions within the word. n represents the total number of characters in the word, i.e., the length of the word. By calculating character weights, we provide a quantitative basis for the subsequent character-word composition relationship and word vector generation, ensuring that key characters are highlighted during word semantic analysis.
[0136] In formula (1), represents the weight of each character in the word, and x represents the position of the character in the word. The later the position of the character, the greater the weight it has.
[0137] Compared with the classical character weights, in this embodiment, a logarithmic function is used as the basis for weight calculation, so that the rate at which the weight of the character increases the further back it is in the position, the slower it gradually increases. This is different from the linear growth of the classical method, which may overemphasize the importance of the character at the very end. In the classical method, characters at the end of a word may receive very high weights, especially when the vocabulary length is long. The logarithmic function can effectively compress the impact of such extreme values and avoid the overall weight imbalance caused by the position of individual characters. For characters whose morpheme positions are relatively close but not exactly the same, the use of a logarithmic function can provide a more detailed distinction. This means that even if two characters are close in absolute position in a word, the difference in relative weights between them will be more obvious than when the classical method is used, thereby helping to improve the ability to distinguish between similar words.
[0138] S404: Generate character nodes corresponding to each character, wherein the character nodes store the character and the weight of the character.
[0139] For each character, a character node is created. In this embodiment, a character node can be understood as a basic unit in the graph data structure, used to store information. The information stored in each character node includes the character itself and the character weight calculated in S402. Each character has an associated numerical value that represents its relative importance within the corresponding word.
[0140] S406: Establish edges between character nodes based on the adjacent relationships between the characters to obtain a first network.
[0141] Edges in the first network represent direct adjacency between characters. If two characters appear consecutively in the text being processed, they are considered adjacent, and an edge forms between their corresponding nodes. In the first network, character nodes represent characters and their weights, while edges represent adjacency between characters.
[0142] By calculating character weights and constructing the first network, the text is converted into structured graph data to capture the adjacent relationships and semantic contributions between characters.
[0143] Based on the above embodiment, this embodiment optimizes the construction method of the second network, such as Figure 5 As shown, the method for constructing the second network provided in this embodiment mainly includes the following steps S502-S506.
[0144] S502: Generate word nodes corresponding to each word.
[0145] A word node refers to abstracting each word in the text to be processed into a node in the second network. Each node represents an independent word unit. For example, if the text to be processed is "Apple is a fruit, banana is also a fruit", then the word nodes include "apple", "is", "fruit", "banana", "also", and "is".
[0146] The text to be processed is segmented to obtain a word list. After removing duplicates, a node is created for each unique word. Node attributes can include the word itself, part of speech, and frequency of occurrence.
[0147] S504: Establish a first edge between each word node according to the co-occurrence relationship between each word.
[0148] The correlation between co-occurrence words in the text being processed is usually measured by the number of co-occurrences or frequency of co-occurrence. For example, "apple" and "fruit" appear multiple times in the same sentence, indicating a strong co-occurrence relationship. "Apple" and "banana" are also likely to co-occur frequently because they both belong to the fruit category.
[0149] The first edge is established between word nodes based on co-occurrence relationships, representing the strength of the co-occurrence relationship between words in the text. The edge type can be determined by indicators such as the number of co-occurrences or the co-occurrence probability. For example, if "apple" and "fruit" co-occur three times, the edge type of the first edge is 3.
[0150] Create a word × word matrix, matrix element M [i,j] Indicates the number of co-occurrences of word i and j. The diagonal is 0 (no co-occurrence). The edge type can use the matrix value directly or normalize it to a probability.
[0151] S506: Determine the semantic relationship between each word.
[0152] Semantic relationship refers to the semantic relevance between words, including synonymy, antonym, hyponym, whole-part, causal relationship, etc.
[0153] Determining the semantic relationship between words includes: using an existing semantic knowledge base to query the relationship between words. For example, "apple" is a hyponym of "fruit" in a hyponymy relationship.
[0154] The vector representation of the word is obtained through the word vector model, and the cosine similarity is calculated as the semantic similarity. A threshold is set. If the similarity between two words is greater than the set threshold, the edge type of the two words is the similarity value.
[0155] Combining part-of-speech tagging and rules to determine relationships. For example, in the structure of "noun + de + noun", the former may be an attribute of the latter. For example, in "the color of apple", "apple" and "color" are in an attribute relationship.
[0156] In a possible implementation, the method for determining the semantic relationship is optimized, such as Figure 6 As shown, the optimized semantic relationship determination process mainly includes S602-S606.
[0157] S602: Count the number of texts to be processed, the frequency of occurrence of each word in each text to be processed, and the number of different words appearing in multiple texts to be processed.
[0158] The frequency of occurrence refers to the number of times a word appears in a single document to be processed, i.e., the term frequency (TF). The number of distinct words refers to the total number of non-repeated words in all documents to be processed, i.e., the vocabulary size.
[0159] S604: Construct a first matrix based on the number of texts to be processed, the frequency of occurrence of each word in each text to be processed, and the number of different words appearing in multiple texts to be processed.
[0160] The first matrix refers to a matrix constructed based on text-word statistical information, usually a two-dimensional matrix of "text×word".
[0161] Compared to other languages, Chinese emphasizes word meaning, has a more complex sentence structure, and contains a large number of words with the same semantics. Based on this characteristic, when processing Chinese text, it is necessary to fully consider the impact of the semantic information of words on Chinese text expression and text classification, so as to achieve better classification results. Specifically, by first reducing the dimension of text features, the text keyword set obtained after word segmentation can more accurately express the meaning of the text at the semantic level. After the word segmentation preprocessing process, the text can be represented as a word × text matrix, which can be expressed by formula (2).
[0162] A=[a ij ] m×n (2)
[0163] Where A represents the text × text matrix; a ij It represents the frequency of the i-th word in the j-th text, which is a non-negative number; m represents the number of different words that appear in multiple texts to be processed; n represents the number of texts to be processed.
[0164] Furthermore, for any text, the matrix is composed of a clear number of keywords rather than all words. Therefore, the word × text matrix can be regarded as a sparse matrix.
[0165] S606: Determine the weight of each word.
[0166] S606a: Calculate the weight of each word based on the number of texts to be processed, the frequency of occurrence of each word in each text to be processed, and a word weight calculation function.
[0167] The word weight consists of the word global weight and the word local weight, which can be expressed by formula (3).
[0168]
[0169] Where L represents the frequency of a word in a text set; N represents the total number of texts in the text set.
[0170] S608: Perform dimensionality reduction processing on the first matrix using the weights of each word.
[0171] Dimensionality reduction refers to the process of reducing the dimension of a matrix, such as compressing a high-dimensional word space to a low-dimensional one.
[0172] The dimensionality reduction of the keyword-text matrix according to the word weight can be expressed by formula (4).
[0173] A * =W(i,j)×A (4)
[0174] Here, A* represents the word × text matrix after dimensionality reduction.
[0175] S610: Calculate the semantic similarity between each word using the first matrix after dimensionality reduction.
[0176] Semantic similarity refers to an indicator that measures the degree of semantic association between words, including but not limited to: cosine similarity, Euclidean distance, etc.
[0177] The semantic similarity of text words is represented by the word semantic relationship matrix, which can be expressed by formula (5).
[0178] η=A * ×R T (5)
[0179] Among them, η represents the semantic similarity of words in the text; R T Represents the word semantic relationship matrix, which is constructed using the Skip-Gram model and can capture the semantic information of words by predicting the central target word through context information.
[0180] In this embodiment, by performing dimensionality reduction on matrix A, the collapse of the original word-text matrix is eliminated, resulting in a significantly smaller approximate matrix. Keywords that best reflect the statistical characteristics of the text category are selected from the numerous original features, reducing the dimensionality of the text space. This is essentially a simplification of the text keywords. By reducing the dimensionality of text features, this embodiment more accurately expresses the meaning of the text at a semantic level, which helps reduce data redundancy and improve processing efficiency.
[0181] By improving the quantification of word semantic similarity, we can solve the problem of quantifying the similarity of a large number of semantically identical words in Chinese, provide accurate weight values for the semantic relationships between word nodes, and make the graph not only contain word co-occurrence information, but also capture deep semantic associations.
[0182] S508: Establish second edges between word nodes based on the semantic relationship between the words to obtain a second network.
[0183] Secondary edges are established between word nodes based on semantic relationships, representing the type and strength of the semantic association between the words. Secondary edges can be labeled with the relationship type, for example, "Synonymous," and the edge weight can be represented by a semantic similarity score. For example, "apple-fruit" is labeled "Hypernymous" with a similarity of 0.8.
[0184] The second network is a heterogeneous graph network composed of word nodes, first edges and second edges, which is used to comprehensively represent the contextual co-occurrence information and semantic information of words to improve the accuracy of semantic analysis.
[0185] Based on the above embodiment, this embodiment optimizes the construction method of the third network, such as Figure 7 As shown, the method for constructing the third network provided in this embodiment mainly includes the following steps S702-S708.
[0186] S702: Generate sentence nodes and domain nodes corresponding to each sentence.
[0187] Sentence nodes refer to sentences in a text abstracted as nodes in a graph structure. Each node represents an independent sentence and can include attributes such as sentence text and length. Domains refer to categories with specific themes, knowledge systems, or application scenarios. Domain nodes refer to the domains or themes corresponding to sentences abstracted as nodes in a graph structure, each representing a topic.
[0188] In one possible implementation, the text to be processed is segmented into sentences by punctuation marks or syntactic rules, and a sentence node is generated for each sentence. The sentence node attributes may include sentence ID, sentence text, and the position of the sentence in the original text.
[0189] In a possible implementation, each sentence is conceptualized, the domain of each sentence is extracted, and a domain node is generated for each domain.
[0190] S704: Establish a third edge between each sentence node according to the contextual relationship and semantic relationship between each sentence.
[0191] Contextual relationships refer to the relationship between sentences in a text and are used to express sentence order, coherence, and other logical connections, such as causality, progression, and transitions. Third edges are edges established between sentence nodes based on sentence contextual relationships and are used to express sentence order or logical coherence.
[0192] Semantic relations refer to the deep semantic connections between sentences, including synonymy, antonymity, implication, contradiction, etc., which reflect the similarity or difference in the meaning expressed by the sentences.
[0193] In one possible implementation, the relationship type is determined by the order of sentences in the text and their connectives. Directed edges are then created for adjacent sentences in that order. Edge attributes can be labeled with the relationship type, such as "sequence" or "continuation." Sentences are converted into vector representations to facilitate the calculation of semantic similarity. Alternatively, the semantic relationship type between sentences can be determined using models or rules.
[0194] S706: Determine the similarity between each sentence and each field.
[0195] The cosine similarity is used to calculate the angle between the sentence vector and the domain feature vector as the similarity between each sentence and the domain.
[0196] S708 : Based on the similarity between each sentence and each field, a fourth edge is established between each sentence node and the field node to obtain a third network.
[0197] The fourth edge refers to the edge established between the sentence node and the field node based on the similarity between each sentence and each field. The edge type is the similarity between the sentence and the field.
[0198] By constructing a sentence graph network that includes context, semantic relations, and domain similarity, it provides richer structured information for text analysis.
[0199] In one possible implementation, the method for determining the semantic relationship between sentences is optimized, such as Figure 8 As shown, the optimized semantic relationship determination process mainly includes S802-S806.
[0200] S802. Determine a sentence vector for each of the sentences and a domain vector for each of the domains, wherein the sentence vector and the domain vector are represented in the form of a two-dimensional matrix, the first row of the two-dimensional matrix is the term features, and the other rows of the two-dimensional matrix except the first row are the feature attributes corresponding to each of the terms.
[0201] A sentence vector represents a sentence's feature structure in a two-dimensional matrix. A domain vector represents a domain's feature structure in a two-dimensional matrix. The first row of the two-dimensional matrix contains the term features. The first row of the sentence vector contains the features extracted from the sentence, while the domain vector contains the features extracted from each domain. Subsequent rows contain the specific feature attribute values corresponding to each term, such as word frequency and part-of-speech tags.
[0202] Both sentence vectors and domain vectors can be expressed by formula (6).
[0203]
[0204] In formula (6), the text vector or domain vector T is stored as a two-dimensional matrix with the terms W in the first row. i , starting from the second line, the entry W i The corresponding feature attributes f mn .
[0205] S804: Setting influence coefficients for each feature attribute in the sentence vector and the domain vector.
[0206] The influence coefficient is a weight assigned to each feature attribute, used to adjust the feature's contribution to semantics. For example, domain features may be assigned a higher coefficient in specialized texts. Multiplying each feature attribute in the first sentence vector by the corresponding influence coefficient enhances the expressive power of the key feature.
[0207] In similarity calculation, it is necessary to integrate multiple feature attributes, so the influencing factor a is introduced. k , influence factor a k It is a constant, and its size needs to be determined according to the different structures of the text to be processed and the differences in the application fields.
[0208] After determining the vector representation, we need to determine the feature attributes of the term to be extracted. First, we use the TF-IDF algorithm to extract term features. Second, to verify the difference between single-feature and multi-feature similarity calculations, we extract several additional features from multi-feature text vectors and domain vectors in addition to the TF-IDF features and add them to the vectors. Considering the impact of data skew on the TF-IDF algorithm, we use term frequency distribution entropy and text distribution entropy to compensate for this deficiency.
[0209] S806. For each feature attribute, calculate the mean of each feature dimension in the sentence vector and the domain vector.
[0210] A feature dimension refers to the attribute row corresponding to each feature attribute. For example, the "word frequency" dimension in the above matrix. The mean refers to the average value of the attribute value for each feature dimension in the second sentence vector, which is used for subsequent normalization.
[0211] The mean of all feature attribute values in each feature dimension is calculated separately, including: extracting all attribute values of a certain dimension from the sentence vector and calculating the mean of the sentence vector in that dimension; extracting all attribute values of the same dimension from the domain vector and calculating the mean of the domain vector in that dimension.
[0212] S808. Modify the feature attribute values in the sentence vector using the mean value of the sentence vector, and modify the feature attributes in the domain vector using the mean value of the domain vector.
[0213] S810: Calculate the similarity between the third sentence vectors of each sentence using a cosine similarity algorithm to obtain the semantic similarity between each sentence.
[0214] Cosine similarity is to evaluate the semantic similarity by calculating the cosine value of the angle between two vectors. The closer the value is to 1, the more similar they are.
[0215] The similarity measure between text and category is calculated using the modified cosine function. and the domain vector The cosine calculation formula between is shown in formula (7).
[0216]
[0217] in, Represents sentence vector and the domain vector The similarity between Represents the mean of the sentence vector in this dimension, Represents the mean of the domain vector in dimension.
[0218] The multi-feature based text similarity judgment method extracts multiple feature attributes of text entries, extracts appropriate feature values according to the data structure and domain application characteristics of the text to be processed, and performs similarity judgment of Chinese texts. This is of great significance for improving the accuracy of the judgment results and the flexible adaptability of the method.
[0219] Improved methods for determining text similarity provide a more accurate measure of semantic relevance for sentence-level graphs, resolving the semantic misjudgment problem of traditional methods in complex text scenarios. By upgrading from single-word similarity to multi-dimensional semantic consistency, the edge relationships in sentence graphs more closely align with real-world semantic logic, laying a foundation for high-precision semantic relevance for subsequent data enhancement and application.
[0220] Based on the above embodiments, Figure 9 As shown, after "establishing an association relationship among the first network, the second network and the third network to obtain the fourth network", the following steps are also included.
[0221] S902: Extract entity information from the text to be processed, and identify attribute relationships between each entity information.
[0222] Named entity recognition (NER) is used to extract specific entities from the text to be processed, such as names of people, places, and organizations. The extracted information is the basis for building the knowledge graph, and the entity information is used to determine which elements should become nodes in the fourth network.
[0223] Natural language processing techniques are used to analyze sentence structure and context in text to identify relationships between entities. For example, relationships such as "belongs to," "is located in," and "acts on" are used to build edges in the graph, connecting different entity nodes.
[0224] S904: The corresponding node of each entity information in the fourth network.
[0225] The nodes of the query entity information in the fourth network are mapped to form a knowledge graph, with attribute relationships as edges. Each entity is a point in the graph, and the relationships between entities are represented by edges. The knowledge graph not only contains information about the entities themselves, but also includes the semantic connections between them, forming a multidimensional information network.
[0226] S906: Establish edges between nodes corresponding to each piece of entity information based on the attribute relationship between each piece of entity information.
[0227] Based on the constructed knowledge graph, we can further mine and correlate information across different dimensions. For example, we can use queries or reasoning to identify other entities related to a particular entity, or discover emerging potential relationships. This correlation helps deepen our understanding of the text content and provides more context for subsequent data enhancement.
[0228] As more text data is added, the knowledge graph needs to be continuously updated to maintain its timeliness and accuracy. At the same time, machine learning algorithms can automatically adjust and optimize the graph structure to improve the quality of information association.
[0229] On the basis of the comprehensive grid, through entity recognition (NER) and relationship extraction, the nodes and edges at the character level, word level, and sentence level are mapped to a broader knowledge system to form structured information in the knowledge base.
[0230] Based on the above embodiment, this embodiment provides an application example of a text processing method. Figure 10 As shown, it mainly includes the following steps.
[0231] A platform needs to analyze users' evaluations of "mobile phone batteries" and enhance text data by generating a grid structure to explore potential needs.
[0232] S1002: Input text to be processed.
[0233] Enter the user review text: "The battery life is too short and the charging speed is very slow. I hope it can be improved!" "The battery is not durable. A full charge only lasts half a day."
[0234] S1004: Data preprocessing.
[0235] After removing redundant symbols such as punctuation and incorrect characters, we can get the following word segmentation: "battery", "life", "too short", "charging", "speed", "very slow", "hope", "improvement", "not durable", "fully charged", "can only", "use", "half a day".
[0236] S1006. Data classification.
[0237] S1008, character level.
[0238] Each character corresponds to a node, such as "electricity", "cell", "continue", "aircraft", etc. The edge type is the order of adjacent characters, such as "electricity" and "cell" are adjacent, the edge is "sequence"; "continue" and "aircraft" are adjacent, the edge is "sequence").
[0239] S1010. Keep meaningful characters.
[0240] S1012. Supplement and expand according to the meaning of the characters.
[0241] Based on the semantic association of Chinese characters, we supplement and expand the related characters of "electricity", such as "energy" and "source".
[0242] S1014, word level.
[0243] After tokenization, each word corresponds to a node, such as "battery," "battery life," and "charging speed." Edge types include: Co-occurrence relationships: "Battery" and "battery life" co-occur in multiple sentences, with the edge weight being the co-occurrence frequency, such as 2. Semantic associations: "Battery life is too short" and "not durable" are synonyms, with the edge type being "synonymous." For example, "battery" and "battery life" are connected by a co-occurrence edge, and "too short" and "not durable" are connected by a synonymous edge.
[0244] S1016. Retain nouns.
[0245] S1018. Add relevant nouns.
[0246] Add relevant nouns such as "lithium battery" and "battery life" through WordNet synonym expansion.
[0247] S1020. Sentence level.
[0248] The segmented sentences correspond to one node. For example, "The battery life is too short and the charging speed is also very slow. I hope to improve it!" is a sentence node. The edge types include: Topic association: Both sentences discuss "battery performance", and the edge type is "topic similarity", and the edge weight is calculated by cosine similarity (such as 0.85). Logical relationship: "I hope to improve it!" is the summary of the previous sentence, and the edge type is "cause - suggestion".
[0249] S1022. Classify by topic.
[0250] S1024. Expand by topic.
[0251] Generate new sentences by expanding according to the topic classification. For example, expand new sentences such as "The battery standby time is short, which affects the user experience" based on the "battery life" topic.
[0252] S1026. Form a network structure.
[0253] The character "electricity" to the word "battery" to the sentence "The battery life is too short" is connected by "composition relationship" and "topic association" edges. The final network structure contains character, word, and sentence nodes, as well as a dynamic grid with various edge types such as sequence, co-occurrence, synonymy, topic, and causality.
[0254] S1028. Output the enhanced text data.
[0255] The expanded text set contains the original evaluation and the generated new sentences, such as "The lithium battery has insufficient battery life and low charging efficiency", which is used to train a customer service robot or a product improvement model.
[0256] It should be noted that in the technical solution of this disclosure, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations. In the embodiments of this disclosure, various types of data such as personal identity data, operation data, and behavior data related to individuals, customers, and groups have been authorized.
[0257] Based on the same inventive concept, the present disclosure also provides a text processing device, such as the following embodiment. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0258] Figure 11 A schematic diagram of a text processing device according to an embodiment of the present disclosure is shown. Figure 11 As shown, the apparatus includes: a text acquisition module 1110 , a determination module 1120 , a first network generation module 1130 , a second network generation module 1140 , a third network generation module 1150 , a fourth network establishment module 1160 and a text generation module 1170 .
[0259] A text acquisition module 1110 is used to acquire a text to be processed; a determination module 1120 is used to determine multiple characters, multiple words, and multiple sentences in the text to be processed; a first network generation module 1130 is used to generate a first network based on the relationship between each character; a second network generation module 1140 is used to generate a second network based on the relationship between each word; a third network generation module 1150 is used to generate a third network based on the relationship between each sentence; a fourth network establishment module 1160 is used to establish an association relationship between the first network, the second network, and the third network to obtain a fourth network; and a text generation module 1170 is used to generate a processed text based on the fourth network.
[0260] In one possible implementation, the fourth network establishing module 1160 is specifically configured to establish an association relationship between the first network and the second network based on the relationship between characters and words; establish an association relationship between the second network and the third network based on the relationship between words and sentences; and establish an association relationship between the first network and the third network based on the relationship between characters and sentences, thereby obtaining a fourth network.
[0261] In one possible implementation, the first network generation module 1130 is specifically used to calculate the weight of each character in the corresponding word; generate a character node corresponding to each character, wherein the character node stores the character and the weight of the character; establish edges of each character node based on the adjacent relationship between each character to obtain the first network.
[0262] In one possible implementation, the weight of a character in its word is determined by the position of the character in the word and the total number of characters in the word.
[0263] In one possible implementation, the weight of a character in the word to which it belongs is positively correlated with a first distance, where the first distance is the distance between the character and the first character in the word to which it belongs.
[0264] In one possible implementation, the second network generation module 1140 is specifically used to generate word nodes corresponding to each word; establish a first edge between each word node based on the co-occurrence relationship between each word; determine the semantic relationship between each word; and establish a second edge between each word node based on the semantic relationship between each word to obtain a second network.
[0265] In one possible implementation, the semantic relationship between words includes the semantic similarity between the words; the text to be processed includes multiple texts; the second network generation module 1140 is specifically used to count the number of texts to be processed, the frequency of occurrence of each word in each text to be processed, and the number of different words appearing in the multiple texts to be processed; based on the number of texts to be processed, the frequency of occurrence of each word in each text to be processed, and the number of different words appearing in the multiple texts to be processed, construct a first matrix; determine the weight of each word; use the weight of each word to reduce the dimension of the first matrix; and use the first matrix after dimension reduction to calculate the semantic similarity between each word.
[0266] In one possible implementation, the second network generation module 1140 is specifically configured to calculate the weight of each word based on the number of texts to be processed, the frequency of occurrence of each word in each text to be processed, and a word weight calculation function.
[0267] In one possible implementation, the third network generation module 1150 is specifically used to generate sentence nodes and domain nodes corresponding to each sentence; establish third edges between each sentence node based on the contextual relationship and semantic relationship between each sentence; determine the similarity between each sentence and each domain; and establish fourth edges between each sentence node and domain node based on the similarity between each sentence and each domain, thereby obtaining a third network.
[0268] In one possible implementation, the third network generation module 1150 is specifically used to determine the sentence vector of each sentence and the domain vector of each domain, wherein the sentence vector and the domain vector are represented in the form of a two-dimensional matrix, the first row of the two-dimensional matrix is the term feature, and the other rows of the two-dimensional matrix except the first row are the feature attributes corresponding to each term; an influence coefficient is set for each feature attribute in the sentence vector and the domain vector; for each feature attribute, the mean of each feature dimension in the sentence vector and the domain vector is calculated; the feature attribute value in the sentence vector is modified using the mean of the sentence vector, and the feature attribute in the domain vector is modified using the mean of the domain vector; and the similarity between the modified sentence vector and the modified domain vector is calculated using the cosine similarity algorithm.
[0269] In one possible implementation, it also includes: a knowledge graph construction module, which is used to extract entity information from the text to be processed and identify the attribute relationship between each entity information; determine the corresponding node of each entity information in the fourth network; and establish edges between the nodes corresponding to each entity information based on the attribute relationship between each entity information.
[0270] In one possible implementation, the multiple characters of the text to be processed include a first character and a second character, the first character is a character extracted from the text to be processed, and the second character includes a character expanded from the first character; the multiple words of the text to be processed include a first word and a second word, the first word is a word extracted from the text to be processed, and the second word includes a word expanded from the first word; the multiple sentences of the text to be processed include a first sentence and a second sentence, the first sentence is a sentence extracted from the text to be processed, and the second sentence includes a sentence expanded from the first sentence based on the topic type.
[0271] It should be noted that the examples and application scenarios implemented by the modules in the above-mentioned apparatus embodiment are the same as those implemented by the corresponding steps in the method embodiment, but are not limited to the contents disclosed in the above-mentioned method embodiment. It should be noted that the above-mentioned modules, as part of the apparatus, can be executed in a computer system, such as a set of computer-executable instructions.
[0272] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation that combines hardware and software aspects, which may be collectively referred to herein as a "circuit," "module," or "system."
[0273] Based on the same inventive concept, an embodiment of the present disclosure further provides an electronic device, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any of the above-mentioned text processing methods by executing the executable instructions. Since the principles for solving the problem in this electronic device embodiment are similar to those in the above-mentioned method embodiment, the implementation of this electronic device embodiment can refer to the implementation of the above-mentioned method embodiment, and the repeated parts are not repeated here.
[0274] Refer to the following Figure 12 1200 according to this embodiment of the present disclosure will be described. Figure 12 The electronic device 1200 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0275] like Figure 12As shown, electronic device 1200 is implemented as a general-purpose computing device. Components of electronic device 1200 may include, but are not limited to, the aforementioned at least one processing unit 1210, the aforementioned at least one storage unit 1220, and a bus 1230 connecting various system components (including storage unit 1220 and processing unit 1210).
[0276] The storage unit stores program code, which can be executed by the processing unit 1210, so that the processing unit 1210 performs the steps described in the "Exemplary Method" section of this specification according to various exemplary embodiments of the present disclosure. For example, the processing unit 1210 can perform the following steps of the above-mentioned method embodiment: obtaining a text to be processed; determining multiple characters, multiple words, and multiple sentences in the text to be processed; generating a first network based on the relationship between each of the characters; generating a second network based on the relationship between each of the words; generating a third network based on the relationship between each of the sentences; establishing an association relationship between the first network, the second network, and the third network to obtain a fourth network; and generating a processed text based on the fourth network.
[0277] The storage unit 1220 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 12201 and / or a cache 12202 , and may further include a read-only memory unit (ROM) 12203 .
[0278] The storage unit 1220 may also include a program / utility 12204 having a set (at least one) of program modules 12205, such program modules 12205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0279] The bus 1230 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0280] Electronic device 1200 can also communicate with one or more external devices 1240 (e.g., a keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with electronic device 1200, and / or any device that enables electronic device 1200 to communicate with one or more other computing devices (e.g., a router, modem, etc.). Such communication can occur via input / output (I / O) interface 1250. Furthermore, electronic device 1200 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via network adapter 1260. As shown, network adapter 1260 communicates with other modules of electronic device 1200 via bus 1230. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with electronic device 1200, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0281] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0282] Based on the same inventive concept, embodiments of the present disclosure also provide a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements any of the aforementioned text processing methods. Because the principles underlying the problems solved by this computer-readable storage medium embodiment are similar to those of the aforementioned method embodiment, the implementation of this computer-readable storage medium embodiment can be referenced to the implementation of the aforementioned method embodiment, and any repetitions will not be repeated.
[0283] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0284] In the present disclosure, a computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0285] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.
[0286] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0287] Based on the same inventive concept, embodiments of the present disclosure further provide a computer program product, including a computer program or instructions, which, when executed by a processor, implements the text processing method of any one of the above-described method embodiments. Because the principles for solving the problems of this computer program product embodiment are similar to those of the above-described method embodiments, the implementation of this computer program product embodiment can refer to the implementation of the above-described method embodiments, and any repetitions will not be repeated.
[0288] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0289] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0290] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0291] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
Claims
1. A text processing method, characterized in that: include: Get the text to be processed; Determining a plurality of characters, a plurality of words, and a plurality of sentences in the text to be processed; generating a first network according to the relationship between the characters; generating a second network based on the relationships between the words; generating a third network based on the relationship between the sentences; Establishing an association relationship among the first network, the second network, and the third network to obtain a fourth network; A processed text is generated based on the fourth network.
2. The text processing method according to claim 1, characterized in that: The establishing of an association relationship among the first network, the second network, and the third network to obtain a fourth network includes: determining a first association relationship between the first network and the second network according to a relationship between the character and the word; determining a second association relationship between the second network and the third network according to the relationship between the word and the sentence; determining a third association relationship between the first network and the third network according to the relationship between the character and the sentence; A fourth network is determined according to the first association relationship, the second association relationship, and the third association relationship.
3. The text processing method according to claim 1, characterized in that: Generating a first network according to the relationship between the characters includes: Calculate the weight of each character in the word to which it belongs; Generating a character node corresponding to each of the characters, wherein the character node stores the character and the weight of the character; The edges of the character nodes are established based on the adjacent relationships between the characters to obtain the first network.
4. The text processing method according to claim 3, characterized in that: The weight of the character in the word to which it belongs is determined by the position of the character in the word to which it belongs and the total number of characters in the word to which it belongs.
5. The text processing method according to claim 4, characterized in that: The weight of the character in the word to which it belongs is positively correlated with the first distance, wherein the first distance is the distance between the character and the first character in the word to which it belongs.
6. The text processing method according to claim 1, characterized in that: Generating a second network according to the relationship between the words includes: Generate word nodes corresponding to each of the words; Establishing a first edge between each of the word nodes according to the co-occurrence relationship between each of the words; determining semantic relationships between each of the words; Second edges between the word nodes are established based on the semantic relationships between the words to obtain the second network.
7. The text processing method according to claim 6, characterized in that: The semantic relationship between the words includes the semantic similarity between the words; the text to be processed includes multiple texts; Determining the semantic relationship between the words includes: Counting the number of the texts to be processed, the frequency of occurrence of each word in each text to be processed, and the number of different words appearing in multiple texts to be processed; constructing a first matrix based on the number of the texts to be processed, the frequency of occurrence of each word in each of the texts to be processed, and the number of different words appearing in the plurality of texts to be processed; Determining the weight of each of the words; Performing dimensionality reduction processing on the first matrix using the weights of the respective words; The semantic similarity between the words is calculated using the first matrix after dimensionality reduction.
8. The text processing method according to claim 7, characterized in that: Determining the weight of each of the words includes: The weight of each of the words is calculated based on the number of the texts to be processed, the frequency of occurrence of each of the words in each of the texts to be processed, and a word weight calculation function.
9. The text processing method according to claim 1, wherein: Generating a third network according to the relationship between the sentences includes: Generate sentence nodes and domain nodes corresponding to each of the sentences; Establishing a third edge between each of the sentence nodes according to the contextual relationship and semantic relationship between each of the sentences; Determining the similarity between each of the sentences and each of the fields; Based on the similarity between each of the sentences and each of the fields, a fourth edge is established between each of the sentence nodes and the field nodes to obtain the third network.
10. The text processing method according to claim 9, characterized in that: Determining the similarity between each of the sentences and each of the fields includes: Determine a sentence vector for each of the sentences and a domain vector for each of the domains, wherein the sentence vector and the domain vector are represented in the form of a two-dimensional matrix, the first row of the two-dimensional matrix is the term features, and the other rows of the two-dimensional matrix except the first row are the feature attributes corresponding to each of the terms; Setting an influence coefficient for each of the feature attributes in the sentence vector and the domain vector respectively; For each feature attribute, calculate the mean of each feature dimension in the sentence vector and the domain vector; Modifying the characteristic attribute values in the sentence vector using the mean of the sentence vector, and modifying the characteristic attributes in the domain vector using the mean of the domain vector; The cosine similarity algorithm is used to calculate the similarity between the modified sentence vector and the modified domain vector.
11. The text processing method according to claim 1, wherein: Also includes: Extracting entity information from the text to be processed and identifying attribute relationships between each entity information; Determining a corresponding node of each entity information in the fourth network; According to the attribute relationship between each entity information, edges between nodes corresponding to each entity information are established.
12. The text processing method according to any one of claims 1 to 5, characterized in that: The plurality of characters of the text to be processed include a first character and a second character, the first character is a character extracted from the text to be processed, and the second character includes a character obtained by expanding the first character; The plurality of words in the text to be processed include a first word and a second word, wherein the first word is a word extracted from the text to be processed, and the second word includes a word expanded from the first word; The multiple sentences of the text to be processed include a first sentence and a second sentence, the first sentence is a sentence extracted from the text to be processed, and the second sentence includes a sentence expanded from the first sentence based on a topic type.
13. A text processing device, characterized in that: include: A text acquisition module is used to obtain the text to be processed; A determination module, configured to determine a plurality of characters, a plurality of words, and a plurality of sentences in the text to be processed; A first network generating module, configured to generate a first network based on the relationship between the characters; A second network generating module, configured to generate a second network based on the relationship between the words; A third network generating module, configured to generate a third network based on the relationship between the sentences; a fourth network establishing module, configured to establish an association relationship among the first network, the second network, and the third network to obtain a fourth network; A text generation module is used to generate processed text based on the fourth network.
14. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the text processing method according to any one of claims 1 to 12 by executing the executable instructions.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the text processing method according to any one of claims 1 to 12 is implemented.
16. A computer program product comprising: A computer program or instruction, characterized in that when the computer program or instruction is executed by a processor, it implements the text processing method described in any one of claims 1 to 12.