Text processing method and apparatus, device, medium

By extracting text and performing semantic analysis on the original corpus, and using hyper- and hypo-cognitive graph linking for word segmentation, the problem of efficiently extracting information from long texts is solved, achieving both high efficiency and accuracy in information retrieval.

CN114996458BActive Publication Date: 2026-04-24CHINA PING AN LIFE INSURANCE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA PING AN LIFE INSURANCE CO LTD
Filing Date
2022-06-28
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently extract useful information from large, lengthy, and complex original corpora.

Method used

By extracting text from the original corpus, preliminary summary text is obtained, and the text is divided, filtered, and semantically analyzed according to the length of the segmented text. The word segments are then linked to the pre-defined hyper- and hypo-cognitive graphs based on their word types.

Benefits of technology

It improves the efficiency of information acquisition, making it easy to know the superordinate and subordinate information of each word segment and the relationship between word segments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114996458B_ABST
    Figure CN114996458B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of text processing method and device, equipment, storage medium, belong to artificial intelligence technical field.The method comprises: by text extraction processing to original corpus, extract the preliminary abstract text that can express the meaning of original corpus most, again according to the preset segmented text length, preliminary abstract text is divided, and target segmented text is obtained.Target segmented text is filtered to determine target candidate word group, and target candidate word group semantic analysis processing is handled, the importance of word segmentation can be determined by the word type of word segmentation, and according to the word type of word segmentation, word segmentation is linked to the preset upper and lower cognitive graph, and target upper and lower cognitive graph is obtained.Because the upper and lower information of each word segmentation and the connection between each word segmentation can be easily obtained through target upper and lower knowledge graph, therefore, the efficiency of obtaining information can be improved by the text processing method of the embodiment of the application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a text processing method, apparatus, device, and medium. Background Technology

[0002] Currently, major news media outlets, public accounts, and bloggers generate a large amount of raw text data daily, including but not limited to news reports, commentary, predictions, and analyses. This raw text data is often lengthy, complex, and contains differing viewpoints, making it difficult to extract useful information directly. Therefore, providing a text processing method that improves the efficiency of information retrieval has become a pressing technical problem. Summary of the Invention

[0003] The main objective of this application is to provide a text processing method, apparatus, device, and medium that can improve the efficiency of information acquisition.

[0004] To achieve the above objectives, a first aspect of this application provides a text processing method, the method comprising:

[0005] Obtain the original corpus;

[0006] The original corpus is subjected to text extraction processing to obtain preliminary summary text;

[0007] The preliminary summary text is divided into multiple preliminary segmented texts according to the preset segmented text length;

[0008] The multiple preliminary segmented texts are filtered to obtain multiple preliminary candidate word groups;

[0009] Calculate the similarity value between each preliminary candidate word group and the original corpus, and take the preliminary candidate word groups that meet the preset similarity conditions as target candidate word groups;

[0010] The target candidate word groups are subjected to semantic parsing to obtain multiple word segments and word types of the word segments;

[0011] Based on the word type of the segmented words, multiple segmented words are linked to a preset hypernym / hypernym cognitive graph to obtain a target hypernym / hypernym cognitive graph.

[0012] In some embodiments, filtering the multiple preliminary segmented texts to obtain multiple preliminary candidate word groups includes:

[0013] Obtain a reference scene classification for the original corpus;

[0014] The multiple preliminary segmented texts are predicted using a preset classification prediction model to obtain the predicted scenario classification for each preliminary segmented text.

[0015] The multiple preliminary segmented texts are filtered based on the matching relationship between the reference scene classification and the predicted scene classification;

[0016] Based on the filtered preliminary segmented texts, several preliminary candidate word groups are determined.

[0017] In some embodiments, filtering the multiple preliminary segmented texts to obtain multiple preliminary candidate word groups includes:

[0018] Obtain a reference scene classification for the original corpus;

[0019] A keyword set is determined based on the reference scenario classification, and the keyword set includes multiple keywords;

[0020] Calculate the word matching value between each of the preliminary segmented texts and the multiple keywords, and obtain multiple preliminary candidate word groups based on the multiple preliminary segmented texts that meet the preset matching conditions.

[0021] In some embodiments, filtering the multiple preliminary segmented texts to obtain multiple preliminary candidate word groups includes:

[0022] Multiple preliminary segmented texts are filtered using a preset BERT quality model to obtain multiple target segmented texts.

[0023] The text similarity between each target segment and the original corpus is calculated using a pre-defined sifRank semantic model.

[0024] Multiple target segmented texts are filtered based on the text similarity to obtain multiple candidate word groups.

[0025] In some embodiments, linking multiple word segments to a preset hypernym / hypernym cognitive graph based on the word type of the word segmentation to obtain a target hypernym / hypernym cognitive graph includes:

[0026] Obtain a reference scene classification for the original corpus;

[0027] The scene node type is determined based on the reference scene classification;

[0028] If the nodes in the hyper- and hyper-level cognitive graph do not include the scene node type, the node to be processed is determined from the hyper- and hyper-level cognitive graph based on the correlation of the scene node type.

[0029] Create scene nodes corresponding to the scene node type under the node to be processed to obtain a preliminary hierarchical cognitive map;

[0030] Based on the word type of the segmented words, multiple segmented words are linked to the preliminary hypernym-hypernym cognitive graph to obtain the target hypernym-hypernym cognitive graph.

[0031] In some embodiments, linking multiple word segments to a preset hypernym / hypernym cognitive graph based on the word type of the word segmentation to obtain a target hypernym / hypernym cognitive graph includes:

[0032] Construct a preliminary word segmentation set based on the multiple word segments described above;

[0033] If it is determined that the word type of the segmented word does not belong to the preset word type set, the word segment corresponding to the word type of the segmented word is deleted from the preliminary word segmentation set to obtain the target word segmentation set, which includes multiple target word segments; wherein, the word type set includes region category, action category, and quantity category;

[0034] Based on the word type of the segmented words, multiple target segmented words are linked to a preset hypernym / hypernym cognitive graph to obtain a target hypernym / hypernym cognitive graph.

[0035] In some embodiments, the text extraction process performed on the original corpus to obtain preliminary summary text includes:

[0036] The original corpus is processed by a preset event classification prediction model to obtain at least two corpus event types.

[0037] Calculate the degree of matching between each of the corpus event types and preset key events to obtain an event matching value;

[0038] The target event type is determined based on the event matching value and at least two of the corpus event types;

[0039] The original corpus is truncated according to a preset truncated text length to obtain at least two original summary texts;

[0040] The preliminary summary text is determined based on the degree of matching between the original summary text and the target event type.

[0041] To achieve the above objectives, a second aspect of this application provides a text processing apparatus, the apparatus comprising:

[0042] The acquisition module is used to acquire the raw corpus;

[0043] The abstract text extraction module is used to extract text from the original corpus to obtain preliminary abstract text.

[0044] The text segmentation module is used to divide the preliminary summary text into multiple preliminary segmented texts according to a preset segmented text length;

[0045] The preliminary candidate word group generation module is used to filter multiple preliminary segmented texts to obtain multiple preliminary candidate word groups;

[0046] The target candidate word group generation module is used to calculate the similarity value between each of the preliminary candidate word groups and the original corpus, and to take the preliminary candidate word groups that meet the preset similarity conditions as target candidate word groups;

[0047] The semantic parsing processing module is used to perform semantic parsing processing on the target candidate word groups to obtain multiple word segments and word types of the word segments;

[0048] The word linking module is used to link multiple word segments to a preset hyper-hyper-cognitive graph according to the word type of the word segmentation, so as to obtain a target hyper-hyper-cognitive graph.

[0049] To achieve the above objectives, a third aspect of this application provides a computer device, the computer device including a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for implementing communication between the processor and the memory, wherein the program, when executed by the processor, implements the method described in the first aspect.

[0050] To achieve the above objectives, a fourth aspect of the present application provides a storage medium, which is a computer-readable storage medium for computer-readable storage, wherein the storage medium stores one or more programs that can be executed by one or more processors to implement the method described in the first aspect.

[0051] The text processing method, apparatus, device, and medium proposed in this application extract text from the original corpus to obtain a preliminary summary text that best expresses the meaning of the original corpus. The preliminary summary text is then divided according to a preset segment length to obtain target segment text. The target segment text is filtered to determine target candidate word groups, and semantic analysis is performed on these candidate word groups. The importance of each word segment is determined by its word type, and the word segments are linked to a preset hyper- and hypo-level cognitive graph based on their word types to obtain the target hyper- and hypo-level cognitive graph. Since the hyper- and hypo-level information of each word segment and the relationships between them can be easily obtained through the target hyper- and hypo-level cognitive graph, the text processing method of this application can improve the efficiency of information acquisition. Attached Figure Description

[0052] Figure 1 This is a flowchart of the text processing method provided in the embodiments of this application;

[0053] Figure 2 yes Figure 1 The flowchart of step S102 in the document;

[0054] Figure 3 yes Figure 1 The flowchart of step S104 in the process;

[0055] Figure 4 yes Figure 1 The flowchart of step S104 in the process;

[0056] Figure 5 yes Figure 1 The flowchart of step S104 in the process;

[0057] Figure 6 yes Figure 1 The flowchart of step S107 in the process;

[0058] Figure 7 yes Figure 1 The flowchart of step S107 in the process;

[0059] Figure 8 This is a block diagram of the module structure of the text processing method apparatus provided in the embodiments of this application;

[0060] Figure 9 This is a schematic diagram of the hardware structure of the computer device provided in the embodiments of this application. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0062] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0064] First, let's analyze some of the terms used in this application:

[0065] Artificial Intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0066] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information and image processing, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.

[0067] Information Extraction (NER) is a text processing technique that extracts factual information such as entities, relationships, and events from natural language text and outputs it as structured data. Information extraction is a technique for extracting specific information from text data. Text data is composed of specific units, such as sentences, paragraphs, and chapters. Text information is composed of smaller, specific units, such as characters, words, phrases, sentences, paragraphs, or combinations of these units. Extracting noun phrases, names of people, and place names from text data is an example of text information extraction. Of course, text information extraction techniques can extract information of various types.

[0068] Corpus: Linguistic material, the basic unit of a corpus, is typically a collection of text resources of a certain quantity and scale. Corpus sizes can vary greatly, ranging from tens of millions, even hundreds of millions of sentences or more, to as few as a few hundred sentences. People simply use text as a substitute, and regard the contextual relationships within the text as a substitute for the contextual relationships in real-world language. A collection of texts can be called a corpus, and when there are several such collections, it can be called a corpus set. The internet itself is a vast and complex corpus. Corpora can be classified in many ways according to different criteria; for example, a corpus can be a monolingual corpus or a multilingual corpus.

[0069] PEGASUS Model: PEGASUS is a pre-trained model tailored for summarizing. It can be used as a general generative pre-training task. PEGASUS is a standard Transformer (preorder encoder-decoder predictor) that has both an encoder and a decoder. Pre-training objectives include GSG (Gap Sentences Generation) and MLM (Masked Language Model).

[0070] The Longest Common Subsequence (LCS) is obtained by taking as many characters as possible from two given sequences X and Y and arranging them in the order they appear in the original sequences. A subsequence is the result of removing zero or more elements from a given sequence (without changing the relative order of the elements). A common subsequence is a sequence Z that is a subsequence of both X and Y, where Z is a common subsequence of X and Y. For example, if X = [A,B,C,B,D,A,B] and Y = [B,D,C,A,B,A], then Z = [B,C,A] is a common subsequence of X and Y with a length of 3. However, Z is not the longest common subsequence of X and Y, while sequences [B,C,B,A] and [B,D,A,B] are also the longest common subsequences of X and Y, with a length of 4. X and Y do not have any common subsequences of length greater than or equal to 5. The only common subsequence for sequences [A,B,C] and [E,F,G] is the empty sequence []. Longest common subsequence: Given sequences X and Y, select the longest one or more of their common subsequences.

[0071] BERT (Bidirectional Encoder Representations from Transformers) is a deep learning model based on the Transformers architecture and encoder. After pre-training on unlabeled training data, BERT only needs to be trained a small amount of data on the specific downstream processing task before being applied to it. This characteristic of BERT makes it well-suited for applications such as Natural Language Processing (NLP).

[0072] Currently, major news media outlets, public accounts, and bloggers generate a large amount of raw text data daily, including but not limited to news reports, commentary, predictions, and analyses. This raw text data is often lengthy, complex, and contains differing viewpoints, making it difficult to extract useful information directly. Therefore, providing a text processing method that improves the efficiency of information retrieval has become a pressing technical problem.

[0073] Based on this, the main objective of this application is to propose a text processing method, apparatus, device, and medium. The aim is to extract a preliminary summary text that best expresses the meaning of the original corpus by performing text extraction processing on the original corpus. Then, semantic analysis is performed on the preliminary summary text. The importance of each word segment can be determined by its word type, and the segments are linked to a preset hyper- and hypo-level cognitive graph based on their word types to obtain a target hyper- and hypo-level cognitive graph. Since the hyper- and hypo-level information of each word segment and the relationships between them can be easily obtained through the target hyper- and hypo-level cognitive graph, the text processing method of this application can improve the efficiency of information acquisition.

[0074] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0075] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0076] The text processing method provided in this application relates to the field of artificial intelligence technology. The text processing method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the text processing method, but is not limited to the above forms.

[0077] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0078] This application provides text processing methods, apparatus, devices, and media, which are specifically described through the following embodiments. First, the text processing method in the embodiments of this application is described.

[0079] Figure 1 This is an optional flowchart of the text processing method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S107.

[0080] Step S101: Obtain the original corpus;

[0081] Step S102: Perform text extraction processing on the original corpus to obtain preliminary summary text;

[0082] Step S103: Divide the preliminary summary text into multiple preliminary segmented texts according to the preset segmented text length;

[0083] Step S104: Filter the multiple preliminary segmented texts to obtain multiple preliminary candidate word groups;

[0084] Step S105: Calculate the similarity value between each preliminary candidate word group and the original corpus, and take the preliminary candidate word groups that meet the preset similarity conditions as target candidate word groups;

[0085] Step S106: Perform semantic parsing on the target candidate word group to obtain multiple word segments and word types of the word segments;

[0086] Step S107: Link multiple word segments to a preset hypernym / hypernym cognitive graph according to the word type of the segmented words, so as to obtain the target hypernym / hypernym cognitive graph.

[0087] Steps S101 to S107 of this embodiment involve extracting text from the original corpus to obtain a preliminary summary text that best expresses the meaning of the original corpus. This preliminary summary text is then divided according to a preset segment length to obtain target segment texts. The target segment texts are filtered to determine target candidate word groups. Semantic parsing of the target candidate word groups is then performed. The importance of each word segment is determined by its word type, and the word segments are linked to a preset hyper- and hypo-level cognitive graph based on their word types to obtain the target hyper- and hypo-level cognitive graph. Since the hyper- and hypo-level information of each word segment and the relationships between them can be easily obtained through the target hyper- and hypo-level cognitive graph, the text processing method of this embodiment can improve the efficiency of information acquisition.

[0088] In step S101 of some embodiments, the raw corpus may refer to news data, including news titles and news content. The raw corpus may also refer to article data, including article titles and article content. Further, embodiments of this application can use a web crawler to extract the raw corpus from web pages; optionally, the web crawler is built based on Node.js technology. Specifically, obtaining the raw corpus using a web crawler includes: using Node.js to crawl the Uniform Resource Locator (URL) address of the raw corpus to be obtained, and characterizing the raw corpus to be obtained; loading the system interface corresponding to the raw corpus to be obtained based on the URL address; and obtaining the corresponding raw corpus from the system interface based on the character identifier.

[0089] It is understandable that the crawled raw corpus contains a large number of raw titles, and the length of the raw content corresponding to each raw title is not uniform, which brings difficulties to subsequent semantic understanding. Therefore, this application embodiment improves the efficiency of subsequent data processing by preprocessing the raw corpus.

[0090] In step S102 of some embodiments, text extraction processing is performed on the original corpus to obtain preliminary summary text. Specifically, a preset character length threshold is obtained, and text extraction processing is performed on the original corpus using the Pagasus model or LCS algorithm to obtain preliminary summary text. The character length of the preliminary summary text is less than the character length threshold. Preprocessing the original corpus aims to extract preliminary summary text that meets the character length threshold, which helps improve the efficiency of subsequent semantic understanding of the preliminary summary text. Specifically, the preset character length threshold can be 512, or other values; this application embodiment does not specifically limit this.

[0091] Specifically, refer to Figure 2 In some embodiments, step S102 includes, but is not limited to, steps S201 to S205:

[0092] Step S201: The original corpus is processed by a preset event classification prediction model to obtain at least two corpus event types;

[0093] Step S202: Calculate the degree of matching between each corpus event type and the preset key events to obtain the event matching value;

[0094] Step S203: Determine the target event type based on the event matching value and at least two corpus event types;

[0095] Step S204: Extract text from the original corpus according to the preset extraction text length to obtain at least two original summary texts;

[0096] Step S205: Determine the preliminary summary text based on the degree of matching between the original summary text and the target event type.

[0097] Steps S201 to S205, as illustrated in this embodiment, aim to determine the event types of the original corpus by predicting the event types. It should be noted that event types often reflect the type of information users most want to obtain. For example, in the event type of travel, insurance is highly likely to be the information users most want to obtain. Therefore, by determining the target event type from the event types of the original corpus and filtering the original summary text according to the target event type, a preliminary summary text that best matches the target event type is obtained. This allows for the subsequent acquisition of the information most desired by the user based on the preliminary summary text. In summary, this embodiment can help improve the efficiency of obtaining effective information.

[0098] In step S103 of some embodiments, the preliminary summary text is divided into multiple preliminary segmented texts according to a preset segmented text length. It should be noted that since the preliminary summary text contains multiple sentences, each with a potentially different meaning, processing the sentences better reflects the central meaning of the original corpus. Therefore, the preliminary summary text is divided into multiple preliminary segmented texts based on the segmented text length.

[0099] In step S104 of some embodiments, multiple preliminary segmented texts can be filtered based on the semantics of the preliminary segmented texts to select the best preliminary candidate word groups that conform to the semantics of the original corpus. The specific value of the segmented text length can be determined according to actual needs. For example, a segmented text length of ten characters can be used to segment the preliminary summary text, delete the preliminary segmented texts that are semantically incoherent, and obtain multiple preliminary segmented texts with complete semantics, thereby determining the preliminary candidate word groups. It should be noted that a complete preliminary segmented text can be directly used as a preliminary candidate word group, or keyword extraction can be performed on the preliminary segmented texts to obtain candidate word groups.

[0100] Specifically, refer to Figure 3 In some embodiments, step S104 includes, but is not limited to, steps S301 to S304:

[0101] Step S301: Obtain the reference scene classification of the original corpus;

[0102] Step S302: Predict multiple preliminary segmented texts using a preset classification prediction model to obtain the predicted scenario classification for each preliminary segmented text.

[0103] Step S303: Filter multiple preliminary segmented texts based on the matching relationship between the reference scene classification and the predicted scene classification;

[0104] Step S304: Determine multiple preliminary candidate word groups based on the multiple preliminary segmented texts after screening.

[0105] Steps S301 to S304, as illustrated in this embodiment, determine the matching relationship between the reference scene classification of the original corpus and the predicted scene classification of each preliminary segmented text. The aim is to determine the preliminary segmented text that matches the reference scene classification of the original corpus, thereby obtaining preliminary candidate word groups. This embodiment aims to maintain consistency between the scene classification of the candidate word groups and the scene classification of the original corpus, thus improving the accuracy of information retrieval by ensuring that the scene classification of the information is not altered when subsequently obtaining information through the candidate word groups.

[0106] Specifically, refer to Figure 4 In some other embodiments, step S104 includes, but is not limited to, steps S401 to S403:

[0107] Step S401: Obtain the reference scene classification of the original corpus;

[0108] Step S402: Determine the keyword set based on the reference scenario classification. The keyword set includes multiple keywords.

[0109] Step S403: Calculate the word matching value between each preliminary segmented text and multiple keywords, and obtain multiple preliminary candidate word groups based on multiple preliminary segmented texts that meet the preset matching conditions.

[0110] Steps S401 to S403, as illustrated in this embodiment, determine a keyword set by referring to scene classification. If the preliminary segmented text includes the keyword, or if the preliminary text includes words whose matching values ​​meet the matching conditions for the keyword, then preliminary candidate word groups are determined based on the preliminary segmented text. This embodiment aims to maintain consistency between the scene classification of the candidate word groups and the scene classification of the original corpus, thereby improving the accuracy of information retrieval without changing the scene classification of the information when subsequently obtaining information through candidate word groups.

[0111] Specifically, refer to Figure 5 In some other embodiments, step S104 includes, but is not limited to, steps S501 to S503:

[0112] Step S501: Filter multiple preliminary segmented texts using a preset BERT quality model to obtain multiple target segmented texts;

[0113] Step S502: Calculate the text similarity between each target segment text and the original corpus using the preset sifRank semantic model;

[0114] Step S503: Filter multiple target segmented texts based on text similarity to obtain multiple candidate word groups.

[0115] Steps S501 to S503 of this embodiment involve a first filtering using the BERT quality model, followed by a second filtering of the resulting target segmented texts using the SIFRank semantic model to obtain candidate word groups. This embodiment aims to ensure the semantic information and quality of the candidate word groups through the BERT quality model and the SIFRank semantic model, while maintaining consistency between the scene classification of the candidate word groups and the scene classification of the original corpus. This improves the accuracy of information retrieval by ensuring that the scene classification of the information is not altered when subsequently obtained through candidate word groups. It is understood that the feature information of each dimension of the words in each preliminary segmented text is obtained based on the selected phrase quality model (BERT quality model), and a quality score is determined based on this feature information. This quality score is then used to filter the preliminary segmented texts to obtain the target segmented texts.

[0116] In step S105 of some embodiments, the similarity value between each preliminary candidate word group and the original corpus is calculated, and the preliminary candidate word groups that meet the preset similarity conditions are taken as target candidate word groups. In order to maximize the similarity to the semantics of the original corpus, the similarity value between the preliminary candidate word groups and the original corpus can be calculated through existing semantic similarity models. The purpose is to screen multiple preliminary candidate word groups to obtain target candidate word groups that meet the preset similarity conditions.

[0117] In step S106 of some embodiments, semantic parsing is performed on the target candidate word group to obtain multiple word segments and word types. Semantic parsing refers to segmenting the text and tagging the word categories. For example, the target candidate word group could be "black rice with apples made into a porridge, more nutritious," which can be semantically parsed using an existing wordtag model. After semantic parsing of the target candidate word group, the results are shown in Table 1.

[0118] Table 1:

[0119]

[0120] Step S107: Link multiple word segments to a preset hyper- and hypo-level cognitive graph based on their word types to obtain a target hyper- and hypo-level cognitive graph. Specifically, the importance of word segments can be determined based on their word types. For example, adverbs and modifiers are considered non-key word types, while food and beverage, scene and event, etc., can be configured as key word types. Linking key word segments to the preset hyper- and hypo-level cognitive graph yields the target hyper- and hypo-level cognitive graph, allowing for clear and convenient information retrieval and improving information acquisition efficiency.

[0121] In a specific example, information recommendation capabilities can be improved through a target hyper- and hypo-cognitive graph. Specifically, the process involves acquiring pending business tasks; predicting the scenario classification of the pending business tasks using a pre-defined classification model; matching the corresponding scenario nodes from the target cognitive graph based on the predicted scenario classification; and determining the target recommendation result based on multiple parent and child scenario nodes corresponding to the scenario node. This example improves information recommendation capabilities through a target hyper- and hypo-cognitive graph. It should be noted that determining the recommendation result based on nodes includes the following steps: running graph embedding algorithms such as Graphsage on the nodes in the target hyper- and hypo-cognitive graph to obtain the eMB representation of each node; representing a word segment as a weighted sum of the eMBs of related nodes (related nodes include parent nodes, child nodes, and sibling nodes); and using the distance between eMBs to determine the target recommendation result.

[0122] Specifically, refer to Figure 6In some other embodiments, step S107 includes, but is not limited to, steps S601 to S604:

[0123] Step S601: Obtain the reference scene classification of the original corpus, and determine the scene node type based on the reference scene classification;

[0124] Step S602: If the nodes in the upper and lower cognitive graphs do not include scene node types, determine the nodes to be processed from the upper and lower cognitive graphs based on the correlation of scene node types.

[0125] Step S603: Create scene nodes corresponding to the scene node type under the node to be processed to obtain a preliminary hierarchical cognitive map;

[0126] Step S604: Link multiple word segments to the preliminary hypernym / hypernym cognitive graph according to the word type of the segmented words, so as to obtain the target hypernym / hypernym cognitive graph.

[0127] Steps S601 to S604, as illustrated in this embodiment, determine the node to be processed in the hyper-hyper-cognitive graph based on the relevance of scene nodes, and create the scene node under the node to be processed, i.e., treat the scene node as a child node of the node to be processed, thus obtaining a preliminary hyper-hyper-cognitive graph. Then, based on the word type of the segmented words, link multiple segmented words to the preliminary hyper-hyper-cognitive graph to obtain the target hyper-hyper-cognitive graph. In this way, while obtaining the corresponding node information through scene nodes, it is also possible to obtain the node information of the parent node and other sibling nodes, further improving the efficiency of information acquisition.

[0128] Specifically, refer to Figure 7 In some other embodiments, step S107 includes, but is not limited to, steps S701 to S703:

[0129] Step S701: Construct a preliminary word segmentation set based on multiple word segments;

[0130] Step S702: If it is determined that the word type of the segmented word does not belong to the preset word type set, the segmented word corresponding to the word type is deleted from the preliminary word segmentation set to obtain the target word segmentation set, which includes multiple target word segments; wherein, the word type set includes region category, action category, and quantity category;

[0131] Step S703: Link multiple target word segments to a preset hyper-hyper-level cognitive graph according to the word type of the segmented words, so as to obtain the target hyper-hyper-level cognitive graph.

[0132] Steps S701 to S703, as illustrated in this embodiment, determine the importance of word segmentation based on its word type. If the word type of a segmented word does not belong to a preset word type set, for example, if the word type is a modifier, the segmented word corresponding to the modifier is deleted from the initial word segmentation set to obtain the target word segmentation set. The word type set includes categories such as region, action, and quantity, and can be configured according to actual needs. Filtering word segments by their word type reduces information redundancy and further improves the efficiency of information retrieval.

[0133] Please see Figure 8 This application also provides a text processing apparatus that can implement the above-described text processing method. Figure 8 The block diagram of the module structure of the text processing device provided in the embodiments of this application is shown. The device includes: an acquisition module 801, a summary text extraction module 802, a text segmentation module 803, a preliminary candidate word group generation module 804, a target candidate word group generation module 805, a semantic parsing and processing module 806, and a word linking module 807. The system comprises the following modules: an acquisition module 801 for acquiring the original corpus; a summary text extraction module 802 for extracting text from the original corpus to obtain preliminary summary text; a text segmentation module 803 for dividing the preliminary summary text into multiple preliminary segment texts according to a preset segment text length; a preliminary candidate word group generation module 804 for filtering the multiple preliminary segment texts to obtain multiple preliminary candidate word groups; a target candidate word group generation module 805 for calculating the similarity value between each preliminary candidate word group and the original corpus, and using the preliminary candidate word groups that meet the preset similarity conditions as target candidate word groups; a semantic parsing processing module 806 for performing semantic parsing processing on the target candidate word groups to obtain multiple word segments and their word types; and a word linking module 807 for linking multiple word segments to a preset hyper-hyper-cognitive graph according to their word types to obtain the target hyper-hyper-cognitive graph.

[0134] The text processing apparatus of this application embodiment can extract preliminary summary text that best expresses the meaning of the original corpus by performing text extraction processing on the original corpus. Then, it divides the preliminary summary text according to a preset segment text length to obtain target segment text. The target segment text is filtered to determine target candidate word groups. Semantic analysis processing is then performed on the target candidate word groups. The importance of each word segment can be determined by its word type, and the word segments are linked to a preset hyper- and hypo-level cognitive graph based on their word types to obtain the target hyper- and hypo-level cognitive graph. Since the hyper- and hypo-level information of each word segment and the relationships between each word segment can be easily obtained through the target hyper- and hypo-level cognitive graph, the text processing method of this application embodiment can improve the efficiency of information acquisition.

[0135] It should be noted that the specific implementation of this text processing device is basically the same as the specific implementation of the text processing method described above, and will not be repeated here.

[0136] This application also provides a computer device, which includes: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for communication between the processor and the memory. When the program is executed by the processor, it implements the aforementioned text processing method. This computer device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0137] Please see Figure 9 , Figure 9 The hardware structure of a computer device according to another embodiment is illustrated. The computer device includes:

[0138] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0139] The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the text processing method of the embodiments of this application.

[0140] The input / output interface 903 is used to implement information input and output;

[0141] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0142] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);

[0143] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0144] This application also provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, which can be executed by one or more processors to implement the above-described text processing method.

[0145] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0146] The text processing method, text processing apparatus, computer equipment, and storage medium provided in this application embodiment extract text from the original corpus to obtain a preliminary summary text that best expresses the meaning of the original corpus. The preliminary summary text is then divided according to a preset segment length to obtain target segment text. The target segment text is filtered to determine target candidate word groups. Semantic analysis of the target candidate word groups is then performed, and the importance of each word segment is determined by its word type. Based on the word type, the word segments are linked to a preset hyper- and hypo-level cognitive graph to obtain the target hyper- and hypo-level cognitive graph. Since the hyper- and hypo-level information of each word segment and the relationships between each word segment can be easily obtained through the target hyper- and hypo-level cognitive graph, the text processing method of this application embodiment can improve the efficiency of information acquisition.

[0147] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0148] It will be understood by those skilled in the art that Figure 1-7 The technical solutions shown do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0149] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0150] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0151] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0152] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0153] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0154] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0155] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0156] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0157] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A text processing method, characterized in that, The method includes: Obtain the original corpus; The original corpus is subjected to text extraction processing to obtain preliminary summary text; wherein, the preliminary summary text is used to express the meaning of the original corpus; The preliminary summary text is divided into multiple preliminary segmented texts according to the preset segmented text length; The multiple preliminary segmented texts are filtered to obtain multiple preliminary candidate word groups; Calculate the similarity value between each preliminary candidate word group and the original corpus, and take the preliminary candidate word groups that meet the preset similarity conditions as target candidate word groups; The target candidate word groups are subjected to semantic parsing to obtain multiple word segments and word types of the word segments; Obtain a reference scene classification for the original corpus; The scene node type is determined based on the reference scene classification; If the nodes in the preset hierarchical cognitive graph do not include the scene node type, the node to be processed is determined from the hierarchical cognitive graph based on the correlation of the scene node type. Create scene nodes corresponding to the scene node type under the node to be processed to obtain a preliminary hierarchical cognitive map; Based on the word type of the segmented words, multiple segmented words are linked to the preliminary hypernym-hypernym cognitive graph to obtain the target hypernym-hypernym cognitive graph; Obtain the business to be processed and predict the scenario classification of the business to be processed using a preset classification model; The corresponding scene nodes are matched from the target hierarchical cognitive graph through the predicted scene classification. The target recommendation result is determined based on the multiple parent scene nodes and child scene nodes corresponding to the scene node.

2. The method according to claim 1, characterized in that, The filtering of multiple preliminary segmented texts to obtain multiple preliminary candidate word groups includes: Obtain a reference scene classification for the original corpus; The multiple preliminary segmented texts are predicted using a preset classification prediction model to obtain the predicted scenario classification for each preliminary segmented text. The multiple preliminary segmented texts are filtered based on the matching relationship between the reference scene classification and the predicted scene classification; Based on the filtered preliminary segmented texts, several preliminary candidate word groups are determined.

3. The method according to claim 1, characterized in that, The filtering of multiple preliminary segmented texts to obtain multiple preliminary candidate word groups includes: Obtain a reference scene classification for the original corpus; A keyword set is determined based on the reference scenario classification, and the keyword set includes multiple keywords; Calculate the word matching value between each of the preliminary segmented texts and the multiple keywords, and obtain multiple preliminary candidate word groups based on the multiple preliminary segmented texts that meet the preset matching conditions.

4. The method according to claim 1, characterized in that, The filtering of multiple preliminary segmented texts to obtain multiple preliminary candidate word groups includes: Multiple preliminary segmented texts are filtered using a preset BERT quality model to obtain multiple target segmented texts. The text similarity between each target segment and the original corpus is calculated using a pre-defined sifRank semantic model. Multiple target segmented texts are filtered based on the text similarity to obtain multiple candidate word groups.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Construct a preliminary word segmentation set based on the multiple word segments described above; If it is determined that the word type of the segmented word does not belong to the preset word type set, the word segment corresponding to the word type of the segmented word is deleted from the preliminary word segmentation set to obtain the target word segmentation set, which includes multiple target word segments; wherein, the word type set includes region category, action category, and quantity category; Based on the word type of the segmented words, multiple target segmented words are linked to a preset hypernym / hypernym cognitive graph to obtain a target hypernym / hypernym cognitive graph.

6. The method according to any one of claims 1 to 4, characterized in that, The text extraction process performed on the original corpus to obtain preliminary summary text includes: The original corpus is processed by a preset event classification prediction model to obtain at least two corpus event types. Calculate the degree of matching between each of the corpus event types and preset key events to obtain an event matching value; The target event type is determined based on the event matching value and at least two of the corpus event types; The original corpus is truncated according to a preset truncation length to obtain at least two original summary texts; The preliminary summary text is determined based on the degree of matching between the original summary text and the target event type.

7. A text processing device, characterized in that, The device includes: The acquisition module is used to acquire the raw corpus; The summary text extraction module is used to extract text from the original corpus to obtain preliminary summary text; wherein, the preliminary summary text is used to express the meaning of the original corpus; The text segmentation module is used to divide the preliminary summary text into multiple preliminary segmented texts according to a preset segmented text length; The preliminary candidate word group generation module is used to filter multiple preliminary segmented texts to obtain multiple preliminary candidate word groups; The target candidate word group generation module is used to calculate the similarity value between each of the preliminary candidate word groups and the original corpus, and to take the preliminary candidate word groups that meet the preset similarity conditions as target candidate word groups; The semantic parsing processing module is used to perform semantic parsing processing on the target candidate word groups to obtain multiple word segments and word types of the word segments; The word linking module is used to obtain the reference scene classification of the original corpus; determine the scene node type according to the reference scene classification; if the nodes of the preset hyper-hyper-cognitive graph do not include the scene node type, determine the node to be processed from the hyper-hyper-cognitive graph according to the relevance of the scene node type; create a scene node corresponding to the scene node type under the node to be processed to obtain a preliminary hyper-hyper-cognitive graph; link multiple word segments to the preliminary hyper-hyper-cognitive graph according to the word type of the word segmentation to obtain a target hyper-hyper-cognitive graph; The device is also used for: Obtain the business to be processed and predict the scenario classification of the business to be processed using a preset classification model; The corresponding scene nodes are matched from the target hierarchical cognitive graph through the predicted scene classification. The target recommendation result is determined based on the multiple parent scene nodes and child scene nodes corresponding to the scene node.

8. A computer device, characterized in that, The computer device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for enabling communication between the processor and the memory, wherein the program, when executed by the processor, implements the steps of the method as described in any one of claims 1 to 6.

9. A storage medium, said storage medium being a computer-readable storage medium for computer-readable storage, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Entity association method and device, electronic equipment and storage medium

    CN113032584A