Text processing method, electronic equipment, storage medium and program product
By performing chunked transformation and semantic retrieval of to be processed text and updating semantic data sets, the context incoherence caused by text segmentation in the prior art is solved, and the efficiency and accuracy of semantic retrieval is improved.
Patent Information
- Application Number
- CN202510660184.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When building a knowledge base, existing search enhancement technologies cause incoherence of context information when semantic segmentation of the to be processed text, resulting in low query efficiency.
By performing block conversion of the text to be processed, text content features and positioning information are obtained, and semantic search is performed in the initial semantic knowledge base, and matching semantic data sets are updated using text content features and positioning information to generate the target semantic knowledge base.
Improve the context coherence of the semantic dataset in the target semantic knowledge base, thereby improving the efficiency and accuracy of semantic retrieval.
Smart Images

Figure CN120179796A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text processing and information retrieval, and more particularly to a text processing method, a text processing device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Retrieval-augmented Generation (RAG) refers to a natural language processing technology that combines retrieval and generation techniques, aiming to solve natural language understanding and generation tasks by integrating information retrieval and text generation methods. With the continuous development of technology, especially in the era of the explosive development of large language models (LLMs), retrieval-augmented technology effectively combines the parametric knowledge of large language models with non-parametric external knowledge bases, and has been proven to significantly improve the accuracy of question answering in specific fields such as computer hardware, finance, and healthcare.
[0003] However, in the process of constructing a knowledge base in related technologies, the context information is incoherent when performing semantic segmentation on the text to be processed, resulting in low subsequent query efficiency. Summary of the Invention
[0004] In view of the above problems, the present invention provides a text processing method, a text processing device, an electronic device, a computer-readable storage medium, and a computer program product.
[0005] According to a first aspect of the present invention, there is provided a text processing method, including: performing chunk conversion processing on the text to be processed to obtain at least one text content feature and the positioning information of the text content feature, where the positioning information represents the position of the text content feature in the text to be processed; for any text content feature, performing semantic retrieval in an initial semantic knowledge base to obtain a semantic retrieval result, where the initial semantic knowledge base includes at least one semantic data set; in the case where the semantic retrieval result indicates that at least one semantic data set includes a matching semantic data set that matches the text content feature, updating the matching semantic data set by using the text content feature and the positioning information to obtain a target semantic knowledge base.
[0006] The second aspect of the present invention provides a text processing device, including: a chunk conversion module for performing chunk conversion processing on the text to be processed to obtain at least one text content feature and the positioning information of the text content feature, where the positioning information represents the position of the text content feature in the text to be processed; a retrieval matching module for performing semantic retrieval in the initial semantic knowledge base for any text content feature to obtain a semantic retrieval result, where the initial semantic knowledge base includes at least one semantic data set; a semantic update module for, in the case that the semantic retrieval result indicates that at least one semantic data set includes a matching semantic data set that matches the text content feature, updating the matching semantic data set by using the text content feature and the positioning information to obtain a target semantic knowledge base.
[0007] The third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, where the above one or more processors execute the above one or more computer programs to implement the steps of the above method.
[0008] The fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the above computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0009] The fifth aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the above computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0010] According to an embodiment of the present invention, by chunk-converting the text to be processed into a text content feature and the positioning information of the text content feature, performing semantic retrieval on the text content feature from the initial semantic knowledge base to obtain a semantic retrieval result, and in the case that the initial semantic knowledge base includes a matching semantic data set that matches the text content feature, updating the matching semantic data set by using the text content feature and the positioning information to obtain a target semantic knowledge base. Since the target semantic knowledge base stores different text content features with similar semantics in the text to be processed, the semantic data sets in the target semantic knowledge base have strong context coherence, so that the context coherence and retrieval efficiency of semantic retrieval can be improved when using the target semantic knowledge base. Description of the Drawings
[0011] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0012] Figure 1 A diagram showing an application scenario of the text processing method according to an embodiment of the present invention;
[0013] Figure 2Shows a flowchart of a text processing method according to an embodiment of the present invention;
[0014] Figure 3 Shows a flowchart of a chunk conversion method for a text to be processed according to an embodiment of the present invention;
[0015] Figure 4 Shows a flowchart of generating a target semantic knowledge base according to an embodiment of the present invention;
[0016] Figure 5 Shows a flowchart of a text processing method according to another embodiment of the present invention;
[0017] Figure 6 Shows a structural block diagram of a text processing apparatus according to an embodiment of the present invention;
[0018] Figure 7 Shows a block diagram of an electronic device suitable for implementing the text processing method according to an embodiment of the present invention. Detailed implementation manners
[0019] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.
[0020] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0022] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0023] With the popularity of large language models, it has become a consensus to use retrieval-enhanced generation technology to solve the "hallucination" problem of large language models. As the use of retrieval-enhanced generation technology deepens, users will find that the chunking effect of the text will seriously affect the final response of the large language model. The traditional chunking method relies on fixed feature length and window length for segmentation, which will destroy the contextual coherence of the text to be processed; semantic segmentation can segment by detecting semantic changes between sentences, but when the topic in the text to be processed switches back and forth, semantic segmentation may divide related content into different chunks, resulting in information incoherence.
[0024] In view of this, an embodiment of the present invention provides a text processing method, an electronic device, a storage medium and a program product, the method comprising performing block conversion processing on a text to be processed to obtain at least one text content feature and positioning information of the text content feature, wherein the positioning information represents the position of the text content feature in the text to be processed; for any text content feature, performing a semantic search in an initial semantic knowledge base to obtain a semantic retrieval result, wherein the initial semantic knowledge base includes at least one semantic data set; when the semantic retrieval result indicates that at least one semantic data set includes a matching semantic data set that matches the text content feature, the matching semantic data set is updated using the text content feature and the positioning information to obtain a target semantic knowledge base.
[0025] In the technical solution of the present invention, the data involved (including but not limited to data used for analysis, data stored, data displayed, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0026] Figure 1 An application scenario diagram of a text processing method according to an embodiment of the present invention is shown.
[0027] like Figure 1 As shown, the application scenario 100 according to this embodiment may include processing the to-be-processed text of the computer hardware. The network 104 is used to provide a medium of a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0028] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0029] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, desktop computers, and so on.
[0030] The server 105 can be a server that provides various services, such as a background management server that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (for example only). The background management server can analyze and process data such as user requests received (such as text to be processed and semantic query information), and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0031] It should be noted that the text processing method provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the text processing device provided by the embodiments of the present invention can generally be set in the server 105. The text processing method provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the text processing device provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0032] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0033] are merely illustrative. According to actual needs, there can be any number of terminal devices, networks, and servers. Figure 1 The following will be based on Figures 2 to 5 the described scenario, and will describe in detail the text processing method of the disclosed embodiments through
[0034] Figure 2 shows a flowchart of the text processing method according to an embodiment of the present invention.
[0035] As shown Figure 2 in the figure, the text processing method of this embodiment includes operations S210 to S230, and this transaction processing method can be executed by an electronic device.
[0036] In operation S210, the text to be processed is subjected to chunk conversion processing to obtain at least one text content feature and the positioning information of the text content feature, where the positioning information represents the position of the text content feature in the text to be processed;
[0037] In operation S220, for any text content feature, semantic retrieval is performed in the initial semantic knowledge base to obtain a semantic retrieval result, where the initial semantic knowledge base includes at least one semantic data set;
[0038] In operation S230, when the semantic retrieval result indicates that at least one semantic data set in the initial semantic knowledge base includes a matching semantic data set that matches the text content feature, the matching semantic data set is updated using the text content feature and the positioning information to obtain a target semantic knowledge base.
[0039] The text to be processed can be text in any field, for example, it can be a product manual or a customer consultation record text, etc.
[0040] For the text to be processed, it is first necessary to perform chunk conversion on the text to be processed to obtain at least one text content feature and the positioning information of the text content feature. For example, the positioning information shows the paragraph identifier of the text content feature in the text to be processed.
[0041] The initial semantic knowledge base includes semantic databases corresponding to different semantics. For example, it can be a semantic database corresponding to a data processor, a semantic database corresponding to a controller, and so on. For each text content feature, it is retrieved from the initial semantic knowledge base whether there is a matching semantic data set in the initial semantic knowledge base that matches the text content feature, so as to generate a semantic retrieval result for each text content feature.
[0042] For each semantic retrieval result, if the semantic retrieval result indicates that the initial semantic knowledge base includes a matching semantic data set that matches the text content feature, then the text content feature and the positioning information can be used to update the matching semantic data set, so as to obtain a target semantic knowledge base, and the target semantic knowledge base includes the updated matching semantic data set.
[0043] According to an embodiment of the present invention, by chunking the text to be processed into text content features and the positioning information of the text content features, performing semantic retrieval on the text content features from the initial semantic knowledge base to obtain a semantic retrieval result, when the initial semantic knowledge base includes a matching semantic data set that matches the text content features, updating the matching semantic data set using the text content features and the positioning information to obtain a target semantic knowledge base. Since the target semantic knowledge base stores different text content features with similar semantics in the text to be processed, the semantic data set in the target semantic knowledge base has strong context coherence, thereby improving the context coherence and retrieval efficiency of semantic retrieval when using the target semantic knowledge base.
[0044] Figure 3 The flowchart of the chunking conversion method for the text to be processed according to an embodiment of the present invention is shown.
[0045] According to an embodiment of the present invention, as Figure 3 shown, performing chunking conversion processing on the text to be processed to obtain at least one text content feature and the positioning information of the text content feature, including: performing chunking processing on the text to be processed to obtain at least one text content information (such as Figure 3 text content information 1, text content information 2, text content information 3... text content information n in) and the positioning information; performing vector conversion processing on at least one text content information to obtain at least one text content feature (such as Figure 3 text content feature 1, text content feature 2, text content feature 3... text content feature n in).
[0046] The specific chunking method can be specifically set according to actual needs. For example, chunking can be performed according to paragraphs, fixed sizes, according to full stops, and recursive segmentation.
[0047] For the text to be processed, first chunk the text to be processed. For example, chunking according to paragraphs can obtain n text content information corresponding to different paragraphs and the paragraph identifiers of the text content information in the text to be processed, that is, the positioning information.
[0048] Since it is impossible to directly perform retrieval using the text content when retrieving in the initial semantic knowledge base, it is necessary to convert the text content information of the text type into a vector to obtain the text content feature corresponding to the text content information.
[0049] Vector conversion can use the bag-of-words model, the term frequency-inverse document frequency (TF-IDF) model, the word embedding model, etc. Among them, the word embedding model includes models such as Word2Vec, GloVe, and FastText.
[0050] According to an embodiment of the present invention, the types of texts to be processed include document types or Q&A types, and the text content information includes at least one of text content, table content, and picture content.
[0051] According to an embodiment of the present invention, the text to be processed is chunked to obtain at least one text content information and positioning information, including: in the case where the text to be processed is of the document type, the text to be processed is chunked based on the paragraph format to obtain at least one text content information and positioning information; in the case where the text to be processed is of the Q&A type, the text to be processed is chunked based on the Q&A format to obtain at least one text content information and positioning information.
[0052] The text to be processed of the document type can be a product analysis report, a product manual, a thesis, etc. The text to be processed of the Q&A type can be a customer consultation log, for example, a log formed by user consultation on the functions and structures of a controller and the content of the customer service reply.
[0053] For the text to be processed of the document type, the text to be processed can be divided into text content information of different paragraphs based on the paragraph format, and the positioning information is marked with paragraph tags at the same time.
[0054] For the text to be processed of the document type, the text to be processed not only includes text content, but some text to be processed also includes table content and / or picture content. For example, a product manual may use pictures to supplement the description of the product and tables to describe the parameters of different components of the product. For such text to be processed, during the chunking process, a paragraph of text content before or after the table content and / or picture content can be divided into a text content information.
[0055] According to an embodiment of the present invention, by dividing the table content and / or picture content into the text content during the chunking of the text to be processed to form text content information, the coherence of the context can be further ensured, thereby improving the accuracy during semantic retrieval.
[0056] Figure 4 Shows a flowchart of generating a target semantic knowledge base according to an embodiment of the present invention.
[0057] According to an embodiment of the present invention, the semantic data set includes at least one semantic summary information.
[0058] According to an embodiment of the present invention, for any text content feature, semantic retrieval is performed in the initial semantic knowledge base to obtain a semantic retrieval result, including: generating a semantic retrieval result when the semantic similarity between the text content feature and the semantic summary information satisfies a preset similarity threshold, where the semantic retrieval result represents that the semantic data set corresponding to the semantic similarity is a matching semantic data set.
[0059] The semantic summary information may refer to the vectorized representation of the semantic summary text. The preset similarity threshold can be specifically set according to actual needs, for example, it can be set to 0.8.
[0060] During the semantic retrieval process, first, the semantic similarity between the text content feature and each semantic summary information needs to be calculated. Among them, the semantic similarity can be calculated using methods such as cosine similarity, Euclidean distance, Manhattan distance, Chebyshev distance, Pearson correlation coefficient, etc.
[0061] See Figure 4 , for each text to be processed, first in operation S401, the text to be processed is subjected to block conversion processing, thereby obtaining at least one text content feature and location information, such as Figure 4 the text content feature 1, text content feature 2, text content feature 3... in , in operation S402, for each text content feature, a matching semantic data set that matches the text content feature is retrieved in the initial semantic knowledge base. Specifically, for the semantic similarity corresponding to each semantic summary information, if the semantic similarity is greater than the preset similarity threshold, a semantic retrieval result can be generated, and the semantic retrieval result can represent that the semantic data set corresponding to the semantic similarity is a matching semantic data set. Thus, the matching data set can be updated subsequently.
[0062] According to an embodiment of the present invention, by calculating the semantic similarity between the text content feature and each semantic summary information to determine a matching semantic data set that matches the text content feature from the initial semantic knowledge base based on the semantic similarity, text content features with similar semantics can be updated to the matching semantic data set, which can ensure the coherence of the context and thus improve the accuracy during semantic retrieval.
[0063] According to an embodiment of the present invention, when the number of text content features is multiple, semantic retrieval of multiple text content features in the initial semantic knowledge base and updating of the initial semantic knowledge base are performed in parallel.
[0064] If the text to be processed obtains multiple text content features after block processing, at this time, semantic retrieval of multiple text content features can be performed in the initial semantic knowledge base in parallel, and thus update operations of different matching semantic data sets can be performed in parallel.
[0065] According to an embodiment of the present invention, by performing semantic retrieval on multiple text content features in parallel and updating the initial semantic knowledge base, the efficiency of text processing can be improved, effectively saving the time of text processing.
[0066] According to an embodiment of the present invention, in the case where the matching semantic data set matches multiple text content features, the initial semantic knowledge base is updated one by one using the multiple text content features.
[0067] If the text to be processed obtains multiple text content features after chunking, and at least two text content features simultaneously match a matching semantic data set in the initial semantic knowledge base, at this time, the initial semantic knowledge base can be updated using the first text content feature, and the second text content feature is paused for processing. After the update using the first text content feature is completed, the second text content feature is used for update.
[0068] According to an embodiment of the present invention, in the process of updating the same matching semantic data set by multiple text content features, the sequential update method can avoid the update task conflict caused by simultaneously using multiple text content features to update the matching speech data set, thereby avoiding the problem that the updated matching semantic data set is inaccurate due to the update task conflict.
[0069] According to an embodiment of the present invention, the matching semantic data set includes at least one semantic content information and the positioning information of the semantic content information.
[0070] According to an embodiment of the present invention, updating the matching semantic data set using the text content feature and the positioning information to obtain a target semantic knowledge base includes: processing the text content information and at least one semantic content information using a preset language model to obtain text summary information; for any content information in the text content information and at least one semantic content information, updating the matching semantic data set based on the summary similarity between the content information and the text summary information to obtain a target semantic knowledge base.
[0071] The preset language model can be a general large language model, a multi-modal large language model, a vertical domain optimization model, such as the BERT (Bidirectional Encoder Representations from Transformers) model, the Transformer network, the GPT-4V model, and models in the LLaMA (Large Language Model Architecture) series, etc.
[0072] See Figure 4, for each successfully matched text content feature, in operation S411, the text content information corresponding to the text content feature and the semantic content information in the matching semantic dataset are simultaneously input into a preset language model so that the preset language model outputs a text summary information about the text content information and the semantic content information.
[0073] Based on the same calculation method as the semantic similarity, the summary similarity between the text content information and the text summary information and the summary similarity between each semantic content information and the text summary information can be calculated.
[0074] Update the matching semantic dataset based on multiple summary similarities. For example, sort the text content information and the semantic content information in the matching semantic dataset according to the size of the summary similarity.
[0075] According to the embodiments of the present invention, by calculating the summary similarity between each content information and the text summary information generated by the preset language model based on the text content information and the semantic content information, and updating the matching semantic dataset based on the summary similarity, the content information in the matching semantic dataset that is more similar in semantics to the text content information is obtained, improving the coherence of the context. Thus, more accurate semantic retrieval results can be provided in subsequent semantic retrievals, thereby improving the retrieval efficiency.
[0076] According to the embodiments of the present invention, update the matching semantic dataset based on the summary similarity between the content information and the text summary information to obtain a target semantic knowledge base, including: sorting multiple content information and multiple positioning information according to multiple summary similarities to obtain sorted multiple content information and sorted multiple positioning information; generating semantic metadata according to the file name of the text to be processed, the sorted multiple positioning information, multiple summary similarities, and feature length information, where the feature length information includes the feature length of the content information arranged according to the summary similarity and the total feature length of multiple content information; generating an updated matching semantic dataset according to the text summary information, the semantic metadata, and the sorted multiple content information, where the target semantic knowledge base includes the updated semantic dataset.
[0077] Refer to Figure 4 , when updating the matching semantic dataset, first sort multiple summary similarities according to the size of the summary similarity to obtain the sorting order of multiple summary similarities, sort multiple content information according to this sorting order, and simultaneously sort the positioning information of multiple content information to obtain sorted multiple content information and sorted multiple positioning information. Among them, the sorting order is preferably sorted in descending order.
[0078] During the update process, the file name of the text to be processed is also required, so that when retrieving later, it can be known which text to be processed the content information in the matching semantic dataset belongs to. Combining the positioning information, the paragraph where the content information belongs to which text to be processed can be quickly located. At the same time, calculate the characteristic length information corresponding to multiple content information, and thus associate and combine the file name, the sorted multiple positioning information, multiple abstract similarities, and characteristic length information into semantic metadata, such as Figure 4 the operation S412 in. Among them, the characteristic length information may include the characteristic length of each content information arranged based on the sorting order and the total characteristic length of multiple content information.
[0079] In operation S413, use the text content information of the text to be processed to update the content information in the matching semantic dataset. The specific update operation includes sorting according to the abstract similarity. At the same time, in operation S414, combine the semantic metadata and the text abstract information output by the preset language model to complete the update of the matching semantic dataset. In operation S415, write the updated matching semantic dataset into the initial semantic knowledge base to generate the target semantic knowledge base.
[0080] In a specific embodiment, when the preset language model processes the text content information and at least one semantic content information, it can not only output the text abstract information, but also output the text title information, and the text title information can be written into the matching semantic dataset together with the text abstract information. The writing of the text title information is convenient for the user to quickly understand the semantic summary content of the matching semantic dataset during subsequent information retrieval.
[0081] According to the embodiment of the present invention, the content information and the positioning information are sorted according to the abstract similarity related to each content information, and thus the file name of the text to be processed, the sorted multiple positioning information, multiple abstract similarities, and characteristic length information are used to generate semantic metadata. Finally, the text abstract information, semantic metadata, and the sorted multiple content information are used to update the matching semantic dataset. When retrieving semantically, the similarity degree between each content information and the retrieved content can be quickly known from the updated matching semantic dataset, thereby improving the retrieval efficiency.
[0082] According to the embodiment of the present invention, the text processing method further includes: when the semantic retrieval result indicates that the matching semantic dataset is not included, generating a semantic dataset according to the text content characteristics and the positioning information, and generating a target semantic knowledge base according to the semantic dataset.
[0083] Refer to Figure 4, in operation S402, if no matching semantic data set that matches the text content feature is found during semantic retrieval of the text content feature in the initial semantic knowledge base, a new semantic data set can be generated based on the current text content feature and location information, so as to facilitate subsequent updates of the initial semantic knowledge base and semantic queries.
[0084] According to an embodiment of the present invention, when no matching semantic data set that matches the text content feature is retrieved in the initial semantic knowledge base, a new semantic data set is generated by using the text content feature and location information to obtain a target semantic knowledge base, realizing real-time update of the semantic knowledge base, thereby providing a more accurate matching semantic data set for subsequent semantic queries, and further improving the accuracy of semantic queries.
[0085] According to an embodiment of the present invention, generating a semantic data set according to the text content feature and location information includes: processing the text content information corresponding to the text content feature by using a preset language model to obtain content summary information; performing a correlation merging operation on the text to be processed, the text content feature, the content summary information, and the location information to obtain a semantic data set.
[0086] Refer to Figure 4 , in the case where no matching semantic data set corresponding to the text content feature is found, in operation S421, the text content information corresponding to the text content feature is input into the preset language model to use the preset language model to extract the content summary information of the text content information, and thus a correlation merging operation is performed on the text to be processed, the text content feature, the content summary information, and the location information to obtain a semantic data set.
[0087] When performing semantic retrieval on the next text content feature, if the next text content feature is relatively similar to the semantic data set obtained in this embodiment, at this time, this semantic data set can be used as a matching semantic data set that matches the next text content feature, and thus an update operation is performed on this semantic data set based on the above update operation.
[0088] According to an embodiment of the present invention, performing a correlation merging operation on the text to be processed, the text content feature, the content summary information, and the location information to obtain a semantic data set includes: performing a correlation merging operation on the text content information, the file name of the text to be processed, the location information, and the text content feature to obtain semantic metadata; combining the content summary information, the semantic metadata, and the text content information into a semantic data set.
[0089] Refer to Figure 4, in operation S422, during the process of generating a semantic data set corresponding to the text content features that do not match successfully, it is first necessary to generate semantic metadata, which can be generated based on the text content information, the file name of the text to be processed corresponding to the text content information, the positioning information of the text content information, and the text content features.
[0090] In operation S423, the generated semantic metadata, the content summary information generated by the preset language model, and the text content information are formed into a semantic data set corresponding to the text content features. In operation S424, the generated semantic data set is written into the initial semantic knowledge base to generate a target semantic knowledge base.
[0091] In a specific embodiment, when the preset language model processes the text content information, it not only outputs content summary information but also can output content title information, and the content title information can be written into the semantic data set together with the content summary information. Among them, the writing of the content title information facilitates users to quickly understand the semantic summary content of the semantic data set through the text title information during subsequent information retrieval.
[0092] According to the embodiments of the present invention, a correlation and merging operation is performed on the text content information, the file name of the text to be processed, the positioning information, and the text content features to obtain semantic metadata, including: performing a correlation and merging on the content similarity between the file name, the positioning information, the text content information, and the content summary information and the feature length of the text content features to obtain semantic metadata.
[0093] Refer to Figure 4 , in operation S422, in the construction of semantic metadata, the file content information is not directly used, but the content similarity between the file content information and the content summary information is used, combined with the file name of the corresponding text to be processed, the positioning information, and the feature length of the text content features, so as to obtain a semantic data set corresponding to the text content features that do not match successfully. Among them, the calculation method of the content similarity is the same as that of the summary similarity, which will not be elaborated here.
[0094] According to the embodiments of the present invention, by generating semantic metadata based on the file name, the positioning information, the content similarity, and the feature length of the text content features, and then based on the semantic data set constructed by the semantic metadata, a relatively accurate semantic data set is provided for the matching or semantic query of the next text content feature.
[0095] Figure 5 Shows a flowchart of a text processing method according to another embodiment of the present invention.
[0096] According to an embodiment of the present invention, the semantic dataset includes semantic metadata, and the semantic metadata includes a plurality of content information sorted by abstract similarity and the positioning information of the content information.
[0097] As Figure 5 shown, the text processing method further includes operations S501 to S503:
[0098] In operation S501, obtain the input semantic query information.
[0099] In operation S502, use a preset language model to process the semantic query information and the target semantic knowledge base, so as to determine a target dataset matching the semantic query information from multiple semantic datasets in the target semantic knowledge base.
[0100] In operation S503, determine the multiple content information and positioning information in the target dataset as the query text content and text positioning information matching the semantic query information respectively.
[0101] The specific type of the semantic query information can be specifically set according to actual needs. For example, it can be "the working parameters of the data processor of model A".
[0102] After the user inputs the semantic query information, input the semantic query information and the target semantic knowledge base into the preset language model, so as to extract a target dataset matching the semantic query information from the target semantic knowledge base. Through this target dataset, the query text content related to the semantic query information and the file name and paragraph identifier to which each query text content belongs, that is, the text positioning information, can be known.
[0103] In a specific embodiment, if the semantic query information is "the working parameters of the data processor of model A", input the semantic query information and the target semantic knowledge base into the preset language model. At this time, the preset language model can output the target dataset corresponding to the working parameters of the data processor. This target dataset includes the descriptions of data processors of multiple models. At this time, the multiple content information and positioning information in this target dataset can be respectively determined as the query text content and text positioning information matching the semantic query information, for the user to consult the content information and positioning information related to model A that the user wants to query.
[0104] According to an embodiment of the present invention, by using a preset language model to query a target dataset related to the semantic query information from the target semantic knowledge base to determine the query text content and text positioning information, since the target semantic knowledge base is updated based on multiple text content features in at least one text to be processed, the query efficiency and query accuracy for the semantic query information are improved.
[0105] According to an embodiment of the present invention, the semantic metadata further includes the total feature length of multiple content information.
[0106] According to an embodiment of the present invention, for any semantic data set, when the total feature length is greater than the processing length threshold of the preset language model, based on the summary similarity ranking, the content information with a summary similarity less than the similarity threshold is excluded from the multiple content information, and the filtered semantic data set is obtained, so as to input the filtered semantic data set and the semantic query information into the preset language model.
[0107] The input of the preset language model has certain limitations. For example, the data capacity of the input of some preset language models cannot be greater than 128KB, that is, the processing length threshold. At this time, the target semantic knowledge base can be split, and the split semantic data sets are sequentially input into the preset language model to determine whether the semantic data set is the target data set.
[0108] It should be noted that taking 128KB as the processing length threshold is only an example for illustration, and it is not limited that the processing length threshold can only be 128KB. The specific value can be determined according to the type of the preset language model or specifically set by the staff based on actual needs.
[0109] If the content information inside the split semantic data set is so large that the total feature length of the semantic data set is greater than the processing length threshold, at this time, the content information in the semantic data set can be deleted so that the total feature length of the filtered semantic data set is not greater than the processing length threshold, thereby completing the query operation on the semantic data set.
[0110] In a specific embodiment, when performing the deletion operation on the semantic data set, based on the summary similarity of different content information, the content information with a summary similarity less than the similarity threshold can be deleted, so as to obtain the filtered semantic data set. Among them, the similarity threshold can be specifically set according to actual needs. For example, it can be 0.8.
[0111] In another specific embodiment, when performing the deletion operation on the semantic data set, based on the summary similarity of different content information, some content information with a higher summary similarity ranking can be used as the content information in the filtered semantic data set. For example, the first 5 content information in the summary similarity ranking can be used as the content information in the filtered semantic data set.
[0112] According to an embodiment of the present invention, by screening and processing the semantic database, it can be ensured that the data length input into the preset language model meets the data processing window of the preset language model, thereby avoiding the problem that the preset language model cannot perform semantic query operations due to too much input data, and indirectly improving the efficiency of semantic query.
[0113] According to an embodiment of the present invention, the text processing method further includes splitting the text to be processed to obtain a plurality of new texts to be processed; and performing chunk conversion processing on the plurality of new texts to be processed in parallel.
[0114] In the case where the content of the text to be processed is large, for example, the size of the text to be processed exceeds a preset size threshold (such as 10 MB), at this time, the text to be processed can be split according to a preset splitting rule to obtain a plurality of new texts to be processed, so as to perform chunk conversion and retrieval operations in the initial semantic knowledge base.
[0115] In a specific embodiment, splitting can be performed according to chapters or page numbers, whereby a plurality of new texts to be processed with smaller capacities can be obtained. At the same time, the plurality of new texts to be processed can perform chunk conversion and retrieval operations in the initial semantic knowledge base in parallel to improve the construction efficiency of the target semantic knowledge base.
[0116] Based on the above text processing method, the present invention also provides a text processing device. The following will be combined with Figure 6 to describe the device in detail.
[0117] Figure 6 The structural block diagram of the text processing device according to an embodiment of the present invention is shown.
[0118] As Figure 6 shown, the text processing device 600 of this embodiment includes a chunk conversion module 610, a retrieval matching module 620, and a semantic update module 630.
[0119] The chunk conversion module 610 is configured to perform chunk conversion processing on the text to be processed to obtain at least one text content feature and the positioning information of the text content feature, where the positioning information represents the position of the text content feature in the text to be processed.
[0120] The retrieval matching module 620 is configured to perform semantic retrieval on the initial semantic knowledge base for any text content feature to obtain a semantic retrieval result, where the initial semantic knowledge base includes at least one semantic data set.
[0121] The semantic update module 630 is configured to, in the case where the semantic retrieval result indicates that at least one semantic data set includes a matching semantic data set that matches the text content feature, update the matching semantic data set by using the text content feature and the positioning information to obtain a target semantic knowledge base.
[0122] According to an embodiment of the present invention, by chunking the text to be processed into text content features and location information of the text content features, performing semantic retrieval on the text content features from the initial semantic knowledge base to obtain a semantic retrieval result, and when the initial semantic knowledge base includes a matching semantic data set that matches the text content features, updating the matching semantic data set using the text content features and location information to obtain a target semantic knowledge base. Since the target semantic knowledge base stores different text content features with similar semantics in the text to be processed, the semantic data set in the target semantic knowledge base has strong context coherence, thereby improving the context coherence and retrieval efficiency of semantic retrieval when using the target semantic knowledge base.
[0123] According to an embodiment of the present invention, the chunking and conversion module 610 includes a chunking unit and a conversion unit.
[0124] The chunking unit is used to perform chunking processing on the text to be processed to obtain at least one text content information and location information.
[0125] The conversion unit is used to perform vector conversion processing on at least one text content information to obtain at least one text content feature.
[0126] According to an embodiment of the present invention, the type of the text to be processed includes document type or question-and-answer type, and the text content information includes at least one of text content, table content, and picture content.
[0127] According to an embodiment of the present invention, the chunking unit includes a first chunking subunit and a second chunking subunit.
[0128] The first chunking subunit is used to perform chunking processing on the text to be processed based on the paragraph format when the text to be processed is of the document type to obtain at least one text content information and location information.
[0129] The second chunking subunit is used to perform chunking processing on the text to be processed based on the question-and-answer format when the text to be processed is of the question-and-answer type to obtain at least one text content information and location information.
[0130] According to an embodiment of the present invention, the semantic data set includes at least one semantic summary information.
[0131] According to an embodiment of the present invention, the retrieval and matching module 620 includes a generation unit.
[0132] The generation unit is used to generate a semantic retrieval result when the semantic similarity between the text content feature and the semantic summary information meets a preset similarity threshold, where the semantic retrieval result represents that the semantic data set corresponding to the semantic similarity is a matching semantic data set.
[0133] According to an embodiment of the present invention, when the number of text content features is multiple, semantic retrieval of the multiple text content features and update of the initial semantic knowledge base are performed in parallel in the initial semantic knowledge base.
[0134] According to an embodiment of the present invention, when the matching semantic data set matches multiple text content features, the initial semantic knowledge base is updated one by one using the multiple text content features.
[0135] According to an embodiment of the present invention, the matching semantic data set includes at least one semantic content information and the positioning information of the semantic content information.
[0136] According to an embodiment of the present invention, the semantic update module 630 includes an abstract generation unit and an update unit.
[0137] The abstract generation unit is used to process the text content information and at least one semantic content information using a preset language model to obtain text abstract information.
[0138] The update unit is used to update the matching semantic data set based on the abstract similarity between the content information and the text abstract information for any one of the text content information and at least one semantic content information to obtain a target semantic knowledge base.
[0139] According to an embodiment of the present invention, the update unit includes a sorting subunit, a first generation subunit, and a second generation subunit.
[0140] The sorting subunit is used to sort multiple content information and multiple positioning information according to multiple abstract similarities to obtain the sorted multiple content information and the sorted multiple positioning information.
[0141] The first generation subunit is used to generate semantic metadata according to the file name of the text to be processed, the sorted multiple positioning information, multiple abstract similarities, and feature length information, where the feature length information includes the feature lengths of the content information arranged according to the abstract similarity and the total feature length of the multiple content information.
[0142] The second generation subunit is used to generate an updated matching semantic data set according to the text abstract information, the semantic metadata, and the sorted multiple content information, where the target semantic knowledge base includes the updated semantic data set.
[0143] According to an embodiment of the present invention, the text processing device 600 further includes a generation module.
[0144] The generation module is used to generate a semantic data set according to the text content features and the positioning information when the semantic retrieval result indicates that the matching semantic data set is not included, and generate a target semantic knowledge base according to the semantic data set.
[0145] According to an embodiment of the present invention, the generation module includes an obtaining sub-module and a merging sub-module.
[0146] The obtaining sub-module is configured to process text content information corresponding to text content features by using a preset language model to obtain content summary information.
[0147] The merging sub-module is configured to perform an associated merging operation on the text to be processed, text content features, content summary information, and positioning information to obtain a semantic data set.
[0148] According to an embodiment of the present invention, the merging sub-module includes an associated merging unit and a combining unit.
[0149] The associated merging unit is configured to perform an associated merging operation on text content information, the file name of the text to be processed, positioning information, and text content features to obtain semantic metadata.
[0150] The combining unit is configured to combine content summary information, semantic metadata, and text content information into a semantic data set.
[0151] According to an embodiment of the present invention, the associated merging unit includes a merging sub-unit.
[0152] The merging sub-unit is configured to perform an associated merging on the content similarity between the file name, positioning information, text content information, and content summary information and the feature length of the text content features to obtain semantic metadata.
[0153] According to an embodiment of the present invention, the semantic data set includes semantic metadata, and the semantic metadata includes multiple content information sorted by summary similarity and the positioning information of the content information.
[0154] According to an embodiment of the present invention, the text processing device further includes an acquisition module, a semantic search module, and a determination module.
[0155] The acquisition module is configured to acquire input semantic query information.
[0156] The semantic search module is configured to process the semantic query information and the target semantic knowledge base by using a preset language model to determine a target data set matching the semantic query information from multiple semantic data sets in the target semantic knowledge base.
[0157] The determination module is configured to respectively determine multiple content information and positioning information in the target data set as query text content and text positioning information matching the semantic query information.
[0158] According to an embodiment of the present invention, the semantic metadata further includes the total feature length of multiple content information.
[0159] According to an embodiment of the present invention, the text processing device further includes a data elimination module.
[0160] The data elimination module is used for any semantic data set. When the total feature length is greater than the processing length threshold of the preset language model, based on the summary similarity ranking, it eliminates the content information with a summary similarity less than the similarity threshold from multiple content information, and obtains the filtered semantic data set, so as to input the filtered semantic data set and the semantic query information into the preset language model.
[0161] According to an embodiment of the present invention, any multiple of the chunk conversion module 610, the retrieval matching module 620, and the semantic update module 630 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the chunk conversion module 610, the retrieval matching module 620, and the semantic update module 630 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits and other hardware or firmware, or can be implemented in any one of the three implementation manners of software, hardware, and firmware or in any suitable combination of several of them. Or, at least one of the chunk conversion module 610, the retrieval matching module 620, and the semantic update module 630 can be at least partially implemented as a computer program module, and when the computer program module runs, it can execute the corresponding functions.
[0162] Figure 7 A block diagram of an electronic device suitable for implementing the text processing method according to an embodiment of the present invention is shown.
[0163] As Figure 7 shown, the electronic device 700 according to an embodiment of the present invention includes a processor 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory 702 or the program loaded from the storage section 708 into the random access memory 703. The processor 701 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 701 can also include on-board memory for caching purposes. The processor 701 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0164] In the random access memory 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the read-only memory 702, and the random access memory 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method flow according to the embodiments of the present invention by executing programs in the read-only memory 702 and / or the random access memory 703. It should be noted that the program can also be stored in one or more memories other than the read-only memory 702 and the random access memory 703. The processor 701 can also perform various operations of the method flow according to the embodiments of the present invention by executing programs stored in one or more memories.
[0165] According to an embodiment of the present invention, the electronic device 700 may further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 708 including a hard disk, etc.; and a communication portion 709 including a network interface card such as a LAN card, a modem, etc. The communication portion 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage portion 708 as needed.
[0166] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0167] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the read-only memory 702 and / or the random access memory 703 described above and / or one or more memories other than the read-only memory 702 and the random access memory 703.
[0168] An embodiment of the present invention further includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiment of the present invention.
[0169] When the computer program is executed by the processor 701, it executes the above functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0170] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 709, and / or be installed from the removable medium 711. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0171] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or be installed from the removable medium 711. When the computer program is executed by the processor 701, it executes the above functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0172] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0173] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0174] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0175] The above describes the embodiments of the present invention. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A text processing method, characterized in that: The method comprises: Performing block conversion processing on the text to be processed to obtain at least one text content feature and positioning information of the text content feature, wherein the positioning information represents the position of the text content feature in the text to be processed; For any of the text content features, a semantic search is performed in an initial semantic knowledge base to obtain a semantic search result, wherein the initial semantic knowledge base includes at least one semantic data set; When the semantic retrieval result indicates that the at least one semantic data set includes a matching semantic data set that matches the text content feature, the matching semantic data set is updated using the text content feature and the positioning information to obtain a target semantic knowledge base.
2. The method according to claim 1, characterized in that The text to be processed is subjected to block conversion processing to obtain at least one text content feature and location information of the text content feature, including: Processing the text to be processed in blocks to obtain at least one text content information and the positioning information; Perform vector conversion processing on at least one of the text content information to obtain at least one of the text content features.
3. The method according to claim 2, characterized in that The type of the text to be processed includes a document type or a question-and-answer type, and the text content information includes at least one of text content, table content, and image content; The text to be processed is processed in blocks to obtain at least one text content information and the positioning information, including: In the case where the text to be processed is of a document type, the text to be processed is divided into blocks based on a paragraph format to obtain at least one of the text content information and the positioning information; In the case that the text to be processed is of question and answer type, the text to be processed is divided into blocks based on the question and answer format to obtain at least one of the text content information and the positioning information.
4. The method according to claim 1, characterized in that The semantic data set includes at least one semantic summary information; Wherein, for any of the text content features, a semantic search is performed in the initial semantic knowledge base to obtain a semantic search result, including: In the case where the semantic similarity between the text content feature and the semantic summary information meets a preset similarity threshold, generating the semantic retrieval result, wherein the semantic retrieval result represents that the semantic data set corresponding to the semantic similarity is the matching semantic data set; Wherein, when there are multiple text content features, semantic retrieval is performed on the multiple text content features in the initial semantic knowledge base in parallel and the initial semantic knowledge base is updated; Wherein, when the matching semantic data set matches a plurality of the text content features, the initial semantic knowledge base is updated one by one using the plurality of the text content features.
5. The method according to claim 2, characterized in that: The matching semantic data set includes at least one semantic content information and positioning information of the semantic content information; The matching semantic data set is updated using the text content features and the positioning information to obtain a target semantic knowledge base, including: Processing the text content information and at least one semantic content information using a preset language model to obtain text summary information; For any one of the text content information and at least one of the semantic content information, based on summary similarity between the content information and the text summary information, the matching semantic data set is updated to obtain the target semantic knowledge base.
6. The method according to claim 5, characterized in that Based on the summary similarity between the content information and the text summary information, the matching semantic data set is updated to obtain the target semantic knowledge base, including: sorting the plurality of content information and the plurality of positioning information according to the plurality of summary similarities to obtain sorted plurality of content information and sorted plurality of positioning information; Generate semantic metadata according to the file name of the text to be processed, the sorted plurality of location information, the plurality of summary similarities, and feature length information, wherein the feature length information includes the feature length of the content information arranged according to the summary similarity and the total feature length of the plurality of content information; An updated matching semantic data set is generated according to the text summary information, the semantic metadata and the sorted multiple content information, wherein the target semantic knowledge base includes the updated semantic data set.
7. The method according to claim 2, characterized in that Also includes: When the semantic retrieval result indicates that the matching semantic data set is not included, the semantic data set is generated according to the text content features and the positioning information, and a target semantic knowledge base is generated according to the semantic data set.
8. The method according to claim 7, characterized in that Generating the semantic data set according to the text content feature and the positioning information includes: Processing text content information corresponding to the text content feature using a preset language model to obtain content summary information; An association and merging operation is performed on the text to be processed, the text content feature, the content summary information and the positioning information to obtain the semantic data set.
9. The method according to claim 8, characterized in that Performing an association and merging operation on the text to be processed, the text content feature, the content summary information, and the positioning information to obtain the semantic data set includes: Performing an association and merging operation on the text content information, the file name of the text to be processed, the positioning information, and the text content features to obtain semantic metadata; The content summary information, the semantic metadata, and the text content information are combined into the semantic data set.
10. The method according to claim 9, characterized in that Performing an association and merging operation on the text content information, the file name of the text to be processed, the positioning information, and the text content features to obtain semantic metadata includes: The file name, the positioning information, the content similarity between the text content information and the content summary information, and the feature length of the text content feature are associated and merged to obtain the semantic metadata.
11. The method according to claim 1, characterized in that: The semantic data set includes semantic metadata, and the semantic metadata includes a plurality of content information sorted according to summary similarity and location information of the content information; Wherein, the method further comprises: Obtain input semantic query information; Processing the semantic query information and the target semantic knowledge base using a preset language model to determine a target data set matching the semantic query information from a plurality of semantic data sets in the target semantic knowledge base; The plurality of content information and positioning information in the target data set are respectively determined as query text content and text positioning information matching the semantic query information.
12. The method according to claim 11, characterized in that The semantic metadata also includes a characteristic total length of a plurality of the content information; Among them, for any of the semantic data sets, when the total length of the features is greater than the processing length threshold of the preset language model, content information with summary similarity less than the similarity threshold is eliminated from the multiple content information based on summary similarity sorting to obtain a filtered semantic data set, so as to input the filtered semantic data set and the semantic query information into the preset language model.
13. An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 12.
14. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the computer program is used to implement the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Text processing method and device for improving information retrieval and generation quality and computer system
CN118535710A
Multi-knowledge granularity text retrieval method and device for RAG
CN119415623A
An enhanced question answering method integrating context awareness
CN119760082A
Retrieval enhancement generation method and device of large language model, equipment and medium
CN119884339A
Multi-modal document retrieval enhancement generation method based on large model
CN119988588A