Information processing method and device, electronic equipment, computer readable storage medium and computer program product

By building a text knowledge base that includes document structure and hierarchical relationships, the problem of insufficient information comprehensiveness in document knowledge base construction is solved, and the accuracy of text processing is improved.

CN120632059APending Publication Date: 2025-09-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410234276.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-29
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In the prior art, the knowledge base of a document is constructed by dividing the document into equal parts, which results in insufficient comprehensiveness and integrity of the text information and affects the accuracy of the text processing results.

Method used

By building a text knowledge base that contains the text knowledge corresponding to each document structure of the basic document and its hierarchical relationship, and using natural language to describe text processing requests, text processing is performed to improve accuracy.

Benefits of technology

The accuracy of text processing results is improved, and the effect of text processing is enhanced through comprehensive and complete text knowledge query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632059A_ABST
    Figure CN120632059A_ABST
Patent Text Reader

Abstract

The invention provides an information processing method and device, electronic equipment, a computer readable storage medium and a computer program product. The information processing method and device are applied to various document-based text processing scenes such as cloud technology, artificial intelligence, intelligent transportation and games. The information processing method comprises the steps that a to-be-queried text is obtained in response to a text processing request, the to-be-queried text is used for information query in a basic document, and the basic document is an information basis of the text processing request; m pieces of text knowledge matched with the to-be-queried text are queried from a text knowledge base, the text knowledge base comprises text knowledge corresponding to each document structure in the basic document and the hierarchical relation of each piece of text knowledge in the basic document, and the document structure is at least one of a title, a paragraph, a table, a table cell and an attached drawing, m is a positive integer; and performing text processing based on the M pieces of text knowledge to obtain a text processing result. Through the method, the accuracy of the text processing result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to information processing technology in the field of computer applications, and in particular to an information processing method, device, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] Document-based text processing refers to the process of obtaining text processing results corresponding to a text processing request using documents as information. In related technologies, to implement document-based text processing, the text information corresponding to the text processing request is typically first queried from a document knowledge base, and then the text processing results are obtained based on the queried text information. However, the document knowledge base is constructed by dividing the documents into equal parts, which affects the comprehensiveness of the text information obtained from the knowledge base and the accuracy of the text processing results. Summary of the Invention

[0003] Embodiments of the present application provide an information processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can improve the accuracy of text processing results.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] This embodiment of the present application provides an information processing method, the method comprising:

[0006] In response to the text processing request, obtaining a text to be queried, wherein the text to be queried is used to perform information query in a basic document, and the basic document is the information basis of the text processing request;

[0007] Querying M text knowledge matching the query text from a text knowledge base, wherein the text knowledge base includes text knowledge corresponding to each document structure in the basic document and a hierarchical relationship between each text knowledge in the basic document, wherein the document structure is at least one of the following: a title, a paragraph, a table, a table cell, and an illustration, and M is a positive integer;

[0008] Describing the text processing request in natural language based on the M text knowledge to obtain a text processing prompt;

[0009] Text processing is performed based on the text processing prompt to obtain a text processing result.

[0010] An embodiment of the present application provides an information processing device, comprising:

[0011] A request response module, configured to respond to a text processing request and obtain a text to be queried, wherein the text to be queried is used to perform information query in a basic document, and the basic document is the information basis for the text processing request;

[0012] a text query module, configured to query a text knowledge base for M text knowledge matching the text to be queried, wherein the text knowledge base includes text knowledge corresponding to each document structure in the basic document and a hierarchical relationship between each text knowledge in the basic document, wherein the document structure is at least one of the following: a title, a paragraph, a table, a table cell, and an illustration, and M is a positive integer;

[0013] A request description module, configured to describe the text processing request in natural language based on the M text knowledge to obtain a text processing prompt;

[0014] The text processing module is used to perform text processing based on the text processing prompt to obtain a text processing result.

[0015] In an embodiment of the present application, the information processing device also includes a knowledge construction module, which is used to obtain the basic document in response to a document analysis request; perform structural analysis on the basic document to obtain the text knowledge and text hierarchical relationship of each document structure, and the text hierarchical relationship represents the hierarchical relationship of each text knowledge in the basic document; and combine each text knowledge based on the text hierarchical relationship to obtain the text knowledge base.

[0016] In an embodiment of the present application, the knowledge construction module is also used to combine each of the text knowledge based on the text hierarchical relationship to obtain an initial knowledge base; combine the basic text knowledge of the lowest level in the initial knowledge base into a basic text knowledge set; count the character string length corresponding to each basic text knowledge in the basic text knowledge set; in the initial knowledge base, balance the basic text knowledge set based on the character string length and the specified length to obtain the text knowledge base, and the balancing process is the splitting of the basic text knowledge or the merging of multiple basic text knowledge.

[0017] In an embodiment of the present application, the knowledge construction module is also used to select at least one text knowledge to be split from the basic text knowledge set, whose string length is greater than the specified difference of the specified length; in the initial knowledge base, based on the rounded result L of the string length of each text knowledge to be split and the specified length, the corresponding text knowledge to be split is split into L of the text knowledge at the lowest level to obtain the text knowledge base, where L is an integer greater than 1.

[0018] In an embodiment of the present application, the L lowest-level text knowledge is independent based on designated punctuation marks, where the designated punctuation marks refer to punctuation marks representing text sentences.

[0019] In an embodiment of the present application, the knowledge construction module is also used to select from the basic text knowledge set, text knowledge belonging to the same parent level, whose string length is less than the specified length and specified difference, and adjacent text knowledge sets to be merged, to obtain at least one text knowledge set to be merged; in the initial knowledge base, multiple text knowledge to be merged in each text knowledge set to be merged are merged to obtain the text knowledge base.

[0020] In an embodiment of the present application, the request description module is further used to use M of the text knowledge to describe the text processing request in natural language to obtain the text processing prompt; or, to combine M of the text knowledge and a simplified document to describe the text processing request in natural language to obtain the text processing prompt, wherein the simplified document is the simplified result of the basic document, and the simplified document is obtained by summarizing the text knowledge in the text knowledge base.

[0021] In an embodiment of the present application, the knowledge construction module is further used to select at least one text knowledge to be extracted corresponding to each text knowledge to be simplified from the text knowledge base, the level at which the text knowledge to be simplified is located is the parent level of the lowest level, and the level at which the text knowledge to be extracted is located is the child level of the text knowledge to be simplified; perform summary extraction on at least one of the text knowledge to be extracted to obtain a simplified text; use the simplified text to replace at least one text knowledge to be extracted corresponding to each of the text knowledge to be simplified from the text knowledge base to obtain a simplified text library; convert the simplified text library into a document to obtain the simplified document.

[0022] In an embodiment of the present application, the text query module is also used to extract features of the text to be queried to obtain features to be queried; determine T text features whose feature similarity with the features to be queried is greater than a similarity threshold from a text feature library, wherein the text feature library includes the text features of each text knowledge in the text knowledge base, and T is a positive integer; obtain T text knowledge to be recalled corresponding to the T text features from the text knowledge base; and determine M text knowledge that matches the text to be queried based on the T text knowledge to be recalled.

[0023] In an embodiment of the present application, the text query module is also used to perform the following processing on each of the T text knowledge to be recalled: from the text knowledge base, the text knowledge to be recalled is expanded and recalled to obtain a text knowledge set, and the expansion recall is implemented based on at least one of the following: expansion coefficient, lowest level text knowledge coverage, and the association relationship between each of the text knowledge; from the text knowledge set of each of the text knowledge to be recalled, T text knowledge sets corresponding to the T text knowledge to be recalled are obtained; at least one of the text knowledge included in the T text knowledge sets is determined as M text knowledge.

[0024] In an embodiment of the present application, the text query module is also used to obtain the target text knowledge corresponding to the parent level of the text knowledge to be recalled from the text knowledge base; obtain the text knowledge sequence corresponding to the child level of the target text knowledge; from the text knowledge sequence, expand and recall S pieces of text knowledge adjacent to the text knowledge to be recalled, where S is a positive integer and S is the expansion coefficient; and obtain the text knowledge set based on the text knowledge to be recalled and the S pieces of text knowledge.

[0025] In an embodiment of the present application, the text query module is also used to obtain the number of common text knowledge of T of the text knowledge to be recalled and the text knowledge sequence; obtain the number of sequence text knowledge of the text knowledge sequence; determine the ratio of the number of common text knowledge to the number of sequence text knowledge as the text knowledge coverage; when the text knowledge coverage is greater than the coverage threshold, determine the text knowledge set based on the text knowledge sequence.

[0026] In an embodiment of the present application, the text query module is also used to obtain target association information associated with the text knowledge to be recalled from an association relationship library, where the association relationship library is information described in the auxiliary documents of the basic document; obtain at least one associated text knowledge corresponding to the target association information from the text knowledge library; and determine the text knowledge set based on the text knowledge to be recalled and at least one associated text knowledge.

[0027] In an embodiment of the present application, the text processing based on the text processing prompt is implemented through a text processing model, and the information processing device also includes a model training module for obtaining text processing training data, wherein the text processing training data includes text processing sample prompts and result sample labels; text processing is performed on the text processing sample prompts based on the model to be trained to obtain a text processing estimation result, and the model to be trained refers to a neural network model to be trained for text processing; based on the difference between the text processing estimation result and the result sample label, the model to be trained is trained to obtain the text processing model.

[0028] In an embodiment of the present application, when the text to be queried is a question of the basic document, the text processing prompt is a question to be answered with M pieces of text knowledge as background knowledge; the text processing module is also used to perform text processing based on the question to be answered to obtain a target answer to the question to be answered; and the target answer is determined as the text processing result.

[0029] In an embodiment of the present application, when the text to be queried is an imitation description based on the basic document, the text processing prompt is the description to be imitated using M of the text knowledge as an imitation method; the text processing module is also used to perform text processing based on the description to be imitated to obtain a target imitation result; and the target imitation result is determined as the text processing result.

[0030] An embodiment of the present application provides an electronic device for information processing, the electronic device comprising:

[0031] a memory for storing computer-executable instructions or computer programs;

[0032] The processor is used to implement the information processing method provided in the embodiment of the present application when executing the computer-executable instructions or computer programs stored in the memory.

[0033] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the information processing method provided in the embodiment of the present application is implemented.

[0034] An embodiment of the present application provides a computer program product, including computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the information processing method provided in the embodiment of the present application is implemented.

[0035] The embodiments of the present application have at least the following beneficial effects: in the text processing process based on document query, since the text knowledge base of the basic document on which the text processing request is based is constructed through text knowledge corresponding to at least one document structure in the title, paragraph, table, table cell and accompanying drawing, the text knowledge base includes the structural information of the basic document, thereby improving the comprehensiveness of the M text knowledge queried from the text knowledge base; furthermore, when text processing is performed based on the M text knowledge, the accuracy of text processing can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a schematic diagram of the architecture of the information processing system provided by an embodiment of the present application;

[0037] Figure 2 This embodiment of the present application provides a Figure 1 A schematic diagram of the server structure in FIG;

[0038] Figure 3 This is a flow diagram of the information processing method provided in the embodiment of the present application. Figure 1 ;

[0039] Figure 4 This is a flow diagram of the information processing method provided in the embodiment of the present application. Figure 2 ;

[0040] Figure 5 This is a flowchart of the information recall provided by the embodiment of the present application;

[0041] Figure 6 This is an exemplary text processing flowchart provided by an embodiment of the present application;

[0042] Figure 7 This is a schematic diagram of an exemplary method for obtaining answer content provided in an embodiment of the present application;

[0043] Figure 8 is another exemplary schematic diagram of obtaining answer content provided in an embodiment of the present application;

[0044] Figure 9 This is an exemplary document processing diagram provided by an embodiment of the present application;

[0045] Figure 10 This is an exemplary information query diagram provided in an embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0047] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0048] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0049] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0050] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant national laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0051] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0052] 1) Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that aims to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. By studying the design principles and implementation methods of various intelligent machines, AI enables them to possess the capabilities of perception, reasoning, and decision-making.

[0053] It should be noted that artificial intelligence technology covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, pre-trained model technologies, operating / interaction systems, and mechatronics. Among them, artificial intelligence software technologies include computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. In the embodiments of this application, the text processing process can be implemented based on AI.

[0054] 2) Machine Learning (ML) is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It is used to study computer simulation or implementation of human learning behavior to acquire new knowledge or skills; reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Machine learning applications are spread across all areas of artificial intelligence. Machine learning / deep learning usually includes technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning. The big model is the latest development in machine learning / deep learning, which integrates the above technologies. Among them, the big model is also called a pre-trained model or a basic model. The big model can be applied directly or after fine-tuning to downstream tasks in various directions of artificial intelligence. In the embodiment of the present application, the text processing model can be a big model or a fine-tuned big model.

[0055] 3) Artificial neural networks are mathematical models that mimic the structure and function of biological neural networks. Examples of artificial neural network structures in the embodiments of this application include graph convolutional networks (GCNs, a type of neural network used to process graph-structured data), deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), neural state machines (NSMs), and phase-functioned neural networks (PFNNs). In the embodiments of this application, text processing can be achieved using artificial neural networks.

[0056] 4) Natural language processing (NLP) is a field of study in computer science and artificial intelligence that focuses on various theories and methods for effective communication between humans and computers using natural language. Natural language processing involves natural language, the language people use in daily life. Natural language processing generally involves technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs. The information processing method provided in the embodiments of this application can be applied in the field of natural language processing.

[0057] It should be noted that in order to realize document-based text processing, it is usually necessary to first query the document's knowledge base for text information corresponding to the text processing request, and then obtain the text processing result based on the queried text information; however, the construction of the document's knowledge base is usually achieved by first extracting the document's text content and then dividing the document's text content into equal parts, so that the constructed document knowledge base loses the document's structural information; thus, the text information queried from the document knowledge base is scattered and missing, which affects the comprehensiveness and completeness of the text information obtained from the knowledge base, and affects the accuracy of the text processing results.

[0058] Based on this, the embodiments of the present application provide an information processing method, device, electronic device, computer-readable storage medium and computer program product, which can improve the accuracy of text processing results. The following describes an exemplary application of an electronic device for information processing (hereinafter referred to as an information processing device) provided by an embodiment of the present application. The information processing device provided by an embodiment of the present application can be implemented as various types of terminals such as robots, smart phones, smart watches, laptops, tablet computers, desktop computers, smart home appliances, set-top boxes, smart car devices, portable music players, personal digital assistants, dedicated messaging devices, intelligent voice interaction devices, portable gaming devices and smart speakers. It can also be implemented as a server, or as a combination of the two. Below, an exemplary application of the information processing device when it is implemented as a server will be described.

[0059] See also Figure 1 , Figure 1 is a schematic diagram of the architecture of the information processing system provided in the embodiment of the present application; Figure 1 As shown, to support an information processing application, in the information processing system 100, the terminal 400 (terminal 400-1 and terminal 400-2 are shown as examples) is connected to the server 200 via the network 300; the network 300 can be a wide area network or a local area network, or a combination of the two. In addition, the information processing system 100 also includes a database 500 for providing data support to the server 200; and Figure 1, which shows a case where the database 500 is independent of the server 200. In addition, the database 500 may also be integrated into the server 200, which is not limited in the embodiment of the present application.

[0060] The terminal 400 is used to send a text processing request to the server 200 via the network 300. The terminal 400 is also used to receive the text processing result sent by the server 200 via the network 300 and present the text processing result (graphic interface 410-1 and graphical interface 410-2 are shown as examples).

[0061] The server 200 is used to receive a text processing request sent by the terminal 400 through the network 300, and obtain a text to be queried in response to the text processing request. The text to be queried is used to perform information query in the basic document, and the basic document is the information basis for the text processing request; M text knowledge matching the text to be queried is searched from the text knowledge base, and the text knowledge base includes the text knowledge corresponding to each document structure in the basic document, and the hierarchical relationship of each text knowledge in the basic document; the text processing request is described in natural language based on the M text knowledge to obtain a text processing prompt; text processing is performed based on the text processing prompt to obtain a text processing result; and the text processing result is sent to the terminal 400 through the network 300.

[0062] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and the server may be connected directly or indirectly via wired or wireless communication, which is not limited in the embodiments of the present application.

[0063] See also Figure 2 , Figure 2 This embodiment of the present application provides a Figure 1 The structural diagram of the server in Figure 2 As shown, the server 200 includes: at least one processor 210, a memory 250, and at least one network interface 220. The various components in the server 200 are coupled together via a bus system 240. It is understood that the bus system 240 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 2 Various buses are labeled as bus system 240 .

[0064] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0065] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 250 may optionally include one or more storage devices that are physically remote from the processor 210.

[0066] The memory 250 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.

[0067] In some embodiments, the memory 250 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0068] Operating system 251, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0069] The network communication module 252 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB).

[0070] In some embodiments, the information processing device provided in the embodiments of the present application can be implemented in software. Figure 2 The information processing device 255 stored in the memory 250 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a request response module 2551, a text query module 2552, a request description module 2553, a text processing module 2554, a knowledge construction module 2555, and a model training module 2556. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0071] In some embodiments, the information processing device provided in the embodiments of the present application can be implemented in hardware. As an example, the information processing device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the information processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.

[0072] In some embodiments, the terminal or server can implement the information processing method provided by the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (Application, APP), that is, a program that needs to be installed in the operating system to run, such as a game APP or a game creation APP; it can also be a small program that can be embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.

[0073] Below, the information processing method provided by the embodiment of the present application will be described in conjunction with the exemplary application and implementation of the information processing device provided by the embodiment of the present application. In addition, the information processing method provided by the embodiment of the present application is applied to various document-based text processing scenarios such as cloud technology, artificial intelligence, smart transportation, and games.

[0074] See also Figure 3 , Figure 3 This is a flow diagram of the information processing method provided in the embodiment of the present application. Figure 1 ;in, Figure 3 The execution subject of each step is the information processing device; Figure 3 The steps shown are explained.

[0075] Step 101: In response to a text processing request, obtain a text to be queried.

[0076] In an embodiment of the present application, when text processing is performed based on a basic document, the information processing device also receives a text processing request; therefore, the text processing request represents a request to perform information query based on the basic document and perform text processing on the information query result; thus, the information processing device responds to the text processing request and can obtain information to be queried in the basic document from the text processing request, and the text form of the information used for information query in the basic document in the text processing request is referred to as the text to be queried. Here, the information used for information query in the basic document in the text processing request can be in text form, voice form, image form, or a combination of the above, etc., and the embodiment of the present application does not limit this; and when the information used for information query in the basic document in the text processing request is in non-text form, the information is converted into information in text form.

[0077] It should be noted that the text to be queried is used to perform information query in the basic document. The basic document is the information basis for the text processing request, such as the traffic encyclopedia, the game world view document, etc.; in addition, the document query in the document-based text processing refers to the information query in the basic document.

[0078] Step 102: Query M text knowledge matching the text to be queried from the text knowledge base.

[0079] In an embodiment of the present application, the information processing device itself includes a text knowledge base, or can obtain a text knowledge base from other devices (for example, a storage device such as a database), and the text knowledge base is constructed based on a basic document, including the text knowledge corresponding to each document structure in the basic document, and the hierarchical relationship of each text knowledge in the basic document. Thus, when the information processing device performs an information query in the basic document based on the text to be queried, it can be achieved by performing a text knowledge query in the text knowledge base based on the text to be queried. Here, the information processing device matches the text to be queried with each text knowledge in the text knowledge base, so that M text knowledge matching the text to be queried can be obtained from the text knowledge base; wherein M is a positive integer, representing the number of text knowledge matching the text to be queried in the text knowledge base; matching means that the text to be queried is similar to the text knowledge, which can be determined by the similarity being greater than the similarity threshold, and the similarity can be the feature similarity of the feature dimension, or it can be the character matching degree, etc., which is not limited in the embodiment of the present application.

[0080] It should be noted that the document structure is at least one of the following: title, paragraph, table, table cell and illustration; wherein the hierarchical relationship of each text knowledge in the basic document refers to the hierarchical relationship of each document structure in the basic document, which can be a parent-child hierarchical relationship, where the parent-level document structure can be a title, and the child-level document structure can be at least one of a title, paragraph, table, table cell and illustration, while the lowest-level document structure is at least one of a paragraph, table, table cell and illustration; the title can be a single-level title or a multi-level title. When the title is a multi-level title, there is a hierarchical relationship between titles at all levels, and both the parent-level document structure and the child-level document structure can be titles. For example, the sub-level of the first-level title includes at least one second-level title, and the sub-level of the second-level title includes at least one third-level title, etc. In addition, the M text knowledge refers to the information query results similar to the text to be queried from the basic document, and each text knowledge in the M text knowledge is similar to the text to be queried.

[0081] It can be understood that in the text knowledge base, each text knowledge is the complete information of a document structure in the basic document; therefore, when information query is performed on the query text in the text knowledge base and information query is performed on the query text in the basic document, the completeness of the obtained information query results can be improved, and then the accuracy of the information query can be improved.

[0082] See also Figure 4 , Figure 4 This is a flow diagram of the information processing method provided in the embodiment of the present application. Figure 2 ;in, Figure 4 The execution subject of each step in is the information processing device; Figure 4 As shown, step 102 can be implemented through steps 1021 to 1024, that is, the information processing device queries M text knowledge matching the text to be queried from the text knowledge base, including steps 1021 to 1024, and each step is explained below.

[0083] Step 1021: extract features from the query text to obtain query features.

[0084] In an embodiment of the present application, the information processing device can determine whether the text to be queried matches the text knowledge in the text knowledge base through the similarity of the feature dimension; at this time, the information processing device performs feature extraction on the text to be queried, and the extracted features are the features of the text to be queried. Here, the features of the text to be queried are referred to as features to be queried.

[0085] It should be noted that the information processing device can use a trained neural network model to extract features, or can use a one-hot code method to extract features, etc., and the embodiments of the present application are not limited to this.

[0086] Step 1022: Determine T text features from the text feature library whose feature similarity with the feature to be queried is greater than a similarity threshold.

[0087] In an embodiment of the present application, the information processing device itself includes a text feature library, or can obtain a text feature library from other devices, and the text feature library includes text features of each text knowledge in the text knowledge base; thus, the information processing device calculates the similarity between the feature to be queried and each text feature in the text feature library one by one, and the calculated similarity is the feature similarity; since the similarity threshold represents the highest feature similarity at which the feature to be queried is not similar to the text feature, when the feature similarity of the feature to be queried is greater than the similarity threshold, the information processing device determines the text feature with a feature similarity greater than the similarity threshold as a text feature similar to the feature to be queried; here, based on the feature similarity between each text feature and the feature to be queried, the information processing device can determine T text features similar to the feature to be queried from the text feature library, where T is a positive integer, representing the number of text features in the text feature library that are similar to the feature to be queried.

[0088] It should be noted that each text feature in the text feature library represents a feature of text knowledge; each text feature in the text feature library has a corresponding text knowledge in the text knowledge library, or in other words, each text knowledge in the text knowledge library has a corresponding text feature in the text feature library; and T text features are at least one text feature in the text feature library that matches (also known as is similar to) the query feature. Furthermore, the method for extracting text features is the same as the method for extracting query features, but can also be different, and this embodiment of the present application does not limit this.

[0089] It can be understood that by pre-extracting features from the text knowledge in the text knowledge base, text features are obtained, and a text feature library corresponding to the text knowledge base is constructed based on the text features, so that when the text knowledge base performs information query on the query text, it can be matched based on the existing text knowledge base, thereby improving the matching efficiency.

[0090] Step 1023: Obtain T pieces of to-be-recalled text knowledge corresponding to the T text features from the text knowledge base.

[0091] It should be noted that since the text knowledge in the text knowledge base corresponds one-to-one with the text features in the text feature base, the information processing device can obtain T pieces of text knowledge corresponding to the T text features from the text knowledge base. Since the T pieces of text knowledge obtained are the text knowledge to be recalled from the text knowledge base, they are called to-be-recalled text knowledge. Therefore, the T pieces of text knowledge obtained are the T to-be-recalled text knowledge. There is a one-to-one correspondence between the T text features and the T to-be-recalled text knowledge.

[0092] Step 1024: Based on the T pieces of text knowledge to be recalled, determine M pieces of text knowledge that match the text to be queried.

[0093] In an embodiment of the present application, the information processing device can determine T pieces of text knowledge to be recalled as M pieces of text knowledge, in which case T is equal to M; or it can perform expansion recall in the text knowledge base based on the T pieces of text knowledge to be recalled to obtain M pieces of text knowledge, in which case T≤M; it is easy to see that the M pieces of text knowledge include at least T pieces of text knowledge to be recalled.

[0094] See also Figure 5 , Figure 5 This is a flow chart of information recall provided by the embodiment of the present application; wherein, Figure 5 The execution subject of each step in is the information processing device; Figure 5 As shown, step 1024 can be implemented through steps 10241 to 10243, that is, the information processing device determines M text knowledge matching the text to be queried based on T text knowledge to be recalled, including steps 10241 to 10243, and each step is explained below.

[0095] In an embodiment of the present application, the information processing device performs the following processing (ie, step 10241 ) on each piece of text knowledge to be recalled among the T pieces of text knowledge to be recalled.

[0096] Step 10241: Perform expansion recall on the text knowledge to be recalled from the text knowledge base to obtain a text knowledge set.

[0097] In an embodiment of the present application, an information processing device performs an expansion recall on each text knowledge to be recalled in a text knowledge base, and combines the expansion recall result and the text knowledge to be recalled into a text knowledge set. Here, the expansion recall result can represent text knowledge that has not been expanded and recalled that is similar to the text knowledge to be recalled. In this case, the text knowledge set is the text knowledge to be recalled. The similarity can be determined by the similarity between the text to be recalled and the text knowledge being greater than an expansion recall threshold, and the expansion recall threshold is greater than or equal to the similarity threshold. In this way, the expansion recall accuracy can be improved. The expansion recall result can represent the expansion recall of at least one text knowledge. In this case, the text knowledge set is the at least one text knowledge that has been expanded and recalled and the text knowledge to be recalled.

[0098] It should be noted that the expansion recall is implemented based on at least one of the following: expansion coefficient, text knowledge coverage at the lowest level, and the association relationship between each text knowledge; wherein the expansion coefficient represents the amount of text knowledge expanded and recalled in a specified expansion direction (for example, in the same level direction as the level where the text knowledge to be recalled is located); the text knowledge coverage at the lowest level refers to the ratio of the number of public text knowledge to the number of sequence text knowledge, the number of sequence text knowledge refers to the number of text knowledge in the text knowledge sequence, the text knowledge sequence refers to the full amount of text knowledge of the sub-level included in the text knowledge of the parent level of the text knowledge to be recalled, and the number of public text knowledge refers to the number of text knowledge to be recalled included in the text knowledge sequence; the association relationship between each text knowledge, for example, the game character a in text knowledge A and the game character b in text knowledge B are brothers.

[0099] In an embodiment of the present application, when the information processing device performs expansion recall based on the expansion coefficient, in step 10241, the information processing device performs expansion recall on the text knowledge to be recalled from the text knowledge base to obtain a text knowledge set, including: the information processing device first obtains the target text knowledge corresponding to the parent level of the text knowledge to be recalled from the text knowledge base; then obtains the text knowledge sequence corresponding to the child level of the target text knowledge; then, from the text knowledge sequence, expands and recalls S text knowledge adjacent to the text knowledge to be recalled; finally, based on the text knowledge to be recalled and the S text knowledge, a text knowledge set is obtained.

[0100] It should be noted that the target text knowledge corresponding to the parent level of the text knowledge to be recalled is the text knowledge of the parent level to which the text knowledge to be recalled belongs; S is a positive integer, and S is the expansion coefficient; the text knowledge to be recalled is similar to the S text knowledge. Here, the information processing device can combine the text knowledge to be recalled and the S text knowledge into a text knowledge set, or can combine the text knowledge to be recalled and the S text knowledge, as well as text knowledge recalled by other expansion recall methods (expansion recall methods based on one or both of text knowledge coverage and association relationships) into a text knowledge set, and the embodiments of the present application do not limit this.

[0101] In an embodiment of the present application, when the information processing device performs expanded recall based on text knowledge coverage, after the information processing device obtains the text knowledge sequence corresponding to the sub-level of the target text knowledge, the information processing method also includes: the information processing device first obtains the number of common text knowledge of T text knowledge to be recalled and the text knowledge sequence; then obtains the number of sequence text knowledge of the text knowledge sequence; then, determines the ratio of the number of common text knowledge to the number of sequence text knowledge as the text knowledge coverage; finally, when the text knowledge coverage is greater than the coverage threshold, determines the text knowledge set based on the text knowledge sequence.

[0102] It should be noted that the number of common text knowledge can be determined by obtaining the number of common text knowledge of T to-be-recalled text knowledge and the text knowledge sequence. Here, when the text knowledge coverage is less than or equal to the coverage threshold, it means that the expansion recall based on the text knowledge coverage has failed. At this time, the information processing device regards the to-be-recalled text knowledge as a text knowledge set; when the text knowledge coverage is greater than the coverage threshold, the information processing device regards the entire text knowledge sequence as a text knowledge set; in addition, the information processing device can also combine the text knowledge sequence with other expansion recall methods (expansion recall based on one or both of the expansion coefficient and the association relationship) to form a text knowledge set, which is not limited in this embodiment of the present application.

[0103] It should also be noted that the priority of the expansion recall based on text knowledge coverage can be higher than the priority of the expansion recall based on the expansion coefficient; in this way, the information processing device can perform the expansion recall based on the expansion coefficient only when the expansion recall based on text knowledge coverage fails, and end the expansion recall processing based on the expansion coefficient when the expansion recall based on text knowledge coverage succeeds.

[0104] In an embodiment of the present application, when the information processing device performs expansion recall based on an association relationship, in step 10241, the information processing device performs expansion recall on the text knowledge to be recalled from the text knowledge base to obtain a text knowledge set, including: the information processing device first obtains the target association information associated with the text knowledge to be recalled from the association relationship library; then obtains at least one associated text knowledge corresponding to the target association information from the text knowledge base; finally, based on the text knowledge to be recalled and at least one associated text knowledge, determines the text knowledge set.

[0105] It should be noted that the association relationship library is the information described in the auxiliary document of the basic document, which is used to describe the association relationship between various text knowledge. The at least one associated text knowledge is at least one text knowledge that matches the target associated information queried from the text knowledge library. Here, the information processing device can combine the text knowledge to be recalled and at least one associated text knowledge into a text knowledge set, or can combine the text knowledge to be recalled and at least one associated text knowledge, as well as text knowledge recalled by other expansion recall methods (expansion recall based on one or both of the expansion coefficient and text knowledge coverage) into a text knowledge set.

[0106] It is understandable that the information processing device can perform expansion recall on the text knowledge to be recalled in the text knowledge base based on at least one of the expansion coefficient, text knowledge coverage and association relationship, which can improve the recall rate of the expansion recall.

[0107] Step 10242: Obtain T text knowledge sets corresponding to the T pieces of text knowledge to be recalled from the text knowledge sets of each piece of text knowledge to be recalled.

[0108] It should be noted that, after the information processing device obtains the text knowledge set corresponding to each of the T pieces of text knowledge to be recalled, it also obtains T text knowledge sets corresponding one-to-one to the T pieces of text knowledge to be recalled.

[0109] Step 10243: Determine at least one text knowledge included in the T text knowledge sets as M text knowledge.

[0110] It should be noted that the T text knowledge sets include at least one text knowledge, and the information processing device determines the at least one text knowledge included in the T text knowledge sets as M text knowledge.

[0111] Step 103: Describe the text processing request in natural language based on the M text knowledge to obtain a text processing prompt.

[0112] It should be noted that the information processing device uses the M textual knowledge as the basis for text processing. Therefore, when the text processing is natural language processing, the M textual knowledge serves as the knowledge background and a natural language description is used to describe the text processing request to obtain a description prompt for the text processing, referred to herein as a text processing description (Prompt). Here, the information processing device can determine a description template based on the type of text processing request, and then describe the text processing request in natural language based on the determined description template and the M textual knowledge.

[0113] In step 103 of the embodiment of the present application, the information processing device describes the text processing request in natural language based on M text knowledge to obtain a text processing prompt, which may include: the information processing device uses M text knowledge to describe the text processing request in natural language to obtain a text processing prompt; it may also include: the information processing device combines M text knowledge and a streamlined document to describe the text processing request in natural language to obtain a text processing prompt; the embodiment of the present application is not limited to this.

[0114] It should be noted that the simplified document is a simplified result of the basic document, and the simplified document is obtained by extracting the summary of the text knowledge in the text knowledge base.

[0115] It can be understood that since the M text knowledge are the complete content of the matching document structure obtained from the basic document, and the simplified document is the summary information of the basic document, when the information processing device combines the M text knowledge and the simplified document to describe the text processing request in natural language, using the M text knowledge and the simplified document as the information basis for text processing can improve the comprehensiveness of the information basis, and thus improve the accuracy of the text processing description.

[0116] Step 104: Perform text processing based on the text processing prompt to obtain a text processing result.

[0117] It should be noted that because the text processing prompt describes a text processing request based on M pieces of text knowledge, the information processing device can perform text processing on the text processing prompt and obtain a text processing result. The text processing result is the result information corresponding to the text processing requested by the text processing request, such as the answer to the question, the result of text enhancement, the result of information imitation, etc.

[0118] In an embodiment of the present application, an information processing device may implement text processing using a neural network model, where the model used to implement text processing is referred to as a text processing model; that is, the information processing device performs text processing based on text processing prompts through a text processing model, wherein the text processing model is trained through the following steps: the information processing device first obtains text processing training data, and the text processing training data includes text processing sample prompts and result sample labels; then, based on the model to be trained, text processing is performed on the text processing sample prompts to obtain a text processing estimated result; finally, based on the difference between the text processing estimated result and the result sample label, the model to be trained is trained to obtain a text processing model.

[0119] It should be noted that the text processing sample prompt is similar to the text processing prompt, and the result sample label, text processing estimated result and text processing result are similar. The embodiments of the present application will not be repeated here. Here, the information processing device calculates the loss function value based on the difference between the text processing estimated result and the result sample label, and performs backpropagation in the model to be trained based on the calculated loss function value to adjust the model parameters of the model to be trained, thereby realizing the training of the model to be trained. The training process of the model to be trained can be iterative. When the iteration end condition is met, the training is terminated, and the model to be trained by the last iteration is determined as the text processing model. Among them, the iteration end condition can be reaching the accuracy index threshold, or reaching the iteration number threshold, or reaching the iteration duration threshold, or a combination of the above, etc., which is not limited in the embodiments of the present application. Among them, the model to be trained refers to the neural network model to be trained for text processing, which can be the original neural network model constructed, or a pre-trained neural network model, or a large model, etc., which is not limited in the embodiments of the present application.

[0120] In an embodiment of the present application, when the text to be queried is a question based on a document, it indicates that the information processing method is applied in a question answering scenario. At this time, the text processing prompt is a question to be answered with M text knowledge as background knowledge; here, the information processing device performs text processing based on the text processing prompt to obtain a text processing result, including: the information processing device first performs text processing based on the question to be answered to obtain a target answer to the question to be answered; and determines the target answer as the text processing result.

[0121] It should be noted that the information processing device performs text processing based on the question to be answered, which refers to the process of obtaining the answer to the question from M text knowledge. The target answer is the answer to the question in the basic document.

[0122] In an embodiment of the present application, when the text to be queried is an imitation description based on a basic document, it indicates that the information processing method is applied in an information imitation scenario. At this time, the text processing prompt is a description to be imitated using M text knowledge as an imitation method; here, the information processing device performs text processing based on the text processing prompt to obtain a text processing result, including: the information processing device first performs text processing based on the description to be imitated to obtain a target imitation result; and determines the target imitation result as the text processing result.

[0123] It should be noted that the information processing device performs text processing based on the description to be imitated, which means imitating the query text based on the writing styles in the M text knowledge to obtain a target imitation result. The target imitation result is the result obtained by imitating the query text based on the writing styles in the M text knowledge.

[0124] It can be understood that in the text processing process based on document query, since the text knowledge base of the basic document on which the text processing request is based is constructed through the text knowledge corresponding to at least one document structure in the title, paragraph, table, table cell and accompanying drawing, the text knowledge base includes the structural information of the basic document, thereby improving the comprehensiveness of the M text knowledge queried from the text knowledge base; furthermore, when text processing is performed based on the M text knowledge, the accuracy of text processing can be improved.

[0125] Continue to see Figure 4 In an embodiment of the present application, steps 105 to 107 are also included before step 102; that is, before the information processing device queries M text knowledge matching the text to be queried from the text knowledge base, the information processing method also includes steps 105 to 107, and each step is explained below.

[0126] Step 105: In response to the document analysis request, obtain the basic document.

[0127] In an embodiment of the present application, when a request is made to perform structural analysis on a basic document, the information processing device also receives a document analysis request; thus, the document analysis request is used to request to perform structural analysis on the basic document; furthermore, the information processing device responds to the document analysis request and can obtain the basic document from the document analysis request, or can obtain the basic document based on a text identifier in the document analysis request.

[0128] Step 106: Perform structural analysis on the basic documents to obtain textual knowledge of each document structure and textual hierarchical relationships.

[0129] In an embodiment of the present application, the information processing device performs structural analysis on the basic document to obtain the various document structures of the basic document, and obtain the document content corresponding to each document structure (called text knowledge), as well as the hierarchical relationship of the various text knowledge corresponding to each document structure in the basic document (called text hierarchical relationship).

[0130] It should be noted that the text hierarchy represents the hierarchical relationship between the various textual knowledge in the base document. Furthermore, in the text hierarchy, the lowest-level textual knowledge does not include sub-level textual knowledge. The lowest-level textual knowledge can be at least one of a paragraph, a table, a table cell, and an illustration, while the textual knowledge above the lowest level is the title.

[0131] Step 107: Combine the various text knowledge based on the text hierarchical relationship to obtain a text knowledge base.

[0132] In an embodiment of the present application, an information processing device combines various textual knowledge based on the hierarchical relationship of texts to obtain a textual knowledge base. The combination may refer to constructing an N-ary tree, where the parent node is the textual knowledge corresponding to each level of title, and the leaf node is the textual knowledge corresponding to at least one of paragraphs, tables, table cells, and figures. N is a positive integer representing the maximum number of child nodes in the tree structure.

[0133] Continue to see Figure 4 In an embodiment of the present application, step 107 can be implemented through steps 1071 to 1074; that is, the information processing device combines various text knowledge based on the text hierarchical relationship to obtain a text knowledge base, including steps 1071 to 1074, and each step is explained below.

[0134] Step 1071: Combine the various text knowledge based on the text hierarchical relationship to obtain an initial knowledge base.

[0135] In the embodiments of the present application, the information processing device combines various text knowledge based on the text hierarchical relationship, and the result obtained is the initial knowledge base. Here, the information processing device can use the initial knowledge base as the text knowledge base, or can balance the string length of the text knowledge at the lowest level in the initial knowledge base and use the balanced initial knowledge base as the text knowledge base, which is not limited in the embodiments of the present application.

[0136] Step 1072: Combine the lowest-level basic text knowledge in the initial knowledge base into a basic text knowledge set.

[0137] It should be noted that in the initial knowledge base, the lowest level of textual knowledge is the main body of the basic document, while the textual knowledge above the lowest level is the title content of the main body. Because the string lengths of the main body of the basic document vary, the information processing device obtains the lowest level of basic textual knowledge in the initial knowledge base to balance the string lengths of the basic textual knowledge. Here, basic textual knowledge refers to the lowest level of textual knowledge in the initial knowledge base, and the basic textual knowledge set is the set consisting of the entire amount of basic textual knowledge in the initial knowledge base.

[0138] Step 1073: Count the length of the character string corresponding to each basic text knowledge in the basic text knowledge set.

[0139] It should be noted that the information processing device counts the number of characters in each basic text knowledge in the basic text knowledge set, and uses the counted number of characters as the string length of the basic text knowledge; here, the information processing device can obtain a string length set corresponding to the basic text knowledge set, and the basic text knowledge in the basic text knowledge set corresponds one-to-one to the string length in the string length set.

[0140] Step 1074: In the initial knowledge base, the basic text knowledge set is balanced based on the string length and the specified length to obtain a text knowledge base.

[0141] It should be noted that the specified length is used to constrain the balance of the string length of each basic text knowledge, which is a preset length threshold. It can be randomly preset or determined based on the mean or mode of the string length, etc., and the embodiments of this application do not limit this. Here, the balancing process is the splitting of basic text knowledge or the merging of multiple basic text knowledge. The information processing device splits the basic text knowledge with a string length greater than the specified length and merges multiple adjacent basic text knowledge with a corresponding string length less than the specified length.

[0142] In step 1074 of the embodiment of the present application, the information processing device performs balanced processing on the basic text knowledge set in the initial knowledge base based on the string length and the specified length to obtain a text knowledge base, including: the information processing device first selects at least one text knowledge to be split from the basic text knowledge set whose string length is greater than the specified difference of the specified length; then in the initial knowledge base, based on the rounded result L of the string length of each text knowledge to be split and the specified length, the corresponding text knowledge to be split is split into L lowest-level text knowledge to obtain a text knowledge base.

[0143] It should be noted that the text knowledge to be split is the basic text knowledge in the basic text knowledge set whose string length is greater than the specified length specified difference, and the specified difference is less than the specified length; and the string length is greater than the specified length specified difference, indicating that the string length of the basic text knowledge is much greater than the specified length and is to be split into multiple text knowledge. L is an integer greater than 1; the L lowest-level text knowledge is independent based on the specified punctuation marks, and the specified punctuation marks refer to punctuation marks that indicate that the text is a sentence, such as a period, an exclamation mark, a question mark, etc.; the rounding result L can be rounded up or rounded down, and this embodiment of the application does not limit this.

[0144] In step 1074 of the embodiment of the present application, the information processing device performs equalization processing on the basic text knowledge set in the initial knowledge base based on the string length and the specified length to obtain a text knowledge base, including: the information processing device first selects from the basic text knowledge set text knowledge belonging to the same parent level and with a string length less than a specified length and a specified difference, as well as adjacent text knowledge sets to be merged, to obtain at least one text knowledge set to be merged; then, in the initial knowledge base, multiple text knowledge to be merged in each text knowledge set to be merged are merged to obtain a text knowledge base.

[0145] It should be noted that if the string length is less than the specified difference of the specified length, it indicates that the string length of the basic text knowledge is much less than the specified length; adjacent means that they are adjacent in the order of appearance in the basic document.

[0146] In the embodiment of the present application, the splitting and merging can be performed in parallel, in series, or one of them can be performed alternatively, etc., and the embodiment of the present application does not limit this.

[0147] It is understandable that by balancing the string lengths of basic text knowledge, the uniformity of the string lengths of the lowest-level text knowledge in the text knowledge base is improved, thereby improving the accuracy of queries.

[0148] In an embodiment of the present application, the information processing device combines M text knowledge and simplified documents to describe the text processing request in natural language. Before obtaining the text processing prompt, the information processing method also includes: the information processing device first selects at least one text knowledge to be extracted corresponding to each text knowledge to be simplified from the text knowledge base; then performs summary extraction on the at least one text knowledge to be extracted to obtain a simplified text; then, from the text knowledge base, uses the simplified text to replace the at least one text knowledge to be extracted corresponding to each text knowledge to be simplified to obtain a simplified text base; finally, the simplified text base is converted into a document to obtain a simplified document.

[0149] It should be noted that the level where the text knowledge to be simplified is located is the parent level of the lowest level, and the level where the text knowledge to be extracted is located is the child level of the text knowledge to be simplified; thus, the text knowledge to be extracted belongs to the main content of the basic document, and at least one main content corresponding to at least one text knowledge to be extracted belongs to the main content under the same title. By performing summary extraction on at least one text knowledge to be extracted, the basic document can be accurately compressed and the accuracy of the simplified document can be improved.

[0150] Below, we will describe an exemplary application of the embodiments of this application in a practical application scenario. This exemplary application describes the process of text processing based on a game worldview document in a game scenario. It is readily apparent that the information processing method provided by the embodiments of this application can be applied to any document-based text processing scenario. This description will be based on the process of text processing based on a game worldview document in a game scenario.

[0151] See also Figure 6 , Figure 6 This is an exemplary text processing flow chart provided in the embodiment of the present application; Figure 6 As shown, the exemplary text processing flow includes steps 601 to 608, and each step is described below.

[0152] Step 601: parse the game world view document (called the basic document) to obtain the document parsing results (called the text knowledge corresponding to each document structure and the hierarchical relationship of each text knowledge in the basic document).

[0153] It should be noted that the document parsing results include title data, paragraph data, table data, table cell data and image data (referred to as various text knowledge), and also include the hierarchical relationship between title data, paragraph data, table data, table cell data and image data (also known as parent-child relationship). When the uploaded game world view document is received, the document parsing module automatically parses the title, paragraph, table and picture of the game world view document to obtain the document parsing results. Among them, the parsing of the picture can be achieved through optical character recognition (OCR) to obtain the text content in the picture, that is, the picture data; the parsing of the table can obtain the content of each cell based on the format of [table header, column header, cell], that is, the table cell data, and the table data includes the full amount of cell data in the table.

[0154] Step 602: Construct an N-ary tree (called a text knowledge base) based on the document parsing results. Then, execute steps 603 and 604.

[0155] It should be noted that the N-ary tree construction module constructs title data, paragraph data, table title data, table cell data and image data into an N-ary tree based on a hierarchical relationship; wherein the text content of all parent nodes is title data, and the text content of the leaf nodes is any one of paragraph data, table data, table cell data and image data.

[0156] As you can understand, the uploaded game worldview document is automatically parsed from top to bottom, from title to body to tables, to construct an ordered and structured N-ary tree. Each node in the N-ary tree corresponds to a text fragment in the game worldview document, and the relationship between parent and child nodes also corresponds to the structural hierarchy of the game worldview document.

[0157] Step 603: Compress the N-ary tree to obtain a simplified N-ary tree, and then execute step 607.

[0158] It should be noted that the N-ary tree compression module realizes compression of the N-ary tree by extracting a summary of a specified number of characters (for example, 20) from the text content of all leaf nodes of the same parent node.

[0159] It can be understood that, through summary extraction, the number of leaf nodes and the text length of the leaf nodes are reduced, the N-ary tree is pruned, and a simplified version of the document content can be obtained.

[0160] It should also be noted that the N-ary tree, also known as the search tree, is an extension of the binary search tree, in which each node includes at most N child nodes; N can be a fixed value or can be flexibly changed according to actual conditions.

[0161] Step 604: Perform vector representation on the N-ary tree to obtain a vector knowledge base (VKB, also called text feature base) of the N-ary tree.

[0162] It should be noted that each node of the N-ary tree is traversed in a depth-first traversal manner, and the text content vector (called text feature) of each traversed node is obtained through a vector model. Among them, the vector model is used to map discrete categorical variables into a continuous vector space to obtain the semantic relationship between categories; the vector knowledge base is a method for representing and storing knowledge. In the vector knowledge base, entities and relationships are mapped into a high-dimensional vector space. In this way, vector similarity calculations can be used to represent the similarity and semantic distance between different entities; and the vector knowledge base is used to represent and store textual knowledge, storing various structural contents in the text and the relationships between various structural contents in the form of encoded results. In addition, the vector knowledge base can be implemented through a database (such as Faiss) for fast similarity search and density clustering in large-scale data sets.

[0163] Step 605: In response to the query request, select M relevant leaf nodes from the vector knowledge base.

[0164] It should be noted that when searching for content from the game world view document, the search module also receives the query request; in response to the query request, the search engine is called to search the vector knowledge base to search for T leaf nodes (called T to-be-recalled texts) whose similarity is greater than the similarity threshold; then the corresponding parent node is determined based on each of the T leaf nodes, and then horizontal and vertical expansion recall is performed based on the determined parent node to recall leaf nodes whose similarity is greater than the recall similarity threshold, and then M leaf nodes are obtained by combining the T leaf nodes.

[0165] Step 606: Obtain M leaf nodes (called M text knowledge) corresponding to the M leaf nodes in the vector knowledge base from the N-ary tree.

[0166] Step 607: Assemble the texts corresponding to the M leaf nodes in the N-ary tree and the text corresponding to the simplified N-ary tree to obtain prompt words (called text processing prompts).

[0167] It should be noted that the prompt word assembly module is used to assemble prompt words from the text corresponding to the M leaf nodes in the N-ary tree and the text corresponding to the simplified N-ary tree. Different application scenarios use different prompt word templates; for example, in a game quiz scenario, the prompt word template is a question-and-answer prompt word template, while in a game creation scenario, the prompt word template is a writing prompt word template.

[0168] Step 608: Obtain the answer content corresponding to the prompt word (called text processing result) through the large language model.

[0169] It should be noted that the vector model and the large language model can be integrated into an open source framework (for example, the Langchain open source framework) to interact with the language model.

[0170] It should also be noted that the hardware environment of the embodiment of the present application includes a CPU environment and a graphics processing unit (GPU) environment; among them, text queries are performed in the GPU environment, and the large language model for obtaining the answer content also runs in the GPU environment.

[0171] It can be understood that by constructing an N-ary tree based on the structure of the document, the accuracy of document understanding is improved, thereby improving the accuracy of document retrieval.

[0172] For example, see Figure 7 , Figure 7 is a schematic diagram of an exemplary method of obtaining answer content provided in an embodiment of the present application; Figure 7 As shown, the interface 7-1 describes a schematic diagram of a game knowledge question and answer in a game scene. By inputting a query request 7-11, relevant text is obtained from the N-ary tree, and then the large language model is used to process the obtained text and the text corresponding to the simplified N-ary tree to obtain the answer content 7-12.

[0173] See also Figure 8 , Figure 8 This is another exemplary schematic diagram for obtaining answer content provided in an embodiment of the present application; the interface 8-1 describes a game creation schematic diagram in a game scene, and through the input query request 8-11, relevant text is obtained from the N-ary tree, and then the large language model is used to process the obtained text and the text corresponding to the simplified N-ary tree, and the answer content 8-12 is obtained.

[0174] The following is based on Figure 6 Describes the process of processing game world view documents.

[0175] See also Figure 9 , Figure 9 This is an exemplary document processing diagram provided by the embodiment of the present application; Figure 9As shown, the game world view document 9-1 is converted into an N-ary tree 9-2 (referred to as the initial knowledge base), wherein the numbers corresponding to the nodes in the N-ary tree 9-2 represent the number of words in the corresponding text content. Thus, the number of words in the leaf nodes represents the number of words in the text content corresponding to at least one of the paragraphs, tables, table cells, and pictures. For example, if a paragraph has 400 words, then the number of the leaf node formed by the paragraph is 400. Since the difference in the number of words of each leaf node in the N-ary tree 9-2 is greater than the specified difference, the word count is unevenly distributed. For example, a paragraph has hundreds of words, while a table cell has dozens of words. In order to reduce the impact of the number of text words on the accuracy of the vector representation and improve the precision and accuracy of the text similarity calculation, the leaf nodes of the N-ary tree 9-2 are merged and split to adjust the text length of the leaf nodes of the N-ary tree to a stable numerical range.

[0176] When merging and splitting, based on the ideal word count, adjacent leaf nodes whose word count is less than the lower word count threshold are merged, and leaf nodes whose word count is greater than the upper word count threshold are split.

[0177] Continue to see Figure 9 In the N-ary tree 9-2, node 9-21 and node 9-22 (called multiple text knowledge to be merged) are adjacent, and the number of words in node 9-21 is 40, the number of words in node 9-22 is 80, and the ideal number of words is 150, so node 9-21 and node 9-22 are merged to obtain node 9-31 in the N-ary tree 9-3; wherein, the N-ary tree 9-3 is obtained by merging the leaf nodes of the N-ary tree 9-2.

[0178] Continue to see Figure 9 In the N-ary tree 9-3, the number of words in node 9-32 (called the text knowledge to be split) is 400, so based on the ideal number of words of 150, node 9-32 is split into three independent leaf nodes, and nodes 9-41 to 9-43 in the N-ary tree 9-4 are obtained, with the number of words being 120, 130 and 150 respectively; among them, the N-ary tree 9-4 is obtained by splitting the leaf nodes of the N-ary tree 9-3.

[0179] It should be noted that the leaf nodes under the same parent node are merged, and the split leaf nodes are independent based on the punctuation marks (for example, periods) that represent single sentences. In this way, the accuracy of text similarity calculation can be improved while maintaining the original structure of the document, thereby improving the effect of knowledge retrieval.

[0180] Continue to see Figure 9, each node of N-ary tree 9-4 is represented by a vector, resulting in a vector knowledge base 9-5. N-ary tree 9-4 is then pruned to obtain a simplified N-ary tree 9-6. Pruning refers to reducing the number of leaf nodes in an N-ary tree. First, the parent node is kept fixed, and all leaf nodes under the same parent node are compressed and summarized: approximately 20 words are used to summarize the text content of all leaf nodes under the same parent node. Here, by converting simplified N-ary tree 9-6 into a document, a simplified document 9-7 of game worldview document 9-1 is obtained.

[0181] It is understandable that the simplified N-ary tree obtained by pruning can reflect the core content of the document in a general way while retaining the key information of the document, which not only improves the quality of the document summary, but also improves the efficiency and accuracy of document retrieval.

[0182] based on Figure 6 , see Figure 10 , Figure 10 This is an exemplary information query diagram provided by the embodiment of the present application; Figure 10 As shown, the root node of the N-ary tree 10-1 includes 15 nodes (nodes 10-11 to 10-15). When the leaf nodes determined to be similar to the query text are expanded horizontally, if the expansion coefficient is 1, the leaf node to be expanded horizontally is node 10-112 (called the text knowledge to be recalled), then the adjacent node 10-111 can be recalled through horizontal expansion. When the leaf nodes determined to be similar to the query text are leaf nodes 10-15 and leaf node 10-17, if the coverage rate of the recalled leaf nodes under the same parent node is greater than the coverage rate threshold (for example, 1 / 2), all leaf nodes under the parent node can be recalled. If leaf nodes 10-14 and 10-115 have an association relationship—for example, if role A in leaf node 10-14 and role B in leaf node 10-115 are parent-child, or if a function in leaf node 10-14 participates in a class definition in leaf node 10-115—then leaf node 10-115 will be retrieved through a cross-node search. This fully utilizes the N-ary tree structure and effectively retrieves knowledge during vector search.

[0183] It is understood that the embodiments of the present application analyze documents based on their structure to obtain the text content and hierarchical relationships corresponding to each document structure, and then construct the document's text content into an N-ary tree based on this hierarchical relationship. This preserves the document's structural information, thereby improving the accuracy and efficiency of document comprehension and the accuracy of knowledge retrieval. In addition, by pruning the N-ary tree, a condensed version of the document is accurately obtained. This, when combined with the condensed version of the document and the knowledge retrieval results to obtain the answer content for the query request, can improve the accuracy of the answer content.

[0184] The following continues to describe the exemplary structure of the information processing device 255 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the information processing device 255 of the memory 250 may include:

[0185] Request response module 2551, for responding to a text processing request and obtaining a text to be queried, wherein the text to be queried is used to perform information query in a basic document, which is the information basis for the text processing request;

[0186] A text query module 2552 is configured to query a text knowledge base for M text knowledge matching the query text, wherein the text knowledge base includes text knowledge corresponding to each document structure in the basic document and the hierarchical relationship between each text knowledge in the basic document, wherein the document structure is at least one of the following: a title, a paragraph, a table, a table cell, and an illustration, and M is a positive integer;

[0187] A request description module 2553 is configured to describe the text processing request in natural language based on the M text knowledge to obtain a text processing prompt;

[0188] The text processing module 2554 is used to perform text processing based on the text processing prompt to obtain a text processing result.

[0189] In an embodiment of the present application, the information processing device 255 also includes a knowledge construction module 2555, which is used to obtain the basic document in response to a document analysis request; perform structural analysis on the basic document to obtain the text knowledge and text hierarchical relationship of each document structure, and the text hierarchical relationship represents the hierarchical relationship of each text knowledge in the basic document; and combine each text knowledge based on the text hierarchical relationship to obtain the text knowledge base.

[0190] In an embodiment of the present application, the knowledge construction module 2555 is also used to combine each of the text knowledge based on the text hierarchical relationship to obtain an initial knowledge base; combine the basic text knowledge of the lowest level in the initial knowledge base into a basic text knowledge set; count the character string length corresponding to each of the basic text knowledge in the basic text knowledge set; in the initial knowledge base, balance the basic text knowledge set based on the character string length and the specified length to obtain the text knowledge base, and the balancing process is the splitting of the basic text knowledge or the merging of multiple basic text knowledge.

[0191] In an embodiment of the present application, the knowledge construction module 2555 is also used to select at least one text knowledge to be split from the basic text knowledge set, whose string length is greater than the specified difference of the specified length; in the initial knowledge base, based on the rounded result L of the string length of each text knowledge to be split and the specified length, the corresponding text knowledge to be split is split into L text knowledge of the lowest level to obtain the text knowledge base, where L is an integer greater than 1.

[0192] In an embodiment of the present application, the L lowest-level text knowledge is independent based on designated punctuation marks, where the designated punctuation marks refer to punctuation marks representing text sentences.

[0193] In an embodiment of the present application, the knowledge construction module 2555 is also used to select from the basic text knowledge set, text knowledge belonging to the same parent level, whose string length is less than the specified length and specified difference, and adjacent text knowledge sets to be merged, to obtain at least one text knowledge set to be merged; in the initial knowledge base, multiple text knowledge to be merged in each text knowledge set to be merged are merged to obtain the text knowledge base.

[0194] In an embodiment of the present application, the request description module 2553 is also used to use M of the text knowledge to describe the text processing request in natural language to obtain the text processing prompt; or, combine M of the text knowledge and a simplified document to describe the text processing request in natural language to obtain the text processing prompt, and the simplified document is the simplified result of the basic document, and the simplified document is obtained by summarizing the text knowledge in the text knowledge base.

[0195] In an embodiment of the present application, the knowledge construction module 2555 is also used to select at least one text knowledge to be extracted corresponding to each text knowledge to be simplified from the text knowledge base, the level where the text knowledge to be simplified is located is the parent level of the lowest level, and the level where the text knowledge to be extracted is located is the child level of the text knowledge to be simplified; perform summary extraction on at least one of the text knowledge to be extracted to obtain a simplified text; use the simplified text to replace at least one text knowledge to be extracted corresponding to each text knowledge to be simplified from the text knowledge base to obtain a simplified text library; convert the simplified text library into a document to obtain the simplified document.

[0196] In an embodiment of the present application, the text query module 2552 is also used to extract features of the text to be queried to obtain features to be queried; determine T text features whose feature similarity with the features to be queried is greater than a similarity threshold from the text feature library, the text feature library includes the text features of each text knowledge in the text knowledge base, and T is a positive integer; obtain T text knowledge to be recalled corresponding to the T text features from the text knowledge base; based on the T text knowledge to be recalled, determine M text knowledge that matches the text to be queried.

[0197] In an embodiment of the present application, the text query module 2552 is also used to perform the following processing on each of the T text knowledge to be recalled: from the text knowledge base, the text knowledge to be recalled is expanded and recalled to obtain a text knowledge set, and the expansion recall is implemented based on at least one of the following: expansion coefficient, lowest level text knowledge coverage, and the association relationship between each of the text knowledge; from the text knowledge set of each of the text knowledge to be recalled, T text knowledge sets corresponding to the T text knowledge to be recalled are obtained; at least one of the text knowledge included in the T text knowledge sets is determined as M text knowledge.

[0198] In an embodiment of the present application, the text query module 2552 is also used to obtain the target text knowledge corresponding to the parent level of the text knowledge to be recalled from the text knowledge base; obtain the text knowledge sequence corresponding to the child level of the target text knowledge; from the text knowledge sequence, expand and recall S pieces of text knowledge adjacent to the text knowledge to be recalled, where S is a positive integer and S is the expansion coefficient; and obtain the text knowledge set based on the text knowledge to be recalled and the S pieces of text knowledge.

[0199] In an embodiment of the present application, the text query module 2552 is also used to obtain the number of common text knowledge of T of the text knowledge to be recalled and the text knowledge sequence; obtain the number of sequence text knowledge of the text knowledge sequence; determine the ratio of the number of common text knowledge to the number of sequence text knowledge as the text knowledge coverage; when the text knowledge coverage is greater than the coverage threshold, determine the text knowledge set based on the text knowledge sequence.

[0200] In an embodiment of the present application, the text query module 2552 is also used to obtain target association information associated with the text knowledge to be recalled from an association relationship library, where the association relationship library is information described in the auxiliary documents of the basic document; obtain at least one associated text knowledge corresponding to the target association information from the text knowledge library; and determine the text knowledge set based on the text knowledge to be recalled and at least one associated text knowledge.

[0201] In an embodiment of the present application, the text processing based on the text processing prompt is implemented through a text processing model, and the information processing device 255 also includes a model training module 2556, which is used to obtain text processing training data, and the text processing training data includes text processing sample prompts and result sample labels; the text processing sample prompts are subjected to text processing based on the model to be trained to obtain a text processing estimation result, and the model to be trained refers to a neural network model to be trained for text processing; based on the difference between the text processing estimation result and the result sample label, the model to be trained is trained to obtain the text processing model.

[0202] In an embodiment of the present application, when the text to be queried is a question of the basic document, the text processing prompt is a question to be answered with M pieces of text knowledge as background knowledge; the text processing module 2554 is also used to perform text processing based on the question to be answered to obtain a target answer to the question to be answered; and the target answer is determined as the text processing result.

[0203] In an embodiment of the present application, when the text to be queried is an imitation description based on the basic document, the text processing prompt is the description to be imitated using M of the text knowledge as an imitation method; the text processing module 2554 is also used to perform text processing based on the description to be imitated to obtain a target imitation result; and the target imitation result is determined as the text processing result.

[0204] The present invention provides a computer program product comprising computer-executable instructions or a computer program stored in a computer-readable storage medium. A processor of an information processing device reads the computer-executable instructions or the computer program from the computer-readable storage medium and executes the computer-executable instructions or the computer program, causing the information processing device to perform the information processing method described in the present invention.

[0205] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the information processing method provided by the embodiment of the present application, for example, Figure 3 The information processing method shown.

[0206] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0207] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0208] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0209] As an example, computer-executable instructions may be deployed to be executed on one electronic device (in which case, this one electronic device is an information processing device), or on multiple electronic devices located at one location (in which case, the multiple electronic devices located at one location are information processing devices), or on multiple electronic devices distributed at multiple locations and interconnected through a communication network (in which case, the multiple electronic devices distributed at multiple locations and interconnected through a communication network are information processing devices).

[0210] It is understandable that in the embodiments of the present application, when data related to documents and the like is involved, when the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. In the present application, when the data capture technical solution involved in model training is implemented, when the above embodiments of the present application are applied to specific products or technologies, the relevant data collection, use and processing process should comply with the requirements of national laws and regulations, comply with the principles of legality, legitimacy and necessity, do not involve obtaining data types prohibited or restricted by laws and regulations, and will not hinder the normal operation of the target website.

[0211] In summary, in the text processing process based on document query in the embodiment of the present application, since the text knowledge base of the basic document on which the text processing request is based is constructed through the text knowledge corresponding to at least one document structure in the title, paragraph, table, table cell and accompanying drawing, the text knowledge base includes the structural information of the basic document, thereby improving the comprehensiveness of the M text knowledge queried from the text knowledge base; and further, when performing text processing based on the M text knowledge, the accuracy of text processing can be improved. In addition, the embodiment of the present application obtains a simplified document of the basic document by performing summary extraction on the text knowledge base, and then performs text processing in combination with the M text knowledge queried and the simplified document, which can improve the accuracy of text processing. In addition, in the process of constructing the text knowledge base, by merging and splitting the basic text knowledge, the uniformity of the string length of the text knowledge is improved, thereby improving the query accuracy of the M text knowledge.

[0212] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. An information processing method, characterized in that: The method comprises: In response to the text processing request, obtaining a text to be queried, wherein the text to be queried is used to perform information query in a basic document, and the basic document is the information basis of the text processing request; Querying M text knowledge matching the query text from a text knowledge base, wherein the text knowledge base includes the text knowledge corresponding to each document structure in the basic document and the hierarchical relationship of each text knowledge in the basic document, wherein the document structure is at least one of the following: a title, a paragraph, a table, a table cell, and an illustration, and M is a positive integer; Describing the text processing request in natural language based on the M text knowledge to obtain a text processing prompt; Text processing is performed based on the text processing prompt to obtain a text processing result.

2. The method according to claim 1, characterized in that Before searching the text knowledge base for M text knowledge matching the text to be queried, the method further includes: In response to the document analysis request, obtaining the basic document; Performing structural analysis on the basic document to obtain the textual knowledge and textual hierarchical relationships of each document structure, wherein the textual hierarchical relationships represent hierarchical relationships between the various textual knowledge in the basic document; The text knowledge base is obtained by combining the various text knowledge based on the text hierarchical relationship.

3. The method according to claim 2, characterized in that The combining of the individual text knowledge based on the text hierarchical relationship to obtain the text knowledge base includes: Combining the text knowledge based on the text hierarchical relationship to obtain an initial knowledge base; Combining the lowest level of basic text knowledge in the initial knowledge base into a basic text knowledge set; Counting the length of the character string corresponding to each basic text knowledge in the basic text knowledge set; In the initial knowledge base, the basic text knowledge set is balanced based on the character string length and the specified length to obtain the text knowledge base, and the balanced processing is the splitting of the basic text knowledge or the merging of multiple basic text knowledge.

4. The method according to claim 3, characterized in that In the initial knowledge base, the basic text knowledge set is balanced based on the character string length and the specified length to obtain the text knowledge base, including: Selecting at least one text knowledge to be split from the basic text knowledge set, wherein the character string length is greater than the specified length and specified difference; In the initial knowledge base, based on the rounded result L of the string length of each text knowledge to be split and the specified length, the corresponding text knowledge to be split is split into L text knowledge of the lowest level to obtain the text knowledge base, where L is an integer greater than 1.

5. The method according to claim 4, characterized in that The L lowest-level text knowledge are independent based on designated punctuation marks, where the designated punctuation marks refer to punctuation marks representing text sentences.

6. The method according to claim 3, characterized in that In the initial knowledge base, the basic text knowledge set is balanced based on the character string length and the specified length to obtain the text knowledge base, including: Selecting, from the basic text knowledge set, text knowledge sets to be merged that belong to the same parent level and whose character string lengths are less than the specified length and specified difference, and are adjacent to each other, to obtain at least one text knowledge set to be merged; In the initial knowledge base, multiple text knowledge to be merged in each of the text knowledge sets to be merged are merged to obtain the text knowledge base.

7. The method according to any one of claims 1 to 6, characterized in that The step of describing the text processing request in natural language based on the M pieces of text knowledge to obtain a text processing prompt includes: Using the M pieces of text knowledge to describe the text processing request in natural language to obtain the text processing prompt; Alternatively, the text processing request is described in natural language in combination with the M text knowledge and a simplified document to obtain the text processing prompt, the simplified document is the simplified result of the basic document, and the simplified document is obtained by abstracting the text knowledge in the text knowledge base.

8. The method according to claim 7, characterized in that Before describing the text processing request in natural language by combining the M pieces of text knowledge and the simplified document to obtain the text processing prompt, the method further includes: Selecting at least one to-be-extracted text knowledge corresponding to each to-be-simplified text knowledge from the text knowledge base, wherein the level where the to-be-simplified text knowledge is located is the parent level of the lowest level, and the level where the to-be-extracted text knowledge is located is the child level of the to-be-simplified text knowledge; Extracting a summary of at least one of the textual knowledge to be extracted to obtain a condensed text; From the text knowledge base, replacing at least one of the to-be-extracted text knowledge corresponding to each of the to-be-simplified text knowledge with the simplified text to obtain a simplified text base; The simplified text library is converted into a document to obtain the simplified document.

9. The method according to any one of claims 1 to 6, characterized in that The step of searching a text knowledge base for M text knowledge matching the text to be queried includes: Extracting features from the text to be queried to obtain features to be queried; Determine T text features whose feature similarity with the feature to be queried is greater than a similarity threshold from a text feature library, wherein the text feature library includes the text features of each text knowledge in the text knowledge library, and T is a positive integer; Obtaining T pieces of to-be-recalled text knowledge corresponding to the T text features from the text knowledge base; Based on the T pieces of text knowledge to be recalled, M pieces of text knowledge matching the text to be queried are determined.

10. The method according to claim 9, characterized in that The determining, based on the T pieces of text knowledge to be recalled, M pieces of text knowledge matching the text to be queried includes: The following processing is performed on each of the T pieces of text knowledge to be recalled: Performing expansion recall on the to-be-recalled text knowledge from the text knowledge base to obtain a text knowledge set, wherein the expansion recall is implemented based on at least one of the following: an expansion coefficient, a coverage rate of text knowledge at the lowest level, and an association relationship between each of the text knowledge; Obtaining T text knowledge sets corresponding to the T pieces of text knowledge to be recalled from the text knowledge set of each piece of text knowledge to be recalled; At least one of the text knowledge included in the T text knowledge sets is determined as M text knowledge.

11. The method according to claim 10, characterized in that The method of performing expansion recall on the text knowledge to be recalled from the text knowledge base to obtain a text knowledge set includes: Obtaining target text knowledge corresponding to the parent level of the text knowledge to be recalled from the text knowledge base; Obtaining a text knowledge sequence corresponding to a sub-level of the target text knowledge; Expand and recall S pieces of text knowledge adjacent to the text knowledge to be recalled from the text knowledge sequence, where S is a positive integer and S is the expansion coefficient; The text knowledge set is obtained based on the to-be-recalled text knowledge and the S pieces of text knowledge.

12. The method according to claim 11, characterized in that After obtaining the text knowledge sequence corresponding to the sub-level of the target text knowledge, the method further includes: Obtaining the number of common text knowledge between the T pieces of text knowledge to be recalled and the text knowledge sequence; Obtaining the amount of sequence text knowledge of the text knowledge sequence; Determine the ratio of the amount of common text knowledge to the amount of sequence text knowledge as the text knowledge coverage; When the text knowledge coverage is greater than a coverage threshold, the text knowledge set is determined based on the text knowledge sequence.

13. The method according to claim 10, characterized in that The method of performing expansion recall on the text knowledge to be recalled from the text knowledge base to obtain a text knowledge set includes: Obtaining target association information associated with the text knowledge to be recalled from an association relationship library, wherein the association relationship library is information described in the auxiliary document of the basic document; Acquiring at least one associated text knowledge corresponding to the target associated information from the text knowledge base; The text knowledge set is determined based on the to-be-recalled text knowledge and the at least one associated text knowledge.

14. The method according to any one of claims 1 to 6, characterized in that The text processing based on the text processing prompt is implemented by a text processing model, and the text processing model is trained by the following steps: Acquire text processing training data, wherein the text processing training data includes text processing sample prompts and result sample labels; Performing text processing on the text processing sample prompt based on a to-be-trained model to obtain a text processing estimation result, wherein the to-be-trained model refers to a to-be-trained neural network model for performing text processing; Based on the difference between the text processing prediction result and the result sample label, the model to be trained is trained to obtain the text processing model.

15. The method according to any one of claims 1 to 6, characterized in that When the text to be queried is a question of the basic document, the text processing prompt is a question to be answered with M pieces of text knowledge as background knowledge; The performing text processing based on the text processing prompt to obtain a text processing result includes: Performing text processing based on the question to be answered to obtain a target answer to the question to be answered; The target answer is determined as the text processing result.

16. The method according to any one of claims 1 to 6, characterized in that When the text to be queried is an imitation description based on the basic document, the text processing prompt is a description to be imitated using M pieces of text knowledge as an imitation method; The performing text processing based on the text processing prompt to obtain a text processing result includes: Performing text processing based on the description to be imitated to obtain a target imitation result; The target imitation result is determined as the text processing result.

17. An information processing device, characterized in that: The information processing device includes: A request response module, configured to respond to a text processing request and obtain a text to be queried, wherein the text to be queried is used to perform information query in a basic document, and the basic document is the information basis for the text processing request; a text query module, configured to query a text knowledge base for M text knowledge matching the text to be queried, wherein the text knowledge base includes text knowledge corresponding to each document structure in the basic document and a hierarchical relationship between each text knowledge in the basic document, wherein the document structure is at least one of the following: a title, a paragraph, a table, a table cell, and an illustration, and M is a positive integer; A request description module, configured to describe the text processing request in natural language based on the M text knowledge to obtain a text processing prompt; The text processing module is used to perform text processing based on the text processing prompt to obtain a text processing result.

18. An electronic device for information processing, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; A processor, configured to implement the information processing method according to any one of claims 1 to 16 when executing the computer-executable instructions or computer programs stored in the memory.

19. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer-executable instructions or computer programs are executed by a processor, the information processing method according to any one of claims 1 to 16 is implemented.

20. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer-executable instructions or computer programs are executed by a processor, the information processing method according to any one of claims 1 to 16 is implemented.