System document processing method and device, storage medium and program product
By splitting according to the document structure of the system document and storing the split results and shared attributes in the knowledge base, the semantic distortion problem caused by fixed-length splitting is solved, and the effect of improving semantic integrity and answer accuracy is achieved.
Patent Information
- Application Number
- CN202510185015.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, the system documents are split according to fixed lengths, resulting in semantic distortion of the split result, affecting the large model's understanding of the system and the accuracy of users' answers to questions.
Split it according to the document structure of the system document, and store the split result and the shared attributes of the system document in the knowledge base to provide users with response services to answer questions.
It improves the semantic integrity of the results of institutional document splitting, avoids the loss of semantic information caused by fixed term length splitting, and enhances the ability of the big model to understand the system and the accuracy of users to answer questions.
Smart Images

Figure CN120045678A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and more specifically, to a method, device, storage medium and program product for processing institutional documents. Background Art
[0002] In the system question and answer system, after the rules and regulations documents are broken down into knowledge points, as the underlying knowledge points, the semantic retrieval capability of vector search or the large model is used to search for knowledge points, making it easier for business personnel to find the rules and regulations needed to handle the business, or the audit points for business processing.
[0003] In the related technology, when the rules and regulations document is disassembled, the document is disassembled according to a fixed number of terms, but this method is not suitable for questions and answers about rules and regulations in the vertical field. Specifically, when the system document is divided according to a fixed number of terms, the system items near the division point will lose semantic coherence, resulting in distortion of the semantic information of the system, which in turn affects the difficulty of the large model in understanding the system. Splitting knowledge points according to a fixed number of terms does not take into account the relationship between systems, and the relationship between systems affects the effectiveness of the business stipulated by the system items, resulting in low accuracy of the generated response answers.
[0004] Currently, no effective solution has been proposed to the problem that in related technologies, policy documents are split according to fixed lengths, resulting in semantic distortion of the split results and inaccurate answers to user questions. Summary of the invention
[0005] The main purpose of the present application is to provide a method, device, storage medium and program product for processing system documents, so as to solve the problem in the related art that system documents are split according to fixed lengths, resulting in semantic distortion of the splitting results and inaccurate answers to user questions.
[0006] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a method for processing system documents is provided. The method comprises: extracting target attributes of a target document, wherein the target document comprises: a system document based on which business is processed in a financial institution, and the target attributes comprise: attributes possessed by all paragraphs in the target document; splitting the target document based on the document structure of the target document to obtain a target splitting result, wherein the target splitting result comprises at least: M entries in the target document, where M is a positive integer; storing the target attributes and the target splitting result in a target knowledge base, wherein the target knowledge base is used to provide users with a service of replying to questions to be answered.
[0007] Furthermore, based on the document structure of the target document, the target document is split to obtain a target splitting result, including: obtaining preset text, wherein the preset text is used to identify the document structure of the target document; identifying paragraphs in the target document containing the preset text to obtain a recognition result; and splitting the target document based on the recognition result to obtain the target splitting result.
[0008] Furthermore, the target document is split based on the recognition result to obtain the target splitting result, including: based on the recognition result, the target document is split with entries as the splitting granularity to obtain an initial splitting result, wherein the initial splitting result includes: M entries in the target document; the M entries are graded according to preset grading rules to obtain a graded result, wherein the preset grading rules include: chapter, section, entry, the level of the chapter is higher than the level of the section, and the level of the section is higher than the level of the entry; based on the initial splitting result and the graded result, the target splitting result is determined.
[0009] Furthermore, after grading the M entries according to preset grading rules and obtaining the grading results, it also includes: determining an analysis strategy, wherein the analysis strategy includes at least one of the following: tracing analysis, cluster analysis, and federated learning; based on the analysis strategy, analyzing the relationship between the M entries and other entries in the target knowledge base except the M entries to obtain an analysis result.
[0010] Furthermore, extracting the target attribute of the target document includes: obtaining a preset prompt word, wherein the preset prompt word is used to control a first language model to extract the attribute of the target document, and the first language model includes: a pre-trained deep learning model; inputting the preset prompt word and the target document into the first language model, and outputting the target attribute.
[0011] Furthermore, after storing the target attributes and the target splitting results in the target knowledge base, it also includes: receiving the question content of the question raised by the target object, preprocessing the question content, and obtaining processed question content, wherein the preprocessing processing method includes at least one of the following: intent analysis, keyword extraction; based on the question content and the processed question content, extracting data in the target knowledge base to obtain knowledge data, wherein the knowledge data includes: entries or attributes in the target document related to the question content; based on the knowledge data and the processed question content, generating a reply answer.
[0012] Furthermore, based on the knowledge data and the processed question content, a reply answer is generated, including: obtaining preset requirement information, wherein the preset requirement information includes: requirements for the reply answer to be generated; based on the preset requirement information and the knowledge data, a target prompt word is generated, and the target prompt word and the processed question content are input into a second language model to generate the reply answer, wherein the target prompt word is used to control the second language model to generate the reply answer, and the second language model includes: a pre-trained deep learning model.
[0013] In order to achieve the above-mentioned purpose, according to another aspect of the present application, a device for processing system documents is provided. The device comprises: an extraction unit, used to extract target attributes of a target document, wherein the target document comprises: a system document based on which business is processed in a financial institution, and the target attributes comprise: attributes possessed by all paragraphs in the target document; a splitting unit, used to split the target document based on the document structure of the target document to obtain a target splitting result, wherein the target splitting result comprises at least: M entries in the target document, where M is a positive integer; a storage unit, used to store the target attributes and the target splitting result in a target knowledge base, wherein the target knowledge base is used to provide users with a service of replying to unanswered questions.
[0014] Furthermore, the splitting unit includes: a first acquisition subunit, used to acquire preset text, wherein the preset text is used to identify the document structure of the target document; an identification subunit, used to identify paragraphs in the target document containing the preset text to obtain an identification result; and a splitting subunit, used to split the target document based on the identification result to obtain the target splitting result.
[0015] Furthermore, the splitting subunit includes: a splitting module, used to split the target document based on the recognition result and with entries as the splitting granularity, to obtain an initial splitting result, wherein the initial splitting result includes: M entries in the target document; a grading module, used to grade the M entries according to preset grading rules to obtain a grading result, wherein the preset grading rules include: chapter, section, entry, the level of the chapter is higher than the level of the section, and the level of the section is higher than the level of the entry; a first determination module, used to determine the target splitting result based on the initial splitting result and the grading result.
[0016] Furthermore, the splitting sub-unit also includes: a second determination module, used to determine the analysis strategy after grading the M entries according to preset grading rules and obtaining the grading results, wherein the analysis strategy includes at least one of the following: tracing analysis, cluster analysis, and federated learning; an analysis module, used to analyze the relationship between the M entries and other entries in the target knowledge base except the M entries based on the analysis strategy to obtain the analysis results.
[0017] Furthermore, the extraction unit includes: a second acquisition subunit, used to acquire preset prompt words, wherein the preset prompt words are used to control the first language model to extract the attributes of the target document, and the first language model includes: a pre-trained deep learning model; a first processing subunit, used to input the preset prompt words and the target document into the first language model, and output the target attributes.
[0018] Furthermore, the processing device for system documents also includes: a processing unit, which is used to receive the question content of the question raised by the target object after storing the target attributes and the target splitting results in the target knowledge base, and pre-process the question content to obtain processed question content, wherein the pre-processing processing method includes at least one of the following: intention analysis, keyword extraction; an extraction unit, which is used to extract data in the target knowledge base based on the question content and the processed question content to obtain knowledge data, wherein the knowledge data includes: entries or attributes in the target document related to the question content; a generation unit, which is used to generate a reply answer based on the knowledge data and the processed question content.
[0019] Furthermore, the generation unit includes: a second acquisition subunit, used to acquire preset demand information, wherein the preset demand information includes: requirements for the reply answer to be generated; a second processing subunit, used to generate a target prompt word based on the preset demand information and the knowledge data, and input the target prompt word and the processed question content into a second language model to generate the reply answer, wherein the target prompt word is used to control the second language model to generate the reply answer, and the second language model includes: a pre-trained deep learning model.
[0020] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the method for processing the system document.
[0021] According to another aspect of the present application, an electronic device is provided, comprising: a memory storing an executable program; and a processor for running the program, wherein the method for processing the system document is executed when the program is running.
[0022] According to another aspect of the present application, a computer program product is provided, comprising computer instructions, which implement the steps of the method for processing institutional documents when executed by a processor.
[0023] In an embodiment of the present application, by extracting the target attributes of the target document, wherein the target document includes: the system document based on which the business is processed in the financial institution, the target attributes include: the attributes possessed by all paragraphs in the target document; based on the document structure of the target document, the target document is split to obtain a target splitting result, wherein the target splitting result includes at least: M entries in the target document, M is a positive integer; the target attributes and the target splitting results are stored in a target knowledge base, wherein the target knowledge base is used to provide users with services for replying to unanswered questions, thereby solving the technical problem in the related art that the system document is split according to a fixed length, resulting in semantic distortion of the splitting results, resulting in inaccurate answers to questions for users. In the present invention, the system document is split according to the document structure of the system document, and the splitting results and the common attributes (target attributes) of the system document are stored in the knowledge base, so as to provide users with services for replying to unanswered questions, thereby avoiding the situation in the related art that the system document is split according to a fixed term length, resulting in the loss of semantic information in the splitting results, thereby achieving the technical effect of improving the semantic integrity of the system document splitting results. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0025] Figure 1 A hardware structure block diagram of a computer terminal for implementing a method for processing system documents is shown;
[0026] Figure 2 is a flowchart of a method for processing system documents provided in an embodiment of the present application;
[0027] Figure 3 is a schematic diagram of a RAG framework provided according to an embodiment of the present application;
[0028] Figure 4 is a schematic diagram of a system document processing device provided according to an embodiment of the present application;
[0029] Figure 5It is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0032] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:
[0033] Large Language Model (LLM): Large Language Model, referred to as LLM, refers to a deep learning model with a huge number of parameters, such as the BERT model, which usually has hundreds of millions of parameters and is trained on large-scale data, so that it can better fit complex data distributions and better capture long-distance dependencies in language, thereby improving the ability to understand and generate language. Large models have a wide range of applications, including but not limited to natural language processing, computer vision, speech recognition and other fields.
[0034] It should be noted that the collected information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data are in compliance with relevant laws, regulations and standards, necessary confidentiality measures are taken, and public order and good customs are not violated, and corresponding operation entrances are provided for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions to provide users with corresponding operation entrances for users to choose to agree or refuse the results of automated decision-making; if the user chooses to refuse, the expert decision-making process will be entered.
[0035] Embodiment 1
[0036] According to an embodiment of the present application, a method embodiment of a method for processing system documents is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0037] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for processing system documents is shown. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0038] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0039] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for processing system documents in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, the method for processing system documents described above is realized. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0040] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0041] The display may be, for example, a touch screen liquid crystal display (LCD), which may enable a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0042] Under the above operating environment, this application provides Figure 2 The processing method of the institutional documents shown. Figure 2 It is a flowchart of a method for processing system documents according to the first embodiment of the present application.
[0043] Step S201, extracting target attributes of a target document, wherein the target document includes: a system document based on which business is processed in a financial institution, and the target attributes include: attributes possessed by all paragraphs in the target document.
[0044] The above-mentioned target document can be a document of rules and regulations (also referred to as a system document) on which financial institutions conduct business, and the above-mentioned target attributes can be attributes possessed by all paragraphs in the target document. For example, common attributes apply to all items in the target document, including but not limited to: applicable fields of the system, issuing agencies, effective time of the system, applicable agencies of the system, abolition time of the system, etc. By extracting common attributes in the target document (for example, the system document), it is possible to prevent data islands formed by the loss of macro attributes of items in the document after the subdivision and splitting of the system items, resulting in even if the relevant items are recalled through semantics, the lack of macro attributes will affect the model's understanding and control of the document in the subsequent process.
[0045] The above-mentioned institutional documents generally have the following characteristics: (1) The writing style is presented in a general-to-specific manner; (2) The document title contains the sub-sectors to which the current system applies; (3) The beginning of the document can summarize the scope of effectiveness of the current institutional document and other details; (4) The institutional details are presented one by one in the form of institutional items; (5) There is a front-end and back-end dependency relationship between different institutional items in the same institutional document.
[0046] Step S202, based on the document structure of the target document, split the target document to obtain a target split result, wherein the target split result at least includes: M entries in the target document, where M is a positive integer.
[0047] The above-mentioned document structure may include the text structure of the target document, for example, chapters, sections, and entries. Each chapter may include: multiple sections, and each section may include multiple entries. In this embodiment, the target document can be split with entries as the minimum splitting granularity. After splitting to obtain multiple entries, each entry can be assigned a globally unique identifier, and the multiple entries can be graded in the form of chapters, sections, and entries. An index or directory can be established for each entry according to the grading rules of chapters, sections, and entries to facilitate searching for each entry. In this way, the above-mentioned target splitting result can be obtained, avoiding the related technology of segmenting the system file according to a fixed length. The system entries near the segmentation point will lose semantic coherence, resulting in distortion of the semantic information of the system, which in turn leads to the extraction of knowledge points related to user questions (for example, system entries), which affects the difficulty of the large model (the model used to generate responses to questions) in understanding the system.
[0048] Step S203, storing the target attributes and the target splitting results into a target knowledge base, wherein the target knowledge base is used to provide users with a service of replying to questions to be answered.
[0049] In this embodiment, the target attributes and target splitting results can be stored in the target knowledge base, so that after the user asks a question, knowledge points (for example, system entries) related to the question raised by the user can be extracted from the target knowledge base, so as to use the big model to respond to the question to be answered.
[0050] In this embodiment, through the above steps, the system document is split according to the document structure of the system document, and the split results and the common attributes (target attributes) of the system document are stored in the knowledge base, so as to provide users with services to reply to unanswered questions, avoiding the situation in the related technology that the system document is split according to fixed word length, resulting in the loss of semantic information in the split results, thereby achieving the technical effect of improving the semantic integrity of the system document splitting results.
[0051] It should be noted that the target knowledge base in this embodiment can be applied to the institutional question-answering system, which is a system built on the RAG (Retrieval-Augmented Generation) framework. RAG is a large model generative design pattern that can combine a large language model (LLM) with external knowledge retrieval, connect real-time data to the large model, understand the questions raised, and generate answers using the large model based on the knowledge points related to the questions. By providing data as context to the LLM during reasoning, the accuracy and quality of the application can be improved.
[0052] Figure 3 is a schematic diagram of a RAG framework provided according to an embodiment of the present application, such as Figure 3 As shown, the RAG framework includes: query preprocessing module, knowledge retrieval module, prompt word processing, large model generation and post-processing module. The RAG framework can interact with the retrieval knowledge base (corresponding to the target knowledge base) and the large language model.
[0053] Optionally, in the method for processing institutional documents provided in an embodiment of the present application, the target document is split based on the document structure of the target document to obtain a target splitting result, including: obtaining preset text, wherein the preset text is used to identify the document structure of the target document; identifying paragraphs in the target document containing the preset text to obtain a recognition result; and splitting the target document based on the recognition result to obtain a target splitting result.
[0054] The above-mentioned preset text may be used to identify the document structure of the target document, and the above-mentioned preset text may include but is not limited to: chapter, section, and article.
[0055] In this embodiment, the natural paragraphs in the target document can be identified by model recognition, and the natural paragraphs containing preset text (for example, "Chapter X / Section / Article") at the beginning of the paragraph can be identified. For example, text processing technology, such as entity recognition in natural language processing, can be used to mark the position of the preset text in the document, or the paragraph containing the preset text can be identified by regular matching to obtain the recognition result. Then, the target document can be split based on the recognition result to obtain the target splitting result.
[0056] Taking the target document as an institutional document as an example, since institutional documents are usually narrated according to the writing habits of "chapter, section, and article", the natural paragraphs in the article can be identified through pattern recognition, and the natural paragraphs with the words "Chapter X / Section / Article" at the beginning of the paragraph can be identified. The paragraphs that do not contain "Chapter X / Section / Article" at the beginning of the paragraph are subordinate to the paragraphs that contain "Chapter X / Section / Article" in the writing order; then, "article" can be used as the minimum splitting granularity, and the sub-items of rules and regulations contained in the current institutional document can be split according to the dimension of "article", and each item can also be given a globally unique identifier.
[0057] The target document's paragraphs are identified by recognizing preset texts, and then the target document is split according to the recognition results, thereby achieving the purpose of splitting the target document according to the document structure and avoiding the situation in related technologies where fixed length is used to split the institutional document, and the contextual semantics in the split results are easily lost.
[0058] Optionally, in the method for processing institutional documents provided in the embodiment of the present application, the target document is split based on the recognition result to obtain a target split result, including: based on the recognition result, the target document is split with entries as the splitting granularity to obtain an initial split result, wherein the initial split result includes: M entries in the target document; the M entries are graded according to preset grading rules to obtain a grading result, wherein the preset grading rules include: chapter, section, entry, the level of the chapter is higher than the level of the section, and the level of the section is higher than the level of the entry; based on the initial split result and the grading result, the target split result is determined.
[0059] In this embodiment, pattern recognition technology can be used to locate item identifiers in the document. Once an identifier, such as "Article X", is found, this identifier and the paragraph content that follows it can be taken as an independent item, and the entire target document can be traversed to identify all such identifiers and extract the corresponding paragraphs therefrom, thereby forming an initial splitting result containing all items in the target document, where M represents the total number of items in the target document, and each item is an independent part of the document structure.
[0060] The process of grading M entries may include: first, identifying the "chapter" and "section" identifiers in the document. The "chapter" and "section" identifiers usually appear in the form of "Chapter X" or "Section X" and are located before the "entry" identifier. The "chapter" is located before the "section" (that is, the level of the chapter is higher than the level of the section, and the level of the section is higher than the level of the entry). In this embodiment, the "chapter" and "section" identifiers can be matched with the corresponding entries to determine the section and chapter to which each entry belongs. Secondly, according to the matching results, the entries can be sorted and grouped according to the chapters and sections to which they belong, and a hierarchical structure of the document can be constructed. This hierarchical structure will facilitate the understanding and processing of the position and relationship of each entry in the document, and improve the efficiency and accuracy of subsequent question-answering processing.
[0061] In this embodiment, after obtaining the initial splitting results and the grading results, these results can also be integrated, each entry can be marked as an independent knowledge unit, and additional metadata can be given, such as the identification of the chapter and section to which it belongs, and then the entry, the identification of the entry, the identification of the chapter and section to which each entry belongs are stored in the target knowledge base, and each entry will contain its content, a globally unique identifier, and the chapter and section information to which it belongs, forming a structured knowledge representation. In this embodiment, the hierarchical structure can also be further optimized, for example, the logical relationship between the entries can be identified, or the correlation between the entries can be updated to ensure that the organizational structure of the knowledge base is optimal and clearest, achieving the technical effect of ensuring the integrity and coherence of the system documents in the knowledge base.
[0062] For example, institutional documents usually specify different institutional details according to different chapters. The systems belonging to the same chapter cover the same fields. At the same time, the institutional documents are written in a general-to-specific manner. The items split in the initial split result can be classified accordingly, including: (1) chapter-item relationship, for example, chapters contain sections, sections contain articles, and articles contain subdivided content in the institutional document; (2) items belonging to the same section cover the same fields, but the details in the items under the section are different.
[0063] Optionally, in the method for processing institutional documents provided in the embodiment of the present application, after grading the M entries according to preset grading rules to obtain the grading results, it also includes: determining an analysis strategy, wherein the analysis strategy includes at least one of the following: traceability analysis, cluster analysis, and federated learning; based on the analysis strategy, analyzing the relationship between the M entries and other entries in the target knowledge base except the M entries to obtain an analysis result.
[0064] The above analysis strategies may include provenance analysis, cluster analysis, federated learning, etc. Provenance analysis can be used to identify causal relationships or information flows between entries, and to track the original source of an entry or other entries it affects. Cluster analysis can be used to group entries according to their content, subject, or attributes, identify sets of similar entries, help organize knowledge bases, and improve retrieval efficiency. Federated learning allows multiple knowledge bases or institutions to collaborate on training models to improve knowledge sharing and understanding while protecting data privacy, especially in scenarios involving multi-institutional collaboration.
[0065] In this embodiment, after obtaining the grading results, the analysis strategy can be further determined to deeply understand the relationship between each entry in the document structure and the entries in the existing knowledge base. The choice of analysis strategy can depend on the specific application scenario, the existing structure of the knowledge base, and the expected analysis depth. The appropriate analysis strategy can be automatically or manually selected based on the status of the current target knowledge base, the preliminary analysis results between the entries, and the expected analysis goals. For example, if there are already a large number of structured entries in the knowledge base, cluster analysis can be selected to further organize and classify the information. If the analysis goal is to enhance knowledge sharing across institutions, federated learning can be selected.
[0066] In this embodiment, based on the analysis strategy, the relationship between the M entries and other entries in the target knowledge base is analyzed to deeply analyze the relationship between the newly added M entries and the entries in the existing knowledge base, including subordinate relationships, derivative relationships, conflict relationships, etc., and then the analysis results can be stored in the target knowledge base to ensure the coherence and consistency of the knowledge base. For example, in the traceability analysis, you can try to identify which existing entries may directly or indirectly affect the M entries, or which entries are the basis of the M entries. In cluster analysis, the entries can be grouped with other entries in the knowledge base according to the similarity of their content to form topic-related knowledge clusters. In federated learning, collaborative learning can be carried out with the knowledge bases of other institutions to identify the relationship between cross-library entries through joint training of models.
[0067] In this embodiment, the analysis is performed according to the selected strategy, which may involve complex algorithms and models, such as machine learning models for clustering, or encrypted computing for data sharing in federated learning. After the analysis is completed, the obtained analysis results can be integrated into the target knowledge base to update the relationship information between the entries, for example, adding new entries in the subordinate relationship, or establishing new links in the derivative relationship. This integration process can also be manually reviewed to ensure the accuracy and applicability of the analysis results.
[0068] For provenance analysis, it is also necessary to build a reference network between entries and identify information flows through network analysis techniques. For cluster analysis, word embedding or text vector representation can be used to group entries through clustering algorithms. For federated learning, a secure model training protocol can be designed to ensure that multi-party data can be used to improve the model without directly sharing the original data.
[0069] The execution of these analysis strategies can provide a deeper understanding of the relationship between the newly added M entries and the entries of other institutional documents in the knowledge base, so as to improve the organizational efficiency and retrieval performance of the knowledge base, and enhance the accuracy and comprehensiveness of the large model when generating answers. The integration of the analysis results further improves the structure of the knowledge base, enabling subsequent question-answering processing based on richer and more coherent background institutional knowledge.
[0070] For example, in addition to grading and classifying the institutional items involved in the current document, other technical means such as traceability analysis, cluster analysis, federated learning, etc. can be used to identify the relationship between sub-items, including but not limited to subordinate relationships, derivative relationships, etc., to improve the comprehensiveness of the target knowledge base.
[0071] Optionally, in the method for processing institutional documents provided in an embodiment of the present application, extracting target attributes of the target document includes: obtaining preset prompt words, wherein the preset prompt words are used to control a first language model to extract the attributes of the target document, and the first language model includes: a pre-trained deep learning model; inputting the preset prompt words and the target document into the first language model, and outputting the target attributes.
[0072] The above-mentioned preset prompt words are pre-designed and used to guide the first language model to focus on and extract key attributes in the target document (for example, applicable field, issuing agency, effective time, scope of application, expiration time, etc.). The preset prompt words can be in the form of questions, such as "What is the effective time of this document?" or "What is the issuing agency of the document?", etc., or they can be descriptive sentences to elicit specific attributes of the document.
[0073] The above-mentioned first language model is a pre-trained deep learning model with powerful natural language understanding and generation capabilities. It can be a large language model. The first language model can be trained on large-scale text data and can capture the complex patterns and contextual relationships of the language.
[0074] In this embodiment, the preset prompt words and the text content of the target document can be input into the first language model together. In the target document, the preset prompt words are used to provide task instructions to the model to help the model focus on the attribute extraction task. Specifically, the text of the target document can be deeply analyzed through the neural network structure inside the first language model to identify the document attributes related to the prompt words. For example, if the prompt word is "the effective time of the document", the first language model can search for time-related sentences or phrases in the document and try to identify the exact effective time information.
[0075] After the first language model completes the analysis, it generates an output that may include the recognition results of the common attributes of the target document. The output may be structured data (for example, JSON format), including the attribute name and the corresponding value; or it may be in free text form, directly describing the target attribute.
[0076] In this embodiment, through the guidance of preset prompt words, the first language model is used to accurately locate and extract common attributes in the target document, ensuring the accuracy and efficiency of attribute extraction, which is the basis for building an efficient institutional question-answering system.
[0077] For example, the common attributes of the current system document can be extracted from the system document (corresponding to the target document). The system document contains some common attributes applicable to all entries, including but not limited to: the scope of application of the system, the issuing agency, the effective time of the system, the applicable agency of the system, the abolition time of the system, etc.
[0078] Since common attributes are generally explained clearly at the beginning and end of a document, in this embodiment, the corpus can be extracted from a specified part of the policy document (for example, the beginning and end paragraphs of the policy document) with the help of a large model to extract common attributes.
[0079] Optionally, in the method for processing system documents provided in the embodiment of the present application, after storing the target attributes and the target splitting results in the target knowledge base, it also includes: receiving the question content of the question raised by the target object, preprocessing the question content, and obtaining processed question content, wherein the preprocessing processing method includes at least one of the following: intent analysis, keyword extraction; based on the question content and the processed question content, extracting data in the target knowledge base to obtain knowledge data, wherein the knowledge data includes: entries or attributes in the target document related to the question content; based on the knowledge data and the processed question content, generating a reply answer.
[0080] In this embodiment, the question content of the question raised by the target object (for example, the user) can be received. In order to improve the accuracy and generation efficiency of the generated reply answer, the question content can be preprocessed to obtain the processed question content. The process of preprocessing the question content can include but is not limited to: performing intention analysis on the question content, keyword extraction, query subject extraction, retrieval planning (for example, converting the spoken content in the question content into written language, etc.
[0081] Then, based on the question content and the processed question content, the data (entries, attributes, etc.) related to the question content in the target knowledge base can be extracted to obtain knowledge data. Then, based on the extracted knowledge data and the processed question content, the large model can be used to generate responses to the questions raised by the target object.
[0082] For example, users can ask questions (queries) through a page in the system question and answer system, and perform a preliminary analysis of the questions through a preprocessing module (corresponding to preprocessing), and perform a series of operations such as intent analysis, keyword extraction, inquiry subject extraction, and retrieval planning. After that, based on the results of the preprocessing module and the questions, relevant knowledge points (corresponding to knowledge data) that can answer the questions can be recalled from the target knowledge base. Then, the knowledge points, requirements for answers and other information can be spliced into prompt words and transmitted to the big model. The big model is used to generate answers, and finally the reply answers to the questions are returned. This avoids the situation in which the accuracy of the answers generated by the big model is not high due to the use of fixed term length splitting in related technologies, thereby achieving the technical effect of improving the accuracy of reply answers generated by the system question and answer system.
[0083] Optionally, in the method for processing system documents provided in the embodiment of the present application, a reply answer is generated based on knowledge data and processed question content, including: obtaining preset demand information, wherein the preset demand information includes: requirements for the reply answer to be generated; generating target prompt words based on the preset demand information and knowledge data, and inputting the target prompt words and the processed question content into a second language model to generate a reply answer, wherein the target prompt words are used to control the second language model to generate a reply answer, and the second language model includes: a pre-trained deep learning model.
[0084] The above-mentioned preset requirement information can be rules or requirements set according to specific scenarios or user needs when the second language model generates a reply answer. The preset requirement information may include but is not limited to: the format, length, completeness requirements, and priority knowledge types of the reply answer, to ensure that the generated reply can meet the specific needs of the user. The above-mentioned preset requirement information can be obtained through user interface settings, historical query analysis, or scenario-based predefined rules. For example, if the user explicitly wants to get a detailed system analysis when querying, the preset requirement information may include requiring the reply to contain the full text of the relevant entries. For automatic settings, the requirement information can be dynamically generated based on the context of the content of the query system document, the user role (for example, front-line business personnel), and the query history.
[0085] The above-mentioned target prompt words can be used as input of the second language model to guide the second language model to generate reply answers that meet specific requirements. The target prompt words can be combined with the preset demand information and the relevant knowledge points retrieved from the knowledge base to ensure that the generated answers are not only accurate but also meet the specific requirements of the user. In this embodiment, one or more target prompt words can be constructed according to the content and format of the preset demand information. For example, if the demand is to obtain a comprehensive analysis containing multiple relevant knowledge points, the prompt words may include references or brief descriptions of different knowledge points, as well as instructions requiring the model to perform a comprehensive analysis. Knowledge data integration: The knowledge data retrieved from the knowledge base will be integrated into the prompt words as background knowledge for the model to generate answers.
[0086] The above-mentioned second language model can be a pre-trained deep learning model, which has the ability to generate text that meets the context and context requirements. It should be noted that in this embodiment, the first language model and the second language model can be the same large model, or they can be different large models. In this embodiment, the target prompt word and the pre-processed question content can be input into the second language model together. After receiving the input, the second language model can use its internal neural network structure to generate text that meets the context and preset requirements. The second language model can consider the intention and scope of the question content, as well as the knowledge data and generation requirements contained in the target prompt word, and generate a reply answer that can both accurately answer the question and meet specific needs.
[0087] In an optional example, the generated reply answer may also be corrected and formatted to ensure the accuracy and readability of the answer, such as grammatical correction, information supplementation or format adjustment, to ensure that the answer can be presented to the user clearly and completely.
[0088] In this embodiment, knowledge points (corresponding to knowledge data), requirements for answers and other information (corresponding to preset demand information) are spliced into target prompt words and transmitted to the big model (i.e., the second language model). The big model is used to generate answers, and the answers generated by the big model are optimized and adjusted by the post-processing module to ensure that the output results meet specific quality standards and format requirements. Finally, the reply answer to the question asked is returned, thereby generating high-quality answers that meet user needs and improving the practicality of the question-answering system and user satisfaction.
[0089] In this embodiment, after pre-analysis and processing of the document structure, it can be ensured that each system entry is semantically complete, avoiding the semantic loss before and after the segmentation point caused by the fixed-length segmentation method in the related technology; the common attributes of the identified entries can be used in the recall process in the subsequent RAG process to conduct preliminary screening of information such as the applicability and effective time of the questions asked, so as to ensure the validity of the recalled entries; in addition, the relationship between knowledge points in the knowledge base that can be maintained can also be used in the recall stage to cooperate with different processing strategies to form customized background knowledge, emphasizing the before-and-after relationship between entries in the prompt words, and providing a more credible and reliable basis for the large model to understand the background knowledge related to the problem.
[0090] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0091] Embodiment 2
[0092] The embodiment of the present application also provides a system document processing device. It should be noted that the system document processing device of the embodiment of the present application can be used to execute the system document processing method provided in the embodiment of the present application. The system document processing device provided in the embodiment of the present application is introduced below.
[0093] According to an embodiment of the present application, a device for implementing the above-mentioned method for processing system documents is also provided, such as Figure 4 As shown, the device includes: an extraction unit 41, a splitting unit 42 and a storage unit 43.
[0094] The extraction unit 41 is used to extract the target attributes of the target document, wherein the target document includes: the system document based on which the business is processed in the financial institution, and the target attributes include: the attributes possessed by all paragraphs in the target document;
[0095] The splitting unit 42 is used to split the target document based on the document structure of the target document to obtain a target splitting result, wherein the target splitting result at least includes: M items in the target document, where M is a positive integer;
[0096] The storage unit 43 is used to store the target attributes and the target splitting results into a target knowledge base, wherein the target knowledge base is used to provide users with a service of replying to unanswered questions.
[0097] In the system document processing device provided in the embodiment of the present application, the target attributes of the target document can be extracted by the extraction unit 41, wherein the target document includes: the system document based on which the business is processed in the financial institution, the target attributes include: the attributes possessed by all paragraphs in the target document, and the target document is split based on the document structure of the target document by the splitting unit 42 to obtain a target splitting result, wherein the target splitting result includes at least: M entries in the target document, M is a positive integer, and the target attributes and the target splitting results are stored in the target knowledge base by the storage unit 43, wherein the target knowledge base is used to provide users with services for replying to questions to be answered. This solves the technical problem in the related art that the system document is split according to a fixed length, which causes the semantic distortion of the splitting results and leads to inaccurate answers to questions for users. In this embodiment, the system document is split according to the document structure of the system document, and the split results and the common attributes (target attributes) of the system document are stored in the knowledge base to provide users with services for replying to questions to be answered. This avoids the situation in the related art that the system document is split according to a fixed term length, which causes the splitting results to lose semantic information, thereby achieving the technical effect of improving the semantic integrity of the system document splitting results.
[0098] Optionally, in the system document processing device provided in the embodiment of the present application, the splitting unit includes: a first acquisition subunit, used to acquire preset text, wherein the preset text is used to identify the document structure of the target document; an identification subunit, used to identify paragraphs in the target document containing the preset text to obtain an identification result; and a splitting subunit, used to split the target document based on the identification result to obtain a target splitting result.
[0099] Optionally, in the processing device for institutional documents provided in the embodiment of the present application, the splitting subunit includes: a splitting module, which is used to split the target document based on the recognition result and the item as the splitting granularity to obtain an initial splitting result, wherein the initial splitting result includes: M items in the target document; a grading module, which is used to grade the M items according to preset grading rules to obtain a grading result, wherein the preset grading rules include: chapter, section, item, the level of the chapter is higher than the level of the section, and the level of the section is higher than the level of the item; a first determination module, which is used to determine the target splitting result based on the initial splitting result and the grading result.
[0100] Optionally, in the processing device for institutional documents provided in the embodiment of the present application, the splitting sub-unit also includes: a second determination module, used to determine the analysis strategy after grading the M items according to preset grading rules and obtaining the grading results, wherein the analysis strategy includes at least one of the following: traceability analysis, cluster analysis, and federated learning; an analysis module, used to analyze the relationship between the M items and other items in the target knowledge base except the M items based on the analysis strategy to obtain the analysis results.
[0101] Optionally, in the system document processing device provided in the embodiment of the present application, the extraction unit includes: a second acquisition sub-unit, used to acquire preset prompt words, wherein the preset prompt words are used to control the first language model to extract the attributes of the target document, and the first language model includes: a pre-trained deep learning model; a first processing sub-unit, used to input the preset prompt words and the target document into the first language model, and output the target attributes.
[0102] Optionally, in the system document processing device provided in the embodiment of the present application, the system document processing device also includes: a processing unit, which is used to receive the question content of the question raised by the target object after storing the target attributes and the target splitting results in the target knowledge base, and preprocess the question content to obtain processed question content, wherein the preprocessing processing method includes at least one of the following: intent analysis, keyword extraction; an extraction unit, which is used to extract data in the target knowledge base based on the question content and the processed question content to obtain knowledge data, wherein the knowledge data includes: entries or attributes in the target document related to the question content; a generation unit, which is used to generate a reply answer based on the knowledge data and the processed question content.
[0103] Optionally, in the system document processing device provided in the embodiment of the present application, the generation unit includes: a second acquisition subunit, used to acquire preset demand information, wherein the preset demand information includes: requirements for the reply answer to be generated; a second processing subunit, used to generate target prompt words based on the preset demand information and knowledge data, and input the target prompt words and the processed question content into the second language model to generate the reply answer, wherein the target prompt words are used to control the second language model to generate the reply answer, and the second language model includes: a pre-trained deep learning model.
[0104] It should be noted that the above-mentioned extraction unit 41, splitting unit 42 and storage unit 43 correspond to steps S201 to S203 in the first embodiment, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the first embodiment. It should be noted that the above-mentioned modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n), and the above-mentioned modules can also be run in the computer terminal 10 provided in the first embodiment as part of the device.
[0105] Embodiment 3
[0106] An embodiment of the present application may provide an electronic device, Figure 5 is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 5 As shown, the electronic device may include: one or more ( Figure 5 (only one is shown) processor 502, memory 504, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0107] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0108] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: extract the target attributes of the target document, wherein the target document includes: the system document based on which the business is processed in the financial institution, and the target attributes include: the attributes possessed by all paragraphs in the target document; based on the document structure of the target document, split the target document to obtain a target splitting result, wherein the target splitting result includes at least: M entries in the target document, M is a positive integer; store the target attributes and the target splitting results in a target knowledge base, wherein the target knowledge base is used to provide users with services for responding to unanswered questions.
[0109] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: based on the document structure of the target document, split the target document to obtain the target splitting result, including: obtaining preset text, wherein the preset text is used to identify the document structure of the target document; identifying the paragraphs in the target document containing the preset text to obtain the recognition result; splitting the target document based on the recognition result to obtain the target splitting result.
[0110] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: In the method for processing institutional documents provided in the embodiment of the present application, the target document is split based on the recognition result to obtain a target split result, including: based on the recognition result, the target document is split with the item as the splitting granularity to obtain an initial split result, wherein the initial split result includes: M items in the target document; the M items are graded according to preset grading rules to obtain a grading result, wherein the preset grading rules include: chapter, section, item, the level of the chapter is higher than the level of the section, and the level of the section is higher than the level of the item; based on the initial split result and the grading result, the target split result is determined.
[0111] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: In the method for processing system documents provided in the embodiment of the present application, after grading the M items according to the preset grading rules and obtaining the grading results, it also includes: determining the analysis strategy, wherein the analysis strategy includes at least one of the following: traceability analysis, cluster analysis, and federated learning; based on the analysis strategy, analyzing the relationship between the M items and other items in the target knowledge base except the M items to obtain the analysis results.
[0112] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: In the method for processing institutional documents provided in the embodiment of the present application, extracting the target attributes of the target document includes: obtaining a preset prompt word, wherein the preset prompt word is used to control the first language model to extract the attributes of the target document, and the first language model includes: a pre-trained deep learning model; inputting the preset prompt word and the target document into the first language model, and outputting the target attribute.
[0113] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: In the method for processing system documents provided in the embodiment of the present application, after storing the target attributes and target splitting results in the target knowledge base, it also includes: receiving the question content of the question raised by the target object, preprocessing the question content, and obtaining the processed question content, wherein the preprocessing processing method includes at least one of the following: intent analysis, keyword extraction; based on the question content and the processed question content, extracting data in the target knowledge base to obtain knowledge data, wherein the knowledge data includes: entries or attributes in the target document related to the question content; generating a reply answer based on the knowledge data and the processed question content.
[0114] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: In the method for processing the system document provided in the embodiment of the present application, a reply answer is generated based on the knowledge data and the processed question content, including: obtaining preset demand information, wherein the preset demand information includes: requirements for the reply answer to be generated; based on the preset demand information and the knowledge data, a target prompt word is generated, and the target prompt word and the processed question content are input into the second language model to generate a reply answer, wherein the target prompt word is used to control the second language model to generate a reply answer, and the second language model includes: a pre-trained deep learning model.
[0115] By adopting the embodiment of the present application, the system document is split according to the document structure of the system document, and the split result and the common attributes (target attributes) of the system document are stored in the knowledge base, so as to provide users with services for answering questions to be answered, avoiding the situation in the related art where the system document is split according to a fixed term length, resulting in the loss of semantic information in the split result, thereby achieving the technical effect of improving the semantic integrity of the system document split result. This solves the technical problem in the related art where the system document is split according to a fixed length, resulting in semantic distortion of the split result, leading to inaccurate answers to questions for users.
[0116] It can be understood by those skilled in the art that Figure 5The structure shown is for illustration only, and the electronic device may also be a smart phone, a tablet computer, a PDA, a mobile Internet device (MID), a PAD or other terminal device. Figure 5 The structure of the electronic device is not limited. Figure 5 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Figure 5 Different configurations shown.
[0117] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0118] Embodiment 4
[0119] The embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the method for processing the system document provided in the first embodiment.
[0120] Optionally, in this embodiment, the above storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0121] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing the program steps of the method for processing institutional documents.
[0122] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0123] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0124] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0125] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0126] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0127] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.
[0128] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for processing system documents, characterized in that: include: Extracting target attributes of a target document, wherein the target document includes: a system document based on which a financial institution processes business, and the target attributes include: attributes possessed by all paragraphs in the target document; Based on the document structure of the target document, the target document is split to obtain a target split result, wherein the target split result at least includes: M entries in the target document, where M is a positive integer; The target attributes and the target splitting results are stored in a target knowledge base, wherein the target knowledge base is used to provide users with a service of replying to questions to be answered.
2. The processing method according to claim 1, characterized in that: Based on the document structure of the target document, the target document is split to obtain a target splitting result, including: Acquire preset text, wherein the preset text is used to identify the document structure of the target document; Identify the paragraphs in the target document that contain the preset text and obtain a recognition result; The target document is split based on the recognition result to obtain the target splitting result.
3. The processing method according to claim 2, characterized in that: The target document is split based on the recognition result to obtain the target split result, including: Based on the recognition result, the target document is split with the item as the splitting granularity to obtain an initial splitting result, wherein the initial splitting result includes: M items in the target document; Classifying the M items according to a preset classification rule to obtain a classification result, wherein the preset classification rule includes: chapter, section, item, the level of the chapter is higher than the level of the section, and the level of the section is higher than the level of the item; Based on the initial segmentation result and the classification result, the target segmentation result is determined.
4. The processing method according to claim 3, characterized in that: After the M items are graded according to the preset grading rules to obtain the grading result, the method further includes: Determining an analysis strategy, wherein the analysis strategy includes at least one of the following: source tracing analysis, cluster analysis, and federated learning; Based on the analysis strategy, the relationship between the M entries and other entries in the target knowledge base except the M entries is analyzed to obtain an analysis result.
5. The processing method according to claim 1, characterized in that: Extract target attributes of target documents, including: Acquire a preset prompt word, wherein the preset prompt word is used to control a first language model to extract the attribute of the target document, and the first language model includes: a pre-trained deep learning model; The preset prompt word and the target document are input into a first language model, and the target attribute is output.
6. The processing method according to claim 1, characterized in that: After storing the target attribute and the target splitting result in the target knowledge base, the method further includes: Receiving the content of the question raised by the target object, preprocessing the content of the question to obtain processed content of the question, wherein the preprocessing method includes at least one of the following: intention analysis, keyword extraction; Based on the question content and the processed question content, extracting data in the target knowledge base to obtain knowledge data, wherein the knowledge data includes: entries or attributes in the target document related to the question content; A reply answer is generated based on the knowledge data and the processed question content.
7. The processing method according to claim 6, characterized in that: Based on the knowledge data and the processed question content, a reply answer is generated, including: Acquire preset requirement information, wherein the preset requirement information includes: requirements for the reply answer to be generated; Based on the preset demand information and the knowledge data, a target prompt word is generated, and the target prompt word and the processed question content are input into a second language model to generate the reply answer, wherein the target prompt word is used to control the second language model to generate the reply answer, and the second language model includes: a pre-trained deep learning model.
8. A system document processing device, characterized in that: include: An extraction unit, used to extract target attributes of a target document, wherein the target document includes: a system document based on which a financial institution processes business, and the target attributes include: attributes possessed by all paragraphs in the target document; A splitting unit, configured to split the target document based on the document structure of the target document to obtain a target splitting result, wherein the target splitting result at least includes: M items in the target document, where M is a positive integer; A storage unit is used to store the target attributes and the target splitting results into a target knowledge base, wherein the target knowledge base is used to provide users with a service of replying to questions to be answered.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the method for processing the system document described in any one of claims 1 to 7.
10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method for processing system documents described in any one of claims 1 to 7 are implemented.