Document processing method and device, equipment and storage medium

By intelligently recommending and adding knowledge units in documents, the problem of low efficiency in knowledge retrieval and use in existing technologies is solved, and more efficient document management and knowledge recommendation are achieved.

CN120706378APending Publication Date: 2025-09-26NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510796493.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the existing technology, document management relies on tree-like hierarchical classification, which leads to low efficiency in knowledge retrieval and use. Users need to manually remember storage paths and retrieve documents to obtain relevant knowledge.

Method used

By obtaining the interactive content of the current document, the knowledge unit to be selected is determined, and the target knowledge unit is added to the current document at a preset position. The knowledge graph is used to display the relationship between the knowledge unit and the document to achieve intelligent recommendation and addition.

Benefits of technology

It improves the efficiency of retrieval and use of knowledge in documents, reduces the complexity of user operations, and realizes knowledge recommendation and addition that better meets user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706378A_ABST
    Figure CN120706378A_ABST
Patent Text Reader

Abstract

The invention provides a document processing method and device, equipment and a storage medium, and the method comprises the steps: obtaining interaction content for a current document, the interaction content comprising at least one of the following items: question content of a user and document content of the current document; receiving and displaying a knowledge unit to be selected; in response to a trigger operation for a target knowledge unit in the to-be-selected knowledge units, displaying a current document in which the target knowledge unit is added at a preset position; wherein the to-be-selected knowledge unit is determined based on the interaction content and a plurality of historical knowledge units obtained by slicing each historical document, so that the retrieval and use efficiency of knowledge in the document is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of data processing technology, and in particular to a document processing method, apparatus, device, and storage medium. Background Art

[0002] Currently, document management relies primarily on tree-like hierarchical classification, such as a nested folder structure. In this nested folder structure, manually pre-set nested directories are required to associate and organize documents according to different folders.

[0003] As personal knowledge management requirements become increasingly complex, users often need to access knowledge associated with a document when editing it. Conventional document management technologies require users to memorize the document's storage path and search through all folders to find the desired document and the knowledge within it. Consequently, knowledge retrieval and utilization remain inefficient. Summary of the Invention

[0004] The present application provides a document processing method, apparatus, device and storage medium, which can improve the efficiency of retrieving and using knowledge in documents.

[0005] In a first aspect, a document processing method is provided, comprising: obtaining interactive content for a current document, the interactive content comprising at least one of the following: the user's question content, the document content of the current document; receiving and displaying knowledge units to be selected; in response to a trigger operation for a target knowledge unit in the knowledge units to be selected, displaying the current document with the target knowledge unit added at a preset position; wherein the knowledge unit to be selected is determined based on the interactive content and a plurality of historical knowledge units obtained by slicing each historical document.

[0006] In a second aspect, a document processing method is provided, comprising: receiving interactive content for a current document, the interactive content comprising at least one of the following: the user's question content, the document content of the current document; determining a knowledge unit to be selected based on the interactive content and a plurality of historical knowledge units obtained by slicing each historical document; and sending the knowledge unit to be selected to a client, so that the client displays the current document with a target knowledge unit in the knowledge unit to be selected added at a preset position.

[0007] In a third aspect, a document processing device is provided, comprising: a content acquisition module for acquiring interactive content for a current document, the interactive content including at least one of the following: the user's question content, the document content of the current document; a receiving and display module for receiving and displaying knowledge units to be selected; a knowledge reference module for displaying the current document with the target knowledge unit added at a preset position in response to a trigger operation for a target knowledge unit in the knowledge units to be selected; wherein the knowledge unit to be selected is determined based on the interactive content and a plurality of historical knowledge units obtained by slicing each historical document.

[0008] In a fourth aspect, a document processing device is provided, including: a content receiving module for receiving interactive content for a current document, the interactive content including at least one of the following: the user's question content, the document content of the current document; a knowledge recommendation module for determining a knowledge unit to be selected based on the interactive content and multiple historical knowledge units obtained by slicing each historical document; a knowledge sending module for sending the knowledge unit to be selected to a client, so that the client displays the current document with the target knowledge unit in the knowledge unit to be selected added at a preset position.

[0009] In a fifth aspect, an electronic device is provided, comprising: a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method as in the first aspect or its various implementations.

[0010] In a sixth aspect, a computer-readable storage medium is provided for storing a computer program, wherein the computer program enables a computer to execute the method as in the first aspect or its various implementations.

[0011] In a seventh aspect, a computer program product is provided, comprising computer program instructions, which enable a computer to execute the method as in the first aspect or its various implementations.

[0012] In an eighth aspect, a computer program is provided, which enables a computer to execute the method in the first aspect or its various implementations.

[0013] In summary, the present application can determine and display the knowledge unit to be selected based on the document content of the current document and / or the user's question content and multiple historical knowledge units obtained by slicing each historical document, and add the target knowledge unit to the current document based on the user's trigger operation on the target knowledge unit in the knowledge unit to be selected, so as to recommend to the user the knowledge needed to edit the current document more intelligently and more in line with user needs. Moreover, the user can select and directly add the required knowledge to the current document through a simple trigger operation, thereby improving the convenience and efficiency of document editing; at the same time, since the knowledge units recommended to the user are after slicing the document, the corresponding content granularity is smaller than the document, therefore, it can also avoid the user from selecting directly from the document, which can improve the efficiency of retrieval and use of knowledge in the document. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The following is an introduction to the drawings required for describing the embodiments.

[0015] Figure 1 A flowchart of a document processing method provided in an embodiment of the present application;

[0016] Figure 2A A schematic diagram of a document processing method provided in an embodiment of the present application;

[0017] Figure 2B A schematic diagram of another document processing method provided in an embodiment of the present application;

[0018] Figure 2C A schematic diagram of another document processing method provided in an embodiment of the present application;

[0019] Figure 3A A schematic diagram of another document processing method provided in an embodiment of the present application;

[0020] Figure 3B A schematic diagram of another document processing method provided in an embodiment of the present application;

[0021] Figure 4A A schematic diagram of another document processing method provided in an embodiment of the present application;

[0022] Figure 4B A schematic diagram of another document processing method provided in an embodiment of the present application;

[0023] Figure 5A A schematic diagram of another document processing method provided in an embodiment of the present application;

[0024] Figure 5B A schematic diagram of another document processing method provided in an embodiment of the present application;

[0025] Figure 5C A schematic diagram of another document processing method provided in an embodiment of the present application;

[0026] Figure 5D A schematic diagram of another document processing method provided in an embodiment of the present application;

[0027] Figure 6 A flowchart of another document processing method provided in an embodiment of the present application;

[0028] Figure 7 A schematic diagram of a document processing device 700 provided in an embodiment of the present application;

[0029] Figure 8 A schematic diagram of another document processing device 800 provided in an embodiment of the present application;

[0030] Figure 9 Schematic diagram of an electronic device 900 provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] The technical solution of this application will be introduced below in conjunction with the drawings in this application.

[0032] It should be noted that the information, data (including, but not limited to: data used for analysis, stored data, displayed data, etc., such as interactive content, knowledge units, documents, knowledge graphs, user preferences, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the interactive content, knowledge units and documents involved in this application and the operations performed on them are all obtained with full authorization.

[0033] In one embodiment, the technical solution of the present application can be used in document management scenarios, but is not limited thereto. Specifically, it can be applied to scenarios such as editing documents, viewing documents, and the relationship between knowledge units in documents.

[0034] In one embodiment, the solution provided by the present application can be executed by any terminal device and server with data processing capabilities, for example: the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services; the terminal device can be a tablet computer, a laptop computer or a desktop computer, etc.

[0035] Among them, a client can be installed on the terminal device, and the corresponding document processing method can be implemented through the client. The client can be a mobile client, a desktop client or a web client, and this application does not limit this.

[0036] The following is an introduction to various embodiments of the technical solution of this application:

[0037] It should be noted that all technical solutions in this application can be combined in any way to form optional embodiments of this application. For example, the embodiments on the terminal device side and the server side can refer to each other and will not be described in detail here.

[0038] The following first introduces the document processing method involved in this application from the client's perspective.

[0039] Figure 1 This is a flowchart of a document processing method provided by an embodiment of the present application. The method can be executed by the terminal device in the above content. The terminal device can be installed with a client, and the client can implement the corresponding document processing method. Figure 1 As shown, the method includes:

[0040] S110: Obtaining interactive content for the current document, where the interactive content includes at least one of the following: user's question content, and document content of the current document;

[0041] S120: receiving and displaying a knowledge unit to be selected, where the knowledge unit to be selected is determined based on the interactive content and multiple historical knowledge units obtained by slicing each historical document;

[0042] S130: In response to a triggering operation on a target knowledge unit among the knowledge units to be selected, displaying the current document with the target knowledge unit added at a preset position.

[0043] It should be noted that the trigger operation, display operation, viewing operation and other operations in this application can be specifically single-click, double-click, press, drag to a designated area or floating touch gesture, etc., and this application does not impose any restrictions on this.

[0044] The documents in this application may be data files, such as, but not limited to, notes, essays, meeting minutes, or files in formats such as .docx, .pdf, and .txt. The document content may be text, tables, images, and the like. Accordingly, the knowledge units obtained by segmenting the document may be in the form of text, tables, images, and the like, and this application does not impose any restrictions on this.

[0045] Regarding S110, in one embodiment, the current document may refer to the document currently being edited by the user. The client may display an editing page, through which the user can edit the current document. Furthermore, the client may display a dialog box on the editing page, in which the user may enter a question, so that the client can obtain the question. The question may be a question used to trigger the acquisition of the knowledge unit to be selected.

[0046] For example, the edit page is as follows Figure 2A As shown, the document content of the current document is "Title 1, Paragraph 1, Paragraph 2", and the dialog box is as follows Figure 2A As shown on the right side of the dialog box, for example, a user can enter a question in the dialog box: "What are the principles of layout design?", and the client can then obtain the question.

[0047] Alternatively, the user may not input the question content in the dialog box, so that the interactive content obtained by the client may only include the document content of the current document.

[0048] Illustratively, the document content of the current document may be the entire content of the current document, or the content within a set range at the current editing position in the current document, such as the content of the currently edited paragraph, but is not limited thereto.

[0049] In the above process, the user's questions and / or the document content of the current document can be obtained through the dialog box, so that the user can be interacted with quickly, conveniently and friendly, and the knowledge units to be selected that better meet the user's needs can be recommended to the user based on the interactive content, thereby improving the user experience.

[0050] For S120, in one embodiment, the client can first send the interactive content to the server; the server can determine the knowledge unit to be selected based on the interactive content and multiple historical knowledge units obtained by slicing each historical document (this process can refer to the corresponding embodiment on the server side), and return the knowledge unit to be selected to the client; thereafter, the client can obtain and display the knowledge unit to be selected.

[0051] For example, the client can display the knowledge unit to be selected in a dialog box. Figure 2A As shown, the knowledge units to be selected may include "knowledge unit 1" and "knowledge unit 2".

[0052] Among them, if the user does not enter the question content, the client can also display the knowledge units to be selected in the dialog box. That is to say, the dialog box can display the knowledge units to be selected in the form of replies to the question content, or it can directly recommend the knowledge units to be selected to the user based on the document content without the user asking questions.

[0053] Of course, the client can also display the knowledge unit to be selected in other display forms, such as a pop-up window. This application does not limit the display form of the knowledge unit to be selected.

[0054] In the above process, knowledge units with smaller granularity can be recommended to users instead of documents with larger granularity. Moreover, the knowledge units to be selected are determined based on the user's interactive content (for example, document content, question content) and historical knowledge. Therefore, not only can more fine-grained knowledge recommendations be achieved to facilitate user selection without the need for users to remember the storage path of knowledge units or documents; but also, implicit relationships between more knowledge units that better meet user needs (for example, relationships between cross-domain knowledge) can be mined and displayed. For example, the implicit relationship between knowledge unit "A project experience" and knowledge unit "B industry trend" can be mined to provide users with more accurate and comprehensive knowledge units.

[0055] For S130, in one embodiment, the user can select the required target knowledge unit from the knowledge units to be selected through the client and trigger the operation on it; then, the client can obtain the trigger operation and, in response to the trigger operation, add the target knowledge unit to the preset position of the current document and display the current document with the added target knowledge unit.

[0056] Exemplarily, the trigger operation may be a drag operation. Then, in response to the drag operation on the target knowledge unit, the client may add the target knowledge unit at the end of the drag operation in the current document and display the current document, that is, the preset position is the position where the drag operation ends.

[0057] For example, Figure 2B As shown, the user can use the client to drag and drop the target knowledge unit "Knowledge Unit 1" displayed in the dialog box to add it between paragraphs 1 and 2 in the current document; after that, when the drag and drop operation or other triggering operation ends, the client can display the current document with the added target knowledge unit, such as Figure 2C shown.

[0058] Exemplarily, the preset position may include any one of the following: a current editing position for editing the current document; a position determined based on a semantic analysis between the document content of the current document and the target knowledge unit.

[0059] The client can directly and automatically display the target knowledge unit at a preset position in the current document after the user selects the target knowledge unit. For example, the user's triggering operation for the target knowledge unit can be a click operation. Accordingly, after receiving the click operation, the client can automatically add the target knowledge document to the preset position sent by the server of the current document (the position determined by semantic analysis between the document content of the current document and the target knowledge unit) or the current editing position.

[0060] For example, the client can add the semantic tag of the target knowledge unit as metadata to the storage unit corresponding to the current document. The semantic tag of the target knowledge unit is used to represent the semantic information of the target knowledge unit. In other words, the client can embed the semantic tag of the target knowledge unit in the current document as metadata, storing it as a hidden attribute, so that data such as the tag of the target knowledge unit can be automatically added to the current document, so as to facilitate the calculation of the association type between the current document and other documents based on the tag.

[0061] In the above process, knowledge units can be quickly added through trigger operations such as drag-and-drop operations. In addition, in subsequent embodiments, the association between the target knowledge unit and the current document can be automatically established based on the trigger operation. Compared with the related technology that requires manual selection, copying, and pasting operations to add and associate knowledge, the cost of adding and associating can be reduced, the interaction efficiency can be improved, and the convenience of operation and the efficiency of knowledge use and reuse can be improved.

[0062] The document processing method in the above embodiment mainly involves a document editing method. The following embodiment will introduce the knowledge graph involved in the document processing method of this application.

[0063] In one embodiment, the client may obtain a display operation for a knowledge graph, and in response to the display operation, receive the knowledge graph and display at least one of a relationship view and an index view in the knowledge graph. The knowledge graph may be constructed by the server and sent to the client.

[0064] Among them, the relationship view represents the association relationship between each knowledge unit and each document based on multiple nodes and edges connecting multiple nodes. Each knowledge unit is obtained by slicing each document; the index view represents the association relationship between each knowledge unit and each document based on the directory display form.

[0065] It should be noted that the knowledge graph is determined based on the association weights, association types, etc. between each knowledge unit and each document, and the association weights change dynamically. Therefore, the knowledge graph is also dynamic. This application will introduce this in subsequent server-side embodiments.

[0066] Among them, the association relationship between each knowledge unit and each document may include at least one of the following: the association relationship between multiple knowledge units, the association relationship between any knowledge unit and any document, the association relationship between multiple documents, the association relationship between multiple knowledge units and at least one document, and the association relationship between multiple documents and at least one knowledge unit.

[0067] For example, the client may display the relationship view and the index view simultaneously on the same page, or may display the relationship view and the index view separately on different pages, which is not limited in this application.

[0068] For example, Figure 3A and Figure 4A As shown, the relationship view and index view can be displayed respectively under different sub-pages "relationship view page" and "index view page" of the same page "knowledge graph page".

[0069] For example, the knowledge graph can be stored in a graph database, such as Neo4j, to enable efficient querying of the knowledge graph. The graph database can store documents, knowledge units, and knowledge graphs.

[0070] In the above process, the association relationship between knowledge units and documents can be established and displayed through relationship views and index views, making the storage and display of knowledge units and documents more structured and intuitive; moreover, there is no need for users to manually establish it, and the two-way association relationship between knowledge units and documents can be established in a more fine-grained, comprehensive and intelligent way, thereby improving the efficiency of knowledge retrieval and reuse.

[0071] Exemplarily, the client can determine a knowledge graph display identifier, which is used to trigger the display of the knowledge graph; then, in the current document, the knowledge graph display identifier is displayed at the identifier display location of the target knowledge unit. Accordingly, the client can obtain a trigger operation for the knowledge graph display identifier at the identifier display location, and display the knowledge graph in response to the trigger operation, wherein the knowledge graph can be a knowledge graph related to the target knowledge unit. In this process, the embedded target knowledge unit is set as an interactive module.

[0072] For example, Figure 2C As shown, the logo display position can be the upper right corner of the target knowledge unit, and the map display logo can be the words "knowledge map" or a five-pointed star-shaped logo. This application does not impose any restrictions on this.

[0073] Alternatively, the user can also access the page displaying the knowledge graph through other channels through the client, and this application does not limit this. For example, the client can respond to the following Figure 2A The trigger operation of the "Knowledge Graph" logo in the lower left corner of the page shown displays the knowledge graph.

[0074] In the above process, the knowledge graph display identifier can be displayed at the identifier display position of the target knowledge unit added in the current document to quickly view the knowledge graph involving the target knowledge unit, which can help users quickly understand the association relationship of the knowledge unit and improve operational convenience.

[0075] The following is an introduction to the relationship view in the knowledge graph:

[0076] Exemplarily, the nodes in the relationship view correspond to individual documents, and the edges in the relationship view correspond to the association types between two documents and the associated knowledge units.

[0077] The association type includes at least one of the following: reference relationship, supplementary relationship, and contradictory relationship. Different association types have different display styles for edges. For example, Figure 3A As shown in the figure, different association types correspond to different lines representing edges.

[0078] For example, the client can determine and display the knowledge units associated with the two documents corresponding to the target edge in response to a trigger operation on the target edge in the relationship view. Figure 3A In the example, the client can display the knowledge unit 1 associated with the document 1 and the document 2 in response to a click operation on the edge connecting the corresponding nodes of the document 1 and the document 2.

[0079] The associated knowledge unit between two documents may refer to the association relationship between the two documents established due to the knowledge unit. For example, the knowledge unit may be a knowledge unit in one document that references or adds to the other document, or it may refer to a knowledge unit shared by the two documents. In this case, the knowledge unit may be the same knowledge unit contained in each of the two documents, or it may refer to a knowledge unit in the other document that is referenced or added by both documents.

[0080] For example, it is also possible to determine the associated document group involving the first knowledge unit in each document, where the first knowledge unit is any knowledge unit in the first document in each document or any knowledge unit in the same knowledge base (the knowledge base will be introduced in the subsequent embodiments); then, determine the document association identifier, which is used to identify that the corresponding document involves the first knowledge unit, that is, to identify that the corresponding document involves the same document or knowledge unit in the knowledge base; use the document association identifier to identify each document in the associated document group. For example, Figure 3A As shown, if document 1, document 2 and other documents all reference a certain knowledge unit in document 3, then document 1, document 2 and other documents can be circled using a document association identifier in the shape of a circle and containing the name of document 3, "Document 3"; or, if document 1, document 2 and other documents all reference a certain knowledge unit in knowledge base A, then document 1, document 2 and other documents can be circled using a document association identifier in the shape of a circle and containing the name of knowledge base A, "Knowledge base A".

[0081] Among them, the knowledge units or documents associated with the document, and the knowledge units or documents associated with the knowledge units can be determined by searching the knowledge graph.

[0082] In addition, in one embodiment, the nodes in the relationship view can also correspond to each knowledge unit, and the edges in the relationship view correspond to the association type between the corresponding two knowledge units and the associated documents. This is similar to the above embodiment and will not be elaborated in this application.

[0083] In the above process, the association relationship between documents and knowledge units can be represented through nodes, edges and document association identifiers, so that users can obtain the association relationship between documents and knowledge units more intuitively and conveniently, and can achieve more efficient retrieval.

[0084] The following is an introduction to the index view in the knowledge graph:

[0085] Exemplarily, the index view includes each second knowledge unit displayed in a directory presentation format and at least one third document in each document associated with each second knowledge unit; wherein the second knowledge unit is a knowledge unit in any second document in each document or any knowledge unit in the same knowledge base.

[0086] For example, Figure 4A As shown, in the index view, documents 2 and 3 and their respective knowledge units can be displayed in a directory format, where the knowledge units of document 2 include knowledge units 1, 2, and 3, which are different paragraphs of document 2. Alternatively, knowledge base A and knowledge base B and their respective knowledge units can be displayed in a directory format, where the knowledge units of knowledge base A include knowledge units 1, 2, and 3.

[0087] For example, the second knowledge unit and the corresponding third document can be connected in the index view based on a preset graphic. The preset graphic can be a line, for example, Figure 4A The line connecting Document 1 and Knowledge Unit 1 in .

[0088] In addition, the client can also receive the knowledge unit attribute data of each second knowledge unit and the document attribute data of each second document; and display the knowledge unit attribute data and document attribute data in the index view. Figure 4A Or as shown in 4B, the knowledge unit attribute data includes: the number of times the corresponding knowledge unit is cited, the corpus information to which it belongs, the corresponding identifier, and the paragraph information in the corresponding document; the document attribute data includes: the number of documents associated with the corresponding document.

[0089] For example, the client can also highlight the third document associated with the second knowledge unit in response to the trigger operation on the second knowledge unit. Figure 4A As shown, the client can highlight or display the document 1, document 4 and document 5 associated with the knowledge unit 1 with bold borders in response to the user's click operation on the knowledge unit 1.

[0090] The document associated with the knowledge unit may refer to a document that adds or includes the knowledge unit, and may be retrieved from the graph database by the server according to the second knowledge unit.

[0091] For example, the client can also highlight the second knowledge unit associated with the third document in response to the trigger operation on the third document. Figure 4B As shown, the client can highlight or display the knowledge unit 1 and the knowledge unit 2 associated with the document 5 with bold borders in response to the user's click operation on the document 5.

[0092] The knowledge unit associated with the document may be a knowledge unit added or included in the document, and may be retrieved from the graph database by the server based on the third document.

[0093] In the above process, the relationships between each document and other documents in the document can be intuitively displayed through a tree structure. In addition, through user trigger operations on documents or knowledge units, two-way mining and path tracing for documents and knowledge units can be achieved. The local network view of the jump knowledge graph can also be highlighted to enhance the user experience.

[0094] Exemplarily, the client can also display a knowledge graph search box; obtain the keywords to be searched in the knowledge graph search box, and the keywords to be searched are used to search the knowledge graph; receive and display the search results, and the search results are determined based on the knowledge graph and the search keywords.

[0095] For example, the knowledge graph search box can be as follows Figure 3A 、 Figure 4A The input box containing the words "Please enter keywords to search" is shown.

[0096] The client can send the search keyword to the server, and the server can determine the search results based on the knowledge graph and the search keyword. For example, the server can use the search keyword to search the knowledge graph.

[0097] Exemplarily, the client may also zoom in or out the knowledge graph in response to a zoom-in display operation or a zoom-out display operation on the knowledge graph.

[0098] For example, the zoom-in display operation and the zoom-out display operation can be respectively Figure 3A 、 Figure 4A The upper right corner contains the trigger operations of the zoom-in and zoom-out icons indicated by “—” and “+”.

[0099] In the above process, retrieval and scaling of the knowledge graph can be realized, which facilitates users to obtain the content they need in a timely and rapid manner, thereby improving user experience.

[0100] In one embodiment, the client may also obtain a slice viewing operation for the fourth document; in response to the slice viewing operation, multiple knowledge units included in the fourth document are displayed on the knowledge slice page.

[0101] Among them, the knowledge slice page can be as follows Figure 3B shown.

[0102] Exemplarily, the slice viewing operation for the fourth document may be a slice viewing operation for the fourth document in the knowledge graph; or it may be a slice viewing operation for the fourth document in a personal knowledge base page.

[0103] For example, the client may respond to Figure 3A The trigger operation of the corresponding node of document 1 in the document 1 is displayed, and the knowledge slice page is displayed on the knowledge slice page, and multiple knowledge units contained in document 1 are displayed. Alternatively, it can be responsive to the trigger operation for the corresponding node of document 1. Figure 2A The trigger operation of the "Personal Knowledge Base" logo in the upper left corner will display the knowledge slice page.

[0104] For example, the client can respond to the trigger operation of the add slice button on the knowledge slice page, obtain the new knowledge unit and add the new knowledge unit to the fourth document. Figure 3B Indicated by the “+” in the upper left corner.

[0105] In the above process, the slicing results of each document, that is, knowledge units, can be displayed, which is convenient for users to view and add them intuitively, thereby improving the user experience.

[0106] The following is a schematic diagram to illustrate the above process:

[0107] In one embodiment, Figure 5AAs shown, a personal knowledge base (which can be a storage unit, such as a database) can include multiple knowledge bases, such as knowledge base A. Each knowledge base can correspond to one document or multiple documents. Each knowledge base can include multiple knowledge units, such as A1, A2, and A3. The knowledge units included in each knowledge base can belong to the same document or to different documents. For example, each knowledge base can store knowledge units of the same type in at least one document, and the type can be determined based on the subject or semantics of the knowledge unit. Document notes can refer to notes edited by users, such as note ①. When a user edits note ①, the user can add knowledge unit A1 in knowledge base A to it through the above embodiment.

[0108] Furthermore, combined with the above content, such as Figure 5B As shown, in the relationship view, Notes ①, Notes ②, and Notes ③ reference / add knowledge units A1 and A2 in knowledge base A. You can use the box corresponding to knowledge base A to circle Notes ①, Notes ②, and Notes ③.

[0109] Furthermore, combined with the above content, such as Figure 5C As shown, in the index view, each knowledge base and the knowledge units it contains can be displayed on the left. For example, knowledge base A and the knowledge units A1 and A2 it contains can be displayed in a tree structure; notes ①, notes ② and notes ③ can be displayed on the right, where notes ① and notes ② reference / add knowledge unit A1 in knowledge base A. The client can use preset graphics in the index view to connect notes ①, notes ② and knowledge unit A1.

[0110] The following is an introduction to the document processing method involved in this application from the server perspective.

[0111] Figure 6 This is a flowchart of another document processing method provided in an embodiment of the present application. This method can be executed by the server in the above content, such as Figure 6 As shown, the method includes:

[0112] S610: Receive interactive content for the current document, where the interactive content includes at least one of the following: user's question content and the document content of the current document;

[0113] S620: Determine a knowledge unit to be selected based on the interactive content and multiple historical knowledge units obtained by slicing each historical document;

[0114] S630: Sending the knowledge unit to be selected to the client, so that the client displays the current document with the target knowledge unit in the knowledge unit to be selected added at the preset position.

[0115] For S610 and S630, reference may be made to the above-mentioned client-side embodiment, which will not be elaborated in this application.

[0116] Regarding S620, in one embodiment, the server may first slice the historical document to obtain multiple historical knowledge units; then, based on the interactive content, the server may search the multiple historical knowledge units to obtain the knowledge units to be selected that are related to the interactive content. The following describes this process:

[0117] For example, historical documents can be segmented according to dimensions such as chapters, paragraphs, or sentences to obtain historical knowledge units. Alternatively, historical documents can be semantically segmented to identify content with the same or similar semantics in the historical documents as the same historical knowledge unit.

[0118] Historical knowledge units can be, but are not limited to, sentences, paragraphs, images, and tags in historical documents. Historical documents can be, but are not limited to, documents used, created, or edited by users, and can be documents in a knowledge graph. After obtaining and slicing historical documents, the server can store them in a knowledge base. Subsequently, the server can first obtain historical knowledge units from the knowledge base and determine the knowledge units to be selected based on these historical knowledge units.

[0119] In addition, a multi-level nested graph database structure can be used to store the association relationship between all knowledge units (including historical knowledge units) and documents, where each knowledge unit can contain metadata (unique identification of the knowledge unit, creation time, modification record, association weight, etc.) and content data (i.e. the specific content of the knowledge unit).

[0120] For example, the knowledge units to be selected may be determined by performing real-time semantic analysis on the interactive content through natural language processing (NLP).

[0121] Exemplarily, the above-mentioned method of determining the knowledge unit to be selected based on the interaction content and multiple historical knowledge units obtained by slicing each historical document includes: converting the interaction content into a word sequence; processing the word sequence through a pre-trained semantic association model to obtain a first sequence containing context information of the interaction content; compressing the first sequence to obtain a text vector; searching multiple historical knowledge units according to the text vector, and determining the knowledge unit to be selected from the multiple historical knowledge units.

[0122] For example, the semantic association model may be a RoBERTa-wwm-ext model, but is not limited thereto. The semantic association model may be fine-tuned using, for example, Lora fine-tuning, using multiple preset documents of the current document's document type (for example, if the current document is a design document, a layout design document or a user experience (UX) document may be used as a preset document).

[0123] For example, a dynamic learning mechanism can be used to update the semantic association model. Specifically, the semantic association model can be updated in real time based on the user's selection result for the knowledge unit to be selected (for example, accepting or rejecting a certain knowledge unit). In addition, incremental training can be performed on the semantic association model every fixed time period, such as 1 hour. Specifically, incremental training can be implemented using partial_fit (a method for incremental learning that allows the model to gradually update model parameters with new data without retraining the entire data set) and new data (new_data), but is not limited to this.

[0124] For example, the compression of the first sequence may specifically be performing pooling processing on the first sequence.

[0125] For example, an index library for approximate nearest neighbor search can be constructed based on multiple historical knowledge units; the number of knowledge units contained in the knowledge unit to be selected is determined; and the number of knowledge units to be selected that are most similar to the text vector is searched in the index library.

[0126] For example, the above process can be achieved through the following steps:

[0127]

[0128] Among them, input_text represents the interactive content, tokenizer represents the word segmenter, tokens represents the word sequence, and return_tensors = 'pt' indicates that the word segmenter returns a tensor (Tensor) in PyTorch format. model represents the semantic association model, .last_hidden_state refers to the hidden state of the last layer of the model, and .mean (dim = 1) refers to taking the mean of the sequence length dimension (seq_length) of the first sequence to obtain the global vector representation embeddings of the sentence, that is, the text vector. annoy_index refers to the vector index based on the Annoy library (Approximate Nearest Neighbors Oh Yeah), which is used for efficient similarity search. get_nns_by_vector represents input vector embeddings and returns the index of the five most similar candidates. filter_by_directory(candidates) indicates that the business logic filtering of the retrieved candidate results is performed to obtain the knowledge unit to be selected.

[0129] In the above process, the server can fully mine and analyze the interactive content through means such as semantic association models and approximate nearest neighbor search, so as to determine the selected knowledge unit that best matches the interactive content in the historical knowledge units, thereby ensuring that the selected knowledge unit recommended to the user is more consistent with the user's question content and document content.

[0130] In one embodiment, before sending the knowledge unit to be selected to the client, the server also includes: performing semantic analysis between the document content of the current document and the target knowledge unit to determine the specified content in the document content that is most semantically similar to the target knowledge unit; determining the adjacent position of the position of the specified content in the current document as a preset position; and sending the preset position to the client so that the client adds and displays the target knowledge unit at the preset position.

[0131] Exemplarily, NLP may be used to perform semantic analysis between the document content of the current document and the target knowledge unit, but is not limited thereto.

[0132] In the above process, the document content of the current document and the semantic analysis of the target knowledge unit can be used to determine the specified content in the document content that is most similar to the semantics of the target knowledge unit, and then the preset position can be determined based on the adjacent position of the position where the specified content is located. This not only allows for the automatic addition of the target knowledge unit, but also ensures that the addition of the target knowledge unit will not disrupt the semantics of the document itself.

[0133] The following is an introduction to the process of determining / building a knowledge graph:

[0134] In one embodiment, the server may determine the knowledge graph by the following steps:

[0135] S640: Determine the association type between the target knowledge unit and the current document;

[0136] S650: Calculate the association weight between the target knowledge unit and the current document;

[0137] S660: Determine a knowledge graph including a relationship view and an index view according to the association type and the association weight, and send the knowledge graph to the client so that the client displays the knowledge graph.

[0138] Among them, the above steps are the process of determining the knowledge graph using the target knowledge unit and the current document as an example. The steps corresponding to establishing the relationship between other documents and knowledge units in the knowledge graph can be referred to here, and this application will not go into details.

[0139] In addition, after determining to add the target knowledge unit to the current document, the server can execute the above steps to achieve real-time establishment of the association relationship between the target knowledge unit and the current document, realize accurate reference and backtracking of paragraph-level knowledge units, and improve the efficiency and timeliness of knowledge graph creation.

[0140] For S640, in one embodiment, the server can determine that within a preset character window (for example, 50 characters), the target knowledge unit and the current document each have associated words representing the association relationship (specifically, the associated words can be detected through a sliding character window); calculate the structural features and interaction features of the target knowledge unit and the current document; and determine the association type between the target knowledge unit and the current document through a preset classifier based on the associated words, structural features, interaction features, and feature weights.

[0141] Among them, the structural features represent the content creation time difference and directory hierarchy distance between the target knowledge unit and the current document; the interactive features represent the user's modification record of the association relationship between the target knowledge unit and the current document; the feature weight represents the influence weight of each feature on the classifier's determination of the association type.

[0142] For example, the associated words that represent the associated relationship may be, but are not limited to, reference, etc.

[0143] For example, the classifier may be a classifier based on the eXtreme Gradient Boosting (XGBoost) algorithm, but is not limited thereto. The association type may be a reference / supplement / contradiction / other relationship, etc.

[0144] Exemplarily, the feature weight is determined based on at least one of the following: semantic similarity between the target knowledge unit and the current document; and user preference information regarding historical association types.

[0145] For example, the cosine similarity between the target knowledge unit and the current document can be determined as the semantic similarity. The preference information for the historical association type can include the historical association types used by the user.

[0146] In the above process, multi-dimensional features can be integrated, such as user preferences, semantic similarity, content differences, user interactions, etc., to determine the association type, which can ensure a more accurate association type.

[0147] Regarding S640, in one embodiment, the server may calculate the label overlap between the target knowledge unit and the current document's respective semantic labels, where the semantic labels are used to represent the semantic information of the corresponding knowledge unit or document; and determine the association weight based on the label overlap.

[0148] Exemplarily, keywords, entities, and subject categories of the target knowledge unit may be extracted; and semantic tags of the target knowledge unit may be generated based on the extraction results.

[0149] For example, NLP can be used to extract keywords, entities (such as names of people, institutions), and subject categories (such as "finance") of the target knowledge unit; then, feature extraction and fusion of keywords, entities, and subject categories are performed to obtain label vectors; finally, the label vectors are converted into semantic labels.

[0150] Exemplarily, the semantic tag of the target knowledge unit includes the document tag of the historical document or the current document corresponding to the target knowledge unit. In other words, the target knowledge unit can inherit the tag of the document.

[0151] The document tag may be, but is not limited to, the title of the document.

[0152] Exemplarily, the above calculation of the label overlap between the semantic labels of the target knowledge unit and the current document includes: calculating the first label quantity of the semantic labels shared by the target knowledge unit and the current document; calculating the second label quantity of the semantic labels of the target knowledge unit and the current document; and determining the label overlap based on the first label quantity and the second label data.

[0153] For example, the ratio of the first tag quantity to the second tag data may be determined as the tag coincidence degree.

[0154] Exemplarily, the above-mentioned determination of the association weight based on the label overlap includes: calculating the semantic similarity between the target knowledge unit and the current document; and determining the association weight based on the label overlap and the semantic similarity.

[0155] For example, the product of label overlap and semantic similarity can be determined as the association weight.

[0156] For example, multi-dimensional feature extraction and deep semantic analysis can be performed on the target knowledge unit and the current document, and they can be converted into vector representations. The similarity of the vector representations can then be calculated using cosine similarity or knowledge graph embedding to obtain the semantic similarity between the target knowledge unit and the current document.

[0157] In the above process, the association weight can be jointly determined based on the label overlap and semantic similarity, and then the association between the target knowledge unit and the current document can be characterized by multiple aspects of different granularities such as labels and semantics, thereby improving the accuracy of the association weight.

[0158] It is understandable that the association between documents and knowledge units may change, so the association weights can be updated in real time to ensure the accuracy of the knowledge graph. The following describes this process:

[0159] In one embodiment, the server can calculate the user behavior weight of the user for the target knowledge unit and the current document; calculate the semantic weight, which represents the semantic similarity between the target knowledge unit and the current document; calculate the structural weight, which represents the importance of the target knowledge unit and the current document in the corresponding topological structure of the knowledge graph; determine the dynamic association weight under each target duration based on the user behavior weight, semantic weight and structural weight under each target duration; and update the association weight accordingly based on the dynamic association weight under each target duration.

[0160] Correspondingly, after the association weight is updated, the server can update the knowledge graph according to the updated association weight.

[0161] Exemplarily, the above calculation of the user behavior weight for the target knowledge unit and the current document includes: calculating the user behavior weight based on the user's triggering behavior data for the target knowledge unit and the current document; the triggering behavior data includes at least one of the following: associated operation data and access data for the target knowledge unit and the current document.

[0162] For example, the association operation data for the target knowledge unit and the current document can be dragging or embedding the target knowledge unit into the current document. A dragging or embedding operation can correspond to a value of 0.3 for the association operation data. The association operation data can also be the user's resident market for the target knowledge unit and the current document. For example, if the user has read the target knowledge unit and the current document for more than 120 seconds, the association operation data can be determined to have a value of 0.15.

[0163] For example, the access data may be the number of accesses or access frequency to the target knowledge unit and the current document, wherein the access frequency may be click frequency, which can be determined by the following formula: click frequency = log (1 + number of weekly visits).

[0164] Exemplarily, the above-mentioned calculation of semantic weight includes: performing vector representation on the semantic information of each of the target knowledge unit and the current document to obtain the semantic vectors of each of the target knowledge unit and the current document; determining a first semantic weight based on the similarity of the two semantic vectors; determining the lexical probability distribution of the vocabulary involved in each of the target knowledge unit and the current document; determining the topic of each of the target knowledge unit and the current document based on the lexical probability distribution; determining a second semantic weight based on the similarity of the two topics; determining the semantic weight based on the first semantic weight and the second semantic weight (for example, weighted sum, average, etc.).

[0165] For example, the similarity between the two semantic vectors is calculated using a sentence-level bidirectional Transformer encoding representation model (Sentence Bidirectional Encoder Representations from Transformers, Sentence-BERT or SBERT) to determine the first semantic weight. The second semantic weight can be determined by detecting the above-mentioned vocabulary probability distribution using Latent Dirichlet Allocation (LDA).

[0166] For example, the structural weight can be calculated based on betweenness centrality and community detection (e.g., Louvain algorithm), thereby determining the importance of the target knowledge unit and the current document in the corresponding topological structure of the knowledge graph. Among them, betweenness centrality can measure the degree to which a node in the network acts as a "hub", that is, how many shortest paths pass through the node. The Louvain algorithm can identify the associated communities of documents or knowledge units in the knowledge graph.

[0167] Exemplarily, the dynamic association weight can be determined according to the following formula: dynamic association weight = α*(user behavior weight) + β*(semantic weight) + γ*(structural weight), where α, β, and γ are parameters, which can be 0.6, 0.3, and 0.1 respectively.

[0168] For example, the server may set a scheduled task to regularly update the associated weights, for example, asynchronously update every hour.

[0169] In the above process, the association weights can be automatically updated based on user behavior data (click frequency, association modification) and the semantics between documents and knowledge, thereby automatically updating the knowledge graph to build a more accurate knowledge graph that is more in line with the current actual situation.

[0170] For S660, in one embodiment, the server can use a graph structure dynamic update engine to determine the node structure and edge structure corresponding to the knowledge unit and document (for example, it can include association type, association weight, metadata and other information), and then build a relationship view based on the node structure and edge structure.

[0171] For example, the node structure may include the following fields:

[0172]

[0173] For example, the edge structure may contain the following fields:

[0174]

[0175] The knowledge association triple is: (source)-[edge]->(target), where source represents the active knowledge unit initiating the association, for example, a paragraph in the current document being edited by the user; target represents the targeted knowledge unit in the current association, i.e., the passive knowledge unit being associated, such as a historical knowledge unit referenced in a historical document, such as a historical paragraph; and edge represents the association relationship and its attributes, such as the association type and weight. In the example "target":"K_20230325_007" above, "K" stands for "Knowledge Unit," the identifier of the knowledge unit; "20230325" indicates that the knowledge unit was created on March 25, 2023; and "007" indicates that it is the seventh knowledge unit created that day. In a real-world scenario, when a user drags a paragraph P from a knowledge base to a new note N, source is the paragraph ID (Identity Document) of the new note N, and target is the ID of the dragged paragraph P (for example, K_20230325_007).

[0176] For example, the server can optimize the rendering of the knowledge graph through the Web Graphics Library (WebGL) and Level of Detail (LOD). For example, the knowledge graph can be dynamically loaded through the real-time rendering function of 100,000 nodes provided by WebGL, and the rendering accuracy can be automatically switched when the knowledge graph is scaled through LOD.

[0177] Exemplarily, the server may determine the node repulsion and edge attraction corresponding to the relationship view based on the dynamic association weight and the semantic similarity between the target knowledge unit and the current document; and determine the relationship view based on the node repulsion and edge attraction.

[0178] For example, node repulsion can be determined as 1 / (semantic similarity^2 + 0.1), and edge attraction can be determined as dynamic association weight * 5. Node repulsion refers to the repulsive force between nodes, simulating the repulsion of like charges in the physical world; edge attraction refers to the attraction of edges to connected nodes, simulating the spring or gravity phenomenon in the physical world. Node repulsion and edge attraction can make graph layouts clearer and more aesthetically pleasing by simulating the interactions in physical systems.

[0179] For example, the server may determine that edges corresponding to different association weights are displayed in different forms. For example, the greater the weight, the darker the color of the edge or the thicker the line.

[0180] For example, the server can implement the interaction logic between the user and the knowledge graph through the following code:

[0181]

[0182] Among them, the function on_node_click can be triggered when the user clicks a node; reverse association mining means that all edges directly connected to the current node can be queried through the graph database; generating a local subgraph means that a local subgraph can be generated based on the clicked node, and the local subgraph can include its neighbor nodes; linked directory indexing means updating the sorting weight of the directory or index based on the subgraph node list of the generated local subgraph.

[0183] In the above process, the server can generate an intuitive and beautiful relationship view based on the node structure and edge structure corresponding to the knowledge unit and document as well as a variety of rendering technologies; moreover, it can also update the relationship view in real time according to the changes in the association weights to ensure that the relationship view is most in line with the current actual situation and improve the accuracy of the information contained in the relationship view.

[0184] In one embodiment, the server can determine the parent-child relationship of the node directory in the document where the knowledge unit is located, and generate a tree-structured index view by hierarchically traversing the graph database (used to store documents, knowledge units, and knowledge graphs). This can be achieved through the following code:

[0185]

[0186] The code indicates that we can first start from the specified root node (root) based on the directory parent-child relationship (CHILD_OF) of the knowledge units in the document, recursively traverse all child nodes, return the path (path) from the root node to each leaf node, and arrange them in descending order by path length to generate a complete hierarchical tree structure and obtain an index view.

[0187] Exemplarily, the server can implement bidirectional path query and reverse association mining for the index view through the following code, that is, in response to a trigger operation on the second knowledge unit, it can determine the third document associated with the second knowledge unit (specifically, it can be matched from a graph database or a knowledge graph) and highlight the third document, or, in response to a trigger operation on the third document, it can determine the second knowledge unit associated with the third document (specifically, it can be matched from a graph database or a knowledge graph) and highlight the second knowledge unit.

[0188]

[0189] Among them, MATCH path = (a)-[*1..2]-(b) means matching all paths from node a to node b; WHERE a.id IN$selected_nodes means limiting the identifier of the starting node a to be in the given node list ($selected_nodes); AND b.tags CONTAINS "color design" means that the target node b contains the tag "color design"; RETURN path means returning all matching paths.

[0190] Exemplarily, the server can determine the ranking score of each third document in the index view based on the dynamic association weight, the topic similarity between each third document and the corresponding knowledge unit in the index view, and at least one of the time decay factors of each third document; thereafter, the display position of each third document in the index view can be determined based on each ranking score and the dynamic association weight to construct the index view, wherein the value of the time decay factor decreases with time, and its initial value can be determined based on the creation time of the document or knowledge unit.

[0191] For example, the topics of the third document and the corresponding knowledge unit can be extracted first; then the topics of the third document and the corresponding knowledge unit can be vectorized to obtain topic vectors; thereafter, the similarity between the topic vectors can be calculated to obtain topic similarity.

[0192] For example, a weighted sum of at least one of the dynamic association weight, the topic similarity, and the time decay factor may be performed to determine the ranking score.

[0193] For example, the ranking score may be determined by the following formula: ranking score=0.65*association strength+0.35*time decay factor.

[0194] Here, you can first use get_edges_by_directory(current_directory) to obtain all edges of the third document in the directory, that is, to obtain all documents related to the third document; then, sum the dynamic association weights (e.dynamicweight) of all related edges of the target node (e.target == node.id), that is, the third document, to obtain the association strength. The specific code is as follows:

[0195]

[0196] For example, you can calculate the time decay factor based on the creation time of a node (knowledge unit or document) to reduce the weight of old nodes. You can first calculate the difference between the current time and the node creation time using delta_days = (now() - node.create_time).days, in days. Then, determine the time decay factor using the decay formula: 1 / (1 + 0.1 * delta_days). The specific code is as follows:

[0197]

[0198] Exemplarily, the above-mentioned determination of the display position of each third document in the index view based on each ranking score and dynamic association weight includes: combining documents in each third document whose corresponding dynamic association weight is greater than a first threshold (for example, 0.7) to obtain a first association group; combining documents in each third document whose corresponding dynamic association weight is less than or equal to the first threshold and greater than a second threshold (for example, 0.4) and whose corresponding semantic similarity is greater than a third threshold (for example, 0.6) to obtain a second association group; combining documents in each third document whose corresponding dynamic association weight is less than or equal to the second threshold and whose corresponding interaction frequency is greater than a fourth threshold to obtain a third association group; determining the corresponding position of each third document in the first association group, the second association group or the third association group according to each ranking score; determining the positions of the first association group, the second association group and the third association group in the index view; and determining the display position of each third document in the index view according to the corresponding position of each third document in the corresponding association group and the position of each association group in the index view.

[0199] For example, Figure 4AAs shown, documents 1, 4, and 8 form the first association group, documents 5 and 7 form the second association group, and document 6 forms the third association group. The server can determine that the first, second, and third association groups are displayed sequentially in the index view from top to bottom. The server can also determine that the order of the corresponding positions of each third document in the first, second, or third association group is consistent with the ranking score of each third document. For example, because the ranking score of document 1 is greater than that of document 4, document 1 can be determined to be ranked before document 4.

[0200] In addition, the server may also divide the above association groups according to each ranking score. This process is similar to the above process of dividing the association groups according to the dynamic association weights, and this application will not elaborate on this.

[0201] In the above process, the server can generate an intuitive index view based on the directory relationship between knowledge units and documents and the retrieval technology; moreover, it can also update the index view in real time according to the changes in the association weights to ensure that the index view is most in line with the current actual situation and improve the accuracy of the information contained in the index view.

[0202] In one embodiment, the architecture diagram corresponding to the technical solution of the present application can be as follows: Figure 5D As shown, it can include an interaction layer, a computing layer, and a storage layer. Among them, for the interaction layer, the drag operation analysis can be used to analyze the user's drag operation and automatically establish an association with the knowledge base content (including historical documents); the visual query engine can be used to provide users with visual query functions (for example, query conversation content, keywords, etc.), and quickly view the association between documents such as notes through a graphical interface. For the computing layer, the graph query engine can be used to process queries on data in the graph database to help identify the relationship between the note content and the knowledge base; the knowledge-enhanced question-and-answer engine can be used to enhance the question-and-answer function in the knowledge base and enhance the intelligent retrieval of the note content; the dynamic index computing service can be used to update the system's index (for example, association weight, index of knowledge units in documents) in real time to optimize query speed and system response. For the storage layer, a graph database (for example, Neo4i) can be used to store the relationships between notes, optimizing content retrieval in a graph-structured manner. A vector database (for example, milvus) can be used to store high-dimensional data and embedded data of various artificial intelligence (AI) models to achieve faster similarity retrieval. Operation logs (for example, Elastic Search (ES)) can be used to record system operations and query logs for subsequent analysis and optimization.

[0203] Figure 7 A schematic diagram of a document processing device 700 provided in an embodiment of the present application is shown as follows: Figure 7As shown, the device 700 includes: a content acquisition module 701, a receiving and display module 702, a knowledge reference module 703, a dialogue display module 704, a data addition module 705, a graph display module 706, an identification determination module 707, a graph retrieval module 708, a graph zoom module 709, and a knowledge slicing module 710.

[0204] In one embodiment, a content acquisition module 701 is used to acquire interactive content for the current document, and the interactive content includes at least one of the following: the user's question content, the document content of the current document; a receiving and displaying module 702 is used to receive and display the knowledge unit to be selected; a knowledge reference module 703 is used to respond to a trigger operation for a target knowledge unit in the knowledge unit to be selected, and display the current document with the target knowledge unit added at a preset position; wherein the knowledge unit to be selected is determined based on the interactive content and multiple historical knowledge units obtained by slicing each historical document.

[0205] Exemplarily, the dialog display module 704 is used to display a dialog box on the editing page for editing the current document; the content acquisition module 701 is specifically used to obtain the question content entered by the user in the dialog box; the receiving display module 702 is specifically used to display the knowledge unit to be selected in the dialog box.

[0206] Exemplarily, the knowledge reference module 703 is specifically configured to respond to a drag operation on a target knowledge unit, add the target knowledge unit at the end of the drag operation in the current document, and display the current document.

[0207] Exemplarily, the preset position includes any one of the following: a current editing position for editing a current document; a position determined based on a semantic analysis between the document content of the current document and a target knowledge unit.

[0208] Exemplarily, the data adding module 705 is used to add the semantic tag of the target knowledge unit as metadata to the storage unit corresponding to the current document; wherein the semantic tag of the target knowledge unit is used to represent the semantic information of the target knowledge unit.

[0209] Exemplarily, the graph display module 706 is used to obtain a display operation for the knowledge graph; in response to the display operation, the knowledge graph is received, and at least one of the relationship view and the index view in the knowledge graph is displayed; wherein the relationship view represents the association relationship between each knowledge unit and each document based on multiple nodes and edges connecting multiple nodes, and each knowledge unit is obtained by slicing each document; the index view represents the association relationship between each knowledge unit and each document based on a directory display format.

[0210] Exemplarily, the identifier determination module 707 is used to determine the knowledge graph display identifier, and the knowledge graph display identifier is used to trigger the display of the knowledge graph; in the current document, the knowledge graph display identifier is displayed at the identifier display position of the target knowledge unit; the graph display module 706 is specifically used to obtain the trigger operation for the knowledge graph display identifier at the identifier display position.

[0211] Exemplarily, the knowledge graph is a knowledge graph involving target knowledge units.

[0212] Exemplarily, the graph retrieval module 708 is used to display a knowledge graph retrieval box; obtain the keywords to be retrieved in the knowledge graph retrieval box, and the keywords to be retrieved are used to search the knowledge graph; receive and display the retrieval results, and the retrieval results are determined based on the knowledge graph and the retrieval keywords.

[0213] Exemplarily, the graph zoom module 709 is used to zoom in or out the knowledge graph in response to a zoom-in display operation or a zoom-out display operation on the knowledge graph.

[0214] Exemplarily, the nodes in the relationship view correspond to individual documents, and the edges in the relationship view correspond to the association types between two documents and the associated knowledge units.

[0215] Exemplarily, the association type includes at least one of the following: a reference relationship, a supplementary relationship, and a contradictory relationship.

[0216] For example, the display styles of edges corresponding to different association types are different.

[0217] Illustratively, the graph display module 706 is configured to determine and display knowledge units associated with two documents corresponding to a target edge in response to a trigger operation on the target edge in the relationship view.

[0218] Exemplarily, the graph display module 706 is used to determine the associated document group involving the first knowledge unit in each document, where the first knowledge unit is any knowledge unit in the first document in each document or any knowledge unit in the same knowledge base; determine the document association identifier, which is used to identify that the corresponding document involves the first knowledge unit; and use the document association identifier to identify each document in the associated document group.

[0219] Exemplarily, the index view includes each second knowledge unit displayed in a directory presentation format and at least one third document in each document associated with each second knowledge unit; wherein the second knowledge unit is a knowledge unit in any second document in each document or any knowledge unit in the same knowledge base.

[0220] Exemplarily, the graph display module 706 is configured to connect the second knowledge unit with the corresponding third document in the index view based on a preset graph.

[0221] Exemplarily, the graph display module 706 is configured to receive knowledge unit attribute data of each second knowledge unit and document attribute data of each second document; and display each knowledge unit attribute data and document attribute data in an index view.

[0222] Exemplary knowledge unit attribute data include: the number of times the corresponding knowledge unit is cited, the corpus information to which it belongs, the corresponding identifier, and the paragraph information in the corresponding document; document attribute data include: the number of documents associated with the corresponding document.

[0223] Illustratively, the graph display module 706 is configured to highlight a third document associated with the second knowledge unit in response to a triggering operation on the second knowledge unit.

[0224] Exemplarily, the graph display module 706 is configured to highlight the second knowledge unit associated with the third document in response to a triggering operation on the third document.

[0225] Exemplarily, the knowledge slicing module 710 is used to obtain a slice viewing operation for the fourth document; in response to the slice viewing operation, multiple knowledge units contained in the fourth document are displayed on the knowledge slicing page.

[0226] Exemplarily, the knowledge slicing module 710 is used to obtain a newly added knowledge unit in response to a triggering operation of an add slice button in a knowledge slicing page; and add the newly added knowledge unit to the fourth document.

[0227] Exemplarily, the knowledge slicing module 710 is used to obtain a slice viewing operation for the fourth document in the knowledge graph.

[0228] Exemplarily, the knowledge slicing module 710 is used to obtain a slice viewing operation for the fourth document in the personal knowledge base page.

[0229] Figure 8 A schematic diagram of another document processing device 800 provided in an embodiment of the present application is shown as follows: Figure 8 As shown, the device 800 includes: a content receiving module 801, a knowledge recommendation module 802, a knowledge sending module 803, a type determination module 804, a weight calculation module 805, a graph construction module 806, and a graph sending module 807.

[0230] In one embodiment, the content receiving module 801 is used to receive interactive content for the current document, and the interactive content includes at least one of the following: the user's question content, the document content of the current document; the knowledge recommendation module 802 is used to determine the knowledge unit to be selected based on the interactive content and multiple historical knowledge units obtained by slicing each historical document; the knowledge sending module 803 is used to send the knowledge unit to be selected to the client, so that the client displays the current document with the target knowledge unit in the knowledge unit to be selected added at a preset position.

[0231] Exemplarily, the knowledge recommendation module 802 is specifically used to: convert the interactive content into a word sequence; process the word sequence through a pre-trained semantic association model to obtain a first sequence containing contextual information of the interactive content; compress the first sequence to obtain a text vector; search multiple historical knowledge units based on the text vector, and determine the knowledge unit to be selected from the multiple historical knowledge units.

[0232] Exemplarily, the knowledge recommendation module 802 is specifically used to: construct an index library for approximate nearest neighbor search based on multiple historical knowledge units; determine the number of knowledge units contained in the knowledge unit to be selected; and search the index library for the number of knowledge units to be selected that are most similar to the text vector.

[0233] Exemplarily, the knowledge recommendation module 802 is specifically configured to fine-tune the semantic association model using a plurality of preset documents under the document type of the current document.

[0234] Exemplarily, the knowledge recommendation module 802 is specifically used to: perform semantic analysis between the document content of the current document and the target knowledge unit, determine the specified content in the document content that is most semantically similar to the target knowledge unit; determine the adjacent position of the position of the specified content in the current document as a preset position; send the preset position to the client, so that the client adds and displays the target knowledge unit at the preset position.

[0235] Exemplarily, the type determination module 804 is used to determine the association type between the target knowledge unit and the current document; the weight calculation module 805 is used to calculate the association weight between the target knowledge unit and the current document; the graph construction module 806 is used to determine the knowledge graph including the relationship view and the index view based on the association type and the association weight; the graph sending module 807 is used to send the knowledge graph to the client so that the client displays the knowledge graph; wherein, the relationship view represents the association relationship between each knowledge unit and each document based on multiple nodes and edges connecting multiple nodes, and each knowledge unit is obtained by slicing each document; the index view represents the association relationship between each knowledge unit and each document based on a directory display form.

[0236] Exemplarily, the type determination module 804 is specifically used to determine the associated words that appear in the target knowledge unit and the current document respectively to represent the association relationship within a preset number of character windows; calculate the structural features and interaction features of the target knowledge unit and the current document; and determine the association type between the target knowledge unit and the current document through a preset classifier based on the associated words, structural features, interaction features, and feature weights; wherein the structural features represent the content creation time difference and the directory level distance between the target knowledge unit and the current document; the interaction features represent the user's modification record of the association relationship between the target knowledge unit and the current document; and the feature weight represents the influence weight of each feature on the classifier's determination of the association type.

[0237] Exemplarily, the feature weight is determined based on at least one of the following: semantic similarity between the target knowledge unit and the current document; and user preference information regarding historical association types.

[0238] Exemplarily, the weight calculation module 805 is specifically used to: calculate the label overlap between the target knowledge unit and the current document's respective semantic labels, where the semantic labels are used to represent the semantic information of the corresponding knowledge unit or document; and determine the association weight based on the label overlap.

[0239] Exemplarily, the weight calculation module 805 is specifically used to: extract keywords, entities, and subject categories of the target knowledge unit; and generate semantic tags of the target knowledge unit based on the extraction results.

[0240] An exemplary semantic tag of a target knowledge unit includes a document tag of a historical document or a current document corresponding to the target knowledge unit.

[0241] Exemplarily, the weight calculation module 805 is specifically used to: calculate the first label quantity of the semantic labels shared between the target knowledge unit and the current document; calculate the second label quantity of the semantic labels of the target knowledge unit and the current document respectively; and determine the label overlap based on the first label quantity and the second label data.

[0242] Exemplarily, the weight calculation module 805 is specifically used to: calculate the semantic similarity between the target knowledge unit and the current document; and determine the association weight according to the label overlap and the semantic similarity.

[0243] Exemplarily, the weight calculation module 805 is specifically used to: calculate the user behavior weight of the user for the target knowledge unit and the current document; calculate the semantic weight, which represents the semantic similarity between the target knowledge unit and the current document; calculate the structural weight, which represents the importance of the target knowledge unit and the current document in the corresponding topological structure of the knowledge graph; determine the dynamic association weight under each target duration based on the user behavior weight, semantic similarity and structural weight under each target duration; and update the association weight accordingly based on the dynamic association weight under each target duration.

[0244] Exemplarily, the weight calculation module 805 is specifically used to calculate the user behavior weight based on the user's trigger behavior data for the target knowledge unit and the current document; the trigger behavior data includes at least one of the following: associated operation data and access data for the target knowledge unit and the current document.

[0245] Exemplarily, the weight calculation module 805 is specifically used to: perform vector representation on the semantic information of the target knowledge unit and the current document to obtain the semantic vectors of the target knowledge unit and the current document; determine the first semantic weight based on the similarity of the two semantic vectors; determine the lexical probability distribution of the vocabulary involved in the target knowledge unit and the current document; determine the topic of the target knowledge unit and the current document based on the lexical probability distribution; determine the second semantic weight based on the similarity of the two topics; determine the semantic weight based on the first semantic weight and the second semantic weight.

[0246] Exemplarily, the graph construction module 806 is specifically used to: determine the node repulsion and edge attraction corresponding to the relationship view based on the dynamic association weight and the semantic similarity between the target knowledge unit and the current document; determine the relationship view based on the node repulsion and edge attraction.

[0247] Exemplarily, the graph construction module 806 is specifically used to: determine the ranking score of each third document in the index view based on the dynamic association weight, the subject similarity between each third document and the corresponding knowledge unit in the index view, and at least one of the time decay factors of each third document; determine the display position of each third document in the index view based on each ranking score and the dynamic association weight to construct the index view, wherein the value of the time decay factor decreases as time increases.

[0248] Exemplarily, the graph construction module 806 is specifically used to: combine the documents in each third document whose corresponding dynamic association weight is greater than the first threshold to obtain a first association group; combine the documents in each third document whose corresponding dynamic association weight is less than or equal to the first threshold and greater than the second threshold, and whose corresponding semantic similarity is greater than the third threshold, to obtain a second association group; combine the documents in each third document whose corresponding dynamic association weight is less than or equal to the second threshold and whose corresponding interaction frequency is greater than the fourth threshold to obtain a third association group; determine the corresponding position of each third document in the first association group, the second association group or the third association group according to each ranking score; determine the positions of the first association group, the second association group and the third association group in the index view; determine the display position of each third document in the index view according to the corresponding position of each third document in the corresponding association group and the position of each association group in the index view.

[0249] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here.

[0250] The above description of the apparatus 800 and 700 of the embodiment of the present application is based on the functional module in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in the form of hardware, can be implemented by instructions in the form of software, or can be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above-mentioned document processing method embodiment in conjunction with its hardware.

[0251] Figure 9 A schematic diagram of an electronic device 900 provided in an embodiment of the present application.

[0252] like Figure 9 As shown, the electronic device 900 may include:

[0253] The memory 910 and the processor 920 are configured to store computer programs and transmit the program code to the processor 920. In other words, the processor 920 can call and run the computer program from the memory 910 to implement the method in the embodiment of the present application.

[0254] For example, the processor 920 may be configured to execute the above method embodiments according to instructions in the computer program.

[0255] In some embodiments of the present application, the processor 920 may include but is not limited to:

[0256] General-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.

[0257] In some embodiments of the present application, the memory 910 includes but is not limited to:

[0258] Volatile memory and / or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAMbus RAM (DR RAM).

[0259] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 910 and executed by the processor 920 to implement the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0260] like Figure 9 As shown, the electronic device may further include:

[0261] The transceiver 930 may be connected to the processor 920 or the memory 910 .

[0262] The processor 920 may control the transceiver 930 to communicate with other devices. Specifically, the processor 920 may send information or data to other devices or receive information or data sent by other devices. The transceiver 930 may include a transmitter and a receiver. The transceiver 930 may further include one or more antennas.

[0263] It should be understood that the various components in the electronic device are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.

[0264] The present application also provides a computer storage medium having a computer program stored thereon, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment. In other words, the present application also provides a computer program product containing instructions, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment.

[0265] When software is used to implement, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instruction is loaded and executed on a computer, the computer can be made to perform the corresponding flow in each method in the embodiment of the present application, generate the function that each method in the embodiment of the present application can realize in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instruction can be stored in a computer-readable storage medium, or transmitted from a computer-readable storage medium to another computer-readable storage medium. For example, the computer instruction can be transmitted from a website, computer, server, or data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website, computer, server, or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital video disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).

[0266] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0267] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the system, device or module can be electrical, mechanical or other forms.

[0268] Modules described as separate components may or may not be physically separate, and components displayed as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected based on actual needs to achieve the purpose of the present embodiment. For example, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module.

Claims

1. A document processing method, characterized in that: include: Acquire interactive content for the current document, where the interactive content includes at least one of the following: a question asked by a user and the document content of the current document; receiving and displaying knowledge units to be selected; In response to a triggering operation on a target knowledge unit among the knowledge units to be selected, displaying a current document with the target knowledge unit added at a preset position; The knowledge unit to be selected is determined based on the interactive content and multiple historical knowledge units obtained by slicing each historical document.

2. The method according to claim 1, characterized in that Before obtaining the interactive content for the current document, the process further includes: Displaying a dialog box on the editing page for editing the current document; Accordingly, obtaining the interactive content for the current document includes: Obtaining the question content entered by the user in the dialog box; Accordingly, the display of the knowledge unit to be selected includes: The knowledge unit to be selected is displayed in the dialog box.

3. The method according to claim 1, characterized in that The step of displaying the current document with the target knowledge unit added at a preset position in response to a triggering operation on the target knowledge unit among the knowledge units to be selected includes: In response to the drag operation on the target knowledge unit, the target knowledge unit is added to the current document at the end of the drag operation, and the current document is displayed.

4. The method according to claim 1, wherein The preset position includes any one of the following: A current editing position for editing the current document; The position is determined based on semantic analysis between the document content of the current document and the target knowledge unit.

5. A document processing method, characterized in that: include: Receiving interactive content for a current document, the interactive content including at least one of the following: a user's question content, and the document content of the current document; Determining a knowledge unit to be selected based on the interactive content and multiple historical knowledge units obtained by slicing each historical document; The knowledge units to be selected are sent to the client, so that the client displays the current document with the target knowledge unit among the knowledge units to be selected added at a preset position.

6. A document processing device, characterized in that: include: A content acquisition module is used to acquire interactive content for the current document, wherein the interactive content includes at least one of the following: a question asked by a user, and the document content of the current document; A receiving and displaying module, configured to receive and display the knowledge unit to be selected; a knowledge reference module, configured to display a current document with the target knowledge unit added at a preset position in response to a triggering operation on a target knowledge unit among the knowledge units to be selected; The knowledge unit to be selected is determined based on the interactive content and multiple historical knowledge units obtained by slicing each historical document.

7. A document processing device, characterized in that: include: A content receiving module, configured to receive interactive content for a current document, wherein the interactive content includes at least one of the following: a question asked by a user, and the document content of the current document; A knowledge recommendation module, configured to determine a knowledge unit to be selected based on the interactive content and a plurality of historical knowledge units obtained by slicing each historical document; The knowledge sending module is used to send the knowledge units to be selected to the client, so that the client displays the current document with the target knowledge unit in the knowledge units to be selected added at a preset position.

8. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to perform the method according to any one of claims 1 to 5 by executing the executable instructions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

10. A computer program product comprising instructions, characterized in that When the computer program product is run on an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 5.