Document updating method and device, equipment and medium

By acquiring and updating the document structure map, identifying and processing the updated data of electronic documents, the problem of global content modification in electronic documents is solved, adaptive updates of global content are achieved, and the accuracy of document updates is improved.

CN121833862APending Publication Date: 2026-04-10INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2025-12-24
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing technologies, electronic documents can only be modified locally when adding content, which makes it impossible to adapt the global content and results in the loss of corresponding content in the electronic document.

Method used

By obtaining the document structure graph of the initial document, identifying the target user's document update data, determining the content and data type to be processed, identifying the target node based on the data type, updating the document structure graph, and generating the target document.

Benefits of technology

It enables adaptive modification of the global content of electronic documents, improving the accuracy of document updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833862A_ABST
    Figure CN121833862A_ABST
Patent Text Reader

Abstract

The invention discloses a document updating method and device, equipment and a medium. The method comprises the steps that an initial document and a document structure atlas corresponding to the initial document are obtained, atlas nodes in the document structure atlas comprise contents and element data types of document elements, and atlas edges comprise structure relations among the document elements; obtaining document updating data corresponding to the target user; according to the document update data, identifying at least one to-be-processed content corresponding to the document update data and to-be-processed data types corresponding to the to-be-processed contents; obtaining an element data type corresponding to each to-be-processed data type; for each to-be-processed data type, determining at least one target node according to the element data type corresponding to the to-be-processed data type; updating data associated with each target node in the document structure atlas; and generating a target document according to the updated document structure map. According to the embodiment of the invention, the accuracy of document updating can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a document updating method, apparatus, device, and medium. Background Technology

[0002] With the rapid development of technology, the frequency of electronic document applications is gradually increasing, and users can store data to be stored in electronic documents.

[0003] Currently, when adding content to electronic documents, only partial content can be added, which prevents the global content of the electronic document from being adapted accordingly, resulting in the loss of corresponding content in the electronic document. Summary of the Invention

[0004] This invention provides a document updating method, apparatus, device, and medium to improve the accuracy of document updates.

[0005] In a first aspect, embodiments of the present invention provide a document updating method, the method comprising:

[0006] Obtain the initial document and its corresponding document structure graph. The graph nodes in the document structure graph include the content and data type of the document elements, and the graph edges include the structural relationships between the document elements.

[0007] Retrieve document update data for the target user;

[0008] Based on the document update data, identify at least one piece of content to be processed corresponding to the document update data, and the data type to be processed corresponding to each piece of content to be processed;

[0009] Retrieve the element data type corresponding to each data type to be processed;

[0010] For each data type to be processed, at least one target node is determined based on the element data type corresponding to the data type to be processed.

[0011] Update the data associated with each target node in the document structure graph;

[0012] Generate the target document based on the updated document structure graph.

[0013] Secondly, embodiments of the present invention also provide a document updating device, the device comprising:

[0014] The graph acquisition module is used to acquire the initial document and the document structure graph corresponding to the initial document. The graph nodes in the document structure graph include the content and data type of the document elements, and the graph edges include the structural relationships between the document elements.

[0015] Update the data acquisition module to acquire document update data for the target user;

[0016] The data recognition module is used to identify at least one piece of content to be processed corresponding to the document update data, and the data type to be processed corresponding to each piece of content, based on the document update data.

[0017] The type determination module is used to obtain the element data type corresponding to each data type to be processed;

[0018] The node determination module is used to determine at least one target node for each data type to be processed, based on the element data type corresponding to the data type to be processed.

[0019] The data update module is used to update the data associated with each target node in the document structure graph;

[0020] The document generation module is used to generate target documents based on the updated document structure graph.

[0021] Thirdly, embodiments of the present invention also provide a document updating device, the document updating device comprising:

[0022] At least one processor; and

[0023] A memory that is communicatively connected to at least one processor; wherein,

[0024] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the document update method of any embodiment of the present invention.

[0025] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the document update method of any embodiment of the present invention.

[0026] The technical solution of this invention involves obtaining an initial document and its corresponding document structure graph. The graph nodes in the document structure graph include the content and data type of document elements, and the graph edges include the structural relationships between document elements. The solution then proceeds to: obtaining document update data corresponding to the target user; identifying at least one content to be processed corresponding to the document update data, and the data type to be processed corresponding to each content; obtaining the data type of each data type to be processed; determining at least one target node for each data type to be processed based on the data type of the corresponding data type; updating the data associated with each target node in the document structure graph; and generating a target document based on the updated document structure graph. Different document structure graph update methods are executed for different types of document update data, thus updating the document structure graph and adaptively modifying the global content of the initial document, thereby improving the accuracy of document updates.

[0027] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a flowchart of a document update method provided in Embodiment 1 of the present invention;

[0030] Figure 2 This is a flowchart of a document update method provided in Embodiment 2 of the present invention;

[0031] Figure 3 This is a structural diagram of a document updating device according to an embodiment of the present invention;

[0032] Figure 4 This is a schematic diagram of the structure of a document update device provided in an embodiment of the present invention. Detailed Implementation

[0033] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or elements is not necessarily limited to those explicitly listed, but may include other steps or elements not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0035] The acquisition, storage, and application of initial documents and other related matters in the technical solutions of this invention comply with relevant laws and regulations and do not violate public order and good morals.

[0036] Example 1

[0037] Figure 1 This is a flowchart illustrating a document update method according to Embodiment 1 of the present invention. This embodiment of the invention is applicable to document update scenarios, and the method can be executed by a document update device, which can be implemented in hardware and / or software.

[0038] See Figure 1 The document update methods shown include:

[0039] S101. Obtain the initial document and the document structure graph corresponding to the initial document. The graph nodes in the document structure graph include the content and data type of the document elements, and the graph edges include the structural relationships between the document elements.

[0040] Among them, the initial document can be the original document to be updated with data. The document structure graph can be a graph that structures the data of the initial document and describes the relationships between document elements. The document structure graph consists of graph nodes and graph edges. A graph node can be the smallest independent content unit in the initial document. Each graph node contains two key pieces of information: the content of the document element (i.e., the specific text, picture, or table corresponding to the graph node), and the element data type (i.e., the category of the content of the document element stored in the graph node). A document element can be a basic component that makes up a document and is the entity corresponding to a graph node. Document elements can include: titles, charts, or text blocks, etc. A graph edge can be a "relationship line" in the document structure graph that connects two graph nodes and is used to describe the logical or structural relationship between document elements.

[0041] Specifically, obtain the initial document, extract the original content and basic format information of the initial document through a document parsing tool (such as font size, paragraph indentation, or picture position, etc.). Based on the basic format information and semantic rules, split the original content into independent document elements. Example: The format rule can be that the font is larger than a preset value and bold, then the document element is a chapter title; the paragraph is indented by a preset number of characters, then the document element is a text block; the icon position area is recognized, then the document element is a chart. The semantic rule can be that the content starting with "1. or 一、" is a first-level chapter title, and the content starting with "2.2 or (2)" is a second-level chapter title; if it contains technical parameters (such as "contract terms"), then the document element is a text block. According to the format recognition result and semantic recognition result corresponding to the text element, establish a mapping relationship between each document element and the element data type. Based on each document element and the mapping relationship between the document element and the element data type, generate a document structure graph. Example: Convert each document element with an element data type label into a graph node (i.e., the node attributes include the content of the document element and the element data type); establish a "graph edge" between the corresponding nodes through the title hierarchy corresponding to the element data type (such as "3.2.1 third-level title" belongs to the "3.2" second-level title); sort the adjacent graph nodes of the same type in order according to the order of the document elements in the document; integrate the graph nodes and graph edges into a visual document structure graph.

[0042] S102. Obtain the document update data corresponding to the target user.

[0043] Among them, the target user can be the user who performs editing operations on the initial document, which can be an individual or a certain editor in a collaborative scenario. The document update data can be all the modification operations and corresponding content initiated by the target user on the initial document, which is the input that triggers the update of the document structure graph and can include the operation type and the modified content.

[0044] Specifically, through interactive monitoring technology in the document editing interface, all operations of the target user are recorded in real time. Monitoring methods include front-end monitoring, i.e., listening for mouse or keyboard events (such as clicking the "Add Title" button, entering text, pasting images, deleting content, dragging and adjusting paragraph positions, etc.) in the web page or client-side editing tool; and operation log monitoring: temporarily storing each user's operation in the format of "timestamp + username + operation behavior" (to avoid operation loss), and then performing structured parsing on the original operation log. Illegal operations can be filtered out; for example, if the format of a newly added title does not conform to the manual specifications or the updated document data contains sensitive words, an alert can be issued.

[0045] S103. Based on the document update data, identify at least one content to be processed corresponding to the document update data, and the data type to be processed corresponding to each content to be processed.

[0046] The content to be processed can be document elements extracted from document update data that require updating the document structure graph. The data type to be processed can be the element category corresponding to the content to be processed.

[0047] Specifically, redundant information (such as timestamps or usernames) in document update data can be filtered out, retaining only the specific content that needs to be modified or added. If the document update data contains multiple document elements (such as adding both a title and a text block simultaneously), it is split into multiple independent pieces of content to be processed. The "formatting rules + semantic rules" generated during the previous document structure graph generation are invoked to perform preliminary type matching on each piece of content to be processed, resulting in a list of content to be processed with preliminary type labels. For example, the content to be processed is 1.3 Bank Contract Terms (only one document element, no splitting required). According to the formatting rule "bold and 14pt indicates a second-level heading," the content to be processed is determined to be a second-level heading. Applying the semantic rule that content starting with "XX" indicates a second-level heading, the content to be processed is determined to be a second-level heading. Based on the results of the formatting and semantic rules, the content to be processed in the document update data is determined to be 1.3 Bank Contract Terms; the data type corresponding to the content to be processed is a second-level heading.

[0048] S104. Obtain the element data type corresponding to each data type to be processed.

[0049] Specifically, for each data type to be processed, the data type to be processed is matched with the element data types of each node in the document structure graph. The element data type in the graph nodes that matches the data type to be processed is identified as the element data type corresponding to the data type to be processed. If a completely identical type is found (e.g., the data type to be processed is "second-level chapter title" = the element data type "second-level chapter title" in the graph node), the match is successful, and the element data type corresponding to the data type to be processed has been found. If no matching type is found after traversing all graph nodes, the match fails (an alarm is triggered). The search method includes, but is not limited to, depth-first traversal algorithms or breadth-first traversal algorithms; this embodiment of the invention does not impose any restrictions on this.

[0050] S105. For each data type to be processed, determine at least one target node based on the element data type corresponding to the data type to be processed.

[0051] The target node can be a node in the document structure graph that corresponds to the element data type and requires an update operation.

[0052] Specifically, based on the structural design of the document structure graph, preset node positioning rules are established for different element data types. For example, the positioning rules are as follows: Chapter titles are located by their parent nodes (e.g., the parent node of a second-level title is the corresponding first-level title node), and sibling nodes at the same level are used to determine the new position of nodes at the same level. Text blocks are located by finding the graph nodes corresponding to the content of document elements in the document structure graph that have semantically similar content to the text block. Charts are located by finding the graph nodes that reference the chart. According to the positioning rules, the nodes corresponding to the element data type of the data to be processed are traversed in the graph, and candidate nodes are filtered out: if it is a title type, the parent-level node of the candidate node is found and determined as the target node; if it is a text block type, the semantically matching candidate node is found and determined as the target node.

[0053] S106. Update the data associated with each target node in the document structure graph.

[0054] Specifically, if the data type to be processed is a chapter title, a new graph node (containing the content and data type of the document element) corresponding to the new title is generated; a graph edge is established between the new node and the target node (parent node); and the nodes associated with the target node and the relationships between the nodes are updated. If the data type to be processed is a chart, a new graph node (containing the content and type of the chart) corresponding to the chart is generated; a "reference edge" is established between the new node and the target node (the graph node that references the chart identifier); and the nodes associated with the target node and the relationships between the nodes are updated. If the data type to be processed is a text block, no new node needs to be generated; the content attributes of the target node are directly updated, the new text block is added to the content of the document element corresponding to the target node, and the content of the document element corresponding to the target node is updated.

[0055] S107. Generate the target document based on the updated document structure graph.

[0056] The target document can be a document generated based on the updated graph, preserving all structural relationships and content in the graph.

[0057] Specifically, the system reads the attributes (content and data type of document elements) and relationships (dependent, referenced, or parallel) of all nodes in the document structure graph. It then sorts the graph nodes according to the document's hierarchical logic (e.g., document title as first level, first-level headings as second level, second-level headings as third level, text blocks as fourth level, and charts as fifth level) while maintaining the order of content, resulting in a hierarchically sorted "node list." Based on this sorted list, the document generation tool is invoked to render each graph node as document content, starting with the document title, then the first-level headings, and so on downwards. Text blocks follow their respective headings, and charts are inserted after the document element content according to their connection positions in the document structure graph, resulting in a target document containing all content and basic formatting.

[0058] The technical solution of this invention involves obtaining an initial document and its corresponding document structure graph. The graph nodes in the document structure graph include the content and data type of document elements, and the graph edges include the structural relationships between document elements. The solution then proceeds to: obtaining document update data corresponding to the target user; identifying at least one content to be processed corresponding to the document update data, and the data type to be processed corresponding to each content; obtaining the data type of each data type to be processed; determining at least one target node for each data type to be processed based on the data type of the corresponding data type; updating the data associated with each target node in the document structure graph; and generating a target document based on the updated document structure graph. Different document structure graph update methods are executed for different types of document update data, thus updating the document structure graph and adaptively modifying the global content of the initial document, thereby improving the accuracy of document updates.

[0059] Example 2

[0060] Figure 2 This is a flowchart illustrating a document update method according to Embodiment 2 of the present invention. Based on the above embodiments, this embodiment optimizes and improves the document update operation.

[0061] Furthermore, the process of "updating the data associated with each target node in the document structure graph" is refined to "when the data type to be processed is a document title, and the document update data includes a new title, generate a graph node corresponding to the new title, and establish a graph edge between the graph node corresponding to the new title and the target node; when the data type to be processed is a chart, and the document update data includes a new chart, generate a graph node corresponding to the new chart, and establish a graph edge between the graph node corresponding to the new chart and the target node; when the data type to be processed is a text block, and the document update data includes new text content, update the content of the document element corresponding to the target node according to the new text content," in order to improve the document update operation.

[0062] It should be noted that for parts not described in detail in the embodiments of the present invention, please refer to the descriptions in other embodiments.

[0063] See Figure 2 The document update methods shown include:

[0064] S201. Obtain the initial document and the document structure graph corresponding to the initial document. The graph nodes in the document structure graph include the content and data type of the document elements, and the graph edges include the structural relationships between the document elements.

[0065] S202. Obtain the document update data corresponding to the target user.

[0066] S203. Based on the document update data, identify at least one content to be processed corresponding to the document update data, and the data type to be processed corresponding to each content to be processed.

[0067] S204. Obtain the element data type corresponding to each data type to be processed.

[0068] S205. For each data type to be processed, determine at least one target node based on the element data type corresponding to the data type to be processed.

[0069] S206. When the data type to be processed is a document title, and the document update data includes the addition of a new title, generate the graph node corresponding to the new title, and establish the graph edge between the graph node corresponding to the new title and the target node.

[0070] Specifically, key information about the newly added title is extracted from the document update data and used as attributes of the new graph node. For example, the content of the new title (e.g., "1.3 Bank Contract Terms") and the element data type (e.g., "Secondary Chapter Title") are extracted from the document update data. A new graph node is created in the document structure graph, and its attributes include the content of the new title and the element data type. The parent node of the graph node with the same element data type as the one to be processed is obtained and identified as the target node. Graph edges are established between the new graph node and the target node. A graph node with the document title type to be processed and its corresponding graph edge are added to the document structure graph.

[0071] S207. When the data type to be processed is a chart, and the document update data includes the addition of a new chart, generate the graph node corresponding to the new chart, and establish the graph edge between the graph node corresponding to the new chart and the target node.

[0072] Specifically, key information about the newly added charts is extracted from the document update data. This key information may include the chart identifier, chart content, and chart format. A new chart node is created in the document structure graph, with attributes including the chart identifier, chart content, and chart format. Based on the rules for finding graph nodes that reference the chart, the identifier of the newly added chart (e.g., "...") is extracted from the document update data. Figure 1-3 The process iterates through all nodes in the document structure graph, filtering out nodes whose content contains the chart identifier. These filtered nodes are the target nodes for the new chart (if multiple text elements reference the same chart, there will be multiple target nodes). Based on the "reference relationship" between the new chart and the target nodes, graph edges are created between the chart node corresponding to the new chart and the target node. Graph nodes of type "chart" to be processed and their corresponding graph edges are added to the document structure graph.

[0073] S208. When the data type to be processed is a text block, and the document update data includes newly added text content, update the content of the document element corresponding to the target node according to the newly added text content.

[0074] Specifically, the process involves retrieving plain text content from document update data. For example, extracting the new content "Bank Contract Terms Case XXX" from the update data. Based on the semantics of the new text content in the document update data, the process finds the corresponding graph node in the document structure graph that has the same semantic meaning as the new text content. This node is designated as the target node, and no new node is generated. All new text content is appended to the content of the target node. For example, if the target node currently contains "Contract Involves 22 Terms," ​​the new content is appended to the end, updating the content to "Contract Involves 22 Terms, Bank Contract Terms Case XXX," thus avoiding redundancy in text block nodes. The updated text content is then checked for grammatical errors, duplicate content, or formatting issues (such as missing punctuation) to ensure logical coherence. For example, the process verifies the connection between the new and existing content using punctuation; if there is no punctuation before the new content, a comma is automatically added.

[0075] S209. Generate the target document based on the updated document structure graph.

[0076] This invention, through its embodiments, acquires an initial document and its corresponding document structure graph. The document structure graph nodes include the content and data type of document elements, and the graph edges include the structural relationships between document elements. It then acquires document update data corresponding to a target user; based on the document update data, identifies at least one content to be processed corresponding to the document update data, and the corresponding data type to be processed for each content; acquires the data type of each data type to be processed; for each data type to be processed, determines at least one target node based on the corresponding data type of the data type; updates the data associated with each target node in the document structure graph; and generates a target document based on the updated document structure graph. Different document structure graph update methods are executed for different types of content to be processed, thus refining the document structure graph update methods and improving the accuracy of document structure graph acquisition.

[0077] Optionally, at least one target node is determined based on the element data type corresponding to the data type to be processed, including: when the data type to be processed is a document title, at least one parent node of the graph node with the same data type as the data type to be processed is found in the document structure graph and determined as the target node; when the data type to be processed is a chart, the graph node in the document structure graph whose content includes the icon corresponding to the chart is found in the document element content and determined as the target node; when the data type to be processed is a text block, the graph node in the document structure graph whose content is semantically similar to that of the text block is determined as the target node.

[0078] Specifically, when the data type to be processed is a document title, find graph nodes with the same data type, then find at least one parent node corresponding to these nodes, extract the parent nodes of the same type of nodes (e.g., the parent node of second-level heading 1.1 / 1.2 is first-level heading 1, and the parent node of first-level heading 1 is the main document title), select the parent node that matches the level of the new heading, and determine it as the target node (the parent node of the new second-level heading is first-level heading 1, not the main document title). When the data type to be processed is a chart, extract the chart identifier corresponding to the chart. For example, the unique identifier of the new chart is "". Figure 1-3 "Traverse the 'document element content' of all nodes in the graph, find nodes whose content contains the identifier," and filter out nodes containing " Figure 1-3 The node with the tag "" is identified as the target node. When the data type to be processed is a text block, the core semantics of the newly added text block are extracted. For example, the core keywords of the newly added content are "risk assessment, application rating and early warning". Semantically similar text block nodes are searched in the graph, and the semantic similarity between the newly added content and the content of existing text block nodes is calculated. The text block node with the highest semantic similarity (usually ≥70%) is selected as the target node. In the document structure graph, for the document title, if there is no node of the same type (such as adding the first second-level heading), the document's overall title is used as the target node; for the chart, if there is no text block referencing the tag, the title node of the chapter to which the newly added chart belongs is used as the target node; for the text block: if there are no semantically similar nodes, the title node of the chapter to which the newly added text belongs is used as the target node.

[0079] By refining the steps for determining target nodes and improving their accuracy, the following approach is adopted: when the data type to be processed is a document title, the target node is identified by searching for at least one parent node of a graph node in the document structure graph that has the same data type as the data type to be processed; when the data type to be processed is a chart, the target node is identified by searching for graph nodes in the document structure graph whose content includes the icon corresponding to the chart; and when the data type to be processed is a text block, the target node is identified by selecting graph nodes in the document structure graph whose content is semantically similar to that of the text block.

[0080] Optionally, update the content of the document element corresponding to the target node based on the newly added text content, including: querying the content of the document element corresponding to the target node for text paragraphs with the same semantic meaning as the newly added text content; and adding the newly added text content to the end of the text paragraphs.

[0081] Specifically, based on the added semantics of the new text content, for example, keywords (such as "risk assessment, credit rating, and early warning") are extracted from the new text content to form a semantic feature set; the complete content of the target node is split into independent paragraphs according to punctuation marks (period, semicolon, and line break); the semantics of each paragraph are obtained, and the similarity between each paragraph and the added semantics is calculated. For example, paragraphs with a similarity of ≥70% to the added semantics are considered text paragraphs consistent with the added semantics; the start and end positions of the paragraph in the target node content are recorded to determine the insertion point at the end of the paragraph. The new text content is added at the end of the matched paragraph (insertion point). If there is no punctuation at the end of the paragraph, punctuation marks (such as commas and periods) are added first before concatenation (ensuring grammatical correctness); the original "content of document element" attribute of the target node is replaced with the concatenated complete content.

[0082] By querying the content of the target node's document element to find text paragraphs with the same semantic meaning as the newly added text content, the new text content is added to the end of the text paragraphs. Texts with similar semantic meanings are stored in the same document location, ensuring the accuracy of document content generation.

[0083] Optionally, based on the newly added semantics of the new text content, the text paragraphs with the same semantics as the newly added text are queried in the content of the document element corresponding to the target node. This includes: obtaining professional term data, which includes at least one professional term and the text category corresponding to each professional term; identifying at least one professional term from the newly added text content; determining the newly added semantics corresponding to the newly added text content based on each professional term and the professional term data; and querying the text paragraphs with the same semantics as the newly added text in the content of the document element corresponding to the target node.

[0084] The specialized terminology data can be a predefined specialized thesaurus, containing core specialized terms within the industry and their corresponding text categories. Specialized terms can be words in the text that represent core domain concepts (distinguished from general vocabulary, and crucial for semantic recognition). Text categories can be the semantic classifications to which the specialized terms belong (used to define the text's thematic direction).

[0085] Specifically, the process involves retrieving "professional term-text category" mapping data from a specialized terminology database (which must match the document's domain, such as a banking and finance terminology database). A key requirement is that the specialized terminology data must cover the document's core domain. The newly added text content is then segmented to identify words that match the specialized terminology data (i.e., specialized terms), excluding generic words (such as "modify," "after," and "30 minutes"). Specialized term identification can be achieved through keyword matching or word segmentation algorithms. Each identified specialized term is matched with its corresponding text category to form the core semantics of the new text. The document element content of the target node is then split into paragraphs (by punctuation or line breaks). For each paragraph, specialized terms are identified, matched with the text category, and the paragraph's text semantics are defined. The similarity between the paragraph semantics and the newly added semantics is compared; higher similarity indicates a more accurate match. Paragraphs that match the core category of the newly added semantics are classified as text paragraphs.

[0086] By acquiring professional terminology data, which includes at least one professional term and the corresponding text category, at least one professional term is identified from the newly added text content. Based on each professional term and the professional terminology data, the newly added semantics corresponding to the newly added text content are determined. Based on the newly added semantics, text paragraphs with the same semantics as the newly added semantics are queried in the content of the document elements corresponding to the target node. By using professional terminology data, the text semantics can be determined, thereby improving the accuracy of text semantic determination.

[0087] Optionally, after generating the target document based on the updated document structure graph, the process further includes: obtaining the operation information corresponding to the target user, including: the current mouse position, document browsing direction, and document page movement speed; determining the prediction time period based on the document browsing direction and document page movement speed; calculating the prediction area based on the document browsing direction, document page movement speed, and prediction time period; determining the data to be rendered based on the current mouse position, prediction area, and target document; and rendering the data to be rendered according to the rendering format to display the target document.

[0088] The operational information can be real-time collected user behavior data while browsing the document. The current mouse position can be the real-time coordinates of the user's mouse pointer on the document page, reflecting the area the user is currently focusing on. The document browsing direction can be the user's movement trend while browsing the document, categorized as up, down, left, and right. The document page movement speed can be the speed at which the user browses the document, measured by "page switching time" or "scrolling distance or time." The predicted time period can be a "future time window" calculated based on browsing speed, used to predict the content the user will browse within that time period. The predicted area can be the document area the user is about to browse, calculated by combining browsing direction, speed, and the predicted time period. The data to be rendered can be the document content extracted from the target document and located within the predicted area. The rendering format can be a preset document content display format (such as font, font size, color, and layout).

[0089] Specifically, the document's front-end monitoring module collects three types of operation data in real time: the current mouse position, i.e., the pixel coordinates (x, y) of the mouse on the document page; the document browsing direction, determined by continuous scrolling or page-turning actions (e.g., continuous downward scrolling indicates a downward browsing direction; upward scrolling indicates an upward browsing direction); and the document page movement speed, i.e., calculating the amount of page movement per unit time. Example: turning from page 3 to page 5 in 3 seconds → moving 2 pages → speed = 2 / 3 pages / second. The length of the predicted time period is negatively correlated with browsing speed; the faster the user browses, the shorter the predicted time period; the slower the browsing, the longer the predicted time period, ensuring that the area within the predicted time period is content the user is "about to see," rather than an excessively distant invalid area. Combining the three parameters (browsing direction, speed, and predicted time period), the document area the user will reach within the predicted time period is calculated using the formula: Predicted area range = Current position + Browsing direction × Speed ​​× Predicted time period. Based on the calculated prediction area, all content data, including text and charts, of the corresponding area is extracted from the target document. A secondary filtering can be performed based on the current mouse position: if the mouse hovers near a title, the sub-content of that title is extracted first as the data to be rendered. The document rendering engine is invoked to process the data to be rendered according to a preset rendering format. The rendered data is seamlessly stitched with the currently displayed document content and displayed to the user in real time. The content of the prediction area can be rendered in advance, so users do not need to wait for loading when scrolling to that area, achieving "seamless browsing." If the prediction is inaccurate (the user has not viewed that area), the rendering resources for that area are released to avoid memory consumption.

[0090] By acquiring the target user's operation information, including the current mouse position, document browsing direction, and document page movement speed, a prediction time period is determined based on the document browsing direction and document page movement speed. A prediction area is calculated based on the document browsing direction, document page movement speed, and prediction time period. The data to be rendered is determined based on the current mouse position, prediction area, and target document. The data to be rendered is then rendered according to the desired rendering format, and the target document is displayed. The rendering area can be determined based on multi-dimensional data, and only the document content corresponding to the rendering area is rendered, reducing memory usage.

[0091] Optionally, after rendering the data to be rendered according to the rendering format and displaying the target document, the method further includes: dividing the target document to determine at least one region to be detected; for each region to be detected, obtaining at least one item to be detected corresponding to the region to be detected, as well as the data content and rendering format corresponding to each item to be detected; performing similarity detection on the data content of each item to be detected to obtain similarity detection results; performing consistency detection on the rendering format of each item to be detected to obtain consistency detection results; and determining the document verification result based on the similarity detection results and the consistency detection results.

[0092] The area to be checked can be an independent verification unit divided according to the document structure logic, and can be divided according to chapters, paragraphs, and the data type of document elements. The item to be checked can be the core object that needs to be verified within the area to be checked. The data content corresponding to the item to be checked can be the actual text information to be checked.

[0093] Specifically, based on the document's structural logic or validation requirements, the target document is divided into several independent areas to be checked, which can be done at the chapter level: for example, Chapter 1 and Chapter 2 each constitute a separate area to be checked. For each area, at least one item to be checked is extracted, and the corresponding data content and rendering format are collected. Within the area to be checked, items with the same data type can undergo format consistency checks to determine whether the same document elements have the same rendering effect. Semantic similarity checks can also be performed on content corresponding to the same data type within the same chapter to determine if the content belongs to the same chapter. If the similarity check result is semantic similarity and the consistency check result is format consistency, the document validation result is considered passed; if the similarity check result is semantic dissimilarity or the consistency check result is format inconsistency, the document validation result is considered failed.

[0094] By segmenting the target document, at least one region to be detected is identified. For each region, at least one item to be detected, along with its corresponding data content and rendering format, is obtained. Similarity detection is performed on the data content of each item to be detected, yielding similarity detection results. Consistency detection is performed on the rendering format of each item to be detected, yielding consistency detection results. Based on the similarity and consistency detection results, the document verification result is determined. Verification of the data content and rendering format of each item to be detected within the region to be detected ensures accuracy of the content and rendering format, thus improving the accuracy of document generation.

[0095] Example 3

[0096] Figure 3 This is a schematic diagram of a document updating device according to Embodiment 3 of the present invention. This embodiment of the present invention is applicable to document updating situations; the device can execute a document updating method and can be implemented in hardware and / or software.

[0097] See Figure 3 The document update device shown includes: a map acquisition module 301, an update data acquisition module 302, a data recognition module 303, a type determination module 304, a node determination module 305, a data update module 306, and a document generation module 307, wherein...

[0098] The graph acquisition module 301 is used to acquire the initial document and the document structure graph corresponding to the initial document. The graph nodes in the document structure graph include the content and data type of the document elements, and the graph edges include the structural relationships between the document elements.

[0099] The data acquisition module 302 is updated to acquire document update data corresponding to the target user.

[0100] The data recognition module 303 is used to identify at least one content to be processed corresponding to the document update data, and the data type to be processed corresponding to each content to be processed, based on the document update data.

[0101] The type determination module 304 is used to obtain the element data type corresponding to each data type to be processed.

[0102] The node determination module 305 is used to determine at least one target node for each data type to be processed, based on the element data type corresponding to the data type to be processed.

[0103] The data update module 306 is used to update the data associated with each target node in the document structure graph;

[0104] Document generation module 307 is used to generate target documents based on the updated document structure graph.

[0105] The technical solution of this invention involves obtaining an initial document and its corresponding document structure graph. The graph nodes in the document structure graph include the content and data type of document elements, and the graph edges include the structural relationships between document elements. The solution then proceeds to: obtaining document update data corresponding to the target user; identifying at least one content to be processed corresponding to the document update data, and the data type to be processed corresponding to each content; obtaining the data type of each data type to be processed; determining at least one target node for each data type to be processed based on the data type of the corresponding data type; updating the data associated with each target node in the document structure graph; and generating a target document based on the updated document structure graph. Different document structure graph update methods are executed for different types of document update data, thus updating the document structure graph and adaptively modifying the global content of the initial document, thereby improving the accuracy of document updates.

[0106] Optionally, the data update module 306 includes:

[0107] The title addition unit is used to generate a graph node corresponding to the new title when the data to be processed is a document title and the document update data includes the addition of a new title, and to establish a graph edge between the graph node corresponding to the new title and the target node.

[0108] The chart addition unit is used to generate the graph node corresponding to the new chart and establish the graph edge between the graph node corresponding to the new chart and the target node when the data to be processed is a chart and the document update data includes the addition of a new chart.

[0109] The text block addition unit is used to update the content of the document element corresponding to the target node based on the newly added text content when the data type to be processed is a text block and the document update data includes newly added text content.

[0110] Optionally, the node determination module 305 is specifically used for:

[0111] When the data type to be processed is a document title, find at least one parent node of the graph node with the same data type as the data type to be processed in the document structure graph and determine it as the target node;

[0112] When the data type to be processed is a chart, search the document structure graph for graph nodes in the document element whose content includes the icon corresponding to the chart, and determine them as target nodes;

[0113] When the data type to be processed is a text block, the target node is determined by the graph node whose content in the document structure graph is similar to that of the text block.

[0114] Optional, text block addition units include:

[0115] The paragraph determination sub-unit is used to search for text paragraphs with the same semantic meaning as the newly added text content in the content of the document element corresponding to the target node, based on the newly added semantic meaning of the newly added text content.

[0116] The data insertion subcell is used to add new text content to the end of a text paragraph.

[0117] Optionally, paragraphs define sub-units, specifically used for:

[0118] Acquire technical terminology data, which includes at least one technical term and the corresponding text category for each technical term;

[0119] At least one specialized term was identified from the newly added text content;

[0120] Based on the data of various professional terms and jargon, the new semantic meaning corresponding to the newly added text content is determined;

[0121] Based on the newly added semantics, search for text paragraphs with the same semantics as the newly added semantics within the content of the document element corresponding to the target node.

[0122] Optionally, the document update device is also specifically used for:

[0123] After generating the target document based on the updated document structure graph, the operation information corresponding to the target user is obtained. The operation information includes: current mouse position, document browsing direction and document page movement speed.

[0124] The predicted time period is determined based on the document browsing direction and the document page scrolling speed;

[0125] The predicted area is calculated based on the document browsing direction, document page movement speed, and predicted time period.

[0126] Determine the data to be rendered based on the current mouse position, the prediction area, and the target document;

[0127] The data to be rendered is rendered according to the rendering format, and the target document is displayed.

[0128] Optionally, the document update device is also specifically used for:

[0129] After rendering the data to be rendered according to the rendering format and displaying the target document, the target document is divided to determine at least one region to be detected.

[0130] For each region to be detected, obtain at least one item to be detected corresponding to the region to be detected, as well as the data content and rendering format of each item to be detected.

[0131] Similarity detection is performed on the data content of each item to be detected to obtain similarity detection results;

[0132] Perform consistency checks on the rendering formats of each item to be tested, and obtain the consistency check results;

[0133] The document verification result is determined based on the similarity test results and the consistency test results.

[0134] The document update apparatus provided in this embodiment of the invention can execute the document update method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the document update method.

[0135] Example 4

[0136] Figure 4 A schematic diagram of the structure of a document update device 400 that can be used to implement an embodiment of the present invention is shown.

[0137] like Figure 4 As shown, the document update device 400 includes at least one processor 401 and a memory, such as a read-only memory (ROM) 402 and a random access memory (RAM) 403, communicatively connected to the at least one processor 401. The memory stores computer programs executable by the at least one processor. The processor 401 can perform various appropriate actions and processes based on the computer program stored in the ROM 402 or loaded into the RAM 403 from storage element 408. The RAM 403 may also store various programs and data required for the operation of the document update device 400. The processor 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0138] Multiple components in the document update device 400 are connected to the I / O interface 405, including: input elements 406, such as a keyboard, mouse, etc.; output elements 407, such as various types of displays, speakers, etc.; storage elements 408, such as disks, optical discs, etc.; and communication elements 409, such as network interface cards, modems, wireless transceivers, etc. Communication element 409 allows the document update device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0139] Processor 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 401 include, but are not limited to, central processing elements (CPUs), graphics processing elements (GPUs), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 401 performs the various methods and processes described above, such as document update methods.

[0140] In some embodiments, the document update method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage element 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the document update device 400 via ROM 402 and / or communication element 409. When the computer program is loaded into RAM 403 and executed by processor 401, one or more steps of the document update method described above may be performed. Alternatively, in other embodiments, processor 401 may be configured to perform the document update method by any other suitable means (e.g., by means of firmware).

[0141] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0142] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0143] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0144] To provide user interaction, the systems and techniques described herein can be implemented on a document updating device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the document updating device. Other types of devices can also be used to provide user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including voice input, speech input, or tactile input).

[0145] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0146] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability.

[0147] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0148] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A document updating method, characterized in that, The method includes: Obtain the initial document and the document structure graph corresponding to the initial document. The graph nodes in the document structure graph include the content and data type of the document elements, and the graph edges include the structural relationships between the document elements. Retrieve document update data for the target user; Based on the document update data, at least one content to be processed corresponding to the document update data is identified, as well as the data type to be processed corresponding to each content to be processed; Obtain the element data type corresponding to each of the data types to be processed; For each of the data types to be processed, at least one target node is determined based on the element data type corresponding to the data type to be processed; Update the data associated with each target node in the document structure graph; Generate the target document based on the updated document structure graph.

2. The method according to claim 1, characterized in that, The step of updating the data associated with each target node in the document structure graph includes: When the data type to be processed is a document title, and the document update data includes a new title, a graph node corresponding to the new title is generated, and a graph edge is established between the graph node corresponding to the new title and the target node. When the data type to be processed is a chart, and the document update data includes newly added charts, a graph node corresponding to the newly added chart is generated, and a graph edge is established between the graph node corresponding to the newly added chart and the target node. When the data type to be processed is a text block, and the document update data includes newly added text content, the content of the document element corresponding to the target node is updated according to the newly added text content.

3. The method according to claim 2, characterized in that, Determining at least one target node based on the element data type corresponding to the data type to be processed includes: When the data type to be processed is a document title, at least one parent node of a graph node with the same data type as the data type to be processed is found in the document structure graph and determined as the target node. When the data type to be processed is a chart, the target node is determined by searching the content of the document element in the document structure graph that includes the icon corresponding to the chart. When the data type to be processed is a text block, the graph nodes whose content in the document structure graph is semantically similar to that of the text block are determined as target nodes.

4. The method according to claim 2, characterized in that, The step of updating the content of the document element corresponding to the target node based on the newly added text content includes: Based on the added semantics of the newly added text content, query the text paragraphs with the same semantics as the newly added semantics in the content of the document element corresponding to the target node; Add the new text content to the end of the text paragraph.

5. The method according to claim 4, characterized in that, The step of querying the content of the document element corresponding to the target node for text paragraphs with the same semantic meaning as the newly added text content, based on the added semantic meaning, includes: Acquire technical terminology data, which includes: at least one technical term and the text category corresponding to each technical term; At least one specialized term was identified from the newly added text content; Based on the aforementioned professional terms and the data of those professional terms, determine the new semantics corresponding to the newly added text content; Based on the newly added semantics, search for text paragraphs with text semantics consistent with the newly added semantics in the content of the document element corresponding to the target node.

6. The method according to claim 1, characterized in that, After generating the target document based on the updated document structure graph, the process further includes: Obtain the operation information corresponding to the target user, including: current mouse position, document browsing direction, and document page movement speed; The predicted time period is determined based on the document browsing direction and document page movement speed. The predicted area is calculated based on the document browsing direction, the document page movement speed, and the predicted time period. The data to be rendered is determined based on the current mouse position, the predicted area, and the target document; The data to be rendered is rendered according to the rendering format, and the target document is displayed.

7. The method according to claim 1, characterized in that, After rendering the data to be rendered according to the rendering format and displaying the target document, the process further includes: The target document is segmented to determine at least one region to be detected; For each of the regions to be detected, at least one item to be detected corresponding to the region to be detected, as well as the data content and rendering format corresponding to each item to be detected, are obtained. Similarity detection is performed on the data content of each item to be detected to obtain similarity detection results; The rendering format of each item to be detected is checked for consistency, and the consistency detection results are obtained. Based on the similarity detection results and the consistency detection results, the document verification result is determined.

8. A document updating device, characterized in that, The device includes: The graph acquisition module is used to acquire an initial document and the document structure graph corresponding to the initial document. The graph nodes in the document structure graph include the content and data type of the document elements, and the graph edges include the structural relationships between the document elements. Update the data acquisition module to acquire document update data for the target user; The data recognition module is used to identify at least one content to be processed corresponding to the document update data, and the data type to be processed corresponding to each content to be processed, based on the document update data. The type determination module is used to obtain the element data type corresponding to each of the data types to be processed; The node determination module is used to determine at least one target node for each of the data types to be processed, based on the element data type corresponding to the data type to be processed. The data update module is used to update the data associated with each target node in the document structure graph; The document generation module is used to generate target documents based on the updated document structure graph.

9. A document update device, characterized in that, The document update device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the document update method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the document update method according to any one of claims 1-7.