Anchor point double-threshold and pointer propulsion robust back-alignment method and device and electronic equipment

By employing a dual-threshold anchoring method and pointer advancement, the problem of low information display effectiveness in long documents is solved, achieving accurate information positioning and preventing model illusion, thereby improving the accuracy and auditability of information extraction.

CN121901419APending Publication Date: 2026-04-21ZHONGJIE TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHONGJIE TELECOMM
Filing Date
2025-12-30
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing methods for displaying information in long documents have low effectiveness, especially in complex documents. The location methods are easily affected by factors such as paragraph boundaries and whitespace characters, resulting in inaccurate information extraction and failing to meet the requirements for high reliability.

Method used

The method employs anchor point dual threshold and pointer advancement, segments documents using a semantic window slicing algorithm, performs precise matching using short and long anchor points, extracts key information using an AI system, and prevents model illusions through multi-dimensional comparison, thereby achieving inconsistency detection.

Benefits of technology

It improves the accuracy and stability of coordinate positioning in long documents, prevents model illusions, enhances the precision and auditability of information extraction, and strengthens robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901419A_ABST
    Figure CN121901419A_ABST
Patent Text Reader

Abstract

The invention provides an anchor point double-threshold and pointer advancing robust back-alignment method and device and electronic equipment, and relates to the field of data processing.The method comprises the steps that a text with the first preset text length at the beginning of each text fragment is extracted to serve as a text threshold short anchor point, matching is conducted in an original document based on the text threshold short anchor point, and the text threshold short anchor point is obtained; determining a first coordinate position of each text fragment in the original document through matching, and determining a matching ending position of each text fragment through a pointer propelling mode; based on the first coordinate position and the matching end position, inputting each text fragment into an AI system large model to extract key information and generate structured data of the text fragment corresponding to the key information, and extracting a text with a second preset text length from the structured data as a text threshold long anchor point; and in response to verification that the text threshold long anchor point does not exist in the text segment corresponding to the text threshold long anchor point, judging that the key information is model illusion and discarding the key information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a robust back-alignment method, apparatus and electronic device for anchor point dual threshold and pointer advancement. Background Technology

[0002] Currently, in the field of long document information extraction, there are complex documents containing multiple types of information points, such as contracts, tenders, and legal documents. From the perspective of coordinate positioning, existing long document extraction technologies are highly susceptible to interference from factors such as paragraph boundaries, whitespace characters, and unstable slice boundaries. Long documents often have complex formatting layouts; for example, contracts may contain chapter titles, clause lists, nested tables, and other structures. Line breaks between paragraphs, differences in indentation between different paragraphs, and the presence of numerous whitespace characters such as spaces and tabs all cause positioning methods based on simple character matching to frequently fail. Therefore, the effectiveness of existing long document information display methods is relatively low. Summary of the Invention

[0003] The purpose of this invention is to provide a robust back-alignment method, apparatus, and electronic device for anchor point dual thresholds and pointer advancement, so as to solve the technical problem of low effectiveness in displaying long document information in existing technologies.

[0004] In a first aspect, this application provides a robust backalignment method based on anchor point dual thresholds and pointer advancement, applied to a terminal providing a graphical user interface, the method comprising: In response to the acquisition of the original document, the original document is divided into multiple semantically complete text segments according to a preset slice length and preset overlap using a semantic window slicing algorithm; The first preset text length of the beginning of each text segment is extracted as a short text threshold anchor point. Matching is performed in the original document based on the short text threshold anchor point. The first coordinate position of each text segment in the original document is determined by the matching. The end position of the matching of each text segment is determined by the pointer advancement method. Based on the first coordinate position and the matching end position, each text segment is input into the AI ​​system's large model to extract key information and generate structured data of the text segment corresponding to the key information. From the structured data, the text with the first second preset text length is extracted as the text threshold long anchor point; the second preset text length is greater than the first preset text length. In response to the verification that the text threshold long anchor point does not exist in the text segment corresponding to the text threshold long anchor point, the key information is determined to be a model illusion and the key information is discarded; All the key information corresponding to the text fragments that were not discarded are summarized, and the summarized key information of the same type is compared in multiple dimensions. Based on the comparison results, document information difference results are generated, and the document information difference results and the position of the target text fragment corresponding to the document information difference results in the original document are displayed in the graphical user interface.

[0005] In one possible implementation, the terminal is equipped with an image acquisition device; after determining the first coordinate position of each text fragment in the original document through matching, the method further includes: Based on the first coordinate position, the target original document corresponding to the first coordinate position is displayed in the graphical user interface, and during the display of the target original document, the facial image of the user corresponding to the graphical user interface is captured by the image acquisition device. Based on the facial image, the facial expression corresponding to the user is identified. Based on the facial expression, the AI ​​system analyzes the user's emotional response data to the target original document and uses the emotional response data as a sensory threshold anchor point. The original document is aligned according to the user's viewing dimension by combining the text threshold short anchor point and the sensory threshold anchor point to obtain the alignment result, and the alignment result is displayed in the graphical user interface.

[0006] In one possible implementation, the alignment of the original document according to the user's reading dimension by combining the text threshold short anchor point and the sensory threshold anchor point to obtain the alignment result includes: Based on the first coordinate position, the text threshold short anchor point in the text segment is semantically associated with the sensory threshold anchor point to obtain a semantically enhanced document structure relationship; For each text segment, the quantized numerical vector corresponding to the sensory threshold anchor point and the semantic vector extended from the text threshold short anchor point are concatenated and fused to form a user text joint feature vector. Based on the joint feature vector of the user text, the semantically enhanced document structure relationship is used to generate a user sentiment change curve as the user views the document; The emotional change curve is analyzed and identified by the clustering and sequence models of the AI ​​system to obtain the user confusion area, the user interest focus area, and the user attention loss area. The user confusion area, the user interest focus area, and the user attention loss area are mapped back to the original document. The original document is then re-divided based on the mapping results. The division results are then aligned and arranged according to the dimensions of the user confusion area, the user interest focus area, and the user attention loss area to obtain a dynamic document with enhanced user experience. This dynamic document is then used as the alignment result after alignment processing according to the user's reading dimension.

[0007] In one possible implementation, after the image acquisition device acquires the facial image of the user corresponding to the graphical user interface during the display of the target original document, the method further includes: The eye perspective of the user is identified based on the facial image; Based on the eye perspective and the currently displayed content in the graphical user interface, the AI ​​system determines the second coordinate position of the line of sight corresponding to the eye perspective in the target original document, and determines the second coordinate position as the user visual anchor point; The original document is rearranged by combining the user's visual anchor points, the text threshold short anchor points, and the sensory threshold anchor points to obtain an optimized document arranged for the user's precise reading method, and the optimized document is displayed in the graphical user interface.

[0008] In one possible implementation, the process of rearranging the original document by combining the user's visual anchors, the text threshold short anchors, and the sensory threshold anchors to obtain an optimized document arranged for the user's precise reading method includes: Based on the first coordinate position and the second coordinate position, the user visual anchor point, the text threshold short anchor point and the sensory threshold anchor point in the text segment are associated in time and position to obtain a spatiotemporally enhanced document structure relationship; The duration of the user's stay and the density of fixation points at each second coordinate position are determined based on the user's visual anchor points and the text threshold short anchor points, and a visual attention distribution heatmap corresponding to the original document is generated based on the duration of the stay and the density of fixation points. Determine the degree of visual attention in the user confusion area, the user interest focus area, and the user attention loss area in the visual attention distribution heatmap, and generate user difficulty and interest insight rules based on the degree of visual attention corresponding to each area. Based on the user difficulty and interest insight rules, the original document is rearranged to obtain an optimized document that is arranged according to the user's precise reading emotional patterns.

[0009] In one possible implementation, after displaying the document information difference results and the location of the corresponding target text fragment in the original document in the graphical user interface, the method further includes: In response to a specified interactive operation on the target text fragment, an interactive object is determined in the target text fragment according to the operation position of the specified interactive operation, and the interactive object is determined as an important object anchor point; By combining the important object anchors, the user visual anchors, the text threshold short anchors, and the sensory threshold anchors, the target text fragments are rearranged according to their positions in the original document to obtain a target interactive document arranged for the user's interactive reading method, and the target interactive document is displayed in the graphical user interface.

[0010] In one possible implementation, the original document is segmented into multiple semantically complete text fragments using a semantic window slicing algorithm according to a preset slice length and a preset overlap, including: Based on a preset slice length, a semantic window slicing algorithm is used to segment the original document into multiple semantically complete text fragments by preserving the contextual semantic relationship in the overlapping areas corresponding to a preset overlap degree.

[0011] Secondly, this application provides a robust back-alignment device for anchor point dual thresholds and pointer advancement, applied to a terminal providing a graphical user interface, comprising: The segmentation module is used to segment the original document into multiple semantically complete text fragments in response to the acquisition of the original document, using a semantic window slicing algorithm according to a preset slice length and a preset overlap. The matching module is used to extract the first preset text length of each text segment as a text threshold short anchor point, perform matching in the original document based on the text threshold short anchor point, determine the first coordinate position of each text segment in the original document through matching, and determine the matching end position of each text segment by pointer advancement. The extraction module is used to input each text segment into the AI ​​system's large model based on the first coordinate position and the matching end position to extract key information and generate structured data of the text segment corresponding to the key information. From the structured data, the module extracts text of the first second preset text length as a text threshold anchor point; the second preset text length is greater than the first preset text length. The determination module is used to determine that the key information is a model illusion and discard the key information in response to the verification that the text threshold long anchor point does not exist in the text segment corresponding to the text threshold long anchor point; The generation module is used to summarize the key information corresponding to all the text fragments that have not been discarded, and to perform multi-dimensional comparison of the summarized key information of the same type. Based on the comparison results, it generates document information difference results and displays the document information difference results and the position of the target text fragment corresponding to the document information difference results in the original document in the graphical user interface.

[0012] Thirdly, this application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the method described in the first aspect above.

[0013] Fourthly, this application also provides a computer-readable storage medium storing computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method described in the first aspect above.

[0014] This application brings the following beneficial effects: This application provides a robust backalignment method, apparatus, and electronic device using anchor point dual thresholding and pointer advancement. In response to the acquisition of an original document, the method segments the original document into multiple semantically complete text segments using a semantic window slicing algorithm according to a preset slice length and preset overlap. It extracts the first preset text length of each text segment as a short anchor point for text thresholding. Based on these short anchor points, it performs matching within the original document to determine the first coordinate position of each text segment within the original document. It then determines the end position of the matching for each text segment using pointer advancement. Based on the first coordinate position and the end position of the matching, it inputs each text segment into a large AI system model to extract key information and generate structured data for the text segments corresponding to the key information. From this structured data, it extracts the first second preset text length as a long anchor point for text thresholding. The second preset text length is greater than the first preset text length. In response to the verification that the long anchor point for text thresholding does not exist in the text segment corresponding to it, it determines that the key information is a model illusion and discards the key information. Finally, it assigns all text segments to... The key information that is not discarded is summarized, and the summarized key information of the same type is compared in multiple dimensions. Based on the comparison results, document information difference results are generated, and the document information difference results and the positions of the target text fragments corresponding to the document information difference results in the original document are displayed in the graphical user interface. In this solution, the semantic fragments of the document are extracted through the long text slicing process, and the coordinates of the fragments in the original text are located by short anchor points and pointer advancement. Then, the structured information of each fragment is obtained through the large model information extraction process. Next, the authenticity of the evidence output by the long anchor point verification model is verified. Finally, the document information inconsistency items are automatically discovered through the inconsistency comparison module. It realizes the accurate detection of inconsistencies of key information in long documents based on anchor point dual thresholds and pointer advancement. By improving the accuracy and stability of coordinate positioning, it ensures that accurate evidence positioning can be obtained in the long document information extraction process, and can prevent model illusion and improve the accuracy of the results. Compared with the existing methods, it better utilizes the fusion capability of anchor point positioning and large model semantic understanding, improves the effectiveness of long document information display, is more robust, and solves the technical problem of low effectiveness of existing long document information display.

[0015] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 A flowchart illustrating the robust backalignment method for anchor point dual thresholds and pointer advancement provided in this application embodiment; Figure 2 Another flowchart illustrating the robust backalignment method for anchor point dual thresholds and pointer advancement provided in this application embodiment; Figure 3 An example of the text slicing process in the robust back alignment method of anchor point dual threshold and pointer advancement provided in the embodiments of this application; Figure 4 An example of the short anchor point and pointer advancement process in the robust back alignment method of anchor point dual threshold and pointer advancement provided in the embodiments of this application; Figure 5 An example of the information extraction process of a large model in the robust back alignment method of anchor point dual threshold and pointer advancement provided in the embodiments of this application; Figure 6 An example of a long anchor point in the robust back alignment method for anchor point dual threshold and pointer advancement provided in the embodiments of this application; Figure 7 An example of the inconsistency comparison process in the robust back alignment method of anchor point dual threshold and pointer advancement provided in the embodiments of this application; Figure 8 An example of a system corresponding to the robust back alignment method for anchor point dual thresholds and pointer advancement provided in the embodiments of this application; Figure 9 A schematic diagram of a robust back-alignment device for anchor point dual thresholds and pointer advancement provided in an embodiment of this application; Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] The terms "comprising" and "having," and any variations thereof, used in the embodiments of this application, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0020] Currently, in the field of long document information extraction, especially when dealing with complex documents containing multiple types of information points such as contracts, tenders, and legal documents, existing technologies have significant limitations in terms of accuracy, reliability, and auditability.

[0021] From a coordinate positioning perspective, existing long document extraction technologies are highly susceptible to interference from factors such as paragraph boundaries, whitespace characters, and unstable slice boundaries. Long documents often have complex formatting layouts; for example, contracts may contain chapter titles, clause lists, nested tables, and other structures. Line breaks between paragraphs, differences in indentation between different paragraphs, and the presence of numerous whitespace characters such as spaces and tabs can all cause positioning methods based on simple character matching to frequently fail. For instance, in an engineering contract, the "contract amount" may appear in the general provisions section of the first chapter or be mentioned multiple times in subsequent payment clauses. The paragraph boundaries of these mentions may change due to different layouts, and the addition or removal of whitespace characters can also shift the original text position. This causes the system to either locate the wrong paragraph or be completely unable to determine its accurate coordinates in the original text when locating these "contract amount" information points. Consequently, the extracted information loses its precise correlation with the original text, severely affecting the credibility of the information.

[0022] In the long document slicing stage, existing methods struggle to guarantee the continuous stability of slice coordinates when dealing with multi-question extraction scenarios. Current mainstream slicing methods either mechanically segment based on fixed character lengths, completely ignoring the semantic coherence of the text (e.g., splitting a complete clause semantic unit into two slices, leading to incomplete information during subsequent extraction); or they rely on simple punctuation or paragraph marks for segmentation. However, in long documents, the same type of information point (such as "performance period" or "liability for breach of contract" in a contract) may span multiple paragraphs, and punctuation usage is inconsistent. This results in sliced ​​fragments either containing too much irrelevant information, increasing the model's processing burden, or omitting key information, leading to incomplete extraction results. In multi-question extraction, such as simultaneously extracting multiple information points like contract amount, term, and the rights and obligations of both parties, these information points may be scattered across different locations in the document, with varying description lengths and contextual complexity. Existing slicing methods cannot dynamically adjust slicing strategies based on the distribution and semantic features of information points, resulting in slice coordinates that sometimes contain valid information, sometimes omit key content, and sometimes include interfering text, significantly reducing the reliability of the extraction results.

[0023] The inaccurate coordinate positioning and unstable slicing coordinates directly affect the auditability of the extraction results. In fields such as finance and law, where the traceability of results is extremely important, users need to know precisely which chapter, paragraph, or even line in the document the extracted information originates from, in order to conduct manual review or trace responsibility. However, due to deficiencies in positioning and slicing, existing technologies cannot provide precise coordinates of the evidence source. When disputes arise regarding the extraction results, it is impossible to quickly find corresponding evidence in the original text, greatly reducing the auditability of the results. Furthermore, large models may exhibit "illusion" phenomena during information extraction, generating information not present in the original text. Existing technologies lack effective verification mechanisms to check the authenticity of the model's output information, further exacerbating the unreliability of the results and creating potential risks for decisions based on these extraction results.

[0024] Therefore, the shortcomings of existing long document extraction methods in terms of coordinate positioning, slicing stability, and result auditability severely restrict their application in complex scenarios. There is an urgent need for an innovative method that can accurately locate information coordinates, stably divide text segments, and effectively verify the authenticity of data to meet the high-precision and high-reliability requirements of various industries for long document information extraction.

[0025] Based on this, embodiments of this application provide a robust back-alignment method, apparatus, and electronic device for anchor point dual thresholds and pointer advancement, which can solve the technical problem of low effectiveness in displaying long document information in existing applications.

[0026] The embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0027] Figure 1 This is a flowchart illustrating a robust backalignment method using anchor point dual thresholds and pointer advancement, provided in an embodiment of this application. The method is applied to a terminal equipped with a graphical user interface. Figure 1 As shown, the method includes: Step S110: In response to the acquisition of the original document, the original document is divided into multiple semantically complete text segments by a semantic window slicing algorithm according to a preset slice length and a preset overlap.

[0028] like Figure 2 As shown, firstly, the system receives and preprocesses original long documents such as contracts and tenders. Then, it uses a semantic window-based slicing algorithm to divide the original documents into multiple semantically complete text segments. The system sets a segment length threshold (customizable number of characters) and an overlap (customizable number of characters). It utilizes overlapping areas to preserve contextual relationships and generates basic data segments with regular structure and semantic continuity, providing data support for subsequent analysis.

[0029] As an optional implementation, the above-mentioned segmentation of the original document into multiple semantically complete text segments using the semantic window slicing algorithm according to a preset slice length and a preset overlap can specifically include the following steps: using the semantic window slicing algorithm according to the preset slice length, the original document is segmented into multiple semantically complete text segments by preserving the contextual semantic relationship in the overlapping area corresponding to the preset overlap.

[0030] For the process of slicing long text, an example is as follows: Figure 3 As shown, the system receives a long original document (such as a contract or tender document) and uses a semantic window-based slicing algorithm to divide it into multiple semantically complete text segments according to a set slice length threshold and overlap.

[0031] Step S120: Extract the first preset text length of the beginning of each text segment as a short anchor point of the text threshold. Based on the short anchor point of the text threshold, perform matching in the original document. Determine the first coordinate position of each text segment in the original document through matching. Determine the end position of the matching of each text segment by pointer advancement.

[0032] In one possible implementation, such as Figure 4 As shown, in the short anchor point localization step, a pre-defined length of text (e.g., 20 characters) is extracted at the beginning of each text segment as a short anchor point. This short anchor point is then used to accurately match within the original document, determining the coordinate position of each segment in the text. For the pointer advancement process, the matching end position of each segment is recorded using pointer advancement technology, ensuring that the subsequent segment localization order is consistent, non-overlapping, and does not go back.

[0033] For example, such as Figure 4 As shown, short anchor point features are generated for each text fragment, and the text at the beginning of the fragment with a preset length (e.g., 20 characters) is extracted as the short anchor point. Based on this short anchor point, precise matching is performed in the original document to determine the coordinate position of each fragment in the original text, solving the problem of positioning failure caused by format differences (line breaks, spaces, symbols, etc.) and achieving a one-to-one mapping between fragments and the original text positions. Furthermore, a pointer advancement technique is used to record the matching end position of each fragment, ensuring that the positioning order of subsequent fragments is consistent, non-overlapping, and does not go back.

[0034] Step S130: Based on the first coordinate position and the matching end position, each text fragment is input into the AI ​​system big model to extract key information and generate structured data of the text fragments corresponding to the key information. The text of the first second preset text length is extracted from the structured data as the text threshold long anchor point.

[0035] Wherein, the second preset text length is greater than the first preset text length. As an optional implementation, such as... Figure 5As shown, in the large model information extraction step, each text fragment is input into the large model, an information extraction request is initiated, and a structured output containing the extracted information and the original text fragment is obtained. For the long anchor point verification step, the first preset length (e.g., 40 characters) of the original text fragment in the model output is extracted as a long anchor point, and the existence of the long anchor point in the corresponding text fragment is verified to prevent the model from "illusioning".

[0036] For example, such as Figure 5 As shown, each text fragment is input into the large model, initiating an information extraction request (such as an instruction to extract key information like contract amount and term). The large model outputs structured data (such as JSON format) containing the extracted information and the corresponding original text fragments, achieving automated extraction of key information from the document. Then, from the original text fragments output by the large model, such as... Figure 6 As shown, extract the text of a preset length (e.g., 40 characters) as the long anchor point.

[0037] In step S140, in response to the verification that the long anchor point of the text threshold does not exist in the text segment corresponding to the long anchor point of the text threshold, the key information is determined to be a model illusion and the key information is discarded.

[0038] As one possible implementation method, such as Figure 6 As shown, the system verifies whether the aforementioned long anchor points of the text threshold exist in the corresponding text segment. If they do not exist, the system determines that the data is a model illusion and discards the result. This ensures the authenticity of the evidence extracted by the model and improves the auditability of the results.

[0039] Step S150: Summarize the key information that has not been discarded for all text fragments, and perform multi-dimensional comparison on the summarized key information of the same type. Generate document information difference results based on the comparison results, and display the document information difference results and the position of the target text fragment corresponding to the document information difference results in the original document in the graphical user interface.

[0040] For inconsistency comparison steps, such as Figure 7 As shown, the extracted information from all segments is summarized, and multi-dimensional comparisons are performed on key information of the same type, automatically generating a discrepancy report. The data processed by each module is summarized, and multi-dimensional comparisons are performed on key information of the same type (such as the description of contract amounts in different chapters). When discrepancies are detected in key information, a discrepancy report is automatically generated, clearly identifying the discrepancies and the location of the corresponding segments in the original text, achieving accurate identification and presentation of document inconsistencies.

[0041] like Figure 8As shown, the robust backalignment system corresponding to the anchor point double threshold and pointer advancement method includes a long text slicing module, a short anchor point module, a large model information extraction module, a long anchor point module, and a comparison inconsistency module.

[0042] In this embodiment, semantic fragments of a document are extracted through a long text slicing process. Short anchor points are used to locate the coordinates of the fragments in the original text and to advance the pointer. Then, the structured information of each fragment is obtained through a large model information extraction process. Next, the authenticity of the evidence is verified by the long anchor point verification model. Finally, the document information inconsistency is automatically detected by the inconsistency comparison module. This achieves accurate detection of inconsistencies in key information in long documents based on anchor point dual thresholds and pointer advancement. By improving the accuracy and stability of coordinate positioning, it ensures accurate evidence positioning during the long document information extraction process and prevents model illusions, thus improving the accuracy of the results. Compared with existing methods, it better utilizes the fusion capability of anchor point positioning and large model semantic understanding, thereby improving the effectiveness of displaying information in long documents and enhancing robustness.

[0043] In some embodiments, the terminal is equipped with an image acquisition device; after determining the first coordinate position of each text fragment in the original document through matching, the method may further include the following steps: Based on the first coordinate position, the target original document corresponding to the first coordinate position is displayed in the graphical user interface, and during the display of the target original document, the facial image of the user corresponding to the graphical user interface is captured by the image acquisition device. Based on facial image recognition, the system analyzes the user's emotional response data to the target original document using an AI system. This emotional response data is then used as a sensory threshold anchor point. By combining the text threshold short anchor point and the sensory threshold anchor point, the original document is aligned according to the user's viewing dimension. The alignment result is then displayed in the graphical user interface.

[0044] The original document is segmented into multiple semantically complete text fragments using a semantic window slicing algorithm. The beginning of each fragment ("text threshold short anchor point") is used to match within the original document, determining its first coordinate position. For target document display and user perception acquisition, the terminal highlights or positions the corresponding target original document content in the graphical user interface (GUI) based on this coordinate position. Simultaneously, an image acquisition device on the terminal (such as a camera) captures real-time facial images of the user viewing the document.

[0045] For emotion recognition and the generation of sensory threshold anchors, AI systems (such as facial expression recognition models) are used to analyze facial images and identify the user's current facial expression. Based on the expression, the user's emotional response data to the current document content (such as confusion, interest, boredom, etc.) is inferred, and this data is used as the sensory threshold anchor.

[0046] For multimodal alignment and dynamic document reconstruction, short text threshold anchors (text side) are correlated and aligned with sensory threshold anchors (user perception side). Based on this, the original document is reorganized or annotated from the user's reading perspective, generating alignment processing results (such as annotating user confusion areas, interest areas, etc.). For result visualization, the alignment processing results enhanced by user perception are displayed in a graphical user interface to achieve personalized reading feedback.

[0047] In this embodiment, cross-modal alignment of document content with real-time user emotional feedback is achieved, thereby dynamically generating "sensory-enhanced dynamic documents" centered on the user's cognitive state, significantly improving the personalization and interactive intelligence of human-computer collaborative reading. Specifically, the system can not only "display documents" but also "understand how users perceive documents"; it automatically identifies areas of confusion, areas of focused interest, and areas of attention loss for users during the reading process. This effect breaks through the traditional static document browsing mode, integrating user physiological / emotional signals into document structure understanding, and is a typical innovation of the fusion of affective computing and intelligent document processing.

[0048] In some embodiments, the original document is aligned according to the user's reading dimension by combining text threshold short anchors and perceptual threshold anchors to obtain the alignment result. The method may further include the following steps: Based on the first coordinate position, the text threshold short anchor points and the sensory threshold anchor points in the text segment are semantically associated to obtain the semantically enhanced document structure relationship; for each text segment, the quantized numerical vector corresponding to the sensory threshold anchor point and the semantic vector extended from the text threshold short anchor point are concatenated and fused into the user text joint feature vector. Based on the joint feature vector of user text, the semantically enhanced document structure relationship is used to generate the user's emotional change curve as the user reads. The clustering model and sequence model of the AI ​​system are used to perform sentiment analysis and sentiment recognition on the emotional change curve to obtain the user's confusion area, user interest focus area and user attention loss area. The user confusion area, user interest focus area, and user attention loss area are mapped back to the original document. The original document is then re-divided based on the mapping results. The division results are then aligned and arranged according to the dimensions of the user confusion area, user interest focus area, and user attention loss area to obtain a dynamic document with enhanced user experience. This dynamic document is then used as the alignment result after alignment processing according to the user's reading dimension.

[0049] The process of concatenating and fusing the quantized numerical vector corresponding to the perceptual threshold anchor point and the semantic vector extended from the short text threshold anchor point into a joint feature vector for each text segment can specifically include the following steps: For each text segment, the quantized numerical vector corresponding to the sensory threshold anchor point and the semantic vector extended from the text threshold short anchor point are concatenated and fused into a user text joint feature vector using the following first calculation formula: =ReLU(W f ·[ ; ]+ ); in, ReLU represents the fused user text joint feature vector; ReLU represents the modified linear unit activation function. This represents the semantic vector extended from the short anchor points of the text threshold; W represents the quantized numerical vector corresponding to the sensory threshold anchor point; f Represents the fusion weight matrix; This represents the fusion bias parameter vector.

[0050] The process described above for generating a user sentiment change curve based on the user text joint feature vector and semantically enhanced document structure relationships can specifically include the following steps: Based on the user text joint feature vector and the semantically enhanced document structure relationships, the user sentiment change curve based on the user text joint feature vector is generated using the following second calculation formula:

[0051] in, This represents the emotional change curve; i Indicates the fragment index; Indicates the total number of text fragments; This represents the normalized exponential function; This represents a function that indicates the similarity between locations and segments. Indicates continuous positional parameters in the document; This represents the context-aware hidden state of the i-th segment; This represents the sentiment mapping weight matrix. In this embodiment of the application, the calculation method of the above calculation formula can make the data spliced ​​and fused into the user text joint feature vector and the generated user emotional change curve as the browsing trajectory more accurate.

[0052] By precisely aligning real-time emotional feedback from users during reading (such as confusion, loss of interest, and attention loss) with the semantic structure of the original document, a dynamically reconstructed "emotionally enhanced document" centered on the user's cognitive experience is automatically generated, achieving a paradigm shift from "static content presentation" to "personalized emotionally perceptive reading." Breaking through the limitations of traditional document processing that relies solely on textual semantics, this approach is the first to integrate users' subjective feelings as a core dimension into document structure modeling. Through multimodal fusion (combining textual semantics with emotional signals) and sequence modeling, it automatically identifies and marks key areas in the document that affect user comprehension and attention. The generated dynamic document not only retains the original information but also overlays a user readability heatmap (such as highlighting confusing paragraphs, folding low-interest content, and recommending rereading paths), significantly improving reading efficiency and interactive intelligence.

[0053] In some embodiments, after capturing the facial image of the user corresponding to the graphical user interface during the display of the target original document using an image acquisition device, the method may further include the following steps: The system identifies the user's eye perspective based on facial images; based on the eye perspective and the currently displayed content in the graphical user interface, the AI ​​system determines the second coordinate position of the line of sight corresponding to the eye perspective in the target original document, and determines the second coordinate position as the user's visual anchor point. The original document is rearranged by combining user visual anchors, text threshold short anchors, and sensory threshold anchors to obtain an optimized document arranged for the user's precise reading method, and the optimized document is displayed in the graphical user interface.

[0054] During the display of the target original document in the graphical user interface (GUI), the terminal captures the user's facial image in real time through a built-in image acquisition device (such as a camera). To identify the eye perspective and locate the gaze point, based on the facial image, an AI vision model (such as an eye-tracking algorithm) is used to identify the user's eye posture and gaze direction (i.e., "eye perspective"). Combined with the layout of the document content displayed in the current GUI, this eye perspective is mapped to a specific location on the target original document to obtain a second coordinate position, which is defined as the user's visual anchor point.

[0055] For multimodal anchor fusion, the following three types of anchors are associated and aligned: text threshold short anchors (from text semantic structure), sensory threshold anchors (from emotional responses in facial expression recognition), and user visual anchors (from eye movement / gaze behavior). For generating optimized documents, based on the collaborative information of these three anchors, the original document undergoes content reorganization, key point highlighting, paragraph sorting, or adaptive interface adjustments to generate an optimized document that better matches the user's actual reading behavior and cognitive state. This optimized document is then dynamically displayed in the graphical user interface, achieving a personalized reading experience where "what you see is what you feel, and what you focus on is what you emphasize."

[0056] By integrating user gaze behavior (visual anchors), emotional feedback (sensory threshold anchors), and text semantic structure (text threshold short anchors), a multimodal document understanding model is constructed, centered on the user's true reading intention and cognitive focus. This enables precise, personalized, and real-time optimized arrangement of the original document. The system upgrades from "passive display" to "active adaptation": it no longer presents documents in a fixed format but dynamically adjusts content organization based on "where the user looks, how they feel, and what they focus on." This improves reading efficiency and comprehension depth: for example, it automatically enlarges paragraphs that the user repeatedly stares at but whose expression is confused, or folds redundant content that the user quickly scans without emotional fluctuation. It also enables fine-grained human-computer collaborative reading: visual anchors provide spatial accuracy (which line is being read), emotional anchors provide cognitive state (whether the user understands), and text anchors provide semantic context; the combination of these three elements forms a high-fidelity user reading profile.

[0057] In some embodiments, the original document is rearranged by combining user visual anchors, text threshold short anchors, and sensory threshold anchors to obtain an optimized document arranged for the user's precise reading method. This may specifically include the following steps: Based on the first and second coordinate positions, the user visual anchor points, text threshold short anchor points and sensory threshold anchor points in the text fragment are associated with time and position to obtain spatiotemporally enhanced document structure relationships; The duration of the user's visual anchor and the short anchor of the text threshold are determined at each second coordinate position, and the visual attention distribution heatmap corresponding to the original document is generated based on the duration of the visual anchor and the short anchor of the text threshold. The visual attention levels of user confusion areas, user interest focus areas, and user attention loss areas are determined in the visual attention distribution heatmap, and user difficulty and interest point insight rules are generated based on the visual attention levels corresponding to each area. The original document is then rearranged based on the user difficulty and interest point insight rules to obtain an optimized document that is arranged according to the user's precise reading emotional patterns.

[0058] The text threshold short anchor points are obtained by matching them in the original document, representing the text position of the semantic fragment. The second coordinate position is obtained by mapping eye viewpoint recognition to GUI content, representing the user's actual visual gaze position (i.e., the user's visual anchor point). The sensory threshold anchor points are based on emotional response data derived from facial expression analysis.

[0059] To construct a spatiotemporally enhanced document structure, the three types of anchors (visual, textual, and emotional) are aligned according to timestamps and spatial locations within the document to establish a spatiotemporally enhanced document structure that includes "when viewed, where viewed, and what feelings were experienced." For generating the visual attention distribution heatmap, the visual attention intensity of each region is calculated using the user's sustained dwell time and fixation density (e.g., the number of fixations or duration within a unit area) at each second coordinate position. Based on this data, a visual attention distribution heatmap covering the entire original document is generated.

[0060] To identify key reading areas and their attention levels, known areas of user confusion, areas of focused interest, and areas of attention loss are located in the heatmap (derived from emotion and behavior fusion analysis). The visual attention level corresponding to each area is quantified (e.g., high dwell time + high density = high attention; low dwell time + low density = attention loss).

[0061] For the generation of rules for identifying user difficulties and points of interest, interpretable rules are extracted based on the correlation pattern between visual attention and emotional labels. For example, "If the gaze time of a certain paragraph is >5 seconds and the expression is confused, it is judged as a cognitive difficulty"; "If the gaze density of a certain paragraph is high and the expression is pleasant, it is judged as a focus of interest".

[0062] For dynamically reconstructed and optimized documents, based on the aforementioned insight rules, the original document undergoes content rearrangement, key point highlighting, difficulty prompts, interest guidance, or redundancy compression to generate an optimized document that aligns with the individual user's reading emotional patterns. This optimized document is then displayed in a graphical user interface, achieving a personalized, context-aware intelligent reading experience.

[0063] This application, for the first time, deeply integrates user visual gaze behavior (spatiotemporal trajectory), emotional feedback, and text semantic structure to construct an intelligent document system with a three-in-one perception capability of "cognition-emotion-behavior," thereby automatically generating "emotionally driven dynamic documents" highly adapted to individual reading habits and comprehension bottlenecks. It upgrades from "coarse-grained emotional feedback" to "fine-grained spatiotemporal behavior + emotion joint modeling": it not only knows "whether the user is confused," but also "how long they were confused about which words," achieving accurate diagnosis. It generates actionable reading insight rules: transforming black-box AI analysis into explainable and interventionist teaching or editing strategies (such as automatically inserting annotations and adjusting paragraph order). It achieves a truly user-centric document evolution: the document is no longer a static carrier, but an intelligent medium that can "perceive the user and self-optimize."

[0064] In some embodiments, after displaying the document information difference results and the location of the target text fragment corresponding to the document information difference results in the original document in the graphical user interface, the method may further include the following steps: In response to a specified interactive operation on the target text fragment, the interactive object is determined in the target text fragment according to the operation position of the specified interactive operation, and the interactive object is identified as an important object anchor point. Combining the important object anchor point, user visual anchor point, text threshold short anchor point and sensory threshold anchor point, the target text fragment is rearranged according to its position in the original document to obtain a target interactive document arranged according to the user's interactive reading method, and the target interactive document is displayed in the graphical user interface.

[0065] Users perform specific interactive operations (such as clicking, highlighting, annotating, asking questions, etc.) on target text fragments in a graphical user interface (GUI). To determine the interactive object and important object anchors, the system precisely locates the manipulated content unit (such as a word, sentence, or paragraph) within the target text fragment based on the operation's location (such as mouse coordinates, touch point, or selected text area), identifying it as the interactive object. This interactive object is marked as an important object anchor, representing key information points subjectively perceived by the user. For multimodal anchor fusion, the following four types of anchors are collaboratively associated: important object anchors (from explicit user interaction), user visual anchors (from eye movement / gaze behavior), text threshold short anchors (from semantic structure slices), and sensory threshold anchors (from facial expression emotion feedback).

[0066] For document rearrangement based on fusion anchors, the presentation or structural order of target text fragments in the original document is dynamically adjusted, with important object anchors as the core and combined with user attention, emotional state, and semantic context reflected by other anchors. For example, this includes enlarging the paragraph containing the interactive object, associating it with relevant context, inserting explanatory content, and collapsing irrelevant parts. For generating and displaying the target interactive document, an optimized "target interactive document" based on the user's current interaction intent and reading state is output and presented in real-time in the GUI, achieving an intelligent response of "interaction is reconstruction."

[0067] By deeply integrating explicit user interactions (such as clicking and selecting) with implicit cognitive signals (gaze and emotion) and text semantic structure, an integrated intelligent document system of "interaction-driven, perception-enhanced, and dynamic reconstruction" is constructed, thereby achieving personalized and contextualized real-time document optimization centered on the user's proactive intent. Upgrading from "passive response" to "proactive collaboration": The system not only responds to user operations but also intelligently expands or reorganizes relevant content based on the user's potential cognitive state, enhancing the depth of interaction. Bridging explicit intent and implicit needs: When a user selects a word (explicit), the system, combined with the user's confused expression and prolonged gaze (implicit), automatically pushes terminology explanations or case studies, achieving precise assistance. Supporting advanced reading tasks: In scenarios such as research, review, and learning, users can trigger system-level document reorganization through simple interactions, greatly improving information acquisition efficiency and comprehension quality.

[0068] Figure 9 A schematic diagram of a robust back-alignment device with anchor point dual thresholds and pointer advancement is provided. This device can be applied to terminals with graphical user interfaces. Figure 9 As shown, the robust realignment device 900 for anchor point dual thresholds and pointer advancement includes: The segmentation module 901 is used to segment the original document into multiple semantically complete text fragments in response to the acquisition of the original document, according to a preset slice length and a preset overlap, using a semantic window slicing algorithm. The matching module 902 is used to extract the first preset text length of the beginning of each text segment as a text threshold short anchor point, perform matching in the original document based on the text threshold short anchor point, determine the first coordinate position of each text segment in the original document through matching, and determine the matching end position of each text segment by pointer advancement. The extraction module 903 is used to input each of the text segments into the AI ​​system's large model based on the first coordinate position and the matching end position to extract key information and generate structured data of the text segments corresponding to the key information. It extracts text of the first second preset text length from the structured data as a text threshold long anchor point. The second preset text length is greater than the first preset text length. The determination module 904 is used to determine that the key information is a model illusion and discard the key information in response to the verification that the text threshold long anchor point does not exist in the text segment corresponding to the text threshold long anchor point; The generation module 905 is used to summarize the key information corresponding to all the text fragments that have not been discarded, and to perform multi-dimensional comparison of the summarized key information of the same type. Based on the comparison results, it generates document information difference results and displays the document information difference results and the position of the target text fragment corresponding to the document information difference results in the original document in the graphical user interface.

[0069] The robust realignment device for anchor point double threshold and pointer advancement provided in this application embodiment has the same technical features as the robust realignment method for anchor point double threshold and pointer advancement provided in the above embodiment, so it can also solve the same technical problems and achieve the same technical effects.

[0070] An electronic device provided in this application embodiment, such as Figure 10 As shown, the electronic device 1000 includes a processor 1002 and a memory 1001. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the steps of the method provided in the above embodiments.

[0071] See Figure 10 The electronic device also includes a bus 1003 and a communication interface 1004. The processor 1002, the communication interface 1004 and the memory 1001 are connected via the bus 1003. The processor 1002 is used to execute executable modules, such as computer programs, stored in the memory 1001.

[0072] The memory 1001 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 1004 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.

[0073] Bus 1003 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 10 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0074] The memory 1001 is used to store programs. After receiving an execution instruction, the processor 1002 executes the program. The method executed by the apparatus defined by the process disclosed in any of the preceding embodiments of this application can be applied to the processor 1002 or implemented by the processor 1002.

[0075] The processor 1002 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 1002 or by instructions in software form. The processor 1002 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 1001. Processor 1002 reads the information in memory 1001 and, in conjunction with its hardware, completes the steps of the above method.

[0076] Corresponding to the robust backalignment method of anchor point double threshold and pointer advancement described above, this application embodiment also provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and run by a processor, the computer-executable instructions cause the processor to perform the steps of the robust backalignment method of anchor point double threshold and pointer advancement described above.

[0077] The robust realignment device for anchor point dual thresholds and pointer advancement provided in this application embodiment can be specific hardware on the device or software or firmware installed on the device. The implementation principle and technical effects of the device provided in this application embodiment are the same as those in the foregoing method embodiments. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the foregoing method embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can all be referred to the corresponding processes in the above method embodiments, and will not be repeated here.

[0078] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0079] For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0080] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0081] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0082] If the aforementioned function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the robust back-alignment method of anchor point dual threshold and pointer advancement described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.

[0083] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0084] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A robust backalignment method using anchor point dual thresholds and pointer advancement, characterized in that, Applied to a terminal providing a graphical user interface, the method includes: In response to the acquisition of the original document, the original document is divided into multiple semantically complete text segments according to a preset slice length and preset overlap using a semantic window slicing algorithm; The first preset text length of the beginning of each text segment is extracted as a short text threshold anchor point. Matching is performed in the original document based on the short text threshold anchor points. The first coordinate position of each text segment in the original document is determined by the matching. The end position of the matching of each text segment is determined by the pointer advancement method. Based on the first coordinate position and the matching end position, each text segment is input into the AI ​​system's large model to extract key information and generate structured data of the text segment corresponding to the key information. From the structured data, the text with the first second preset text length is extracted as the text threshold long anchor point; the second preset text length is greater than the first preset text length. In response to the verification that the text threshold long anchor point does not exist in the text segment corresponding to the text threshold long anchor point, the key information is determined to be a model illusion and the key information is discarded; All the key information corresponding to the text fragments that were not discarded are summarized, and the summarized key information of the same type is compared in multiple dimensions. Based on the comparison results, document information difference results are generated, and the document information difference results and the position of the target text fragment corresponding to the document information difference results in the original document are displayed in the graphical user interface.

2. The method according to claim 1, characterized in that, The terminal is equipped with an image acquisition device; After determining the first coordinate position of each text fragment in the original document through matching, the method further includes: Based on the first coordinate position, the target original document corresponding to the first coordinate position is displayed in the graphical user interface, and during the display of the target original document, the facial image of the user corresponding to the graphical user interface is captured by the image acquisition device. Based on the facial image, the facial expression corresponding to the user is identified. Based on the facial expression, the AI ​​system analyzes the user's emotional response data to the target original document and uses the emotional response data as a sensory threshold anchor point. The original document is aligned according to the user's viewing dimension by combining the text threshold short anchor point and the sensory threshold anchor point to obtain the alignment result, and the alignment result is displayed in the graphical user interface.

3. The method according to claim 2, characterized in that, The process of aligning the original document according to the user's reading dimension by combining the text threshold short anchor point and the perceptual threshold anchor point to obtain the alignment result includes: Based on the first coordinate position, the text threshold short anchor point in the text segment is semantically associated with the sensory threshold anchor point to obtain a semantically enhanced document structure relationship; For each text segment, the quantized numerical vector corresponding to the sensory threshold anchor point and the semantic vector extended from the text threshold short anchor point are concatenated and fused to form a user text joint feature vector. Based on the joint feature vector of the user text, the semantically enhanced document structure relationship is used to generate a user sentiment change curve as the user views the document; The emotional change curve is analyzed and identified by the clustering and sequence models of the AI ​​system to obtain the user confusion area, the user interest focus area, and the user attention loss area. The user confusion area, the user interest focus area, and the user attention loss area are mapped back to the original document. The original document is then re-divided based on the mapping results. The division results are then aligned and arranged according to the dimensions of the user confusion area, the user interest focus area, and the user attention loss area to obtain a dynamic document with enhanced user experience. This dynamic document is then used as the alignment result after alignment processing according to the user's reading dimension.

4. The method according to claim 3, characterized in that, After the image acquisition device captures the facial image of the user corresponding to the graphical user interface during the display of the target original document, the process further includes: The eye perspective of the user is identified based on the facial image; Based on the eye perspective and the currently displayed content in the graphical user interface, the AI ​​system determines the second coordinate position of the line of sight corresponding to the eye perspective in the target original document, and determines the second coordinate position as the user visual anchor point. The original document is rearranged by combining the user visual anchor point, the text threshold short anchor point, and the sensory threshold anchor point to obtain an optimized document arranged for the user's precise reading method, and the optimized document is displayed in the graphical user interface.

5. The method according to claim 4, characterized in that, The process of rearranging the original document by combining the user's visual anchor points, the text threshold short anchor points, and the sensory threshold anchor points to obtain an optimized document arranged for the user's precise reading method includes: Based on the first coordinate position and the second coordinate position, the user visual anchor point, the text threshold short anchor point and the sensory threshold anchor point in the text segment are associated in time and position to obtain a spatiotemporally enhanced document structure relationship; The duration of the user's stay and the density of fixation points at each second coordinate position are determined based on the user's visual anchor points and the text threshold short anchor points, and a heatmap of visual attention distribution corresponding to the original document is generated based on the duration of the stay and the density of fixation points. Determine the degree of visual attention in the user confusion area, the user interest focus area, and the user attention loss area in the visual attention distribution heatmap, and generate user difficulty and interest insight rules based on the degree of visual attention corresponding to each area. Based on the user difficulty and interest insight rules, the original document is rearranged to obtain an optimized document that is arranged according to the user's precise reading emotional patterns.

6. The method according to claim 5, characterized in that, After displaying the document information difference results and the location of the corresponding target text fragment in the original document in the graphical user interface, the method further includes: In response to a specified interactive operation on the target text fragment, an interactive object is determined in the target text fragment according to the operation position of the specified interactive operation, and the interactive object is determined as an important object anchor point; By combining the important object anchors, the user visual anchors, the text threshold short anchors, and the sensory threshold anchors, the target text fragments are rearranged according to their positions in the original document to obtain a target interactive document arranged for the user's interactive reading method, and the target interactive document is displayed in the graphical user interface.

7. The method according to claim 1, characterized in that, The original document is segmented into multiple semantically complete text fragments using a semantic window slicing algorithm according to a preset slice length and preset overlap, including: Based on a preset slice length, a semantic window slicing algorithm is used to segment the original document into multiple semantically complete text fragments by preserving the contextual semantic relationship in the overlapping areas corresponding to a preset overlap degree.

8. A robust realignment device for anchor point dual thresholds and pointer advancement, characterized in that, Applied to terminals that provide a graphical user interface, including: The segmentation module is used to segment the original document into multiple semantically complete text fragments in response to the acquisition of the original document, using a semantic window slicing algorithm according to a preset slice length and a preset overlap. The matching module is used to extract the first preset text length of each text segment as a text threshold short anchor point, perform matching in the original document based on the text threshold short anchor point, determine the first coordinate position of each text segment in the original document through matching, and determine the matching end position of each text segment by pointer advancement. The extraction module is used to input each text segment into the AI ​​system's large model based on the first coordinate position and the matching end position to extract key information and generate structured data of the text segment corresponding to the key information. From the structured data, the module extracts text of the first second preset text length as a text threshold anchor point; the second preset text length is greater than the first preset text length. The determination module is used to determine that the key information is a model illusion and discard the key information in response to the verification that the text threshold long anchor point does not exist in the text segment corresponding to the text threshold long anchor point; The generation module is used to summarize the key information corresponding to all the text fragments that have not been discarded, and to perform multi-dimensional comparison of the summarized key information of the same type. Based on the comparison results, it generates document information difference results and displays the document information difference results and the position of the target text fragment corresponding to the document information difference results in the original document in the graphical user interface.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method according to any one of claims 1 to 7.