Processing method and device for carrying out element labeling on PDF (Portable Document Format) file

By combining pre-parsing and real-time processing with a behavior prediction model, efficient, accurate, and consistent annotation of PDF file elements is achieved, solving the problems of low efficiency, inaccurate cross-page recognition, and poor consistency in traditional annotation methods.

CN121009880APending Publication Date: 2025-11-25BEIJING DP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511177369.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Traditional PDF file element annotation methods are inefficient, lack cross-page recognition accuracy, and have poor overall consistency.

Method used

By employing pre-parsing, real-time processing, behavior prediction, multimodal feature recognition, and asynchronous processing mechanisms, combined with human-computer collaboration, the system achieves annotation trajectory refresh, candidate target reference, multimodal feature addition, and associated target matching, thereby improving annotation efficiency and cross-page recognition accuracy and ensuring full-text consistency.

Benefits of technology

It improves the efficiency and accuracy of PDF file element annotation, enhances the recognition and fusion efficiency of cross-page elements, and ensures the consistency of full-text annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009880A_ABST
    Figure CN121009880A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a processing method and device for performing element labeling on a PDF (Portable Document Format) file. The method comprises the following steps: performing image conversion and basic element analysis on the PDF file input by a labeling person; in the labeling process, the labeling track and the target element set are refreshed by recording the labeling behavior of a labeling person; providing a candidate element set for the next step of labeling by the behavior prediction model according to the labeling track; the performance of the prediction model is improved based on candidate feedback of the labeler; multi-modal element features are added to the target elements based on the multi-modal feature recognition model; the associated target trajectory is refreshed through a target matching and trajectory tracking processing mechanism; after labeling is finished, cross-page element fusion and labeling consistency check are carried out; and finally, feeding back the target set completing the consistency check to the annotator. According to the method, the labeling efficiency can be improved, the recognition accuracy and fusion efficiency of the cross-page elements are improved, and the labeling consistency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to a processing method and device for element annotation of a PDF file. BACKGROUND

[0002] The annotation of file elements (such as text, table, image, etc.) of a Portable Document Format (PDF) file refers to that an annotator selects target elements of the PDF file, sets an annotation box for the selected target, and annotates the annotation box based on a predetermined label category. The traditional annotation method is a pure manual annotation method. Limited by the manual experience and working state of the annotator, this traditional method has some problems: 1) low annotation efficiency; 2) insufficient recognition accuracy and low fusion efficiency for cross-page elements; and 3) poor annotation consistency of the entire text. SUMMARY

[0003] The present application is aimed at the defects of the prior art, and provides a processing method and device for element annotation of a PDF file, an electronic device, and a computer readable storage medium. Before starting an annotation task, the present application performs pre-analysis on a PDF file based on a PDF element analysis tool. During the annotation process: 1) based on a real-time processing method, the annotation trajectory (i.e., a first annotation trajectory) and the annotation element set (i.e., a first target element set) of the annotator are refreshed in real time according to the pre-analysis information; 2) based on the real-time processing method, a behavior prediction model is used to provide a candidate target reference (i.e., a candidate element set) for the annotator after each single-step annotation behavior according to the real-time annotation trajectory; 3) based on an asynchronous processing method, the model performance of the prediction model is continuously improved according to the candidate target feedback of the annotator; 4) based on the asynchronous processing mechanism, a multi-modal feature recognition model is used to add multi-modal (visual + text) element features to the annotated target elements; 5) based on the asynchronous processing mechanism, the associated target matching and associated target trajectory tracking processing are performed according to the annotation box geometric features (annotation box area, annotation box aspect ratio) and the multi-modal element features of the annotated elements, and the tracking trajectory set (i.e., a first trajectory set) is refreshed based on the processing result. After the annotation is completed, the cross-page elements in the annotation element set are first fused according to the tracking trajectory; then the fused element set (analysis target set) is subjected to annotation consistency checking, and the consistency modification of the same type of element label is completed through a man-machine cooperation method during the checking process; finally, the analysis target set that has passed the consistency checking is fed back as the final annotation result. The present application can improve the annotation efficiency through the candidate element prediction mechanism, improve the recognition accuracy and fusion efficiency of the cross-page elements through the associated element tracking mechanism, and improve the annotation consistency of the entire text through the consistency checking.

[0004] To achieve the above objectives, a first aspect of the present invention provides a method for annotating elements in a PDF file, the method comprising:

[0005] Receive a first PDF file input by the first annotator; the first PDF file is the annotation object of the first annotator.

[0006] Based on a preset PDF element parsing tool, the first PDF file is processed for image conversion and basic element parsing to obtain the corresponding first image sequence and first parsing sequence;

[0007] When the first annotator selects and annotates target elements in the first PDF file according to a preset tag type set, the annotation behavior of the first annotator is recorded by combining the first image sequence and the first parsing sequence, and the corresponding first annotation trajectory and first target element set are refreshed based on the recorded information; and a preset behavior prediction model provides a candidate element set for the next annotation target after each single-step annotation behavior of the first annotator according to the first annotation trajectory; and the behavior prediction model is improved and trained based on the first annotator's selection feedback of candidate elements.

[0008] During the annotation process, corresponding multimodal element features are added to the target elements of the first target element set based on a preset multimodal feature recognition model; and target matching and trajectory tracking are performed on the target elements according to the first target element set, and the corresponding first trajectory set is refreshed based on the tracking results.

[0009] After annotation is completed, the first target element set is subjected to cross-page element fusion processing based on the first trajectory set to obtain the corresponding parsed target set; and the annotation consistency check is performed on the parsed target set.

[0010] The parsed target set that has completed the consistency check is fed back to the first annotator.

[0011] Preferably, the first PDF file includes multiple first file pages;

[0012] The first image sequence is formed by sequentially sorting multiple first frame images; each first frame image corresponds one-to-one with the first file page.

[0013] The first parsing sequence is formed by sequentially sorting multiple first frame element sets; each first frame element set corresponds one-to-one with the first file page.

[0014] The first frame element set includes multiple first basic elements; the first basic elements include a first element identifier, a first page number identifier, a first paragraph identifier, a first start position, a first end position, a first element type, and a first element attribute set; the first element type includes article title, chapter title, abstract, paragraph, table, and image;

[0015] The tag type set includes a basic type subset and a domain type subset; both the basic type subset and the domain type subset consist of one or more tag types; the tag types of the basic type subset include article title, chapter title, abstract, paragraph, table, and image; the tag types of the domain type subset correspond to the file type of the first PDF file; if the file type of the first PDF file is an academic paper, then the tag types of the domain type subset include mathematical symbols, physical symbols, mathematical formulas, and chemical formulas; if the file type of the first PDF file is a technical report, then the tag types of the domain type subset include formulas, technical terms, and technical parameters; if the file type of the first PDF file is a legal document, then the tag types of the domain type subset include legal terms and legal clauses.

[0016] The first annotation trajectory is formed by sequentially sorting one or more first trajectory points; the trajectory point attributes of each first trajectory point include a first timestamp, a first mouse sliding trajectory, a first mouse click position, a first annotation box drag direction, a first annotation box size, a first annotation box center point, and a first annotation type;

[0017] The first target element set consists of one or more first target elements; each first target element corresponds to a selected target element, a first trajectory point, and a parent basic element; the parent basic element is a first basic element; the element content of each first target element is part or all of the element content of the parent basic element; the first target element includes a first target identifier, a second timestamp, a second element identifier, a second page number identifier, a second paragraph identifier, a first annotation box, a first target type, and a first multimodal feature; the first annotation box includes a first starting point, a first center point, a first height, a first width, a first diagram, and first text;

[0018] The candidate element set consists of at most M first candidate elements; the number M is a preset positive integer; the first candidate element includes the candidate box starting point, candidate box height, candidate box width, candidate box annotation type, and candidate box confidence.

[0019] The first trajectory set consists of one or more first tracking trajectories; each first tracking trajectory corresponds to a first trajectory identifier; each first tracking trajectory is formed by sequentially sorting one or more second trajectory points; the trajectory point attributes of each second trajectory point include a second target identifier, a third timestamp, a third element identifier, a third page number identifier, a third paragraph identifier, a second annotation box, a second target type, and a second multimodal feature; the second annotation box includes a second starting point, a second center point, a second height, a second width, a second bounding box, and second text;

[0020] The behavior prediction model is used to predict candidate elements based on the historical behavior sequence X input to the model and output the corresponding set of candidate elements; wherein, the historical behavior sequence X consists of N behavior feature vectors x i The data is sorted chronologically, with 1 ≤ index i ≤ N, and the number N is a preset positive integer; each behavioral feature vector x i Corresponding to one of the first trajectory points; each of the behavioral feature vectors x i The vector feature data consists of the first timestamp of the corresponding first trajectory point, the first mouse sliding trajectory, the first mouse click position, the drag direction of the first annotation box, the size of the first annotation box, and the first annotation type;

[0021] The multimodal feature recognition model is used to extract visual basic features, visual layout features, image features, and text features based on the element annotation box information input to the model. It then fuses the three types of visual features using an attention-weighted fusion mechanism and performs feature dimensionality reduction on the concatenated features of the visual fusion features and text features to obtain the corresponding multimodal element features. The element annotation box information corresponds to one of the first target elements in the first target element set. The element annotation box information includes the parent page image, annotation box center point, annotation box height, annotation box width, annotation box image, and annotation box text.

[0022] The parsing target set consists of one or more first parsing targets; the first parsing target includes a parsing target type, a parsing target identifier group, a parsing target graph, and a parsing target text; the parsing target type is one of the tag types in the tag type set; the parsing target identifier group consists of one or two second target identifiers; the parsing target graph consists of one or two second block graphs corresponding to the parsing target identifier group; the parsing target text consists of one or two second texts corresponding to the parsing target identifier group.

[0023] Preferably, the behavior prediction model is composed of an embedding coding layer, a behavior feature encoder, a candidate box prediction module, and a candidate box filtering module connected in sequence;

[0024] The embedding coding layer is used to perform actions based on the behavior feature vector x. i The embedding encoding rules corresponding to each vector feature data are used for each behavioral feature vector x. i Each vector feature data is embedded and encoded to obtain a corresponding feature encoding vector; and each behavioral feature vector x is then used to generate a feature encoding vector. i All the aforementioned feature encoding vectors are sequentially concatenated into a corresponding embedding encoding vector; and a corresponding embedding encoding sequence E composed of N such embedding encoding vectors is sent to the behavior feature encoder;

[0025] The behavior feature encoder is implemented based on an LSTM model or a Transformer architecture encoder model; the behavior feature encoder is used to perform feature encoding processing on the embedded encoding sequence E to obtain the corresponding feature vector H and send it to the candidate box prediction module;

[0026] The candidate box prediction module consists of M ’ It consists of M parallel first prediction units; the number is M ’ M is a preset positive integer. ’ >M; The candidate box prediction module is used to use each of the first prediction units to predict a single candidate box based on the feature vector H to obtain a corresponding predicted candidate box y. j 1 ≤ index j ≤ M ’ ; and from the obtained M ’ The predicted candidate boxes y j The predicted candidate boxes are sequentially sorted to form a corresponding candidate box sequence Y and sent to the candidate box filtering module; each predicted candidate box y j Each includes a corresponding set of candidate box starting point, candidate box height, candidate box width, candidate box annotation type, and candidate box confidence level;

[0027] The prediction method for the j-th first prediction unit is:

[0028]

[0029] The four weight vector parameters are the j-th first prediction unit. These are the four offset vector parameters corresponding to the j-th first prediction unit; The candidate box position prediction vector is the candidate box position prediction vector corresponding to the j-th first prediction unit. The vector length is 3, and it consists of the starting point of the candidate box corresponding to the j-th first prediction unit, the height of the candidate box, and the width of the candidate box; The candidate box type prediction vector corresponding to the j-th first prediction unit is composed of multiple type prediction probabilities; the total number of the type prediction probabilities is consistent with the total number of label types in the label type set; the candidate box type prediction vector The label type corresponding to the highest predicted probability of the type is the candidate box label type corresponding to the j-th first prediction unit; The confidence level of the candidate box corresponding to the j-th first prediction unit;

[0030] The candidate box prediction module is used to select up to M corresponding first candidate elements from the candidate box sequence Y according to the non-maximum suppression mechanism to form the corresponding candidate element set and output it.

[0031] Preferably, the step of recording the annotation behavior of the first annotator by combining the first image sequence and the first parsing sequence, and refreshing the corresponding first annotation trajectory and first target element set based on the recorded information, specifically includes:

[0032] When the first annotator selects a target element, the currently selected target element is taken as the corresponding currently selected target element; the current time is taken as the corresponding first timestamp and second timestamp; the global page number of the file page where the currently selected target element is located is taken as the current page number identifier; the first file page corresponding to the current page number identifier is taken as the current file page; and the first frame image and the first frame element set corresponding to the current file page in the first image sequence and the first parsing sequence are taken as the corresponding current frame image and current frame element set.

[0033] It also confirms whether the currently selected target element is the first selected target element of the first PDF file; if it is confirmed, the first annotation trajectory and the first target element set are initialized to empty.

[0034] The system confirms whether the currently selected target element is the first selected target element on the current file page. If yes, the system tracks and records the mouse movement of the first annotator on the current file page before the first annotator sets a label box for the currently selected target element by long-pressing the mouse. If no, the system tracks and records the mouse movement of the first annotator on the current file page from the previous target element with a completed label box to the currently selected target element before the first annotator sets a label box for the currently selected target element by long-pressing the mouse. During this tracking and recording process, a corresponding mouse trajectory point is periodically formed by the current page number identifier and the image coordinates of the mouse position on the corresponding current frame image at a preset sampling frequency. The tracking and recording process stops when the first annotator starts setting a label box by long-pressing the mouse. At the end of this tracking and recording process, all the mouse trajectory points obtained in this tracking and recording process are sampled at equal intervals according to the preset total number of mouse trajectory points, and the corresponding first mouse movement trajectory is formed by all the trajectory points obtained in this sampling process.

[0035] The image coordinates of the start and end positions of the first annotator's operation when setting the annotation box for the currently selected target element by long-pressing the mouse are recorded on the current frame image as the corresponding current start coordinates and current end coordinates; the current start coordinates are used as the corresponding first mouse click position; the horizontal and vertical displacement vectors from the current start coordinates to the current end coordinates form the corresponding first annotation box drag direction; the height and width of the annotation box of the currently selected target element in the current frame image are identified, and the corresponding first annotation box size is set based on the identification results; the image coordinates of the center point of the current annotation box in the current frame image are identified based on the current first mouse click position and the first annotation box size, and the corresponding center point of the first annotation box is set based on the identification results; and the label type set by the first annotator for the current annotation box is used as the corresponding first annotation type.

[0036] The first trajectory point is composed of the first timestamp corresponding to the currently selected target element, the first mouse sliding trajectory, the first mouse click position, the first annotation box drag direction, the first annotation box size, the first annotation box center point, and the first annotation type, and is added to the first annotation trajectory; and the first trajectory point added this time is used as the corresponding current trajectory point;

[0037] And assign a unique target identifier to the currently selected target element as the corresponding first target identifier;

[0038] Based on the first start position and the first end position of each of the first basic elements in the current frame element set, a corresponding rectangle is drawn on the current frame image as the corresponding basic element box; and the first basic element that intersects with the annotation box of the currently selected target element is taken as the corresponding parent basic element; and the first element identifier, the first page number identifier, and the first paragraph identifier of the current parent basic element are taken as the corresponding second element identifier, the second page number identifier, and the second paragraph identifier.

[0039] The first mouse click position, the center point of the first annotation box, and the first annotation type of the current trajectory point are taken as the corresponding first starting point, the first center point, and the first target type; the height and width of the first annotation box of the current trajectory point are taken as the corresponding first height and the first width; the annotation box area sub-image circled by the annotation box of the currently selected target element on the current frame image is taken as the corresponding first frame image; the sub-text information in the text attribute of the current parent basic element that is within the annotation box of the currently selected target element is taken as the corresponding first text; and the first annotation box is composed of the first starting point, the first center point, the first height, the first width, the first frame image, and the first text corresponding to the currently selected target element.

[0040] An empty multimodal feature is initialized for the currently selected target element as the corresponding first multimodal feature; and a corresponding first target element is added to the first target element set, consisting of the first target identifier, the second timestamp, the second element identifier, the second page number identifier, the second paragraph identifier, the first annotation box, the first target type, and the first multimodal feature corresponding to the currently selected target element.

[0041] Preferably, the provision of a candidate element set by a preset behavior prediction model for the next annotation target after each single-step annotation action of the first annotator based on the first annotation trajectory specifically includes:

[0042] When the first annotation trajectory completes a trajectory point, the first file page that the first annotator is currently annotating is taken as the current file page; and the first trajectory points that match the single page number identifier of the first mouse sliding trajectory in the first annotation trajectory with the current file page are extracted to form the corresponding first sub-trajectory;

[0043] The total number of trajectory points in the first sub-trajectory is taken as the current total number; the previous total number is identified; if the previous total number is greater than or equal to N, the sub-trajectory composed of the N nearest first trajectory points in the first sub-trajectory is taken as the corresponding second sub-trajectory; if the previous total number is less than N but greater than 0, the total number of trajectory points in the first sub-trajectory is increased to N by adding one or more preset zero-value trajectory points to the head of the first sub-trajectory, and the first sub-trajectory with completed trajectory point addition is taken as the corresponding second sub-trajectory; if the previous total number is 0, the corresponding second sub-trajectory is set to empty; the data format of the zero-value trajectory points is consistent with that of the first trajectory points, but all trajectory point attributes inside are 0;

[0044] When the second sub-trajectory is not empty, a corresponding behavior feature vector x is formed by the first timestamp of each first trajectory point of the second sub-trajectory, the first mouse sliding trajectory, the first mouse click position, the first annotation box drag direction, the first annotation box size, and the first annotation type. i ; and from the N behavioral feature vectors x obtained this time i A corresponding historical behavior sequence X is formed; and the current historical behavior sequence X is input into the behavior prediction model for processing to obtain the corresponding candidate element set;

[0045] When the candidate element set is not empty, the candidate box area corresponding to each of the first candidate elements in the candidate element set is highlighted on the current annotation interface of the first annotator; and the candidate box annotation type and the candidate box confidence of the current first candidate element are displayed synchronously in each displayed candidate box area.

[0046] Preferably, the step of improving the behavior prediction model based on the first annotator's selection feedback on candidate elements specifically includes:

[0047] When the candidate element set is first displayed, the consecutive failure counter is initialized to 0; and the boosted training dataset is initialized to empty.

[0048] Each time the candidate element set is displayed, it is determined whether the first annotator has selected a first candidate element from the current candidate element set as the corresponding current target element; if yes, the consecutive failure counter is cleared to zero; if no, the consecutive failure counter is incremented by 1.

[0049] And each time the first annotator fails to successfully select the current target element from the candidate element set displayed at that time, the target element selected by the first annotator is taken as the corresponding current manually selected element; and when the first target element corresponding to the current manually selected element is successfully added to the first target element set, the first starting point, first height, first width, and first target type corresponding to the current manually selected element are taken as a set of corresponding candidate box starting point, candidate box height, candidate box width, and candidate box label type; and the confidence level of the candidate box corresponding to the current manually selected element is set to 1; and a corresponding first label candidate box is formed by the candidate box starting point, candidate box height, candidate box width, candidate box label type, and candidate box confidence level corresponding to the current manually selected element; and M ’ Each identical first label candidate box forms a corresponding first label vector; and the historical behavior sequence X corresponding to the candidate element set displayed at the current time is used as a corresponding first training behavior sequence; and the first training behavior sequence corresponding to the currently selected element and the first label vector form a corresponding first data record which is added to the improvement training dataset;

[0050] Each time a record is added to the training dataset, the system checks whether the consecutive failure counter exceeds a preset failure threshold. If it does, the consecutive failure counter is reset to zero, and the first number of most recently added first data records in the training dataset are extracted to form the current dataset, where the first number is a preset positive integer. The system then copies the model structure and parameters of the embedding encoding layer, the behavior feature encoder, and the candidate box prediction module of the current behavior prediction model to obtain the corresponding first embedding layer, first encoder, and first prediction module. These copied first embedding layer, first encoder, and first prediction module are then sequentially connected to form the current training framework. The current training framework is then fine-tuned based on the current dataset. During this fine-tuning training, only the model parameters of the first prediction module are fine-tuned. At the end of this fine-tuning training, the model parameters of the candidate box prediction module of the current behavior prediction model are reset based on the current model parameters of the first prediction module.

[0051] Preferably, the multimodal feature recognition model includes a visual basic feature encoder, a visual layout feature encoder, an image feature encoder, a text feature encoder, a first feature mapping network, a second feature mapping network, a third feature mapping network, a fourth feature mapping network, an attention weighted fusion layer, a visual text feature splicing layer, and a feature dimensionality reduction layer.

[0052] The inputs of the visual basic feature encoder, the visual layout feature encoder, the image feature encoder, and the text feature encoder are all connected to the model input. The outputs of the visual basic feature encoder, the visual layout feature encoder, the image feature encoder, and the text feature encoder are respectively connected to the inputs of the corresponding first, second, third, and fourth feature mapping networks. The outputs of the first, second, and third feature mapping networks are respectively connected to the first, second, and third inputs of the attention-weighted fusion layer. The outputs of the attention-weighted fusion layer and the fourth feature mapping network are respectively connected to the first and second inputs of the visual-text feature splicing layer. The output of the visual-text feature splicing layer is connected to the input of the feature dimensionality reduction layer. The output of the feature dimensionality reduction layer is connected to the model output.

[0053] The visual fundamental feature encoder is used to extract the corresponding parent page image and the annotation box image from the element annotation box information; and to identify the color histogram of the annotation box image; and to identify the texture features of the annotation box image; and to identify the edge features of the annotation box image on the parent page image; and to send the corresponding visual fundamental feature tensor composed of the obtained color histogram, texture features and edge features to the first feature mapping network;

[0054] The visual layout feature encoder is used to extract the center point, height and width of the label box of the element label box information to form a corresponding visual layout feature vector and send it to the second feature mapping network;

[0055] The image feature encoder is implemented based on a CNN network model, a residual neural network model, or a Transformer framework encoder model; the image feature encoder is used to perform high-dimensional feature encoding processing on the image of the labeled bounding box information of the element to obtain the corresponding high-dimensional image feature tensor and send it to the third feature mapping network;

[0056] The text feature encoder is implemented based on the BERT series model or the encoder model of the Transformer framework; the text feature encoder is used to perform high-dimensional feature encoding on the text of the labeled box information of the element to obtain the corresponding high-dimensional text feature tensor and send it to the fourth feature mapping network;

[0057] The first feature mapping network is implemented based on a fully connected network, which is composed of one or more fully connected layers connected sequentially; the first feature mapping network is used to map a preset first feature vector space. As the corresponding target vector space, the visual basic feature tensor is mapped to the feature vector of the target vector space to obtain the corresponding first feature vector, which is then sent to the attention weighted fusion layer; d1 is the feature dimension of the first feature vector space, and the feature dimension d1 is a preset positive integer;

[0058] The second feature mapping network is implemented based on a fully connected network, which is composed of one or more fully connected layers connected sequentially; the second feature mapping network is used to map the first feature vector space. As the corresponding target vector space, the visual layout feature vector is mapped to the feature vector of the target vector space to obtain the corresponding second feature vector, which is then sent to the attention weighted fusion layer; the vector shapes of the first and second feature vectors are consistent.

[0059] The third feature mapping network is implemented based on a fully connected network, which is composed of one or more fully connected layers connected sequentially. The third feature mapping network is used to map the first feature vector space. As the corresponding target vector space, the image high-dimensional feature tensor is mapped to the feature vector of the target vector space to obtain the corresponding third feature vector, which is then sent to the attention weighted fusion layer; the vector shapes of the first and third feature vectors are consistent.

[0060] The fourth feature mapping network is implemented based on a fully connected network, which is composed of one or more fully connected layers connected sequentially. The fourth feature mapping network is used to map the first feature vector space. As the corresponding target vector space, the text high-dimensional feature tensor is mapped to the feature vector of the target vector space to obtain the corresponding fourth feature vector; and the fourth feature vector is sent to the visual text feature splicing layer as the corresponding annotation box text feature vector; the vector shapes of the first and fourth feature vectors are consistent;

[0061] The attention-weighted fusion layer consists of a first feature tensor composed of the received first, second, and third feature vectors; and performs corresponding query, key, and value vector transformations on the first feature tensor based on preset query, key, and value weight vectors to obtain corresponding query vector Q, key vector K, and value vector V; then performs attention operations based on the query vector Q, the key vector K, and the value vector V, and sends the operation result as the corresponding bounding box visual feature vector to the visual text feature splicing layer; the shape of the bounding box visual feature vector and the bounding box text feature vector are consistent;

[0062] The visual text feature concatenation layer is used to concatenate the visual feature vector of the annotation box and the text feature vector of the annotation box according to the vector concatenation method to obtain the corresponding concatenated feature vector and send it to the feature dimensionality reduction layer;

[0063] The feature reduction layer is implemented based on an MLP model; the feature reduction layer is used to reduce the preset second feature vector space. As the corresponding target vector space, the concatenated feature vector is mapped to the feature vector of the target vector space, and the resulting mapped feature vector is used as the corresponding multimodal element feature and output; d2 is the feature dimension of the second feature vector space, and the feature dimension d2 is a preset positive integer, d2 < d1.

[0064] Preferably, the step of adding corresponding multimodal element features to the target elements of the first target element set based on the preset multimodal feature recognition model specifically includes:

[0065] When the first annotator begins to select target elements in the first PDF file, an empty queue is initialized as the corresponding first identifier queue; and each time a first target element is added to the first target element set, the first target identifier of the first target element added at that time is added to the first identifier queue.

[0066] When the total number of identifiers in the first identifier queue is not zero, the earliest added first target identifier is taken as the current identifier; the first target element in the first target element set whose first target identifier matches the current identifier is taken as the current element; the first center point, first height, first width, first frame image, and first text of the current element are taken as a set of corresponding annotation box center point, annotation box height, annotation box width, annotation box image, and annotation box text; the first frame image corresponding to the second page number identifier of the current element is taken as the corresponding parent page image; and the parent page image, annotation box center point, annotation box height, annotation box width, annotation box image, and annotation box text corresponding to the current element form a corresponding element annotation box information; the current element annotation box information is input into the multimodal feature recognition model for processing to obtain the corresponding multimodal element features; and the first multimodal features of the current element are set based on the current multimodal element features; and when this setting ends, the current identifier is removed from the first identifier queue.

[0067] Preferably, the step of performing target matching and trajectory tracking processing on target elements according to the first target element set and refreshing the corresponding first trajectory set based on the tracking results specifically includes:

[0068] When the first annotator begins to select target elements in the first PDF file, an empty queue is initialized as the corresponding second identifier queue, and the first trajectory set is initialized to empty; and each time a multimodal element feature setting is completed for the first target element set, the first target identifier of the first target element corresponding to the first multimodal feature set in that current setting is added to the second identifier queue.

[0069] When the total number of identifiers in the second identifier queue is not zero, the earliest added first target identifier is taken as the current identifier; the total number of trajectories in the current first trajectory set is counted to obtain the corresponding current trajectory total number; the first target element in the first target element set whose first target identifier matches the current identifier is taken as the current element; and it is determined whether the current trajectory total number is zero; if yes, a corresponding first tracking trajectory is initialized in the first target element set based on the current element; if no, it is confirmed whether there is a first tracking trajectory in the current first target element set that matches the current element; if it is confirmed that there is, the trajectory is updated based on the first tracking trajectory that matches the current element; if it is confirmed that there is no, a corresponding first tracking trajectory is initialized in the first target element set based on the current element; and when the trajectory refresh operation ends, the current identifier is removed from the second identifier queue.

[0070] Furthermore, confirming whether there exists a first tracking trajectory matching the current element in the current first target element set specifically includes:

[0071] Step 101: Take the current element as the corresponding current observation point;

[0072] Step 102: Each of the first tracking trajectories in the first target element set is taken as the corresponding current tracking trajectory; the last second trajectory point of the current tracking trajectory is taken as the previous trajectory point; the area and aspect ratio of the corresponding bounding box are calculated based on the second height and second width of the previous trajectory point to obtain the corresponding previous area and aspect ratio; the area and aspect ratio of the corresponding bounding box are calculated based on the first height and first width of the current observation point to obtain the corresponding current area and aspect ratio; the shape similarity is calculated based on the previous area, the previous aspect ratio, the current area, and the current aspect ratio; the feature similarity between the second multimodal feature of the previous trajectory point and the first multimodal feature of the current observation point is calculated based on the cosine vector similarity algorithm; and the weighted sum of the shape similarity and the feature similarity is obtained to obtain the corresponding first similarity.

[0073] The shape similarity and the first similarity are calculated as follows:

[0074]

[0075] First similarity = ω3 × shape similarity + ω4 × feature similarity

[0076] ω1, ω2, ω3, and ω4 are four preset similarity weighting parameters;

[0077] Step 103: Take the first similarity that exceeds the preset first similarity threshold as the corresponding second similarity; and identify whether the number of second similarities is zero; if yes, confirm that there is no first tracking trajectory in the current first target element set that matches the current element; if no, confirm that there is a first tracking trajectory in the current first target element set that matches the current element, and take the first tracking trajectory corresponding to the largest second similarity as the first tracking trajectory that matches the current element.

[0078] Preferably, the step of performing cross-page element fusion processing on the first target element set based on the first trajectory set to obtain the corresponding parsed target set specifically includes:

[0079] Each of the first tracking trajectories in the first trajectory set is taken as the current trajectory; and all the second trajectory points of the current trajectory are sequentially traversed once; during this traversal, the currently traversed second trajectory point is taken as the current trajectory point, and the next second trajectory point of the current trajectory point is taken as the next trajectory point; and it is confirmed whether there is a continuous cross-page and continuous content association between the current trajectory point and the next trajectory point; if a continuous cross-page and continuous content association is confirmed, a corresponding continuous trajectory point mark is added to the current trajectory point and the next trajectory point; and at the end of this traversal, the corresponding parsing target subset is obtained by analyzing all the continuous trajectory point marks of the current trajectory; and all the parsing target subsets corresponding to the first trajectory set are merged to obtain the corresponding parsing target set; the parsing target subset consists of one or more first parsing targets.

[0080] Furthermore, confirming whether there is a continuous cross-page and continuous content association between the current trajectory point and the next trajectory point specifically includes:

[0081] Step 121: Calculate the difference between the third page number identifier of the next trajectory point and the current trajectory point to obtain the corresponding first page number difference; and identify whether the first page number difference is 1; if yes, proceed to step 122; if no, set the corresponding first confirmation status to no and proceed to step 126.

[0082] Step 122: Identify whether the two second target types of the current trajectory point and the next trajectory point match; if yes, proceed to step 123; if no, set the corresponding first confirmation status to no and proceed to step 126.

[0083] Step 123: Calculate the similarity between the two second multimodal features of the current trajectory point and the next trajectory point to the corresponding third similarity based on the cosine vector similarity algorithm; and identify whether the third similarity exceeds the preset second similarity threshold; if yes, proceed to step 124; if no, set the corresponding first confirmation state to no and proceed to step 126.

[0084] Step 124: Based on the second annotation box of the current trajectory point, confirm the image coordinates of the start and end positions of the element corresponding to the current trajectory point to obtain the corresponding first start coordinates and first end coordinates; and based on the second annotation box of the next trajectory point, confirm the image coordinates of the start and end positions of the element corresponding to the next trajectory point to obtain the corresponding second start coordinates and second end coordinates; calculate the horizontal spacing between the first and second start coordinates to obtain the corresponding first horizontal spacing; calculate the horizontal spacing between the first and second end coordinates to obtain the corresponding second horizontal spacing; calculate the linear distance between the first end coordinate and the preset single-page end position image coordinates to obtain the corresponding first coordinate spacing; and calculate the linear distance between the second start coordinate and the preset single-page start position image coordinates to obtain the corresponding second coordinate spacing.

[0085] Step 125: Take the second target type of the current trajectory point as the current type; and identify the current type; if the current type is a table or an image, identify whether the first and second horizontal spacings are both less than a preset first spacing threshold; if yes, set the corresponding first confirmation state to yes; otherwise, set the corresponding first confirmation state to no; if the current type is not a table or an image, identify whether the first and second coordinate spacings are both less than a preset second spacing threshold; if yes, set the corresponding first confirmation state to yes; otherwise, set the corresponding first confirmation state to no.

[0086] Step 126: Identify the first confirmation state; if the first confirmation state is yes, then confirm that there is a continuous cross-page and continuous content association between the current trajectory point and the next trajectory point; if the first confirmation state is no, then confirm that there is no continuous cross-page and continuous content association between the current trajectory point and the next trajectory point.

[0087] Preferably, the step of performing a label consistency check on the parsed target set specifically includes:

[0088] Each of the first parsing targets in the parsing target set is taken as the current target; and it is identified whether the fusion element group of the current target contains two second target identifiers; if so, the average feature of the two second multimodal features of the two second trajectory points corresponding to the current target is calculated to obtain the corresponding third multimodal feature; if not, the second multimodal feature of one second trajectory point corresponding to the current target is taken as the corresponding third multimodal feature.

[0089] The similarity between any two third multimodal features is calculated using a cosine similarity algorithm to obtain a corresponding fourth similarity. Based on the fourth similarity, all first parsed targets in the parsed target set are clustered to obtain multiple corresponding first target clusters. The first target clusters with a total number of targets greater than 1 are taken as the corresponding current clusters. The parsed target types of all first parsed targets within the current cluster are identified; otherwise, all first parsed targets in the current cluster form a corresponding secondary confirmation target set. Each first target cluster consists of one or more first parsed targets. In each first target cluster, if the total number of targets within the cluster is greater than 1, the fourth similarity between any two first parsed targets within the cluster exceeds a preset third similarity threshold.

[0090] When the total number of all the obtained secondary confirmation target sets is not zero, all the secondary confirmation target sets are output to the first annotator; and the first annotator is prompted to perform consistency checks and secondary type annotations on all the parsed target types in each of the secondary confirmation target sets; and the parsed target sets are corrected accordingly based on the feedback results of the first annotator on the consistency checks and secondary type annotations of each of the secondary confirmation target sets.

[0091] A second aspect of the present invention provides an apparatus for implementing the processing method for element annotation of PDF files as described in the first aspect above. The apparatus includes: an annotation object receiving module, an annotation preprocessing module, a first annotation tracking module, a second annotation tracking module, an annotation postprocessing module, and an annotation feedback module.

[0092] The annotation object receiving module is used to receive a first PDF file input by a first annotator; the first PDF file is the annotation object of the first annotator.

[0093] The annotation preprocessing module performs image conversion and basic element parsing processing on the first PDF file based on a preset PDF element parsing tool to obtain the corresponding first image sequence and first parsing sequence;

[0094] The first annotation tracking module is used to record the annotation behavior of the first annotator when the first annotator selects and annotates target elements in the first PDF file according to a preset tag type set, combining the first image sequence and the first parsing sequence, and to refresh the corresponding first annotation trajectory and first target element set based on the recorded information; and a preset behavior prediction model provides a candidate element set for the next annotation target after each single-step annotation behavior of the first annotator according to the first annotation trajectory; and to improve and train the behavior prediction model based on the first annotator's selection feedback of candidate elements.

[0095] The second annotation and tracking module is used to add corresponding multimodal element features to the target elements of the first target element set based on a preset multimodal feature recognition model during the annotation process; and to perform target matching and trajectory tracking processing on the target elements according to the first target element set and refresh the corresponding first trajectory set based on the tracking results;

[0096] The post-annotation processing module is used to perform cross-page element fusion processing on the first target element set according to the first trajectory set after the annotation is completed to obtain the corresponding parsed target set; and to perform annotation consistency check on the parsed target set.

[0097] The annotation feedback module is used to provide feedback on the parsed target set that has completed the consistency check to the first annotator.

[0098] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;

[0099] The processor is used to couple with the memory, read and execute instructions in the memory to implement the steps of the method described in the first aspect above;

[0100] The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

[0101] A fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions that, when executed by a computer, cause the computer to perform the instructions described in the first aspect.

[0102] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for annotating elements in PDF files. As described above, before initiating the annotation task, this invention pre-parses the PDF file using a PDF element parsing tool. Then, during the annotation process: 1) Based on real-time processing, the annotator's annotation trajectory (i.e., the first annotation trajectory) and the set of annotated elements (i.e., the first target element set) are updated in real-time according to the pre-parsed information; 2) Based on real-time processing, after each single-step annotation action, a behavior prediction model provides the annotator with a reference for the next candidate target (i.e., a candidate element set) based on the real-time annotation trajectory; 3) Based on asynchronous processing, the performance of the prediction model is continuously improved based on the annotator's candidate target feedback; 4) Based on an asynchronous processing mechanism, a multimodal feature recognition model is used to add multimodal (visual + text) element features to the annotated target elements; 5) Based on an asynchronous processing mechanism, associated target matching and associated target trajectory tracking are performed based on the geometric features of the annotation boxes (annotation box area, annotation box aspect ratio) and multimodal element features of the annotated elements, and the tracking trajectory set (i.e., the first trajectory set) is updated based on the processing results. After annotation, the cross-page elements in the annotated element set are first fused according to the tracking trajectory; then, the fused element set (parsing target set) undergoes an annotation consistency check, and during the check, the labels of similar elements are modified for consistency through human-computer collaboration; finally, the parsing target set that has completed the consistency check is fed back as the final annotation result. This embodiment of the invention improves annotation efficiency through a candidate element prediction mechanism, improves the recognition accuracy and fusion efficiency of cross-page elements through a related element tracking mechanism, and improves the annotation consistency of the entire text through a consistency check. Attached Figure Description

[0103] Figure 1 This is a schematic diagram of a method for annotating elements in a PDF file according to Embodiment 1 of the present invention;

[0104] Figure 2 A schematic diagram of the modules of the behavior prediction model provided in Embodiment 1 of the present invention;

[0105] Figure 3 This is a schematic diagram of the modules of the multimodal feature recognition model provided in Embodiment 1 of the present invention;

[0106] Figure 4 This is a module structure diagram of a processing device for annotating elements in a PDF file, provided in Embodiment 2 of the present invention;

[0107] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation

[0108] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0109] Embodiment 1 of the present invention provides a method for annotating elements in PDF files, such as... Figure 1 The diagram illustrates a method for annotating elements in a PDF file according to Embodiment 1 of the present invention. This method mainly includes the following steps:

[0110] Step 1: Receive the first PDF file input by the first annotator.

[0111] The first PDF file is the annotation object of the first annotator; the first PDF file includes multiple first file pages.

[0112] Step 2: Based on the preset PDF element parsing tool, perform image conversion and basic element parsing processing on the first PDF file to obtain the corresponding first image sequence and first parsing sequence.

[0113] Here, the PDF element parsing tools in this embodiment of the invention include software such as PyMuPDF, PDFMiner, and PDF-Extract-Kit.

[0114] The first image sequence is composed of multiple first frame images arranged in sequence; each first frame image corresponds one-to-one with a first file page.

[0115] The first parsing sequence is formed by sequentially sorting multiple first frame element sets; each first frame element set corresponds one-to-one with a first file page.

[0116] The first frame element set includes multiple first basic elements; each first basic element includes a first element identifier, a first page number identifier, a first paragraph identifier, a first start position, a first end position, a first element type, and a first element attribute set. Among them:

[0117] 1) The first element identifier is the unique element identifier of the current basic element;

[0118] 2) The first page number is the global page number of the file page where the current basic element is located;

[0119] 3) The first paragraph identifier is the global paragraph number of the paragraph containing the current basic element;

[0120] 4) The first start position and the first end position are the pixel coordinates of the start and end positions of the current basic element in the corresponding first frame image, respectively;

[0121] 5) The first element type includes article title, chapter title, abstract, paragraph, table, and image;

[0122] 6) The attribute types of the first element attribute set correspond to the first element type, specifically: when the first element type is an article title, the first element attribute set includes the first title text; when the first element type is a chapter title, the first element attribute set includes the first chapter level and the first title text; when the first element type is an abstract, the first element attribute set includes the first abstract text; when the first element type is a paragraph, the first element attribute set includes the first paragraph text; when the first element type is a table, the first element attribute set includes the first table height, the first table width, and the first table text, where the first table text is a CSV text file; when the first element type is an image, the first element attribute set includes the first image height, the first image width, the first image file, and the first image description text.

[0123] Step 3: When the first annotator selects and annotates target elements in the first PDF file according to the preset tag type set, the annotation behavior of the first annotator is recorded by combining the first image sequence and the first parsing sequence, and the corresponding first annotation trajectory and first target element set are refreshed based on the recorded information; and the preset behavior prediction model provides a candidate element set for the next annotation target after each single-step annotation behavior of the first annotator according to the first annotation trajectory; and the behavior prediction model is improved and trained based on the first annotator's selection feedback of candidate elements.

[0124] Here, the tag type set in this embodiment of the invention includes a basic type subset and a domain type subset. Both the basic type subset and the domain type subset consist of one or more tag types. The tag types of the basic type subset include article titles, chapter titles, abstracts, paragraphs, tables, and images. The tag types of the domain type subset correspond to the file type of the first PDF file. Specifically: if the file type of the first PDF file is an academic paper, then the tag types of the domain type subset include mathematical symbols, physical symbols, mathematical formulas, and chemical formulas; if the file type of the first PDF file is a technical report, then the tag types of the domain type subset include formulas, technical terms, and technical parameters; if the file type of the first PDF file is a legal document, then the tag types of the domain type subset include legal terms and legal clauses.

[0125] The first annotation trajectory in this embodiment of the invention is formed by sequentially sorting one or more first trajectory points. The trajectory point attributes of each first trajectory point include a first timestamp, a first mouse sliding trajectory, a first mouse click position, a first annotation box drag direction, a first annotation box size, a first annotation box center point, and a first annotation type. The first mouse movement trajectory is a single-page mouse movement trajectory. The trajectory point attribute of each point in the first mouse movement trajectory consists of a single-page page number identifier and single-page image coordinates of the mouse position. The first mouse click position is the image coordinates of the mouse's starting position when the first annotator sets the annotation box for the currently selected target element by long-pressing the mouse, specifically a pixel coordinate on the corresponding first frame image. The first annotation box drag direction is the direction in which the first annotator drags the mouse when setting the annotation box for the currently selected target element by long-pressing the mouse, specifically composed of two components: horizontal drag direction and vertical drag direction. The first annotation box size is the size of the annotation box corresponding to the currently selected target element, specifically composed of two components: height and width. The first annotation box center point is the center point of the annotation box corresponding to the currently selected target element, specifically a pixel coordinate on the corresponding first frame image. The first annotation type is a type of label in the label type set.

[0126] The first target element set in this embodiment of the invention consists of one or more first target elements. Each first target element corresponds to a selected target element, a first trajectory point, and a parent basic element; the parent basic element is actually a corresponding first basic element. The element content of each first target element is part or all of the element content of the parent basic element.

[0127] The first target element in this embodiment of the invention includes a first target identifier, a second timestamp, a second element identifier, a second page number identifier, a second paragraph identifier, a first annotation box, a first target type, and a first multimodal feature. Wherein:

[0128] The first target identifier is the unique target identifier of the current target element;

[0129] The second timestamp is the first timestamp of the corresponding first trajectory point;

[0130] The second element identifier, the second page number identifier, and the second paragraph identifier are respectively the first element identifier, the first page number identifier, and the first paragraph identifier of the corresponding first basic element;

[0131] The first annotation box includes a first starting point, a first center point, a first height, a first width, a first frame diagram, and first text; the first starting point and the first center point are respectively the first mouse click position of the corresponding first trajectory point and the center point of the first annotation box; the first height and the first width are the height and width of the first annotation box size of the corresponding first trajectory point; the first frame diagram is the annotation box area sub-image on the first frame image corresponding to the corresponding second page number identifier, defined by the current annotation box; the first text is the sub-text information within the current annotation box range or corresponding to the current annotation box in the text attribute set of the first element of the corresponding first basic element.

[0132] The first target type is the first annotation type of the corresponding first trajectory point;

[0133] The first multimodal feature is the multimodal element feature of the current target element; this multimodal element feature is formed by fusing the corresponding bounding box visual features and bounding box text features; the bounding box visual features are formed by fusing the corresponding visual basic features, visual layout features and image high-dimensional features through attention weighting; at the initial creation time of the first target element, the first multimodal feature is set to empty.

[0134] like Figure 2 The diagram shows a module of the behavior prediction model provided in Embodiment 1 of the present invention. The behavior prediction model of the present invention is used to perform candidate element prediction processing based on the historical behavior sequence X input to the model and output the corresponding candidate element set.

[0135] Wherein, the historical behavior sequence X consists of N behavior feature vectors x i Sort chronologically, where 1 ≤ index i ≤ N, and the number N is a pre-defined positive integer; each behavior is a feature vector x. i Corresponding to a first trajectory point; each behavior feature vector x i The vector feature data consists of the first timestamp of the corresponding first trajectory point, the first mouse movement trajectory, the first mouse click position, the first annotation box drag direction, the first annotation box size, and the first annotation type. The candidate element set consists of at most M first candidate elements; the number M is a preset positive integer; the first candidate element includes the candidate box starting point, candidate box height, candidate box width, candidate box annotation type, and candidate box confidence; the candidate box annotation type is a label type from the label type set; the candidate box confidence is a value between 0 and 1; each first candidate element corresponds to one candidate annotation box.

[0136] like Figure 2 As shown, the behavior prediction model consists of an embedding encoding layer, a behavior feature encoder, a candidate box prediction module, and a candidate box filtering module connected sequentially. The functionalities of the model components of the behavior prediction model are shown below.

[0137] 1) Embedded coding layer:

[0138] Embedded coding layers are used to define the behavior feature vector x i The embedding encoding rules corresponding to each vector feature data, for each behavioral feature vector x i Each vector feature data is embedded and encoded to obtain the corresponding feature encoding vector; and each behavioral feature vector x is used to... i All feature encoding vectors are sequentially concatenated into a corresponding embedding encoding vector; and N embedding encoding vectors form a corresponding embedding encoding sequence E, which is sent to the behavior feature encoder.

[0139] Here, the behavioral feature vector x i The embedding encoding rules corresponding to each vector feature data can be customized based on actual application needs, and the embodiments of the present invention do not impose specific technical limitations on them.

[0140] 2) Behavioral feature encoder:

[0141] The behavioral feature encoder is implemented based on an LSTM model or a Transformer architecture encoder model. The behavioral feature encoder is used to perform feature encoding processing on the embedded encoding sequence E to obtain the corresponding feature vector H, which is then sent to the candidate box prediction module.

[0142] 3) Candidate box prediction module:

[0143] The candidate box prediction module consists of M ’ It consists of M parallel first prediction units. ’ M is a preset positive integer. ’ >M.

[0144] The candidate box prediction module is used to predict a single candidate box y based on the feature vector H using each first prediction unit. j ; and from the obtained M ’ y prediction candidate boxes j The candidate boxes are sequentially sorted to form a corresponding candidate box sequence Y, which is then sent to the candidate box filtering module. Where 1 ≤ index j ≤ M ’ Each predicted candidate box y j Each includes a corresponding set of candidate box starting point, candidate box height, candidate box width, candidate box annotation type, and candidate box confidence level.

[0145] The prediction method for the j-th first prediction unit is:

[0146]

[0147] in, The four weight vector parameters are the j-th first prediction unit. These are the four offset vector parameters corresponding to the j-th first prediction unit; Let $\mathbf{j}$ be the candidate box position prediction vector corresponding to the first prediction unit. The vector length is 3, and it consists of the starting point, height and width of the candidate box corresponding to the j-th first prediction unit; The candidate box type prediction vector corresponding to the j-th first prediction unit is composed of multiple type prediction probabilities; the total number of type prediction probabilities is consistent with the total number of label types in the label type set; the candidate box type prediction vector... The label type corresponding to the highest type prediction probability is the candidate box label type corresponding to the j-th first prediction unit; is the confidence score of the candidate box corresponding to the j-th first prediction unit.

[0148] 4) Candidate box prediction module:

[0149] The candidate box prediction module is used to select at most M corresponding first candidate elements from the candidate box sequence Y according to the non-maximum suppression mechanism, form the corresponding candidate element set, and output it. Specifically:

[0150] Step A1: When the candidate box prediction module receives the candidate box sequence Y, it initializes an empty sequence as the corresponding current target sequence; and uses the current candidate box sequence Y as the corresponding current source sequence; and initializes the candidate element set to empty.

[0151] Step A2: Select the predicted candidate box y with the highest confidence in the current source sequence. j As the current candidate box; and identify whether the confidence of the current candidate box is less than the preset confidence threshold; if yes, proceed to step A7; if no, add the current candidate box to the current target sequence and delete the current candidate box from the current source sequence, and proceed to step A3;

[0152] Here, the confidence threshold is a pre-set threshold parameter;

[0153] Step A3: Identify whether the total number of candidate boxes in the current source sequence is 0; if yes, proceed to step A7; if no, proceed to step A4.

[0154] Step A4, select each predicted candidate box y in the current source sequence. j As the corresponding current comparison box; and calculate the corresponding comparison-candidate box intersection-union ratio (IOU) based on the candidate box starting point, candidate box height, and candidate box width of the current comparison box and the current candidate box;

[0155] Step A5: In the current source sequence, select the predicted candidate boxes y whose alignment-candidate box Intersection over Union (IOU) is greater than a preset IOU threshold and whose candidate box confidence is less than a confidence threshold. j delete;

[0156] Here, the intersection-union ratio threshold is a pre-set threshold parameter;

[0157] Step A6: Identify whether the total number of candidate boxes in the current source sequence is 0; if yes, proceed to step A7; if no, return to step A2.

[0158] Step A7: Identify whether the current target sequence is empty; if yes, proceed to step A9; if no, proceed to step A8.

[0159] Step A8: First, identify the predicted candidate boxes whose starting points in the current target sequence exceed the preset single-page coordinate range. j Delete; then sort all predicted candidate boxes y for the current target sequence in descending order of candidate box confidence. j The candidate boxes are reordered; the total number of candidate boxes in the reordered current target sequence is identified; if the total number of candidate boxes is greater than or equal to M, then the top M predicted candidate boxes y in the current target sequence are selected. j Add the corresponding M first candidate elements to the candidate element set; if the total number of current candidate boxes is less than M but greater than 0, then add each predicted candidate box y in the current target sequence. j As a corresponding first candidate element, all first candidate elements obtained this time are added to the candidate element set;

[0160] Here, in this embodiment of the invention, after performing image conversion on the first PDF file, multiple first frame images of the same size and resolution are obtained. By analyzing the image coordinates of the first frame images, the coordinate range of a single page can be obtained.

[0161] Step A9: Output the candidate element set obtained this time.

[0162] The specific steps for implementing step 3 include:

[0163] Step 31: Combine the first image sequence and the first parsing sequence to record the annotation behavior of the first annotator and refresh the corresponding first annotation trajectory and first target element set based on the recorded information;

[0164] Specifically, this includes: step 311, when the first annotator selects a target element, the currently selected target element is taken as the corresponding currently selected target element; the current time is taken as the corresponding first timestamp and second timestamp; the global page number of the file page where the currently selected target element is located is taken as the current page number identifier; the first file page corresponding to the current page number identifier is taken as the current file page; and the first frame image and the first frame element set corresponding to the current file page in the first image sequence and the first parsing sequence are taken as the corresponding current frame image and the current frame element set.

[0165] Step 312, and confirm whether the currently selected target element is the first selected target element of the first PDF file; if confirmed, initialize the first annotation trajectory and the first target element set to empty;

[0166] Step 313, and confirm whether the currently selected target element is the first selected target element on the current file page; if confirmed, track and record the first annotator's mouse movement on the current file page before the first annotator sets the annotation box for the currently selected target element by long-pressing the mouse; if not confirmed, track and record the first annotator's mouse movement from the previous target element with an annotation box set to the currently selected target element on the current file page before the first annotator sets the annotation box for the currently selected target element by long-pressing the mouse; during this tracking and recording process, based on a preset sampling frequency, periodically form a corresponding mouse trajectory point by the current page number identifier and the image coordinates of the mouse position on the corresponding current frame image at the current moment; stop this tracking and recording when the first annotator starts setting the annotation box by long-pressing the mouse; and at the end of this tracking and recording, sample all the mouse trajectory points obtained in this tracking and recording at equal intervals according to the preset total number of mouse trajectory points, and form the corresponding first mouse movement trajectory by all the trajectory points obtained in this sampling.

[0167] Here, the sampling frequency is a preset time frequency; the total number of mouse trajectory points is a preset positive integer;

[0168] Step 314: Record the image coordinates of the start and end positions of the first annotator's operation when setting the annotation box for the currently selected target element by long-pressing the mouse, as the corresponding current start coordinates and current end coordinates; use the current start coordinates as the corresponding first mouse click position; and form the corresponding first annotation box drag direction by the horizontal and vertical displacement vectors from the current start coordinates to the current end coordinates; identify the height and width of the annotation box of the currently selected target element in the current frame image and set the corresponding first annotation box size based on the identification results; identify the image coordinates of the center point of the current annotation box in the current frame image based on the current first mouse click position and the first annotation box size, and set the corresponding first annotation box center point based on the identification results; and use the label type set by the first annotator for the current annotation box as the corresponding first annotation type.

[0169] Step 315: A first trajectory point is added to the first annotation trajectory, consisting of the first timestamp, first mouse movement trajectory, first mouse click position, first annotation box drag direction, first annotation box size, first annotation box center point, and first annotation type corresponding to the currently selected target element; and the first trajectory point added this time is used as the corresponding current trajectory point.

[0170] Step 316, and assign a unique target identifier to the currently selected target element as the corresponding first target identifier;

[0171] Step 317: Based on the first start position and the first end position of each first basic element in the current frame element set, draw a corresponding rectangle on the current frame image as the corresponding basic element box; and take the first basic element that intersects with the annotation box of the currently selected target element as the corresponding parent basic element; and take the first element identifier, first page number identifier, and first paragraph identifier of the current parent basic element as the corresponding second element identifier, second page number identifier, and second paragraph identifier.

[0172] Step 318: The first mouse click position, the center point of the first annotation box, and the first annotation type of the current trajectory point are taken as the corresponding first starting point, first center point, and first target type; the height and width of the first annotation box of the current trajectory point are taken as the corresponding first height and first width; the annotation box area sub-image defined by the annotation box of the currently selected target element on the current frame image is taken as the corresponding first frame image; the sub-text information in the text attribute of the current parent basic element that is within the annotation box of the currently selected target element is taken as the corresponding first text; and the first annotation box is composed of the first starting point, first center point, first height, first width, first frame image, and first text corresponding to the currently selected target element.

[0173] Step 319: Initialize an empty multimodal feature for the currently selected target element as the corresponding first multimodal feature; and add a corresponding first target element to the first target element set, which is composed of the first target identifier, second timestamp, second element identifier, second page number identifier, second paragraph identifier, first annotation box, first target type, and first multimodal feature corresponding to the currently selected target element.

[0174] Step 32, and the preset behavior prediction model provides a set of candidate elements for the next annotation target after each single-step annotation behavior of the first annotator based on the first annotation trajectory;

[0175] Specifically, it includes: step 321, when the first annotation trajectory completes each trajectory point, the first file page that the first annotator is currently annotating is taken as the current file page; and the first trajectory point that matches the single page number identifier of the first mouse sliding trajectory in the first annotation trajectory with the current file page is extracted to form the corresponding first sub-trajectory;

[0176] Step 322: The total number of trajectory points in the first sub-trajectory is taken as the current total number; the previous total number is identified; if the previous total number is greater than or equal to N, the sub-trajectory composed of the N nearest first trajectory points in the first sub-trajectory is taken as the corresponding second sub-trajectory; if the previous total number is less than N but greater than 0, the total number of trajectory points in the first sub-trajectory is increased to N by adding one or more preset zero-value trajectory points to the head of the first sub-trajectory, and the first sub-trajectory with the completed trajectory point addition is taken as the corresponding second sub-trajectory; if the previous total number is 0, the corresponding second sub-trajectory is set to empty;

[0177] Here, the data format of the zero-value trajectory point in this embodiment of the invention is consistent with that of the first trajectory point, but all the trajectory point attributes inside are 0;

[0178] Step 323, and when the second sub-trajectory is not empty, a corresponding behavior feature vector x is formed by the first timestamp of each first trajectory point of the second sub-trajectory, the first mouse sliding trajectory, the first mouse click position, the first annotation box drag direction, the first annotation box size, and the first annotation type. i ; and from the N behavioral feature vectors x obtained this time i A corresponding historical behavior sequence X is formed; and the current historical behavior sequence X is input into the behavior prediction model for processing to obtain the corresponding candidate element set.

[0179] Step 324: When the candidate element set is not empty, highlight the candidate box area corresponding to each first candidate element in the candidate element set on the current annotation interface of the first annotator; and synchronously display the candidate box annotation type and candidate box confidence of the current first candidate element in each displayed candidate box area.

[0180] Step 33, and improve the behavior prediction model based on the first annotator's feedback on the selection of candidate elements;

[0181] Specifically, this includes: step 331, when displaying the candidate element set for the first time, initializing the consecutive failure counter to 0; and initializing the boosted training dataset to be empty;

[0182] Step 332: Each time the candidate element set is displayed, identify whether the first annotator has selected a first candidate element from the current candidate element set as the corresponding current target element; if yes, clear the consecutive failure counter; if no, increment the consecutive failure counter by 1.

[0183] Step 333: Each time the first annotator fails to successfully select the current target element from the currently displayed candidate element set, the target element selected by the first annotator is taken as the corresponding currently selected element. When the first target element corresponding to the currently selected element is successfully added to the first target element set, the first starting point, first height, first width, and first target type corresponding to the currently selected element are taken as a set of corresponding candidate box starting point, candidate box height, candidate box width, and candidate box label type. The confidence level of the candidate box corresponding to the currently selected element is set to 1. A corresponding first label candidate box is formed by the candidate box starting point, candidate box height, candidate box width, candidate box label type, and candidate box confidence level corresponding to the currently selected element. And M... ’ Each identical first label candidate box forms a corresponding first label vector; and the historical behavior sequence X corresponding to the candidate element set displayed at this time is used as a corresponding first training behavior sequence; and the first training behavior sequence corresponding to the currently selected element and the first label vector form a corresponding first data record, which is added to the improved training dataset.

[0184] Step 334: When adding a record to the training dataset, identify whether the current consecutive failure counter exceeds a preset failure count threshold. If it does, clear the consecutive failure counter and extract the first data record with the first number of most recently added records from the training dataset to form the corresponding current dataset. Copy the model structure and model parameters of the embedding encoding layer, behavior feature encoder, and candidate box prediction module of the current behavior prediction model to obtain the corresponding first embedding layer, first encoder, and first prediction module. The copied first embedding layer, first encoder, and first prediction module are sequentially connected to form the corresponding current training framework. Perform a round of fine-tuning training on the current training framework based on the current dataset. During this round of fine-tuning training, only the model parameters of the first prediction module are fine-tuned. At the end of this round of fine-tuning training, reset the model parameters of the candidate box prediction module of the current behavior prediction model based on the current model parameters of the first prediction module.

[0185] Among them, the failure count threshold and the first quantity are two pre-set positive integers.

[0186] It should be noted that the behavior prediction model in this embodiment of the invention has completed one round of model training based on a pre-collected dataset before its first use; however, during use, the candidate box prediction module of the behavior prediction model can be fine-tuned based on the annotator's personalized data (i.e., improving the training dataset).

[0187] Furthermore, this embodiment of the invention does not impose specific technical limitations on the training and fine-tuning process of the behavior prediction model; users can customize it based on specific application needs. Generally, supervised training / fine-tuning is used for training / fine-tuning.

[0188] For example, during fine-tuning, the first training action sequence of each first data record in the training dataset is input into the current training framework for processing to obtain the corresponding candidate box sequence Y as the corresponding first prediction vector; each first prediction vector and its corresponding first label vector form a first prediction-label pair; all the obtained first prediction-label pairs are then fed into a preset first model loss function; and a preset first model optimizer (such as the Adam optimizer, SGD optimizer, etc.) is used to fine-tune the model parameters of the first prediction module of the current training framework in the direction that minimizes the first model loss function. Here, the first model loss function can be composed of three parts: candidate box position loss, candidate box classification loss, and candidate box confidence loss; the candidate box position loss is used to calculate the candidate box position prediction vector. The prediction-label loss and the candidate box classification loss are used to calculate the candidate box type prediction vector. The prediction-label loss and the candidate box confidence loss are used to calculate the candidate box confidence. The prediction-label loss; candidate box position loss and candidate box confidence loss can be implemented based on L1 loss function, L2 loss function or RMSE loss function; candidate box classification loss can be implemented based on cross-entropy loss function.

[0189] Step 4: During the annotation process, based on the preset multimodal feature recognition model, add corresponding multimodal element features to the target elements of the first target element set; and perform target matching and trajectory tracking processing on the target elements according to the first target element set, and refresh the corresponding first trajectory set based on the tracking results.

[0190] like Figure 3 As shown in the schematic diagram of the multimodal feature recognition model provided in Embodiment 1 of the present invention, the multimodal feature recognition model of the present invention is used to extract visual basic features, visual layout features, image features and text features based on the element annotation box information input by the model, and to fuse the three types of visual features based on the attention weighted fusion mechanism, and to perform feature dimensionality reduction on the spliced ​​features of visual fusion features and text features to obtain the corresponding multimodal element features.

[0191] The element annotation box information corresponds to a first target element in the first target element set; the element annotation box information includes the parent page image, annotation box center point, annotation box height, annotation box width, annotation box image, and annotation box text; the parent page image is the first frame image of the file page where the corresponding target element is located; the annotation box center point, annotation box height, annotation box width, annotation box image, and annotation box text are the first center point, first height, first width, first frame image, and first text of the corresponding target element.

[0192] like Figure 3 As shown, the model components of the multimodal feature recognition model include: a visual basic feature encoder, a visual layout feature encoder, an image feature encoder, a text feature encoder, a first feature mapping network, a second feature mapping network, a third feature mapping network, a fourth feature mapping network, an attention-weighted fusion layer, a visual-text feature splicing layer, and a feature dimensionality reduction layer.

[0193] The connection relationships of the model components in the multimodal feature recognition model are as follows: the inputs of the visual basic feature encoder, visual layout feature encoder, image feature encoder, and text feature encoder are all connected to the model input; the outputs of the visual basic feature encoder, visual layout feature encoder, image feature encoder, and text feature encoder are connected to the inputs of the corresponding first, second, third, and fourth feature mapping networks, respectively; the outputs of the first, second, and third feature mapping networks are connected to the first, second, and third inputs of the attention-weighted fusion layer, respectively; the outputs of the attention-weighted fusion layer and the fourth feature mapping network are connected to the first and second inputs of the visual-text feature concatenation layer, respectively; the output of the visual-text feature concatenation layer is connected to the input of the feature dimensionality reduction layer; and the output of the feature dimensionality reduction layer is connected to the model output.

[0194] The model component functions of the multimodal feature recognition model are shown below.

[0195] 1) Visual basic feature encoder:

[0196] The visual basic feature encoder is used to extract the corresponding parent page image and annotation box image from the element annotation box information; and to identify the color histogram of the annotation box image; to identify the texture features of the annotation box image; and to identify the edge features of the annotation box image on the parent page image; and to send the corresponding visual basic feature tensor composed of the obtained color histogram, texture features and edge features to the first feature mapping network.

[0197] Here, the visual basic feature encoder uses existing image processing tools such as OpenCV, scikit-image, and MATLAB to perform color histogram, texture feature, and edge feature recognition.

[0198] 2) Visual layout feature encoder:

[0199] The visual layout feature encoder is used to extract the center point, height, and width of the element annotation box information to form the corresponding visual layout feature vector, which is then sent to the second feature mapping network.

[0200] 3) Image feature encoder:

[0201] Image feature encoders are implemented based on CNN network models, residual neural network models, or Transformer framework encoder models. The image feature encoder performs high-dimensional feature encoding on the bounding box images of element annotation information to obtain the corresponding high-dimensional image feature tensors, which are then sent to a third feature mapping network.

[0202] 4) Text Feature Encoder:

[0203] The text feature encoder is implemented based on the BERT series models or the encoder model of the Transformer framework. The text feature encoder is used to perform high-dimensional feature encoding on the text of the bounding box information of the element annotation box to obtain the corresponding high-dimensional text feature tensor, which is then sent to the fourth feature mapping network.

[0204] 5) First Feature Mapping Network:

[0205] The first feature mapping network is implemented based on a fully connected network, which is composed of one or more fully connected layers connected sequentially. The first feature mapping network is used to map a preset first feature vector space. As the corresponding target vector space, the first feature vector is obtained by mapping the visual basic feature tensor to the feature vector of the target vector space and sending it to the attention weighted fusion layer.

[0206] Where d1 is the feature dimension of the first feature vector space, and the feature dimension d1 is a preset positive integer.

[0207] 6) Second Feature Mapping Network:

[0208] The second feature mapping network is implemented based on a fully connected network, which is composed of one or more fully connected layers connected sequentially. The second feature mapping network is used to map the first feature vector space. As the corresponding target vector space, the visual layout feature vector is mapped to the feature vector of the target vector space to obtain the corresponding second feature vector, which is then sent to the attention weighted fusion layer.

[0209] Here, the vector shapes of the first and second feature vectors are kept consistent.

[0210] 7) Third Feature Mapping Network:

[0211] The third feature mapping network is implemented based on a fully connected network, which is composed of one or more fully connected layers connected sequentially. The third feature mapping network is used to map the first feature vector space. As the corresponding target vector space, the image high-dimensional feature tensor is mapped to the feature vector of the target vector space to obtain the corresponding third feature vector, which is then sent to the attention weighted fusion layer.

[0212] Here, the vector shapes of the first, second, and third eigenvectors are all kept consistent.

[0213] 8) Fourth Feature Mapping Network:

[0214] The fourth feature mapping network is implemented based on a fully connected network, which is composed of one or more fully connected layers connected sequentially. The fourth feature mapping network is used to map the first feature vector space... As the corresponding target vector space, the feature vector of the high-dimensional feature tensor of the text is mapped to the feature vector of the target vector space to obtain the corresponding fourth feature vector; and the fourth feature vector is sent to the visual text feature splicing layer as the corresponding annotation box text feature vector.

[0215] Here, the vector shapes of the first, second, third, and fourth eigenvectors are all kept consistent.

[0216] 9) Attention-weighted fusion layer:

[0217] The attention-weighted fusion layer consists of the first, second, and third feature vectors received, forming a corresponding first feature tensor. Based on the preset query, key, and value weight vectors, the first feature tensor is transformed into corresponding query vectors Q, key vector K, and value vector V. Attention operations are performed based on the query vector Q, key vector K, and value vector V, and the results are sent to the visual text feature splicing layer as the corresponding bounding box visual feature vectors.

[0218] Here, the preset query, key, and value weight vectors are three pre-defined weight vectors W. Q W K W V Let Z be the first feature tensor. Then the query vector Q, key vector K, and value vector V are: Q = W Q Z, K = W K Z, V = W V Z; Let Z be the visual feature vector of the annotation box. * ,So:

[0219] d k Let K be the feature dimension of the key vector.

[0220] It should be noted that the shape of the visual feature vector of the annotation box and the text feature vector of the annotation box (that is, the fourth feature vector) in the embodiments of the present invention is consistent.

[0221] 10) Visual text feature splicing layer:

[0222] The visual text feature concatenation layer is used to concatenate the visual feature vector and the text feature vector of the annotation box according to the vector concatenation method to obtain the corresponding concatenated feature vector, which is then sent to the feature dimensionality reduction layer.

[0223] 11) Feature dimensionality reduction layer:

[0224] The feature reduction layer is implemented based on an MLP model. The feature reduction layer is used to reduce the size of the predefined second feature vector space. As the corresponding target vector space, the concatenated feature vectors are mapped to the feature vectors of the target vector space, and the resulting mapped feature vectors are used as the corresponding multimodal element features and output.

[0225] Where d2 is the feature dimension of the second feature vector space, and the feature dimension d2 is a preset positive integer, d2 < d1.

[0226] As can be seen from the above, the multimodal feature recognition model in this embodiment of the invention is actually a pre-trained model of a multimodal feature encoder.

[0227] It should be noted that the multimodal feature recognition model in this embodiment of the invention has been trained before its first use and will not be adjusted during use. This embodiment of the invention does not impose specific technical limitations on the model training process of the multimodal feature recognition model; users can customize it based on specific application needs. Generally, there are two types of training methods for such pre-trained models: supervised training and unsupervised training. Supervised training typically involves connecting the multimodal feature recognition model with multiple downstream image and text processing task models to form a unified multi-task training framework. The framework is then trained based on the label datasets of all downstream image and text processing tasks, and the model parameters of the multimodal feature recognition model within the framework are used as the final training result at the end of training. Unsupervised training generally involves treating the multimodal feature recognition model as a generator model and connecting it to a discriminator model. The generator model and the discriminator model together form a generative adversarial network (GAN), which is then trained using unsupervised training methods. At the end of training, the model parameters of the generator model are used as the final training result.

[0228] The first trajectory set in this embodiment of the invention consists of one or more first tracking trajectories. Each first tracking trajectory corresponds to a first trajectory identifier. Each first tracking trajectory is formed by sequentially sorting one or more second trajectory points. The trajectory point attributes of each second trajectory point include a second target identifier, a third timestamp, a third element identifier, a third page number identifier, a third paragraph identifier, a second annotation box, a second target type, and a second multimodal feature; the second annotation box includes a second starting point, a second center point, a second height, a second width, a second diagram, and second text.

[0229] The specific steps for implementing step 4 include:

[0230] Step 41: Add corresponding multimodal element features to the target elements of the first target element set based on the preset multimodal feature recognition model;

[0231] Specifically, this includes: Step 411, when the first annotator begins to select target elements in the first PDF file, an empty queue is initialized as the corresponding first identifier queue; and each time a first target element is added to the first target element set, the first target identifier of the first target element added at that time is added to the first identifier queue;

[0232] Step 412: When the total number of identifiers in the first identifier queue is not zero, the earliest added first target identifier is taken as the current identifier; the first target element in the first target element set whose first target identifier matches the current identifier is taken as the current element; the first center point, first height, first width, first frame, and first text of the current element are taken as a set of corresponding annotation box center point, annotation box height, annotation box width, annotation box image, and annotation box text; the first frame image corresponding to the second page number identifier of the current element is taken as the corresponding parent page image; and the parent page image, annotation box center point, annotation box height, annotation box width, annotation box image, and annotation box text of the current element are combined to form a corresponding element annotation box information; the current element annotation box information is input into the multimodal feature recognition model for processing to obtain the corresponding multimodal element features; and the first multimodal feature of the current element is set based on the current multimodal element features; and when this setting is completed, the current identifier is removed from the first identifier queue.

[0233] Step 42, and perform target matching and trajectory tracking processing on the target elements according to the first target element set, and refresh the corresponding first trajectory set based on the tracking results;

[0234] Specifically, this includes: Step 421, when the first annotator starts selecting target elements in the first PDF file, an empty queue is initialized as the corresponding second identifier queue, and the first trajectory set is initialized to empty; and each time a multimodal element feature setting is completed for the first target element set, the first target identifier of the first target element corresponding to the first multimodal feature set at that time is added to the second identifier queue;

[0235] Step 422: When the total number of identifiers in the second identifier queue is not zero, the earliest added first target identifier is taken as the current identifier; the total number of trajectories in the current first trajectory set is counted to obtain the corresponding current trajectory total number; the first target element in the first target element set whose first target identifier matches the current identifier is taken as the current element; and it is determined whether the total number of current trajectories is zero; if yes, a corresponding first tracking trajectory is initialized in the first target element set based on the current element; if no, it is confirmed whether there is a first tracking trajectory in the current first target element set that matches the current element; if it is confirmed that there is, the trajectory is updated based on the first tracking trajectory that matches the current element; if it is confirmed that there is no, a corresponding first tracking trajectory is initialized in the first target element set based on the current element; and when the trajectory refresh operation ends, the current identifier is removed from the second identifier queue.

[0236] Specifically, initializing a corresponding first tracking trajectory based on the current element in the first target element set includes:

[0237] Initialize an empty first tracking trajectory in the first target element set as the currently added trajectory; assign a unique trajectory identifier to the currently added trajectory as the corresponding first trajectory identifier; and add a corresponding second trajectory point to the currently added trajectory by combining the first target identifier, second timestamp, second element identifier, second page number identifier, second paragraph identifier, first annotation box, first target type, and first multimodal feature of the current element as a set of corresponding second target identifier, third timestamp, third element identifier, third page number identifier, third paragraph identifier, second annotation box, second target type, and second multimodal feature.

[0238] Confirmation is performed on whether a first tracking trajectory matching the current element exists in the current first target element set, specifically including:

[0239] Step B1: Set the current element as the corresponding current observation point;

[0240] Step B2: Each first tracking trajectory in the first target element set is taken as the corresponding current tracking trajectory; the last second trajectory point of the current tracking trajectory is taken as the previous trajectory point; the area and aspect ratio of the corresponding bounding box are calculated based on the second height and second width of the previous trajectory point to obtain the corresponding previous area and aspect ratio; the area and aspect ratio of the corresponding bounding box are calculated based on the first height and first width of the current observation point to obtain the corresponding current area and aspect ratio; the shape similarity is calculated based on the previous area, previous aspect ratio, current area, and current aspect ratio; the feature similarity between the second multimodal features of the previous trajectory point and the first multimodal features of the current observation point is calculated based on the cosine vector similarity algorithm; and the shape similarity and feature similarity are weighted and summed to obtain the corresponding first similarity.

[0241] Here, the shape similarity and first similarity are calculated as follows:

[0242]

[0243] First similarity = ω3 × shape similarity + ω4 × feature similarity

[0244] Among them, ω1, ω2, ω3, and ω4 are four preset similarity weighting parameters;

[0245] Step B3: The first similarity exceeding the preset first similarity threshold is taken as the corresponding second similarity; and the number of second similarities is identified as zero; if yes, it is confirmed that there is no first tracking trajectory matching the current element in the current first target element set; if no, it is confirmed that there is a first tracking trajectory matching the current element in the current first target element set, and the first tracking trajectory corresponding to the largest second similarity is taken as the first tracking trajectory matching the current element.

[0246] Here, the first similarity threshold is a pre-set threshold parameter.

[0247] The trajectory is updated based on the first tracking trajectory matched by the current element, specifically including:

[0248] The first tracking trajectory matched by the current element is taken as the current trajectory; and the first target identifier, second timestamp, second element identifier, second page number identifier, second paragraph identifier, first annotation box, first target type, and first multimodal feature of the current element are combined into a corresponding second target identifier, third timestamp, third element identifier, third page number identifier, third paragraph identifier, second annotation box, second target type, and second multimodal feature to form a corresponding second trajectory point and add it to the current trajectory.

[0249] Step 5: After the annotation is completed, the first target element set is fused across pages according to the first trajectory set to obtain the corresponding parsed target set; and the annotation consistency is checked on the parsed target set.

[0250] Here, the parsing target set in this embodiment of the invention consists of one or more first parsing targets. Each first parsing target includes a parsing target type, a parsing target identifier group, a parsing target graph, and parsing target text. Specifically, the parsing target type is a tag type from a tag type set; the parsing target identifier group consists of one or two second target identifiers; the parsing target graph consists of one or two second block graphs corresponding to the parsing target identifier group; and the parsing target text consists of one or two second texts corresponding to the parsing target identifier group.

[0251] The specific steps for implementing step 5 include:

[0252] Step 51: Perform cross-page element fusion processing on the first target element set based on the first trajectory set to obtain the corresponding parsed target set;

[0253] Specifically, this includes: taking each first tracking trajectory in the first trajectory set as the current trajectory; performing a sequential traversal of all second trajectory points of the current trajectory; during this traversal, taking the currently traversed second trajectory point as the current trajectory point and the next second trajectory point as the next trajectory point; confirming whether there is a continuous cross-page and continuous content relationship between the current trajectory point and the next trajectory point; if a continuous cross-page and continuous content relationship is confirmed, adding a corresponding continuous trajectory point marker to the current trajectory point and the next trajectory point; at the end of this traversal, analyzing all continuous trajectory point markers of the current trajectory to obtain the corresponding parsing target subset; and merging all parsing target subsets corresponding to the first trajectory set to obtain the corresponding parsing target set; the parsing target subset consists of one or more first parsing targets;

[0254] This includes confirming whether there is a continuous cross-page relationship between the current trajectory point and the next trajectory point, specifically including:

[0255] Step C1: Calculate the difference between the third page number identifier of the next trajectory point and the current trajectory point to obtain the corresponding first page number difference; and identify whether the first page number difference is 1; if yes, proceed to step C2; if no, set the corresponding first confirmation status to no and proceed to step C6.

[0256] Here, a first page number difference of 0 indicates that the two do not cross pages, and a first page number difference > 1 indicates that although the two cross pages, the page numbers are not consecutive. In both cases, there is no consecutive page crossing relationship between the current trajectory point and the next trajectory point.

[0257] Step C2: Identify whether the two second target types of the current trajectory point and the next trajectory point match; if yes, proceed to step C3; if no, set the corresponding first confirmation status to no and proceed to step C6.

[0258] Here, if the two second target types of the current trajectory point and the next trajectory point are not the same, then there is no content continuity relationship between them;

[0259] Step C3: Calculate the similarity between the two second multimodal features of the current trajectory point and the next trajectory point based on the cosine vector similarity algorithm to the corresponding third similarity; and identify whether the third similarity exceeds the preset second similarity threshold; if yes, proceed to step C4; if no, set the corresponding first confirmation state to no and proceed to step C6.

[0260] Here, the second similarity threshold is a pre-set threshold parameter;

[0261] Here, if the third similarity does not exceed the second similarity threshold, it means that the content similarity between the current trajectory point and the next trajectory point is insufficient, and it is also considered that there is no content continuity between the two.

[0262] Step C4: Based on the second annotation box of the current trajectory point, confirm the image coordinates of the start and end positions of the element corresponding to the current trajectory point to obtain the corresponding first start coordinates and first end coordinates; and based on the second annotation box of the next trajectory point, confirm the image coordinates of the start and end positions of the element corresponding to the next trajectory point to obtain the corresponding second start coordinates and second end coordinates; calculate the horizontal spacing between the first and second start coordinates to obtain the corresponding first horizontal spacing; calculate the horizontal spacing between the first and second end coordinates to obtain the corresponding second horizontal spacing; calculate the linear distance between the first end coordinate and the preset single-page end position image coordinates to obtain the corresponding first coordinate spacing; and calculate the linear distance between the second start coordinate and the preset single-page start position image coordinates to obtain the corresponding second coordinate spacing.

[0263] As mentioned earlier, by analyzing the image coordinates of the first frame, we can obtain the coordinate range of a single page, and then obtain the corresponding image coordinates of the start and end positions of the single page.

[0264] Step C5: Take the second target type of the current trajectory point as the current type; and identify the current type; if the current type is a table or image, identify whether the first and second horizontal spacings are both less than the preset first spacing threshold. If yes, set the corresponding first confirmation state to yes; otherwise, set the corresponding first confirmation state to no. If the current type is not a table or image, identify whether the first and second coordinate spacings are both less than the preset second spacing threshold. If yes, set the corresponding first confirmation state to yes; otherwise, set the corresponding first confirmation state to no.

[0265] Here, the first spacing threshold and the second spacing threshold are two preset threshold parameters;

[0266] It should be noted that when the current type is a table or an image, if the two table / image elements corresponding to the current trajectory point and the next trajectory point are two sub-elements generated by the same table / image being truncated due to pagination, then the start and end positions of the annotation boxes of these two sub-tables / images should have the characteristic of left and right horizontal alignment. That is to say, the horizontal error of the start / end coordinates, i.e. the first and second horizontal spacing, should be within a very small error range (< the first spacing threshold). Therefore, in this embodiment of the invention, the current trajectory point and the next trajectory point are confirmed to be two consecutive trajectory points by judging whether the first and second horizontal spacings are both less than the first spacing threshold.

[0267] It should also be noted that when the current type is a text object (such as article title, chapter title, abstract, paragraph, mathematical symbol, physical symbol, mathematical formula / formula, chemical formula, technical terminology, technical parameter, legal terminology, legal clause, etc.), if the two text objects corresponding to the current trajectory point and the next trajectory point are two child elements generated by the same text object being truncated due to pagination, then the ending position of the previous text object should be aligned with the ending position of the previous page, the next text object should not have any text line indentation, and its starting position should be aligned with the starting position of the next page. In other words, the ending position of the previous text object... The end coordinates of the annotation box should be aligned with the end position image coordinates of the single page, and the start coordinates of the annotation box of the next text object should be aligned with the start position image coordinates of the single page. In other words, the error between the end position coordinates of the previous annotation box and the end position image coordinates of the single page (i.e., the first coordinate spacing) and the error between the start position coordinates of the next annotation box and the start position image coordinates of the single page (i.e., the second coordinate spacing) should be within a very small error range (< the second spacing threshold). Therefore, in this embodiment of the invention, the current trajectory point and the next trajectory point are confirmed to be two consecutive trajectory points by judging whether the first and second coordinate spacings are both less than the second spacing threshold.

[0268] Step C6: Identify the first confirmation state; if the first confirmation state is yes, then confirm that there is a continuous cross-page and continuous content relationship between the current trajectory point and the next trajectory point; if the first confirmation state is no, then confirm that there is no continuous cross-page and continuous content relationship between the current trajectory point and the next trajectory point.

[0269] The corresponding analytical target subset is obtained by analyzing all continuous trajectory point markers of the current trajectory, specifically including:

[0270] Step D1: Record the second trajectory point in the current trajectory that is not marked as a continuous trajectory point as an independent trajectory point; and take the second target type of each independent trajectory point as a corresponding parsing target type; and take the second target identifier of each independent trajectory point as a corresponding parsing target identifier group; take the second block diagram and second text of each independent trajectory point as a corresponding parsing target diagram and parsing target text; and take the parsing target type, parsing target identifier group, parsing target diagram and parsing target text corresponding to each independent trajectory point as a corresponding first parsing target.

[0271] Step D2: For each consecutive trajectory point in the current trajectory, mark the two corresponding second trajectory points to form a corresponding first trajectory point group; and use the second target type corresponding to each first trajectory point group as a corresponding parsing target type; and use the two corresponding second target identifiers of each first trajectory point group to form a corresponding parsing target identifier group; and use the two corresponding second block diagrams of each first trajectory point group to form a corresponding parsing target diagram; and use the two corresponding second texts of each first trajectory point group to form a corresponding parsing target text; and use a set of parsing target type, parsing target identifier group, parsing target diagram, and parsing target text corresponding to each first trajectory point group to form a corresponding first parsing target.

[0272] It should be noted that, in the embodiments of the present invention, when a corresponding analytical target image is composed of two second block images corresponding to each first trajectory point group: the two second block images are first aligned horizontally in the width direction, and then spliced ​​vertically in the height direction, and the resulting spliced ​​image is used as the corresponding analytical target image;

[0273] It should also be noted that, in this embodiment of the invention, when two second texts corresponding to each first trajectory point group are combined to form a corresponding parsing target text: firstly, the parsing target type corresponding to the current first trajectory point group is identified; if the current parsing target type is a type of text object (such as article title, chapter title, abstract, paragraph, mathematical symbol, physical symbol, mathematical formula / formula, chemical formula, technical terminology, technical parameter, legal terminology, legal clause, etc.), then the two second texts are concatenated in a sequential text concatenation manner, and the resulting concatenated text is used as the corresponding parsing target text; if the current parsing target type is a table, then the two second texts are actually two CSV text files, and in this case, the two second texts are concatenated in a sequential table concatenation manner, and the resulting new CSV text file is used as the corresponding parsing target text; if the current parsing target type is an image, then the two second texts are actually two first image description texts, and in this case, the two second texts are concatenated in a sequential text concatenation manner, and the resulting new image description text is used as the corresponding parsing target text;

[0274] Step D3, and the obtained first parsing targets form the corresponding parsing target subset;

[0275] Step 52, and perform a label consistency check on the parsed target set;

[0276] Specifically, this includes: step 521, taking each of the first parsing targets in the parsing target set as the current target; and identifying whether the fusion element group of the current target contains two second target identifiers; if so, calculating the corresponding third multimodal feature by averaging the two second multimodal features of the two second trajectory points corresponding to the current target; if not, taking the second multimodal feature of one second trajectory point corresponding to the current target as the corresponding third multimodal feature;

[0277] Step 522: Calculate the similarity between any two third multimodal features using the cosine similarity algorithm to obtain the corresponding fourth similarity; cluster all first parsing targets in the parsing target set based on the fourth similarity to obtain multiple corresponding first target clusters; and take the first target cluster with a total number of targets greater than 1 in each cluster as the corresponding current cluster, and identify whether the parsing target types of all first parsing targets in the current cluster are all the same. If not, form a corresponding secondary confirmation target set by combining all first parsing targets in the current cluster.

[0278] Here, each first target cluster in this embodiment of the invention consists of one or more first parsed targets; in each first target cluster, if the total number of targets in the cluster is greater than 1, then the fourth similarity between any two first parsed targets in the cluster exceeds the preset third similarity threshold; wherein, the third similarity threshold is a preset threshold parameter;

[0279] Step 523: When the total number of all secondary confirmed target sets is not zero, output all secondary confirmed target sets to the first annotator; prompt the first annotator to perform consistency checks and secondary type annotations on all parsed target types in each secondary confirmed target set; and make corresponding corrections to the parsed target sets based on the feedback results of the first annotator's consistency checks and secondary type annotations on each secondary confirmed target set.

[0280] Step 6: Feed back the parsed target set that has completed the consistency check to the first annotator.

[0281] Figure 4 This is a module structure diagram of a processing device for element annotation of PDF files provided in Embodiment 2 of the present invention. This device can be a terminal device or server implementing the aforementioned method embodiment, or it can be a device that enables the aforementioned terminal device or server to implement the aforementioned method embodiment. For example, the device can be a device or chip system of the aforementioned terminal device or server. Figure 4 As shown, the device includes: a labeling object receiving module 201, a labeling preprocessing module 202, a first labeling tracking module 203, a second labeling tracking module 204, a labeling postprocessing module 205, and a labeling feedback module 206.

[0282] The annotation object receiving module 201 is used to receive the first PDF file input by the first annotator; the first PDF file is the annotation object of the first annotator.

[0283] The annotation preprocessing module 202 performs image conversion and basic element parsing processing on the first PDF file based on a preset PDF element parsing tool to obtain the corresponding first image sequence and first parsing sequence.

[0284] The first annotation tracking module 203 is used to record the annotation behavior of the first annotator in combination with the first image sequence and the first parsing sequence when the first annotator selects and annotates target elements in the first PDF file according to the preset label type set, and refreshes the corresponding first annotation trajectory and first target element set based on the recorded information; and the preset behavior prediction model provides a candidate element set for the next annotation target after each single-step annotation behavior of the first annotator according to the first annotation trajectory; and improves the behavior prediction model based on the first annotator's selection feedback of candidate elements.

[0285] The second annotation and tracking module 204 is used to add corresponding multimodal element features to the target elements of the first target element set based on the preset multimodal feature recognition model during the annotation process; and to perform target matching and trajectory tracking processing on the target elements according to the first target element set and refresh the corresponding first trajectory set based on the tracking results.

[0286] The post-annotation processing module 205 is used to perform cross-page element fusion processing on the first target element set according to the first trajectory set after the annotation is completed to obtain the corresponding parsed target set; and to perform annotation consistency check on the parsed target set.

[0287] The annotation feedback module 206 is used to provide feedback on the parsed target set that has completed the consistency check to the first annotator.

[0288] The present invention provides a processing device for annotating elements in PDF files, which can execute the method steps in the above method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0289] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented by processing element calls to software, while others are implemented in hardware. For example, the labeling object receiving module can be a separate processing element, or it can be integrated into a chip in the above device. Alternatively, it can be stored as program code in the memory of the above device, and called and executed by a processing element of the device. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.

[0290] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a System-on-a-Chip (SOC).

[0291] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the foregoing method embodiments are generated. The computer described above can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The aforementioned computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the aforementioned computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, Bluetooth, microwave, etc.) means. The aforementioned computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).

[0292] Figure 5 This is a schematic diagram of an electronic device provided in Embodiment 3 of the present invention. This electronic device can be a terminal device or server implementing the methods of the aforementioned embodiments, or it can be a terminal device or server connected to the aforementioned terminal device or server implementing the methods of the aforementioned embodiments. Figure 5 As shown, the electronic device may include: a processor 301 (e.g., CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transmission and reception operations of the transceiver 303. The memory 302 may store various instructions for performing various processing functions and implementing the processing steps described in the foregoing embodiments. Preferably, the electronic device involved in the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize communication connections between components. The communication port 306 is used for communication between the electronic device and other peripherals.

[0293] exist Figure 5The system bus 305 mentioned can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, it is represented by only one thick line in the diagram, but this does not indicate that there is only one bus or one type of bus. The communication interface is used to enable communication between the database access device and other devices (e.g., clients, read-write libraries, and read-only libraries). Memory may include Random Access Memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.

[0294] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), graphics processing units (GPUs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0295] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to perform the methods and processes provided in the above embodiments.

[0296] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for annotating elements in PDF files. As described above, before initiating the annotation task, this invention pre-parses the PDF file using a PDF element parsing tool. Then, during the annotation process: 1) Based on real-time processing, the annotator's annotation trajectory (i.e., the first annotation trajectory) and the set of annotated elements (i.e., the first target element set) are updated in real-time according to the pre-parsed information; 2) Based on real-time processing, after each single-step annotation action, a behavior prediction model provides the annotator with a reference for the next candidate target (i.e., a candidate element set) based on the real-time annotation trajectory; 3) Based on asynchronous processing, the performance of the prediction model is continuously improved based on the annotator's candidate target feedback; 4) Based on an asynchronous processing mechanism, a multimodal feature recognition model is used to add multimodal (visual + text) element features to the annotated target elements; 5) Based on an asynchronous processing mechanism, associated target matching and associated target trajectory tracking are performed based on the geometric features of the annotation boxes (annotation box area, annotation box aspect ratio) and multimodal element features of the annotated elements, and the tracking trajectory set (i.e., the first trajectory set) is updated based on the processing results. After annotation, the cross-page elements in the annotated element set are first fused according to the tracking trajectory; then, the fused element set (parsing target set) undergoes an annotation consistency check, and during the check, the labels of similar elements are modified for consistency through human-computer collaboration; finally, the parsing target set that has completed the consistency check is fed back as the final annotation result. This embodiment of the invention improves annotation efficiency through a candidate element prediction mechanism, improves the recognition accuracy and fusion efficiency of cross-page elements through a related element tracking mechanism, and improves the annotation consistency of the entire text through a consistency check.

[0297] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0298] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for annotating elements in a PDF file, characterized in that, The method includes: Receive a first PDF file input by the first annotator; the first PDF file is the annotation object of the first annotator. Based on a preset PDF element parsing tool, the first PDF file is processed for image conversion and basic element parsing to obtain the corresponding first image sequence and first parsing sequence; When the first annotator selects and annotates target elements in the first PDF file according to a preset set of tag types, the annotation behavior of the first annotator is recorded by combining the first image sequence and the first parsing sequence, and the corresponding first annotation trajectory and first target element set are refreshed based on the recorded information; and a preset behavior prediction model provides a candidate element set for the next annotation target after each single-step annotation behavior of the first annotator according to the first annotation trajectory; and the behavior prediction model is improved and trained based on the first annotator's selection feedback of candidate elements. During the annotation process, corresponding multimodal element features are added to the target elements of the first target element set based on a preset multimodal feature recognition model; and target matching and trajectory tracking are performed on the target elements according to the first target element set, and the corresponding first trajectory set is refreshed based on the tracking results. After annotation is completed, the first target element set is subjected to cross-page element fusion processing based on the first trajectory set to obtain the corresponding parsed target set; and the annotation consistency check is performed on the parsed target set. The parsed target set that has completed the consistency check is fed back to the first annotator.

2. The method for annotating elements in a PDF file according to claim 1, characterized in that, The first PDF file includes multiple first file pages; The first image sequence is formed by sequentially sorting multiple first frame images; each first frame image corresponds one-to-one with the first file page. The first parsing sequence is formed by sequentially sorting multiple first frame element sets; each first frame element set corresponds one-to-one with the first file page. The first frame element set includes multiple first basic elements; the first basic elements include a first element identifier, a first page number identifier, a first paragraph identifier, a first start position, a first end position, a first element type, and a first element attribute set; the first element type includes article title, chapter title, abstract, paragraph, table, and image; The tag type set includes a basic type subset and a domain type subset; both the basic type subset and the domain type subset consist of one or more tag types; the tag types of the basic type subset include article title, chapter title, abstract, paragraph, table, and image; the tag types of the domain type subset correspond to the file type of the first PDF file; if the file type of the first PDF file is an academic paper, then the tag types of the domain type subset include mathematical symbols, physical symbols, mathematical formulas, and chemical formulas; if the file type of the first PDF file is a technical report, then the tag types of the domain type subset include formulas, technical terms, and technical parameters; if the file type of the first PDF file is a legal document, then the tag types of the domain type subset include legal terms and legal clauses. The first annotation trajectory is formed by sequentially sorting one or more first trajectory points; the trajectory point attributes of each first trajectory point include a first timestamp, a first mouse sliding trajectory, a first mouse click position, a first annotation box drag direction, a first annotation box size, a first annotation box center point, and a first annotation type; The first target element set consists of one or more first target elements; each first target element corresponds to a selected target element, a first trajectory point, and a parent basic element; the parent basic element is a first basic element; the element content of each first target element is part or all of the element content of the parent basic element; the first target element includes a first target identifier, a second timestamp, a second element identifier, a second page number identifier, a second paragraph identifier, a first annotation box, a first target type, and a first multimodal feature; The first annotation box includes a first starting point, a first center point, a first height, a first width, a first frame, and first text; The candidate element set consists of at most M first candidate elements; the number M is a preset positive integer; the first candidate element includes the candidate box starting point, candidate box height, candidate box width, candidate box annotation type, and candidate box confidence. The first trajectory set consists of one or more first tracking trajectories; each first tracking trajectory corresponds to a first trajectory identifier; each first tracking trajectory is formed by sequentially sorting one or more second trajectory points; the trajectory point attributes of each second trajectory point include a second target identifier, a third timestamp, a third element identifier, a third page number identifier, a third paragraph identifier, a second annotation box, a second target type, and a second multimodal feature; the second annotation box includes a second starting point, a second center point, a second height, a second width, a second bounding box, and second text; The behavior prediction model is used to predict candidate elements based on the historical behavior sequence X input to the model and output the corresponding set of candidate elements; wherein, the historical behavior sequence X consists of N behavior feature vectors x i The data is sorted chronologically, with 1 ≤ index i ≤ N, and the number N is a preset positive integer; each behavioral feature vector x i Corresponding to one of the first trajectory points; each of the behavioral feature vectors x i The vector feature data consists of the first timestamp of the corresponding first trajectory point, the first mouse sliding trajectory, the first mouse click position, the drag direction of the first annotation box, the size of the first annotation box, and the first annotation type; The multimodal feature recognition model is used to extract visual basic features, visual layout features, image features, and text features based on the element annotation box information input to the model. It then fuses the three types of visual features using an attention-weighted fusion mechanism and performs feature dimensionality reduction on the concatenated features of the visual fusion features and text features to obtain the corresponding multimodal element features. The element annotation box information corresponds to one of the first target elements in the first target element set. The element annotation box information includes the parent page image, annotation box center point, annotation box height, annotation box width, annotation box image, and annotation box text. The parsing target set consists of one or more first parsing targets; the first parsing target includes a parsing target type, a parsing target identifier group, a parsing target graph, and a parsing target text; the parsing target type is one of the tag types in the tag type set; the parsing target identifier group consists of one or two second target identifiers; the parsing target graph consists of one or two second block graphs corresponding to the parsing target identifier group; the parsing target text consists of one or two second texts corresponding to the parsing target identifier group.

3. The method for annotating elements in a PDF file according to claim 2, characterized in that, The behavior prediction model is composed of an embedding coding layer, a behavior feature encoder, a candidate box prediction module, and a candidate box filtering module connected sequentially. The embedding coding layer is used to perform actions based on the behavior feature vector x. i The embedding encoding rules corresponding to each vector feature data are used for each behavioral feature vector x. i Each vector feature data is embedded and encoded to obtain a corresponding feature encoding vector; and each behavioral feature vector x is then used to generate a feature encoding vector. i All the aforementioned feature encoding vectors are sequentially concatenated into a corresponding embedding encoding vector; and a corresponding embedding encoding sequence E, composed of N such embedding encoding vectors, is sent to the behavior feature encoder; The behavior feature encoder is implemented based on an LSTM model or a Transformer architecture encoder model; the behavior feature encoder is used to perform feature encoding processing on the embedded encoding sequence E to obtain the corresponding feature vector H and send it to the candidate box prediction module; The candidate box prediction module consists of M ’ It consists of M parallel first prediction units; the number is M ’ M is a preset positive integer. ’ >M; The candidate box prediction module is used to use each of the first prediction units to predict a single candidate box based on the feature vector H to obtain a corresponding predicted candidate box y. j 1 ≤ index j ≤ M ’ ; and from the obtained M ’ The predicted candidate boxes y j The predicted candidate box y is sequentially sorted to form a corresponding candidate box sequence Y and sent to the candidate box filtering module; each predicted candidate box y j Each includes a corresponding set of candidate box starting point, candidate box height, candidate box width, candidate box annotation type, and candidate box confidence level; The prediction method for the j-th first prediction unit is: The four weight vector parameters are the j-th first prediction unit. These are the four offset vector parameters corresponding to the j-th first prediction unit; The candidate box position prediction vector is the candidate box position prediction vector corresponding to the j-th first prediction unit. The vector length is 3, and it consists of the starting point of the candidate box corresponding to the j-th first prediction unit, the height of the candidate box, and the width of the candidate box; The candidate box type prediction vector corresponding to the j-th first prediction unit is composed of multiple type prediction probabilities; the total number of the type prediction probabilities is consistent with the total number of label types in the label type set; the candidate box type prediction vector The label type corresponding to the highest predicted probability of the type is the candidate box label type corresponding to the j-th first prediction unit; The confidence level of the candidate box corresponding to the j-th first prediction unit; The candidate box prediction module is used to select up to M corresponding first candidate elements from the candidate box sequence Y according to the non-maximum suppression mechanism to form the corresponding candidate element set and output it.

4. The method for annotating elements in a PDF file according to claim 2, characterized in that, The step of recording the annotation behavior of the first annotator by combining the first image sequence and the first parsing sequence, and refreshing the corresponding first annotation trajectory and first target element set based on the recorded information, specifically includes: When the first annotator selects a target element, the currently selected target element is taken as the corresponding currently selected target element; the current time is taken as the corresponding first timestamp and second timestamp; the global page number of the file page where the currently selected target element is located is taken as the current page number identifier; the first file page corresponding to the current page number identifier is taken as the current file page; and the first frame image and the first frame element set corresponding to the current file page in the first image sequence and the first parsing sequence are taken as the corresponding current frame image and current frame element set. It also confirms whether the currently selected target element is the first selected target element of the first PDF file; if it is confirmed, the first annotation trajectory and the first target element set are initialized to empty. The system confirms whether the currently selected target element is the first selected target element on the current file page. If yes, the system tracks and records the mouse movement of the first annotator on the current file page before the first annotator sets a label box for the currently selected target element by long-pressing the mouse. If no, the system tracks and records the mouse movement of the first annotator on the current file page from the previous target element with a completed label box to the currently selected target element before the first annotator sets a label box for the currently selected target element by long-pressing the mouse. During this tracking and recording process, a corresponding mouse trajectory point is periodically formed by the current page number identifier and the image coordinates of the mouse position on the corresponding current frame image at a preset sampling frequency. The tracking and recording process stops when the first annotator starts setting a label box by long-pressing the mouse. At the end of this tracking and recording process, all the mouse trajectory points obtained in this tracking and recording process are sampled at equal intervals according to the preset total number of mouse trajectory points, and the corresponding first mouse movement trajectory is formed by all the trajectory points obtained in this sampling process. The image coordinates of the start and end positions of the first annotator's operation when setting the annotation box for the currently selected target element by long-pressing the mouse are recorded on the current frame image as the corresponding current start coordinates and current end coordinates; the current start coordinates are used as the corresponding first mouse click position; the horizontal and vertical displacement vectors from the current start coordinates to the current end coordinates form the corresponding first annotation box drag direction; the height and width of the annotation box of the currently selected target element in the current frame image are identified, and the corresponding first annotation box size is set based on the identification results; the image coordinates of the center point of the current annotation box in the current frame image are identified based on the current first mouse click position and the first annotation box size, and the corresponding center point of the first annotation box is set based on the identification results; and the label type set by the first annotator for the current annotation box is used as the corresponding first annotation type. The first trajectory point is composed of the first timestamp corresponding to the currently selected target element, the first mouse sliding trajectory, the first mouse click position, the first annotation box drag direction, the first annotation box size, the first annotation box center point, and the first annotation type, and is added to the first annotation trajectory; and the first trajectory point added this time is used as the corresponding current trajectory point; And assign a unique target identifier to the currently selected target element as the corresponding first target identifier; Based on the first start position and the first end position of each of the first basic elements in the current frame element set, a corresponding rectangle is drawn on the current frame image as the corresponding basic element box; and the first basic element that intersects with the annotation box of the currently selected target element is taken as the corresponding parent basic element; and the first element identifier, the first page number identifier, and the first paragraph identifier of the current parent basic element are taken as the corresponding second element identifier, the second page number identifier, and the second paragraph identifier. The first mouse click position, the center point of the first annotation box, and the first annotation type of the current trajectory point are taken as the corresponding first starting point, the first center point, and the first target type; the height and width of the first annotation box of the current trajectory point are taken as the corresponding first height and the first width; the annotation box area sub-image circled by the annotation box of the currently selected target element on the current frame image is taken as the corresponding first frame image; the sub-text information in the text attribute of the current parent basic element that is within the annotation box of the currently selected target element is taken as the corresponding first text; and the first annotation box is composed of the first starting point, the first center point, the first height, the first width, the first frame image, and the first text corresponding to the currently selected target element. An empty multimodal feature is initialized for the currently selected target element as the corresponding first multimodal feature; and a corresponding first target element is added to the first target element set, consisting of the first target identifier, the second timestamp, the second element identifier, the second page number identifier, the second paragraph identifier, the first annotation box, the first target type, and the first multimodal feature corresponding to the currently selected target element.

5. The method for annotating elements in a PDF file according to claim 2, characterized in that, The provision of a candidate element set by a preset behavior prediction model for the next annotation target after each single-step annotation action of the first annotator based on the first annotation trajectory specifically includes: When the first annotation trajectory completes a trajectory point, the first file page that the first annotator is currently annotating is taken as the current file page; and the first trajectory points that match the single page number identifier of the first mouse sliding trajectory in the first annotation trajectory with the current file page are extracted to form the corresponding first sub-trajectory; The total number of trajectory points in the first sub-trajectory is taken as the current total number; the previous total number is identified; if the previous total number is greater than or equal to N, the sub-trajectory composed of the N nearest first trajectory points in the first sub-trajectory is taken as the corresponding second sub-trajectory; if the previous total number is less than N but greater than 0, the total number of trajectory points in the first sub-trajectory is increased to N by adding one or more preset zero-value trajectory points to the head of the first sub-trajectory, and the first sub-trajectory with completed trajectory point addition is taken as the corresponding second sub-trajectory; if the previous total number is 0, the corresponding second sub-trajectory is set to empty; the data format of the zero-value trajectory points is consistent with that of the first trajectory points, but all trajectory point attributes inside are 0; When the second sub-trajectory is not empty, a corresponding behavior feature vector x is formed by the first timestamp of each first trajectory point of the second sub-trajectory, the first mouse sliding trajectory, the first mouse click position, the first annotation box drag direction, the first annotation box size, and the first annotation type. i ; and from the N behavioral feature vectors x obtained this time i A corresponding historical behavior sequence X is formed; and the current historical behavior sequence X is input into the behavior prediction model for processing to obtain the corresponding candidate element set; When the candidate element set is not empty, the candidate box area corresponding to each of the first candidate elements in the candidate element set is highlighted on the current annotation interface of the first annotator; and the candidate box annotation type and the candidate box confidence of the current first candidate element are displayed synchronously in each displayed candidate box area.

6. The method for annotating elements in a PDF file according to claim 3, characterized in that, The step of improving the behavior prediction model based on the first annotator's selection feedback on candidate elements specifically includes: When the candidate element set is first displayed, the consecutive failure counter is initialized to 0; and the boosted training dataset is initialized to empty. Each time the candidate element set is displayed, it is determined whether the first annotator has selected a first candidate element from the current candidate element set as the corresponding current target element; if yes, the consecutive failure counter is cleared to zero; if no, the consecutive failure counter is incremented by 1. And each time the first annotator fails to successfully select the current target element from the candidate element set displayed at that time, the target element selected by the first annotator is taken as the corresponding current manually selected element; and when the first target element corresponding to the current manually selected element is successfully added to the first target element set, the first starting point, first height, first width, and first target type corresponding to the current manually selected element are taken as a set of corresponding candidate box starting point, candidate box height, candidate box width, and candidate box label type; and the confidence level of the candidate box corresponding to the current manually selected element is set to 1; and a corresponding first label candidate box is formed by the candidate box starting point, candidate box height, candidate box width, candidate box label type, and candidate box confidence level corresponding to the current manually selected element; and M ’ Each identical first label candidate box forms a corresponding first label vector; and the historical behavior sequence X corresponding to the candidate element set displayed at the current time is used as a corresponding first training behavior sequence; and the first training behavior sequence corresponding to the currently selected element and the first label vector form a corresponding first data record which is added to the improvement training dataset; Each time a record is added to the training dataset, the system checks whether the consecutive failure counter exceeds a preset failure threshold. If it does, the consecutive failure counter is reset to zero, and the first number of most recently added first data records in the training dataset are extracted to form the current dataset, where the first number is a preset positive integer. The system then copies the model structure and parameters of the embedding encoding layer, the behavior feature encoder, and the candidate box prediction module of the current behavior prediction model to obtain the corresponding first embedding layer, first encoder, and first prediction module. These copied first embedding layer, first encoder, and first prediction module are then sequentially connected to form the current training framework. The current training framework is then fine-tuned based on the current dataset. During this fine-tuning training, only the model parameters of the first prediction module are fine-tuned. At the end of this fine-tuning training, the model parameters of the candidate box prediction module of the current behavior prediction model are reset based on the current model parameters of the first prediction module.

7. The method for annotating elements in a PDF file according to claim 2, characterized in that, The multimodal feature recognition model includes a visual basic feature encoder, a visual layout feature encoder, an image feature encoder, a text feature encoder, a first feature mapping network, a second feature mapping network, a third feature mapping network, a fourth feature mapping network, an attention weighted fusion layer, a visual text feature splicing layer, and a feature dimensionality reduction layer. The inputs of the visual basic feature encoder, the visual layout feature encoder, the image feature encoder, and the text feature encoder are all connected to the model input. The outputs of the visual basic feature encoder, the visual layout feature encoder, the image feature encoder, and the text feature encoder are respectively connected to the inputs of the corresponding first, second, third, and fourth feature mapping networks. The outputs of the first, second, and third feature mapping networks are respectively connected to the first, second, and third inputs of the attention-weighted fusion layer. The outputs of the attention-weighted fusion layer and the fourth feature mapping network are respectively connected to the first and second inputs of the visual-text feature splicing layer. The output of the visual-text feature splicing layer is connected to the input of the feature dimensionality reduction layer. The output of the feature dimensionality reduction layer is connected to the model output. The visual fundamental feature encoder is used to extract the corresponding parent page image and the annotation box image from the element annotation box information; and to identify the color histogram of the annotation box image; and to identify the texture features of the annotation box image; and to identify the edge features of the annotation box image on the parent page image; and to send the corresponding visual fundamental feature tensor composed of the obtained color histogram, texture features and edge features to the first feature mapping network; The visual layout feature encoder is used to extract the center point, height and width of the label box of the element label box information to form a corresponding visual layout feature vector and send it to the second feature mapping network; The image feature encoder is implemented based on a CNN network model, a residual neural network model, or a Transformer framework encoder model; the image feature encoder is used to perform high-dimensional feature encoding processing on the image of the labeled bounding box information of the element to obtain the corresponding high-dimensional image feature tensor and send it to the third feature mapping network; The text feature encoder is implemented based on the BERT series model or the encoder model of the Transformer framework; the text feature encoder is used to perform high-dimensional feature encoding on the text of the labeled box information of the element to obtain the corresponding high-dimensional text feature tensor and send it to the fourth feature mapping network; The first feature mapping network is implemented based on a fully connected network, which is composed of one or more fully connected layers connected sequentially; the first feature mapping network is used to map a preset first feature vector space. As the corresponding target vector space, the visual basic feature tensor is mapped to the feature vector of the target vector space to obtain the corresponding first feature vector, which is then sent to the attention weighted fusion layer; d1 is the feature dimension of the first feature vector space, and the feature dimension d1 is a preset positive integer; The second feature mapping network is implemented based on a fully connected network, which is composed of one or more fully connected layers connected sequentially; the second feature mapping network is used to map the first feature vector space. As the corresponding target vector space, the visual layout feature vector is mapped to the feature vector of the target vector space to obtain the corresponding second feature vector, which is then sent to the attention weighted fusion layer; the vector shapes of the first and second feature vectors are consistent. The third feature mapping network is implemented based on a fully connected network, which is composed of one or more fully connected layers connected sequentially. The third feature mapping network is used to map the first feature vector space... As the corresponding target vector space, the image high-dimensional feature tensor is mapped to the feature vector of the target vector space to obtain the corresponding third feature vector, which is then sent to the attention weighted fusion layer; the vector shapes of the first and third feature vectors are consistent. The fourth feature mapping network is implemented based on a fully connected network, which is composed of one or more fully connected layers connected sequentially. The fourth feature mapping network is used to map the first feature vector space. As the corresponding target vector space, the text high-dimensional feature tensor is mapped to the feature vector of the target vector space to obtain the corresponding fourth feature vector; and the fourth feature vector is sent to the visual text feature splicing layer as the corresponding annotation box text feature vector; the vector shapes of the first and fourth feature vectors are consistent; The attention-weighted fusion layer consists of a first feature tensor composed of the received first, second, and third feature vectors; and performs corresponding query, key, and value vector transformations on the first feature tensor based on preset query, key, and value weight vectors to obtain corresponding query vector Q, key vector K, and value vector V; then performs attention operations based on the query vector Q, the key vector K, and the value vector V, and sends the operation result as the corresponding bounding box visual feature vector to the visual text feature splicing layer; the shape of the bounding box visual feature vector and the bounding box text feature vector are consistent; The visual text feature concatenation layer is used to concatenate the visual feature vector of the annotation box and the text feature vector of the annotation box in a vector concatenation manner to obtain a corresponding concatenated feature vector, which is then sent to the feature dimensionality reduction layer. The feature reduction layer is implemented based on an MLP model; the feature reduction layer is used to reduce the preset second feature vector space. As the corresponding target vector space, the concatenated feature vector is mapped to the feature vector of the target vector space, and the resulting mapped feature vector is used as the corresponding multimodal element feature and output. d2 is the feature dimension of the second feature vector space. The feature dimension d2 is a preset positive integer, and d2 < d1.

8. The method for annotating elements in a PDF file according to claim 2, characterized in that, The preset multimodal feature recognition model adds corresponding multimodal element features to the target elements of the first target element set, specifically including: When the first annotator begins to select target elements in the first PDF file, an empty queue is initialized as the corresponding first identifier queue; and each time a first target element is added to the first target element set, the first target identifier of the first target element added at that time is added to the first identifier queue. When the total number of identifiers in the first identifier queue is not zero, the earliest added first target identifier is taken as the current identifier; the first target element in the first target element set whose first target identifier matches the current identifier is taken as the current element; the first center point, first height, first width, first frame image, and first text of the current element are taken as a set of corresponding annotation box center point, annotation box height, annotation box width, annotation box image, and annotation box text; the first frame image corresponding to the second page number identifier of the current element is taken as the corresponding parent page image; and the parent page image, annotation box center point, annotation box height, annotation box width, annotation box image, and annotation box text corresponding to the current element form a corresponding element annotation box information; the current element annotation box information is input into the multimodal feature recognition model for processing to obtain the corresponding multimodal element features; and the first multimodal features of the current element are set based on the current multimodal element features; and when this setting ends, the current identifier is removed from the first identifier queue.

9. The method for annotating elements in a PDF file according to claim 2, characterized in that, The step of performing target matching and trajectory tracking processing on target elements according to the first target element set and refreshing the corresponding first trajectory set based on the tracking results specifically includes: When the first annotator begins to select target elements in the first PDF file, an empty queue is initialized as the corresponding second identifier queue, and the first trajectory set is initialized to empty; and each time a multimodal element feature setting is completed for the first target element set, the first target identifier of the first target element corresponding to the first multimodal feature set in that current setting is added to the second identifier queue. When the total number of identifiers in the second identifier queue is not zero, the earliest added first target identifier is taken as the current identifier; the total number of trajectories in the current first trajectory set is counted to obtain the corresponding current trajectory total number; the first target element in the first target element set whose first target identifier matches the current identifier is taken as the current element; and it is determined whether the current trajectory total number is zero; if yes, a corresponding first tracking trajectory is initialized in the first target element set based on the current element; if no, it is confirmed whether there is a first tracking trajectory in the current first target element set that matches the current element; if it is confirmed that there is, the trajectory is updated based on the first tracking trajectory that matches the current element; if it is confirmed that there is no, a corresponding first tracking trajectory is initialized in the first target element set based on the current element; and when the trajectory refresh operation ends, the current identifier is removed from the second identifier queue.

10. The method for annotating elements in a PDF file according to claim 9, characterized in that, The step of confirming whether there exists a first tracking trajectory matching the current element in the current first target element set specifically includes: Step 101: Take the current element as the corresponding current observation point; Step 102: Each of the first tracking trajectories in the first target element set is taken as the corresponding current tracking trajectory; the last second trajectory point of the current tracking trajectory is taken as the previous trajectory point; the area and aspect ratio of the corresponding bounding box are calculated based on the second height and second width of the previous trajectory point to obtain the corresponding previous area and aspect ratio; the area and aspect ratio of the corresponding bounding box are calculated based on the first height and first width of the current observation point to obtain the corresponding current area and aspect ratio; the shape similarity is calculated based on the previous area, the previous aspect ratio, the current area, and the current aspect ratio; the feature similarity between the second multimodal feature of the previous trajectory point and the first multimodal feature of the current observation point is calculated based on the cosine vector similarity algorithm; and the weighted sum of the shape similarity and the feature similarity is obtained to obtain the corresponding first similarity. The shape similarity and the first similarity are calculated as follows: First similarity = ω3 × shape similarity + ω4 × feature similarity ω1, ω2, ω3, and ω4 are four preset similarity weighting parameters; Step 103: Take the first similarity that exceeds the preset first similarity threshold as the corresponding second similarity; and identify whether the number of second similarities is zero; if yes, confirm that there is no first tracking trajectory in the current first target element set that matches the current element; if no, confirm that there is a first tracking trajectory in the current first target element set that matches the current element, and take the first tracking trajectory corresponding to the largest second similarity as the first tracking trajectory that matches the current element.

11. The method for annotating elements in a PDF file according to claim 2, characterized in that, The step of performing cross-page element fusion processing on the first target element set based on the first trajectory set to obtain the corresponding parsed target set specifically includes: Each of the first tracking trajectories in the first trajectory set is taken as the current trajectory; and all the second trajectory points of the current trajectory are sequentially traversed once; during this traversal, the currently traversed second trajectory point is taken as the current trajectory point, and the next second trajectory point of the current trajectory point is taken as the next trajectory point; and it is confirmed whether there is a continuous cross-page and continuous content association between the current trajectory point and the next trajectory point; if a continuous cross-page and continuous content association is confirmed, a corresponding continuous trajectory point mark is added to the current trajectory point and the next trajectory point; and at the end of this traversal, the corresponding parsing target subset is obtained by analyzing all the continuous trajectory point marks of the current trajectory; and all the parsing target subsets corresponding to the first trajectory set are merged to obtain the corresponding parsing target set; the parsing target subset consists of one or more first parsing targets.

12. The method for annotating elements in a PDF file according to claim 11, characterized in that, The confirmation of whether there is a continuous cross-page and continuous content relationship between the current trajectory point and the next trajectory point specifically includes: Step 121: Calculate the difference between the third page number identifier of the next trajectory point and the current trajectory point to obtain the corresponding first page number difference; and identify whether the first page number difference is 1; if yes, proceed to step 122; if no, set the corresponding first confirmation status to no and proceed to step 126. Step 122: Identify whether the two second target types of the current trajectory point and the next trajectory point match; if yes, proceed to step 123; if no, set the corresponding first confirmation status to no and proceed to step 126. Step 123: Calculate the similarity between the two second multimodal features of the current trajectory point and the next trajectory point to the corresponding third similarity based on the cosine vector similarity algorithm; and identify whether the third similarity exceeds the preset second similarity threshold; if yes, proceed to step 124; if no, set the corresponding first confirmation state to no and proceed to step 126. Step 124: Based on the second annotation box of the current trajectory point, confirm the image coordinates of the start and end positions of the element corresponding to the current trajectory point to obtain the corresponding first start coordinates and first end coordinates; and based on the second annotation box of the next trajectory point, confirm the image coordinates of the start and end positions of the element corresponding to the next trajectory point to obtain the corresponding second start coordinates and second end coordinates; calculate the horizontal spacing between the first and second start coordinates to obtain the corresponding first horizontal spacing; calculate the horizontal spacing between the first and second end coordinates to obtain the corresponding second horizontal spacing; calculate the linear distance between the first end coordinate and the preset single-page end position image coordinates to obtain the corresponding first coordinate spacing; and calculate the linear distance between the second start coordinate and the preset single-page start position image coordinates to obtain the corresponding second coordinate spacing. Step 125: Take the second target type of the current trajectory point as the current type; and identify the current type; if the current type is a table or an image, identify whether the first and second horizontal spacings are both less than a preset first spacing threshold; if yes, set the corresponding first confirmation state to yes; otherwise, set the corresponding first confirmation state to no; if the current type is not a table or an image, identify whether the first and second coordinate spacings are both less than a preset second spacing threshold; if yes, set the corresponding first confirmation state to yes; otherwise, set the corresponding first confirmation state to no. Step 126: Identify the first confirmation state; if the first confirmation state is yes, then confirm that there is a continuous cross-page and continuous content association between the current trajectory point and the next trajectory point; if the first confirmation state is no, then confirm that there is no continuous cross-page and continuous content association between the current trajectory point and the next trajectory point.

13. The method for annotating elements in a PDF file according to claim 2, characterized in that, The step of performing a label consistency check on the parsed target set specifically includes: Each of the first parsing targets in the parsing target set is taken as the current target; and it is identified whether the fusion element group of the current target contains two second target identifiers; if so, the average feature of the two second multimodal features of the two second trajectory points corresponding to the current target is calculated to obtain the corresponding third multimodal feature; if not, the second multimodal feature of one second trajectory point corresponding to the current target is taken as the corresponding third multimodal feature. The similarity between any two third multimodal features is calculated using a cosine similarity algorithm to obtain a corresponding fourth similarity. Based on the fourth similarity, all first parsed targets in the parsed target set are clustered to obtain multiple corresponding first target clusters. The first target clusters with a total number of targets greater than 1 are taken as the corresponding current clusters. The parsed target types of all first parsed targets within the current cluster are identified; otherwise, all first parsed targets in the current cluster form a corresponding secondary confirmation target set. Each first target cluster consists of one or more first parsed targets. In each first target cluster, if the total number of targets within the cluster is greater than 1, the fourth similarity between any two first parsed targets within the cluster exceeds a preset third similarity threshold. When the total number of all the obtained secondary confirmation target sets is not zero, all the secondary confirmation target sets are output to the first annotator; and the first annotator is prompted to perform consistency checks and secondary type annotations on all the parsed target types in each of the secondary confirmation target sets; and the parsed target sets are corrected accordingly based on the feedback results of the first annotator on the consistency checks and secondary type annotations of each of the secondary confirmation target sets.

14. An apparatus for performing the processing method for element annotation of a PDF file according to any one of claims 1-13, characterized in that, The device includes: a labeled object receiving module, a labeled preprocessing module, a first labeled tracking module, a second labeled tracking module, a labeled postprocessing module, and a labeled feedback module; The annotation object receiving module is used to receive a first PDF file input by a first annotator; the first PDF file is the annotation object of the first annotator. The annotation preprocessing module performs image conversion and basic element parsing processing on the first PDF file based on a preset PDF element parsing tool to obtain the corresponding first image sequence and first parsing sequence; The first annotation tracking module is used to record the annotation behavior of the first annotator when the first annotator selects and annotates target elements in the first PDF file according to a preset tag type set, combining the first image sequence and the first parsing sequence, and to refresh the corresponding first annotation trajectory and first target element set based on the recorded information; and a preset behavior prediction model provides a candidate element set for the next annotation target after each single-step annotation behavior of the first annotator according to the first annotation trajectory; and to improve and train the behavior prediction model based on the first annotator's selection feedback of candidate elements. The second annotation and tracking module is used to add corresponding multimodal element features to the target elements of the first target element set based on a preset multimodal feature recognition model during the annotation process; and to perform target matching and trajectory tracking processing on the target elements according to the first target element set and refresh the corresponding first trajectory set based on the tracking results; The post-annotation processing module is used to perform cross-page element fusion processing on the first target element set according to the first trajectory set after the annotation is completed to obtain the corresponding parsed target set; and to perform annotation consistency check on the parsed target set. The annotation feedback module is used to provide feedback on the parsed target set that has completed the consistency check to the first annotator.

15. An electronic device, characterized in that, include: Memory, processor, and transceiver; The processor is configured to be coupled to the memory, read and execute instructions in the memory to implement the method according to any one of claims 1-13; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1-13.