A literature data security labeling method and system based on confidential computing

CN122778445APending Publication Date: 2026-09-18NANHU LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611257346.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-19
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

然而,这种方式仍保留了文档的整体版式、段落结构以及前后文的语义关联,标注人员依然有可能通过上下文推断出被遮盖的内容,或者对文档类型、业务性质等元信息进行推测,从而造成间接泄密

Benefits of technology

[0054] This solution establishes a coordinate system and a restoration key based on image and text coordinate encapsulation. Combined with single-word segmentation to generate single-word task packages without context, it completely separates the anonymization and annotation process from the typesetting information. After proofreading, the key is used to restore the document, thereby achieving a complete manually proofread result while protecting privacy, thus balancing security and restoration accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122778445A_ABST
    Figure CN122778445A_ABST
Patent Text Reader

Abstract

This invention discloses a secure annotation method and system for document data based on confidential computing. It separates image and text regions by establishing a planar coordinate system and analyzing page layout. The text region is then identified and segmented using OCR to create individual character images. A coordinate mapping table is established, containing text logical coordinates, physical pixel coordinates, unique identifiers, and identified characters. These individual character images are randomly assigned to multiple de-identification task packages. Based on the coordinate mapping table, a local restoration key is generated for each task package to restore the context information of the target character object. In a trusted execution environment, the local restoration key is used to restore the context and generate annotation suggestions, which are then provided to the annotation end. After receiving the annotation results, the management end reconstructs the complete text consistent with the original document layout based on the coordinate mapping table and merges it with the image region to generate the final document. This invention achieves privacy protection by keeping document data invisible during OCR error correction, balancing data security and restoration accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of document data processing and information security technology, and in particular to a document data security annotation method and system based on confidential computing. Background Technology

[0002] Optical Character Recognition (OCR) technology has been widely used in the digitization of various documents, such as scanned contracts, invoices, and historical document research. However, errors often occur in OCR results due to factors such as scan quality, font distortion, and interference from similar-looking characters. Therefore, manual proofreading and correction are necessary. Traditional proofreading methods typically involve having a third-party annotator correct all or part of the document. This annotator has access to all or part of the original text information, posing a risk of leakage of sensitive information within the document.

[0003] To mitigate privacy risks, existing technologies have proposed several anonymization and proofreading schemes. For example, when displaying documents, key fields (such as names and amounts) are masked or blurred, allowing annotators to observe only non-sensitive areas. However, this method still preserves the overall document layout, paragraph structure, and semantic connections between contexts. Annotators may still infer the masked content from the context or speculate on metadata such as document type and business nature, leading to indirect leaks. Another approach is to send the complete document to a trusted, closed environment (such as an internal corporate server) for processing by designated personnel. However, this requires annotators with extremely high security qualifications and is difficult to implement on a large scale, in a distributed, and low-cost crowdsourced annotation basis. Summary of the Invention

[0004] The purpose of this invention is to propose a secure annotation method and system for document data based on confidential computing to address the problems existing in the prior art. This method and system can achieve privacy protection of document data by making it invisible during the OCR error correction process, enabling annotators to complete character-level error correction efficiently and accurately without access to any complete document or character context, and to subsequently restore the correct text containing the original typesetting format with high precision.

[0005] To achieve the above objectives, the present invention adopts the following technical solutions:

[0006] A method for secure annotation of documentary data based on confidential computing includes the following steps:

[0007] Establish a planar coordinate system for the original document image, separate the image area from the text area through layout analysis, record the boundary coordinates of each area, and save the image area;

[0008] The text region is subjected to OCR recognition to obtain the recognized characters, physical pixel coordinates, and text logical coordinates allocated by row and column for each character object;

[0009] Cut out individual character images from each character object and assign them unique identifiers;

[0010] Establish a mapping table containing the text logical coordinates, physical pixel coordinates, unique identifiers, and coordinates of the identified characters for each character object;

[0011] Each individual character image, along with its unique identifier and recognition character, is randomly assigned to multiple de-identification task packages. Based on the coordinate mapping table, a local restoration key is generated for each de-identification task package to restore the context information of several target character objects.

[0012] The de-identification task package and its local restoration key are distributed to a trusted execution environment. In the trusted execution environment, the local restoration key is used to restore the context information of the target character object, and annotation suggestions are generated based on the context information. The de-identification task package and its annotation suggestions are then provided to the annotation end.

[0013] Receive annotation results returned from multiple annotation terminals, and reconstruct the annotated characters into complete text consistent with the original document layout based on the coordinate mapping table;

[0014] Based on the boundary coordinates, the image region is merged with the complete text to generate the final document.

[0015] In the above-mentioned document data security annotation method based on confidential computing, the character objects include text, letters, numbers, punctuation marks, and various symbols;

[0016] The reconstructed complete text contains character objects and spaces.

[0017] In the above-mentioned document data security annotation method based on confidential computing, the coordinate mapping table is constructed using text logical coordinates or unique identifiers as keys;

[0018] The local restoration key consists of the target character object and the mapping terms corresponding to a predetermined number of adjacent character objects in the text logical coordinates.

[0019] In the above-mentioned document data security annotation method based on confidential computing, the annotation result includes the unique identifier of the target character object and the annotation mapping relationship between the identified character and the annotated character;

[0020] The annotation mapping relationship is generated by annotating the single character images based on the target character object displayed by the annotation terminal user, the recognized characters, and annotation suggestions.

[0021] After generating the annotation mapping relationship, the annotation mapping relationship is digitally signed within the trusted execution environment and then returned to the management terminal.

[0022] In the aforementioned method for secure annotation of document data based on confidential computing, this method also includes a dispute arbitration process:

[0023] After reconstructing the complete text, the complete text is compared character by character with the OCR recognition results of each text region in the original document image;

[0024] If there are inconsistent character objects in the comparison results, the inconsistent character object is extracted and used as the target character object in the arbitration stage.

[0025] The physical pixel coordinates of the target character object are obtained according to the coordinate mapping table, and a portion of the document image containing the target character object and its context is cropped with the physical pixel coordinates as the center.

[0026] All captured partial document images are packaged into an arbitration package, which is then sent to a trusted execution environment. The arbitration end performs arbitration annotation on the target character objects in the partial document images in the trusted execution environment, generates arbitration annotation results, and returns them to the management end.

[0027] The management system updates the complete text based on the arbitration annotation results.

[0028] In the above-mentioned document data security annotation method based on confidential computing, the recognition results obtained by performing OCR recognition on each text region in the original document image during the dispute arbitration process and the recognition results obtained by performing OCR recognition on the text region during the annotation process are the same OCR recognition results, or they are recognized separately by mutually independent OCR modules.

[0029] In the above-mentioned document data security annotation method based on confidential computing, the management end updates the corresponding identification character in the coordinate mapping table using the annotation character based on the received annotation results;

[0030] Based on the updated coordinate mapping table, the context mapping items are extracted again to regenerate the local restoration key for the next round of de-identification task package;

[0031] The next round of de-identification task package and its local restoration key are distributed to the trusted execution environment, and the process of context restoration, generating annotation suggestions, distributing the de-identification task package and its annotation suggestions to the annotation end, and obtaining the annotation results are re-executed.

[0032] Repeat the above process until any of the following conditions are met, then perform the complete text reconstruction and image-text merging based on the updated coordinate mapping table:

[0033] The annotation results for all character objects remained consistent for two consecutive rounds.

[0034] Reach the preset maximum number of iterations;

[0035] The administrator can manually terminate the iteration.

[0036] A secure annotation system for document data based on confidential computing, comprising a management terminal, a trusted execution environment, and an annotation terminal;

[0037] The management terminal is configured as follows:

[0038] Establish a planar coordinate system for the original document image, separate the image area from the text area through layout analysis, record the boundary coordinates of each area, and save the image area;

[0039] The text region is subjected to OCR recognition to obtain the recognized characters, physical pixel coordinates, and text logical coordinates allocated by rows and columns for each character object; individual character images of each character object are cut out and assigned a unique identifier;

[0040] Establish a mapping table containing the text logical coordinates, physical pixel coordinates, unique identifiers, and coordinates of the identified characters for each character object;

[0041] Each individual character image, along with its unique identifier and recognition character, is randomly assigned to multiple de-identification task packages. Based on the coordinate mapping table, a local restoration key is generated for each de-identification task package to restore the context information of its target character object.

[0042] Distribute the de-identification task package and its local restoration key to a trusted execution environment;

[0043] Receive annotation results returned from multiple annotation terminals, and reconstruct the annotated characters into complete text consistent with the original document layout based on the coordinate mapping table;

[0044] Based on the boundary coordinates, the image region is merged with the complete text to generate the final document;

[0045] A trusted execution environment is used to restore the context information of the target character object using the local restoration key, generate annotation suggestions based on the context information, provide the de-identification task package and its annotation suggestions to the annotation end, and generate, according to the annotation operation of the annotation user, a unique identifier of the target character object, the annotation mapping relationship between the identified character and the annotated character, and return it to the management end.

[0046] The annotation end is used to display small images of individual characters of the target character object, the recognized characters and annotation suggestions, and to receive annotation operations from the annotation end user;

[0047] The trusted execution environment is located at the annotation end or at the server end accessible to the annotation end.

[0048] The aforementioned document data security annotation system based on confidential computing also includes an arbitration end;

[0049] The management terminal is further configured to: after reconstructing the complete text, compare the complete text with the OCR recognition results of each text region in the original document image character by character; if there are inconsistent character objects in the comparison results, extract the inconsistent character objects as the target character objects in the arbitration stage; obtain the physical pixel coordinates of the inconsistent character objects according to the coordinate mapping table, and, with the physical pixel coordinates as the center, extract a portion of the document image containing the target character object and its context from the original document image; package all the extracted partial document images into an arbitration package, and send the arbitration package to the trusted execution environment;

[0050] The arbitration end is configured to display the annotation screen output by the trusted execution environment, so that the arbitrator can perform arbitration annotation on target character objects in some document images, generate arbitration annotation results and return them to the management end.

[0051] In the aforementioned document data security annotation system based on confidential computing, the management terminal includes a first OCR recognition module for the annotation process and a second OCR recognition module for the arbitration process.

[0052] The first OCR recognition module and the second OCR recognition module are either the same module or independent modules.

[0053] The advantages of this invention are:

[0054] This solution establishes a coordinate system and a restoration key based on image and text coordinate encapsulation. Combined with single-word segmentation to generate single-word task packages without context, it completely separates the anonymization and annotation process from the typesetting information. After proofreading, the key is used to restore the document, thereby achieving a complete manually proofread result while protecting privacy, thus balancing security and restoration accuracy.

[0055] This solution combines single-word segmentation with distributed annotation in a trusted execution environment, allowing annotators to process isolated single words without accessing context. At the same time, it utilizes the trusted execution environment for context-based secure reasoning to assist in proofreading, thereby improving the proofreading efficiency and accuracy of annotators while protecting document privacy.

[0056] Since each annotator sees a scrambled image of a single character, even if all these annotators share all the task packages, they cannot piece together and restore the complete document without the restoration key, which effectively ensures the information security of the annotation process.

[0057] By separating text from images, annotators only process the separated text characters without having to access the image area. This avoids the leakage of sensitive information in the image, makes the annotation task more focused, reduces interference, and improves annotation concentration and efficiency. Attached Figure Description

[0058] Figure 1 This is a flowchart of a document data security annotation method based on confidential computing, according to Embodiment 1 of the present invention.

[0059] Figure 2 This is a system architecture diagram of the document data security annotation system based on confidential computing, according to Embodiment 1 of the present invention.

[0060] Figure 3 This is a schematic diagram of establishing a planar coordinate system for a document image in Embodiment 1 of the present invention;

[0061] Figure 4 This is a schematic diagram illustrating the separation of image and text regions in a document image after layout analysis in Embodiment 1 of the present invention.

[0062] Figure 5 This is a schematic diagram of the annotation terminal REE interface displaying single-character images, character recognition, and annotation suggestions in Embodiment 1 of the present invention;

[0063] Figure 6 This is a schematic diagram of the document reconstruction process in Embodiment 1 of the present invention;

[0064] Figure 7 This is a flowchart illustrating the dispute arbitration process in Embodiment 2 of the present invention;

[0065] Figure 8 This is a flowchart of the iterative annotation process in Embodiment 3 of the present invention. Detailed Implementation

[0066] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0067] Example 1

[0068] like Figure 1 and Figure 2As shown, this invention provides a method and system for secure annotation of document data based on confidential computing and image-text separation. The method first performs layout analysis and image-text separation on the original document image, using OCR technology to identify the text content and physical location of each character in the text region, and then cuts out isolated single-character images for each character. A coordinate mapping table is established to associate the text logical coordinates, physical pixel coordinates, unique identifiers, and identification characters of each character. Subsequently, the single-character images are randomly assigned to multiple de-identification task packages. Based on the coordinate mapping table, a local restoration key is generated for each de-identification task package to restore the context information of several target character objects. The local restoration key for each de-identification task package contains only the target character objects belonging to that de-identification task package and their limited adjacent character information.

[0069] Each de-identification task package and its local restoration key are provided to the Trusted Execution Environment (TEE). An auxiliary proofreading model runs within the TEE, using the local restoration key to restore the context information of the target character objects in each de-identification task package, and the auxiliary proofreading model generates annotation suggestions based on this.

[0070] Here, the TEE can be deployed at each annotation endpoint. In this case, based on the assigned de-identification task package, the de-identification task package and its local restoration key are provided to the corresponding annotation endpoint's TEE. After the TEE generates annotation suggestions for all target character objects in the de-identification task package, it sequentially displays the individual character images, recognized characters, and annotation suggestions of the target character objects in the de-identification task package within the annotation endpoint's Rich Execution Environment (REE) for user annotation. Alternatively, the TEE can be deployed on a server. This TEE sequentially receives each de-identification task package and its local restoration key, restores the target character objects, generates annotation suggestions, and then distributes the de-identification task package and its annotation suggestions to the corresponding annotation endpoint. The annotation endpoint will then display the individual character images, recognized characters, and annotation suggestions of the target character objects in its de-identification task package for user annotation. The following description in this embodiment uses the former as an example.

[0071] The management system collects annotation results from multiple sources, reconstructs complete text with a layout completely consistent with the original document using a coordinate mapping table, and merges it with the image area to generate the final document.

[0072] The following section uses a scanned copy of a corporate purchase contract as the original input and provides a detailed explanation of each step with a specific example.

[0073] Preprocessing stage

[0074] The management system first receives the original document image, such as a high-resolution scan of a corporate procurement contract, records its basic parameters: pixel size, resolution, file format, etc., and generates an original document image identifier ID. For documents containing multiple pages, each page corresponds to an original document image, and each page has an original document image identifier ID.

[0075] Subsequently, as Figure 3 As shown, a planar coordinate system is established for the image. In order to facilitate the description of the positional relationships in the image, this embodiment takes the lower left corner of the image as the origin, the horizontal direction to the right as the positive X-axis, and the vertical direction upward as the positive Y-axis. Each pixel point corresponds to a unique pixel coordinate (x, y). This coordinate system covers the entire image and provides a unified benchmark for the subsequent recording of the position of all areas and characters.

[0076] This solution utilizes layout analysis technology to automatically identify and distinguish image and text regions in the original document image, segments these regions, stores the image regions on a server, and transmits the text regions to a secure preprocessing layer on the management end for the next stage of OCR recognition. Layout analysis can employ deep learning-based detection models, such as YOLO or Mask R-CNN, or traditional image processing methods; this solution does not impose any restrictions on either approach.

[0077] like Figure 4 As shown, a page of the scanned contract in this embodiment contains a company logo image area and two text areas. Based on layout analysis, three rectangular bounding boxes are output. Each bounding box contains the area type (image or text) and its coordinates (x1, y1, x2, y2) in the original image, where (x1, y1) is the coordinate of the top-left corner of the area, and (x2, y2) is the coordinate of the bottom-right corner. The boundary coordinates of the original document image and the boundary coordinates of each area are recorded. The image area is stored on the backend processing server and does not participate in the subsequent character-level annotation process. The text area is transmitted to the security and processing layer for subsequent OCR recognition.

[0078] In the security preprocessing layer, the OCR engine is invoked to perform character recognition on the text area. OCR processing identifies each character object within the text area, including visible text, letters, numbers, punctuation, and various symbols. Based on the OCR recognition results and layout analysis, logical text coordinates are assigned to each character object by row and column. For example, row numbers range from 0 to 4, corresponding to 5 lines of text, and column numbers start from 0 and increment character by character. These logical text coordinates preserve the original document's line breaks and paragraph relationships.

[0079] Based on the established planar coordinate system, the physical pixel coordinates (x1, y1, x2, y2) of each character object in the original document image are locked, where (x1, y1) is the top-left corner of the character object and (x2, y2) is the bottom-right corner. By using pixel size, the full-width and half-width characters of a symbol can be determined, thus enabling high-precision restoration of symbol types, such as full-width colons and half-width colons, during subsequent text restoration. Furthermore, OCR cannot recognize character content in blank areas such as first-line indentation and placeholders in the original document. This embodiment, combined with physical pixel coordinates, preserves these spaces during later text restoration, ensuring that text indentation, alignment, and other details are accurately recorded and restored.

[0080] At this point, the logical text coordinates and physical pixel coordinates of each character object, as well as the recognized characters obtained based on OCR recognition, have been obtained.

[0081] Based on the physical pixel coordinates of each character object, individual character images are independently cut out and assigned a unique identifier. Specifically, based on the character bounding box (x1, y1, x2, y2), a preset safety margin is extended outwards, preferably by 2-5 pixels, to crop an image region containing only that single character object. Extending the safety margin ensures that character edges are not accidentally truncated, while avoiding the introduction of image information from adjacent characters and the inclusion of spaces, thus preventing the loss of space information. After cutting, a globally unique identifier is generated for each character image, such as img_003, img_004, etc. Then, based on the previously obtained text logical coordinates and physical pixel coordinates, complete information for each character object (character image) can be obtained. This complete information for all character objects is then integrated and combined to obtain a coordinate mapping table for the current original document image.

[0082] The coordinate mapping table can use text logical coordinates or a unique identifier as the key; this embodiment prefers the latter. The values ​​corresponding to each key include: text logical coordinates, physical pixel coordinates, and a recognized character.

[0083]

[0084] For example, the correct original text of a company's purchase contract is:

[0085] Purchase Contract

[0086] Party A: XX Technology Co., Ltd.

[0087] Party B: XY Supply Co., Ltd.

[0088] Contract No.: 20240306

[0089] Procurement Items: 10 servers

[0090] OCR recognition result:

[0091] Procurement with Company

[0092] Party A: XX Technology Co., Ltd.

[0093] Party B: XY Supply Co., Ltd.

[0094] With Company Code Error: 20240306

[0095] Procurement Content: 10 servers

[0096] An example of the obtained coordinate mapping is as follows:

[0097]

[0098] The management end randomly allocates all cut single-character small images (together with their corresponding unique identifiers and recognition characters) to a plurality of desensitization task packages, each desensitization task package contains single-character small images of a plurality of character objects, and distributes these task packages to different labeling ends.

[0099] In this embodiment, there are 3 annotators, that is, 3 labeling ends. The management end randomly allocates the character objects in the above contract, and obtains the following three task packages:

[0100] Task Package A:

[0101] ["采","司","乙","亏","4","容","器","限","司"……] (each element represents the recognition character of a character object and the corresponding single-character small image and unique identifier)

[0102] Task Package B:

[0103] ["购","甲","编",":","3","服","1","公"……]

[0104] Task Package C:

[0105] ["含","方","2","0","6","内","务","台","有"……]

[0106] Each task package only contains several isolated and random single-character small images and their recognition characters, and does not contain context information or the original image. After the annotator opens the task package, he can only see a series of shuffled single-character small images and the corresponding recognition character for each small image. For example, for a small image displaying the character "合", the recognition character marked above is "含".

[0107] For each desensitization task package, the management end generates a local restoration key for it, which is used to restore the context information within a limited range of the target character object in the TEE, so as to generate high-quality annotation suggestions. The local restoration key is composed of the mapping items corresponding to the target character object and a predetermined number of adjacent character objects thereof under the text logical coordinates, and the specific generation process is as follows:

[0108] (1) For each target character object in the desensitization task package (that is, the character object that the task package requires the corresponding annotator to annotate), the management end searches for the mapping item of the character object from the coordinate mapping table according to its unique identifier.

[0109] (2) Based on the text logical coordinates in the mapping item of the target character object, extract its context in the original document: a predetermined number of adjacent characters, and the predetermined number can be N characters before and after, for example, 1, 2, 3 or 10. The specific value can be configured according to the balance between security and annotation accuracy.

[0110] (3) Combine the mapping items of these context character objects with the mapping item of the target character itself to form a "local mapping item set", which is the local restoration key fragment of the target character. It contains the text logical coordinates, physical pixel coordinates and recognized characters of these character objects, based which the position and text logical relationship can be restored.

[0111] (4) If the desensitization task package contains multiple target character objects, package the local restoration key fragments of each target character object into a complete local restoration key. The key can be encoded in a structured data format (such as JSON), then converted to binary via UTF-8, and then encoded by Base64 to generate the final string for easy transmission.

[0112] For example, for the target character object "合" (text logical coordinate (0,2)) in task package C, the predetermined number of adjacent characters is 1 character before and after. According to the text logical coordinates, the recognized adjacent character in the upper context is "购" at (0,1), and the recognized adjacent character in the lower context is "司" at (0,3). Therefore, the local restoration key fragment of this target character object includes the following mapping items:

[0113] Target character: unique identifier—(0,2)—physical pixel coordinates—contain

[0114] Left character: unique identifier—(0,1)—physical pixel coordinates—购

[0115] Right character: unique identifier—(0,3)—physical pixel coordinates—司

[0116] If task package C further includes another target character object "fang" (text logical coordinate (2,1)), the mapping items of its adjacent characters shall also be extracted for it. Finally, the local restoration key of the entire task package C contains mapping information of all target character objects and their respective adjacent characters.

[0117] The management end distributes each desensitization task package and its corresponding local restoration key to the TEE of each annotation end through a secure channel. TEE is a CPU hardware-level secure isolation area, which can ensure that the code and data running inside TEE are invisible to the operating system and other processes. Before distribution, TEE will establish a trust relationship with the management end through a remote attestation mechanism to ensure that the TEE environment is authentic and has not been tampered with.

[0118] It is worth noting that the generation process of the local restoration key does not rely on encryption algorithms to hide the content, because the key itself will be passed into TEE, which has the security feature of hardware isolation. However, to prevent eavesdropping during transmission, a TLS encryption channel will be established between all parties. Meanwhile, since each task package is independent and the target character objects in each task package are random, the local restoration key only contains limited context. Therefore, even if the transmission is eavesdropped, illegal personnel cannot obtain context beyond the scope, so it can still achieve a high-level security protection effect.

[0119] After receiving the task package and the local restoration key, the TEE of each annotation end performs the following operations:

[0120] (1) Parse the local restoration key; for each target character object in the task package, according to the adjacent character mapping items provided in the key, splice the recognition characters therein into a section of context text with limited length based on the text logical coordinates and physical pixel coordinates. For example, for the Chinese character "han", the context may be "gou han si" or a longer "XX gou han si XX", which depends on the predetermined number. Moreover, the context only exists inside TEE and will not be transmitted to the ordinary execution environment.

[0121] (2) Obtain the single-character small image of the target character object from the task package, and input the unique identifier of the single-character small image, the recognition character, and the above context text in character form into the auxiliary proofreading model pre-deployed in TEE. This model is a lightweight masked pre-training model with semantic understanding capability, which can judge wrong characters based on context text. The model performs inference inside TEE and outputs annotation suggestions for the target character object. For example, combining the context "XX gou han si XX", the model judges that the recognition character "han" corresponding to the target character object is most likely "he", so it outputs the annotation suggestion "he".

[0122] Preferably, for each target character object, after obtaining its annotation suggestion, the TEE automatically deletes the local restoration key of the corresponding target character object. It should be noted that the TEE has been confirmed to be untampered through remote trusted authentication, which can ensure that the pre-deployed application with the aforementioned deletion function is executed as deployed.

[0123] Spaces existing in the original document image will also be included in the context text. Spaces are obtained based on text logical coordinates and physical pixel coordinates. The annotation end determines whether there is a space between characters according to the spacing between physical pixel coordinates of adjacent characters based on text logical coordinates and a preset spacing threshold, and determines the number of spaces to be inserted according to the ratio of the spacing to the threshold, thereby accurately restoring the space width in the original typesetting in the context text. For example, between "cai" and "gou", if there is no space, the text logical coordinates of the two characters are adjacent, and their physical pixel coordinates are also relatively adjacent; for example, the X-coordinate difference is usually within 30 pixels. When there is a space, the text logical coordinates are adjacent, but the physical pixel coordinates are obviously separated; for example, the X-coordinate difference usually exceeds 60 pixels, and the spacing threshold can be set to 50 pixels.

[0124] (3) Packing the annotation suggestion, the single-character small image of the target character object, and the recognition character of the character object together, and transmitting them to the REE of the annotation end through the secure output interface of the TEE.

[0125] The REE interface of the annotation end displays a series of character object entries in the current task package. For each character object, if the annotation suggestion is inconsistent with the OCR recognition character result, the interface display is as Figure 5 shown. If they are consistent, only the recognition character, that is, the current recognition character in the mapping item of the corresponding target character object, can be displayed.

[0126] For the case where the recognition character is inconsistent with the annotation suggestion, the annotator selects or inputs the correct character by clicking one of the two characters on the interface or through handwritten input based on the actual glyph of the single-character small image and in combination with the annotation suggestion. For the case where the recognition character is consistent with the annotation suggestion, the annotator judges whether the recognition is accurate according to the single-character small image. If it is not accurate, the correct character is input; if it is accurate, the recognition character is adopted by clicking.

[0127] The annotation result submitted by the annotator is transmitted back from the REE to the TEE. After the TEE receives the result, it generates an annotation mapping relationship, that is, the mapping from the original recognition character to the annotated character. Preferably, when the annotation result is inconsistent with the original character, the mapping relationship includes two characters at the same time, such as { img_003, "含": "合"}; when the annotation result is consistent with the original recognition character, one character is retained in the mapping relationship, such as { img_001, "采"}.

[0128] After the annotation operation is completed, the TEE digitally signs the annotation result: the private key of the TEE is used to sign data such as the annotation result, task ID, and timestamp to generate signature data, and the signed result package is uploaded to the management terminal through a secure channel.

[0129] The management terminal receives annotation results from multiple annotation terminals, each result is attached with a TEE signature, and the management terminal can verify the validity of the signature, thereby confirming that the result indeed comes from a trusted environment and has not been tampered with.

[0130] After the management terminal collects all annotation results returned by the annotation terminals, it reconstructs the annotated characters into a complete text with the same layout as the original document based on the coordinate mapping table. The specific process is as follows:

[0131] (1) Read the coordinate mapping table and the original recognized character version from the storage;

[0132] (2) For each annotation result package, extract the annotation mapping relationship therein. For example, annotator A returns {img_003, "含": "合"}, annotator B returns {img_004, "司": "同"}, annotator C returns {img_023, "亏": "号"}, the recognition character corresponding to the unique identifier in the coordinate mapping table is updated to the annotated character. Of course, if the annotated character is consistent with the original recognized character, no update is required.

[0133] (3) After updating all characters, sort according to the text logical coordinates (line number, column number) in the coordinate mapping table: sort by line number from small to large, and within the same line, sort by column number from small to large. Take out characters one by one in this order and splice them into a character string. At the same time, correct physical pixel coordinates, insert spaces, adjust full-width / half-width symbols, and finally obtain layout text that is highly consistent with the original document, including line breaks, spaces, and full-width / half-width symbols.

[0134] As Figure 6 shown, the management terminal has stored the boundary coordinates of the image area and its corresponding image data in the preprocessing stage. After the complete text is reconstructed, the image area is merged with the complete text according to the pixel proportion. The merging method is: calculate the occupation interval of each area on the target blank document for reconstructing the original document according to the boundary coordinates of the original document image and the boundary coordinates of each area, adjust the size of the image area and the text size according to the occupation interval, place each area at the corresponding occupation interval on the target blank document, and the finally generated document is consistent with the content of the original document image.

[0135] This solution establishes a coordinate system, separates text and images, segments individual characters to assign annotation tasks, and creates a coordinate mapping table. It breaks down the original document into isolated, context-free small images of individual characters, randomly assigning them to multiple anonymized task packages. This prevents annotators from accessing the complete document, complete text lines, or any semantic context. Simultaneously, it securely restores a limited number of neighboring characters within the TEE using a local restoration key, combined with a lightweight auxiliary proofreading model to generate intelligent annotation suggestions. This significantly improves annotation efficiency and accuracy while ensuring privacy. Finally, the management end, through the coordinate mapping table, can losslessly reconstruct the complete text, including original line breaks, spaces, and full-width / half-width symbols, and merge it with the image region to generate the final document. This achieves secure, efficient, and high-quality annotation results with invisible data, making it particularly suitable for OCR error correction and data annotation scenarios in highly sensitive documents such as contracts, medical records, and confidential documents.

[0136] Example 2

[0137] This embodiment is basically the same as Embodiment 1, except that this embodiment further includes an arbitration process, which is executed after the complete text reconstruction is completed.

[0138] As described in Example 1, after collecting the annotation results returned by each annotation terminal, the management terminal reconstructs the complete text based on the coordinate mapping table. Subsequently, as... Figure 7 As shown, the management end compares the reconstructed complete text with the OCR recognition results of each text region in the original document image character by character. It should be noted that the OCR recognition results used for this comparison can be the same as the results obtained from the OCR recognition of the text regions during the annotation process, or they can be obtained by independent OCR modules. When different OCR modules are used, cross-validation can be achieved to further improve the credibility of the arbitration benchmark.

[0139] If the comparison results are completely consistent, it means that the text after being annotated and corrected by the user on the annotation end is completely consistent with the original OCR recognition result. There is no need to trigger arbitration, and the text and image merging stage can be directly entered.

[0140] If the comparison results show inconsistent character objects, then the inconsistent character object is extracted and used as the target character object in the arbitration phase.

[0141] The physical pixel coordinates of the target character object are obtained from the coordinate mapping table. Using these physical pixel coordinates as the center, a portion of the original document image containing the target character object and its context is cropped. The cropping range can be set to a predetermined window size centered on the target character to provide the arbitrator with sufficient visual reference.

[0142] In this embodiment, it is preferable to extract each page containing disputed characters.

[0143] Pack all intercepted partial document images into an arbitration package, deliver the arbitration package to a trusted execution environment, and operate it by an arbitrator with arbitration authority.

[0144] What the arbitrator sees is the labeled picture output by the TEE and displayed on the arbitration end—the partial document image and the prominently marked target character objects. Based on the complete partial document image, the arbitrator labels one or more target character objects marked as needing arbitration in the image one by one, independently gives the recognition result of the target characters, and finally obtains the arbitration labeling result.

[0145] Return the arbitration labeling result to the management end, and the management end uses the arbitration labeling result to update the corresponding recognized characters in the coordinate mapping table.

[0146] Embodiment 3

[0147] As Figure 8 shown, this embodiment is basically the same as Embodiment 1 or Embodiment 2, the difference is that this embodiment takes into account that the labeling suggestion quality of the auxiliary proofreading model is affected by the quality of the recognized characters in the coordinate mapping table, and further designs an iterative optimization labeling process to ensure labeling accuracy through at least two rounds of interaction.

[0148] The management end executes the following iterative optimization process based on the received labeling results:

[0149] Update the corresponding recognized characters in the coordinate mapping table using the labeled characters, for example, update "han" corresponding to img_003 to "he", update "si" corresponding to img_004 to "tong", and update "kui" corresponding to img_030 to "hao".

[0150] Based on the updated mapping table, re-extract context mapping items for the target character objects of each desensitization task package and generate a local restoration key. At this time, the recognized characters in the mapping table are already characters corrected after one round of labeling, so the context generated in the next round will be more accurate.

[0151] Distribute the desensitization task package and its new local restoration key to the TEE, and re-execute the processes of context restoration, distributing the desensitization task package and its labeling suggestions to the labeling end, generating labeling suggestions by the auxiliary proofreading model, REE display and user labeling. In this round of labeling, it is preferable to prominently display labeling suggestions inconsistent with the characters recorded in the current coordinate mapping table. Specifically, if the labeling suggestion generated by the TEE is different from the current character in the mapping table, the labeling personnel are reminded by means of highlight and color marking, which is convenient for them to quickly locate the label.

[0152] The above iterative process can be repeated, and the stop condition can be preset as any of the following situations:

[0153] The annotation results of all character objects remain consistent for two consecutive rounds;

[0154] The preset maximum number of iterations is reached, for example, 3 rounds;

[0155] The administrator user manually terminates the iteration.

[0156] When the stop condition is satisfied, complete text reconstruction and image-text combination are performed based on the finally updated coordinate mapping table, and the final document is output.

[0157] In this embodiment, through at least two rounds of iteration, the result corrected in the previous round provides more accurate context for subsequent rounds, enabling the auxiliary proofreading model to continuously generate more reliable annotation suggestions. Taking the character "亏" as an example, it may fail to be correctly recognized, inferred and corrected in the first round; when entering the second round, its context "合同编" has been corrected to the correct text, and the model can more easily infer that "亏" corresponds to "号" based on this, thereby improving the overall annotation accuracy.

[0158] It should be noted that although multi-round iteration is adopted, in each round, TEE generates annotation suggestions for each character in the current task package, and displays them side by side with the single-character thumbnail and OCR recognition characters on the REE interface. Annotators can perform rapid annotation by scanning, and can accept annotation suggestions consistent with the current recognition characters in one-click batch. For characters whose annotation suggestions are inconsistent with the recognition characters, they can manually click or input the correct characters. Therefore, in this solution, the efficiency of each round of annotation is very high. Moreover, after the first round, most of the character annotation suggestions are consistent with the correct results, and the number of characters requiring manual intervention in subsequent rounds is very small. Therefore, the total workload of multi-round iteration is basically the same as or even lower than that of traditional single-round character-by-character check, and both text privacy protection and final annotation quality have been significantly improved.

[0159] The above description is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any equivalent substitution or modification made by a person skilled in the art within the technical scope disclosed by the present invention based on the technical solution of the present invention and the inventive concept thereof shall be covered within the protection scope of the present invention.

Claims

1. A method for secure annotation of documentary data based on confidential computing, characterized in that, Includes the following steps: Establish a planar coordinate system for the original document image, separate the image area from the text area through layout analysis, record the boundary coordinates of each area, and save the image area; The text region is subjected to OCR recognition to obtain the recognized characters, physical pixel coordinates, and text logical coordinates allocated by row and column for each character object; Cut out individual character images from each character object and assign them unique identifiers; Establish a mapping table containing the text logical coordinates, physical pixel coordinates, unique identifiers, and coordinates of the identified characters for each character object; Each individual character image, along with its unique identifier and recognition character, is randomly assigned to multiple de-identification task packages. Based on the coordinate mapping table, a local restoration key is generated for each de-identification task package to restore the context information of several target character objects. The de-identification task package and its local restoration key are distributed to a trusted execution environment. In the trusted execution environment, the local restoration key is used to restore the context information of the target character object, and annotation suggestions are generated based on the context information. The de-identification task package and its annotation suggestions are then provided to the annotation end. Receive annotation results returned from multiple annotation terminals, and reconstruct the annotated characters into complete text consistent with the original document layout based on the coordinate mapping table; Based on the boundary coordinates, the image region is merged with the complete text to generate the final document.

2. The document data security annotation method based on confidential computing according to claim 1, characterized in that, The character objects include text, letters, numbers, punctuation marks, and various symbols; The reconstructed complete text contains character objects and spaces.

3. The document data security annotation method based on confidential computing according to claim 1, characterized in that, Construct the coordinate mapping table using text logical coordinates or unique identifiers as keys; The local restoration key consists of the target character object and the mapping terms corresponding to a predetermined number of adjacent character objects in the text logical coordinates.

4. The document data security annotation method based on confidential computing according to claim 1, characterized in that, The annotation results include the unique identifier of the target character object and the annotation mapping relationship between the identified character and the annotated character; The annotation mapping relationship is generated by annotating the single character images based on the target character object displayed by the annotation terminal user, the recognized characters, and annotation suggestions. After generating the annotation mapping relationship, the annotation mapping relationship is digitally signed within the trusted execution environment and then returned to the management terminal.

5. The document data security annotation method based on confidential computing according to claim 1, characterized in that, This method also includes a dispute arbitration process: After reconstructing the complete text, the complete text is compared character by character with the OCR recognition results of each text region in the original document image; If there are inconsistent character objects in the comparison results, the inconsistent character object is extracted and used as the target character object in the arbitration stage. The physical pixel coordinates of the target character object are obtained according to the coordinate mapping table, and a portion of the document image containing the target character object and its context is cropped with the physical pixel coordinates as the center. All captured partial document images are packaged into an arbitration package, which is then sent to a trusted execution environment. The arbitration end performs arbitration annotation on the target character objects in the partial document images in the trusted execution environment, generates arbitration annotation results, and returns them to the management end. The management system updates the complete text based on the arbitration annotation results.

6. The document data security annotation method based on confidential computing according to claim 5, characterized in that, The recognition results obtained by performing OCR recognition on each text region in the original document image during the dispute arbitration process and the recognition results obtained by performing OCR recognition on the text region during the annotation process are the same OCR recognition results, or they are recognized separately by independent OCR modules.

7. The document data security annotation method based on confidential computing according to claim 1, characterized in that, Based on the received annotation results, the management terminal updates the corresponding recognition characters in the coordinate mapping table using the annotation characters; Based on the updated coordinate mapping table, the context mapping items are extracted again to regenerate the local restoration key for the next round of de-identification task package; The next round of de-identification task package and its local restoration key are distributed to the trusted execution environment, and the process of context restoration, generating annotation suggestions, distributing the de-identification task package and its annotation suggestions to the annotation end, and obtaining the annotation results are re-executed. Repeat the above process until any of the following conditions are met, then perform the complete text reconstruction and image-text merging based on the updated coordinate mapping table: The annotation results for all character objects remained consistent for two consecutive rounds. Reach the preset maximum number of iterations; The administrator can manually terminate the iteration.

8. A document data security annotation system based on confidential computing, characterized in that, This includes the management interface, trusted execution environment, and annotation interface; The management terminal is configured as follows: Establish a planar coordinate system for the original document image, separate the image area from the text area through layout analysis, record the boundary coordinates of each area, and save the image area; The text region is subjected to OCR recognition to obtain the recognized characters, physical pixel coordinates, and text logical coordinates allocated by rows and columns for each character object; individual character images of each character object are cut out and assigned a unique identifier; Establish a mapping table containing the text logical coordinates, physical pixel coordinates, unique identifiers, and coordinates of the identified characters for each character object; Each individual character image, along with its unique identifier and recognition character, is randomly assigned to multiple de-identification task packages. Based on the coordinate mapping table, a local restoration key is generated for each de-identification task package to restore the context information of its target character object. Distribute the de-identification task package and its local restoration key to a trusted execution environment; Receive annotation results returned from multiple annotation terminals, and reconstruct the annotated characters into complete text consistent with the original document layout based on the coordinate mapping table; Based on the boundary coordinates, the image region is merged with the complete text to generate the final document; A trusted execution environment is used to restore the context information of the target character object using the local restoration key, generate annotation suggestions based on the context information, provide the de-identification task package and its annotation suggestions to the annotation end, and generate, according to the annotation operation of the annotation user, a unique identifier of the target character object, the annotation mapping relationship between the identified character and the annotated character, and return it to the management end. The annotation end is used to display small images of individual characters of the target character object, the recognized characters and annotation suggestions, and to receive annotation operations from the annotation end user; The trusted execution environment is located at the annotation end or at the server end accessible to the annotation end.

9. The document data security annotation system based on confidential computing according to claim 8, characterized in that, It also includes the arbitration side; The management terminal is further configured to: after reconstructing the complete text, compare the complete text with the OCR recognition results of each text region in the original document image character by character; if there are inconsistent character objects in the comparison results, extract the inconsistent character objects as the target character objects in the arbitration stage; obtain the physical pixel coordinates of the inconsistent character objects according to the coordinate mapping table, and extract a portion of the document image containing the target character objects and their context from the original document image with the physical pixel coordinates as the center; All captured document images are packaged into an arbitration package, and the arbitration package is sent to the trusted execution environment. The arbitration end is configured to display the annotation screen output by the trusted execution environment, so that the arbitrator can perform arbitration annotation on target character objects in some document images, generate arbitration annotation results and return them to the management end.

10. The document data security annotation system based on confidential computing according to claim 8, characterized in that, The management terminal includes a first OCR recognition module for the annotation process and a second OCR recognition module for the arbitration process; The first OCR recognition module and the second OCR recognition module are either the same module or independent modules.