Multi-language alignment method and system for rich text document analysis
Through data augmentation and multilingual dictionary construction, multilingual document data generating alignment semantics, and aligning document representations of different languages in the encoding space, solving the problem of cross-language rich text document parsing and achieving efficient multilingual alignment and general capabilities.
Patent Information
- Application Number
- CN202311544121.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to effectively parse rich text documents across languages, especially when it does not rely on large amounts of labeled data and retrained models, and the multilingual alignment capabilities are insufficient.
The data enhancement steps generate document data in different languages, build multilingual dictionaries and use vocabulary replacement to generate multilingual document data with aligned semantics. Then, document representations with the same semantics and different languages are aligned in the encoding space to achieve multilingual semantic alignment.
Unsupervised multilingual alignment is realized, reducing the need to rely on labeled data, improving the efficiency of cross-language document parsing and multilingual common capabilities, and reducing data labeling costs.
Smart Images

Figure CN120046597A_ABST
Abstract
Claims
1. A multi - language alignment method for rich text document parsing, characterized in that, it includes: Data enhancement step: Without changing the semantic content of the target document, generate document data in different languages. First, construct a multi - language dictionary based on the target document corpus, and then use the method of word - list replacement based on the multi - language dictionary to perform data enhancement and generate multi - language document data with aligned semantics; Multi - language alignment training step: Based on the existing document encoder, align the representations of documents in different languages with the same semantics in the encoding space, so as to achieve multi - language semantic alignment of documents.
2. The multi - language alignment method for rich text document parsing according to claim 1, characterized in that, The process of constructing the multi - language dictionary is as follows: The input of the dictionary construction process is the target rich text document corpus for data enhancement. This kind of document data is stored in the form of pictures or PDFs. Use the OCR tool to extract the text information and layout information in this kind of document in units of bounding boxes. A document consists of multiple bounding boxes, and each bounding box records a text field and a layout field. The layout field is the position coordinate information of the bounding box in the two - dimensional plane; Classify the languages according to the extracted text fields; Merge the texts into paragraph - level texts through string splicing operations, and use regular expressions to filter and clean the special characters in the texts; Use different tools for word segmentation for different languages to obtain entity words with continuous meanings; Adopt TF and Textrank algorithms to extract keywords from numerous words. The TF algorithm calculates the word frequency in the corpus and sorts them to obtain common words; The Textrank algorithm extracts keywords using the co - occurrence information between words in the document and judges the dependency relationship between words; Match and screen according to the open - source stop - word list, and perform part - of - speech filtering based on open - source NLTK or CoreNLP; Translate the keywords of the obtained target rich text document corpus using open - source translation tools or APIs to obtain a multi - language dictionary; The process of constructing the general dictionary is as follows: After passing through the same data cleaning, stop - word filtering, and part - of - speech filtering links for the existing open - source bilingual dictionaries, perform multi - language processing using English bridging.
3. The multi - language alignment method for rich text document parsing according to claim 1, characterized in that, Data enhancement is performed in units of bounding boxes. For the text modality of document data, first extract and clean it, and then pass it through a multi - language tokenizer, maintaining the same tokenization strategy and granularity. At this time, the text in the bounding box is in the form of a list of words. Randomly select words in the list according to the preset ratio threshold, query the replacement words in the multi - language dictionary. During the query process, give priority to the in - domain dictionary. If the query fails in the in - domain dictionary, then query in the general dictionary. If the query still fails in the general dictionary, continue to randomly select from the word list until the ratio threshold requirement is met or all the successfully queried words in the word list are taken out; After querying in the word list, multiple multi - language synonyms will be obtained for the replacement words, and randomly select and replace the original words among these synonyms; After completing the same-semantic multi-language replacement of the text, the layout information is changed by means of mask perturbation. Specifically, for all bounding boxes globally, 20% of the bounding boxes are selected for change, 80% remain unchanged, and the following change methods are randomly adopted for the layout information of 20% of the bounding boxes: set the coordinates to zero, that is, consider that the bounding box is at the origin position in the two-dimensional plane; replace the two-dimensional coordinates with random values; perturb the bounding box, including enlarging, shrinking, and rotating.
4. The multi-language alignment method for rich text document parsing according to claim 1, characterized in that, For the document data after data augmentation, a general document encoder is used to encode the bounding boxes. For each bounding box, the encoding result is a 1×768-dimensional vector representation, denoted as h i , and the expression is: h i = Normalize(LayoutXLM(x i )) ∈ R 768 Among them, Normalize is a regularization operation; LayoutXLM is a document encoder; x i is the bounding box input; R is the set of real numbers.
5. The multi-language alignment method for rich text document parsing according to claim 4, characterized in that, positive and negative examples are constructed through contrastive learning, a contrastive learning loss function is designed, the representations of positive example pairs are pulled closer, and the representations of negative example pairs are pulled farther apart, so as to achieve cross-language representation alignment. The expression is: Among them, B is the training batch size, that is, the number of samples in this training batch; exp is the exponential function with the natural constant e as the base; h p is the positive example representation; h j is the negative example representation; τ is the temperature coefficient, which is used to adjust the distribution shape of the representation. The smaller the temperature coefficient, the more the loss function will focus on the negative sample pairs.
6. A multi-language alignment system for rich text document parsing, characterized in that, comprising: Data augmentation module: Without changing the semantic content of the target document, generate document data in different languages. First, construct a multi-language dictionary according to the target document corpus, and then use the method of word list replacement based on the multi-language dictionary to perform data augmentation and generate multi-language document data with aligned semantics; Multi-language alignment training module: Based on the existing document encoder, align the representations of documents in different languages with the same semantics in the encoding space, so as to achieve multi-language semantic alignment of documents.
7. The multi-language alignment system for rich text document parsing according to claim 6, characterized in that, The process of constructing the multi-language dictionary is as follows: The input of the dictionary construction process is the target rich text document corpus for data augmentation. Such document data is stored in the form of pictures or PDFs. Use the OCR tool to extract the text information and layout information in such documents in units of bounding boxes. A document consists of multiple bounding boxes, and each bounding box records a text field and a layout field. The layout field is the position coordinate information of the bounding box in the two-dimensional plane; Perform language classification according to the extracted text fields; Merge the texts into paragraph-level texts through string concatenation operations, and use regular expressions to filter and clean the special characters in the texts; Use different tools for word segmentation for different languages to obtain entity words with continuous meanings; Adopt the TF and Textrank algorithms to extract keywords from numerous words. The TF algorithm calculates the word frequency in the corpus and sorts them to obtain common words; the Textrank algorithm extracts keywords using the co-occurrence information between words in the document and judges the dependency relationship between words; Perform matching and screening according to the open-source stop word list, and perform part-of-speech filtering based on the open-source NLTK or CoreNLP; Translate the keywords of the obtained target rich text document corpus using open-source translation tools or APIs to obtain a multi-language dictionary; The process of constructing the general dictionary is as follows: After passing through the same data cleaning, stop word filtering, and part-of-speech filtering processes for the existing open-source bilingual dictionaries, perform multi-lingualization using English bridging.
8. The multi - language alignment system for rich text document parsing according to claim 6, characterized in that, data augmentation is performed in units of bounding boxes. For the text modality of document data, it is first extracted and cleaned, and then passed through a multi - language tokenizer, maintaining the same tokenization strategy and granularity. At this time, the text in the bounding box is in the form of a list of words. Randomly select words in the list according to a preset proportional threshold, query the replacement words in the multi - language dictionary, and give priority to the in - domain dictionary during the query process. If the query fails in the in - domain dictionary, then query in the general dictionary. If the query still fails in the general dictionary, then continue to randomly select from the word list until the proportional threshold requirement is met or all the words that are successfully queried in the word list are taken out; after querying in the word list, multiple multi - language synonyms will be obtained for the replacement words, and randomly select and replace the original words among these synonyms; after completing the same - semantic multi - language replacement of the text, the layout information is changed by means of mask perturbation. Specifically: for all bounding boxes globally, 20% of the bounding boxes are selected for change, and 80% remain unchanged. For the layout information of the 20% bounding boxes, the following change methods are randomly adopted: set the coordinates to zero, that is, consider that the bounding box is at the origin position in the two - dimensional plane; replace the two - dimensional coordinates with random values; perturb the border, including magnification, reduction, and rotation.
9. The multi - language alignment system for rich text document parsing according to claim 6, characterized in that, For the document data after data augmentation, a general document encoder is used to encode the bounding boxes. For each bounding box, the encoding result is a 1×768-dimensional vector representation, denoted as h i , and the expression is: h i = Normalize(LayoutXLM(x i )) ∈ ℝ 768 Among them, Normalize is the regularization operation; LayoutXLM is the document encoder; x i is the bounding box input; R is the set of real numbers.
10. The multi - language alignment system for rich text document parsing according to claim 9, characterized in that, positive and negative examples are constructed through contrastive learning, a contrastive learning loss function is designed, the representations of positive example pairs are pulled closer, and the representations of negative example pairs are pulled farther apart, so as to achieve cross - language representation alignment. The expression is: where B is the training batch size, i.e., the number of samples in the training batch; exp is the exponential function with the natural constant e as the base; h p is the positive example representation; h j is the negative example representation; τ is the temperature coefficient used to adjust the distribution shape of the representation. The smaller the temperature coefficient, the more the loss function will focus on negative sample pairs.
Citation Information
Patent Citations
Cross-language information retrieval training method based on aligned query entity pairs
CN116757188A