Model Pretraining Method and Device for Multilingual Tasks

Through multi-stage training methods and target lexicon processing, the adaptability and performance of multi-modal models in multi-language tasks are improved, the problem of poor adaptability in multi-language and multi-modal tasks is solved, and the transfer of multi-modal understanding ability across languages and across fields is realized.

CN119293514BActive Publication Date: 2025-08-01LIANLIAN YINTONG ELECTRONIC PAYMENT CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411831139.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-08-01
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

The existing multimodal models are poorly adaptable in multilingual and multimodal tasks, especially in the field of visual text comprehension that includes multiple languages. They cannot effectively adapt to task processing in the field, and data support is mainly concentrated in English, resulting in cross-language and cross-domain adaptation difficulties.

Method used

The multi-stage training method is adopted, first the comparative learning training of the alignment of visual features and text features is carried out, the parameters of the decoding module are frozen, and then the constraint training of content understanding is carried out, and word segmentation is processed in combination with the target lexicon in the preset business field to construct a model pre-training method and device for multilingual tasks.

Benefits of technology

It improves the multi-language task adaptability of multi-modal models in specific business fields, reduces data demand, shortens training time, improves model performance and universality, and realizes multi-modal capability migration from high-resource languages to low-resource languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119293514B_ABST
    Figure CN119293514B_ABST
Patent Text Reader

Abstract

The present application provides a model pre-training method and apparatus for multilingual tasks, relating to the field of artificial intelligence technology. The method includes: obtaining a multimodal training data set, where the training data set includes multiple sample text data with multilingual content and multiple sample text-image pair data, covering general fields and preset business fields; based on the multiple sample text-image pair data, performing contrastive learning training on the initial model for visual feature and text feature alignment, freezing the model parameters of the decoding module and adjusting the model parameters of the visual encoder and the projection module during the training process until the first end condition is met; based on the multiple sample text-image pair data and the multiple sample text data, performing constraint training on the initial model that meets the first end condition for content understanding, and adjusting the model parameters of the visual encoder, the projection module, and the decoding module during the training process until the second end condition is met to obtain the target model; the present application can significantly improve the information extraction ability of the model in specific fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model pre-training method and device for multi-language tasks. Background Art

[0002] The advancement of the globalization and digitalization waves has made multi-language and multi-modal tasks increasingly prominent. Although pre-trained models have achieved remarkable achievements in the field of natural language processing, there are still many problems to be solved in multi-language and multi-modal tasks. Currently, in the field of multi-modal pre-training, data support mainly focuses on English, resulting in poor adaptability of existing multi-modal models in the field of visual text understanding that includes multiple languages (such as Chinese, German, French, Japanese, etc.). This challenge is more prominent in specific business fields of different scenarios, involving not only multi-language issues but also dealing with the differences between document images and natural images, making multi-modal models unable to adapt to task processing within the field. Summary of the Invention

[0003] This application provides a model pre-training method and device for multi-language tasks, which can significantly improve the model pre-training effect for multi-language tasks and enhance the task processing effect of the model in a preset business field.

[0004] On the one hand, this application provides a model pre-training method for multi-language tasks, and the method includes:

[0005] Obtain a multi-modal training data set and an initial model. The training data set includes multiple sample text data and multiple sample text-image pair data. The multiple sample text-image pair data and the multiple sample text data include multiple language contents, and the multiple sample text-image pair data include sample text-image pair data in the general field and sample text-image pair data in a preset business field in the target scenario. The multiple sample text data include the text data of the preset business field. The initial model includes a visual encoder, a projection module, and a decoding module connected in sequence, and the decoding module is constructed based on a large language model;

[0006] Based on the multiple sample text-image pair data, perform contrastive learning training on the initial model for visual feature and text feature alignment. During the training process, freeze the model parameters of the decoding module and adjust the model parameters of the visual encoder and the projection module until the first end condition is met;

[0007] Based on the multiple sample text-image pair data and the multiple sample text data, perform constraint training on the initial model that meets the first end condition for content understanding. During the training process, adjust the model parameters of the visual encoder, the projection module, and the decoding module until the second end condition is met;

[0008] Determine the initial model that meets the second end condition as the target model; during the training process, perform word segmentation on the text in the sample text-image pair data and the sample text data in the preset business domain in combination with the target word library corresponding to the preset business domain, so as to be used as the input of the projection module.

[0009] On the other hand, a model pre-training device for multilingual tasks is provided, and the device includes:

[0010] An acquisition module: used to acquire a multimodal training data set and an initial model, the training data set includes a plurality of sample text data and a plurality of sample text-image pair data, the plurality of sample text-image pair data and the plurality of sample text data include multiple language contents, and the plurality of sample text-image pair data include sample text-image pair data in the general domain and sample text-image pair data in the preset business domain in the target scenario, the plurality of sample text data include the text data in the preset business domain, the initial model includes a visual encoder, a projection module and a decoding module connected in sequence, and the decoding module is constructed based on a large language model;

[0011] A first training module: used to perform contrastive learning training on the initial model for visual feature and text feature alignment based on the plurality of sample text-image pair data, freeze the model parameters of the decoding module and adjust the model parameters of the visual encoder and the projection module during the training process until the first end condition is met;

[0012] A second training module: used to perform constraint training on the content understanding of the initial model that meets the first end condition based on the plurality of sample text-image pair data and the plurality of sample text data, and adjust the model parameters of the visual encoder, the projection module and the decoding module during the training process until the second end condition is met;

[0013] A model generation module: used to determine the initial model that meets the second end condition as the target model; during the training process, perform word segmentation on the text in the sample text-image pair data and the sample text data in the preset business domain in combination with the target word library corresponding to the preset business domain, so as to be used as the input of the projection module.

[0014] On the other hand, a computer device is provided, the device includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the model pre-training method for multilingual tasks as described above.

[0015] On the other hand, a computer-readable storage medium is provided, in which at least one instruction or at least one program segment is stored, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the model pre-training method for multilingual tasks as described above.

[0016] On the other hand, a server is provided, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the model pre-training method for multilingual tasks as described above.

[0017] On the other hand, a terminal is provided, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the model pre-training method for multilingual tasks as described above.

[0018] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and when the computer instructions are executed by a processor, the model pre-training method for multilingual tasks as described above is implemented.

[0019] The model pre-training method, device, equipment, storage medium, server, terminal, computer program and computer program product for multilingual tasks provided by this application have the following technical effects:

[0020] The multi-modal training data set adopted in this application includes multiple sample text data and multiple sample text-image pair data. The multiple sample text-image pair data and the multiple sample text data include multiple language contents, and the multiple sample text-image pair data include sample text-image pair data in the general field and sample text-image pair data in the preset business field in the target scenario. The multiple sample text data include text data in the preset business field, so as to provide multi-language graphic and text data and pure text data covering the preset business field, improve the adaptability of the multi-modal model to multilingual tasks in a specific business field, and have the understanding ability of both multi-modal data and text data. And during the pre-training process, a two-stage training method is adopted to perform visual feature and text feature alignment training and content understanding constraint training respectively, so as to realize the transfer of multi-modal capabilities from high-resource languages to low-resource languages through multi-language and multi-modal pre-training, reduce data requirements, shorten the training time, and improve the model performance and generality at the same time. Description of the Drawings

[0021] To more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0022] Figure 1 It is a schematic diagram of an application environment provided by an embodiment of the present application;

[0023] Figure 2 It is a schematic flowchart of a model pre-training method for multi-language tasks provided by an embodiment of the present application;

[0024] Figure 3 It is a schematic framework diagram of a model pre-training device for multi-language tasks provided by an embodiment of the present application;

[0025] Figure 4 It is a hardware structure block diagram of an electronic device that executes a model pre-training method for multi-language tasks provided by an embodiment of the present application. Detailed implementation manners

[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or sub-modules does not necessarily have to be limited to those clearly listed steps or sub-modules, but may include other steps or sub-modules that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0028] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0029] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.

[0030] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0031] Computer Vision Technology (CV) Computer vision is a science that studies how to make machines "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as target recognition, tracking, and measurement in machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0032] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; at the same time, it involves computer science and mathematics. An important technology for model training in the field of artificial intelligence, the pre-trained model, is developed from the large language model in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs and other technologies.

[0033] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. The solution provided in the embodiments of this application involves technologies such as machine learning / deep learning, computer vision technology, and natural language processing of artificial intelligence, which will be specifically described through the following embodiments.

[0034] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an application environment provided by the embodiments of this application. As Figure 1 shown, the application environment may include a terminal 01 and a server 02. In actual applications, the terminal 01 and the server 02 can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this here.

[0035] The server 02 in the embodiments of this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0036] Specifically, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing. Cloud technology can be applied in various fields, such as medical cloud, cloud Internet of Things, cloud security, cloud education, cloud conferencing, artificial intelligence cloud services, cloud applications, cloud calling, and cloud social networking. Cloud technology is based on the cloud computing business model. It distributes computing tasks across a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to users to be infinitely expandable and can be obtained at any time, used on demand, expanded at any time, and paid according to usage. As a basic capability provider of cloud computing, a cloud computing resource pool (abbreviated as a cloud platform, generally referred to as IaaS (Infrastructure as a Service)) platform will be established, and various types of virtual resources will be deployed in the resource pool for external customers to select and use. The cloud computing resource pool mainly includes: computing devices (virtual machines, including operating systems), storage devices, and network devices.

[0037] According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. The SaaS layer can also be directly deployed on the IaaS. PaaS is the platform for software operation, such as databases, web containers, etc. SaaS is various business software, such as web portals, mass text message senders, etc. Generally speaking, SaaS and PaaS are the upper layers relative to IaaS.

[0038] Specifically, the above-mentioned server 02 can include physical devices, which can specifically include network communication sub-modules, processors, memories, etc., or can include software running on physical devices, which can specifically include application programs, etc.

[0039] Specifically, the terminal 01 can include physical devices such as smartphones, desktop computers, tablets, laptops, digital assistants, augmented reality (AR) / virtual reality (VR) devices, intelligent voice interaction devices, smart home appliances, smart wearable devices, in-vehicle terminal devices, etc., or can include software running on physical devices, such as application programs, etc.

[0040] In an embodiment of the present application, the terminal 01 can be used to obtain a multi-modal training data set and an initial model, and send them to the server 02, so that the server 02 performs contrastive learning training for aligning visual features and text features on the initial model based on the multiple sample image-text pair data, and performs constraint training for content understanding on the initial model that meets the first end condition based on the multiple sample image-text pair data and the multiple sample text data, to obtain a target model.

[0041] Specifically, the model pre-training method for multi-language tasks of the present application can be applied to image content understanding in a preset business field in a target scenario to perform tasks such as image classification and image content description, such as payment content review services, transaction contract reviews, or services such as visual question answering (VQA), image captioning, object detection, localization, and image classification.

[0042] In addition, it can be understood that Figure 1 The shown is only an application environment of a model pre-training method for multi-language tasks, and this application environment can include more or fewer nodes, and the present application does not make any restrictions here.

[0043] The application environment involved in the embodiments of the present application, or the terminal 01 and the server 02 in the application environment, etc. can be a distributed system formed by connecting a client and multiple nodes (any form of computing device accessing the network, such as a server, a user terminal) in a network communication manner. The distributed system can be a blockchain system, and this blockchain system can provide the above-mentioned model pre-training services, neural network training services, data storage services, etc. for multi-language tasks.

[0044] The technical solutions of the present application are introduced below based on the above application environment. The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc. Please refer to Figure 2 , Figure 2 is a flowchart of a model pre-training method for multi-language tasks provided by an embodiment of the present application. The present specification provides method operation steps such as in the embodiment or the flowchart, but based on routine or non-creative labor, it can include more or fewer operation steps. The step order listed in the embodiment is only one way among the execution orders of numerous steps, and does not represent the only execution order. When the actual system or server product executes, it can be executed in the order of the embodiment or as shown in the drawings, or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). Specifically, as Figure 2 shown, the method can include the following steps S201-S207:

[0045] S201: Obtain a multi-modal training data set and an initial model.

[0046] Specifically, the training dataset includes multiple sample text data and multiple sample image-text pair data. The multiple sample image-text pair data and the multiple sample text data include content in multiple languages. Sample image-text pair data refers to sample data that includes images and text, and sample text data refers to pure text data. There can be sample image-text pair data and sample text data in multiple languages in the training dataset. A single sample image-text pair data can also include image content or text in multiple languages, and a single sample text data can also include text content in multiple languages.

[0047] Specifically, the multiple sample image-text pair data includes sample image-text pair data in the general domain and sample image-text pair data in the preset business domain in the target scenario. The multiple sample text data includes text data in the preset business domain. The target scenario can be a specific business scenario, such as a cross-border business scenario, and the preset business domain can be a specific business domain under the target scenario, such as a multilingual business domain, etc. Correspondingly, a multilingual image-text multimodal training dataset including the cross-border domain is constructed through the training set, including multilingual sample image-text pair data and multilingual sample text data. Exemplarily, the sample image-text pair data in the general domain can be the image-text pairs of an existing public training set or the image-text pairs obtained from the network, etc., without restrictions on the source and content here. The sample image-text pair data in the preset domain can include a multilingual cross-border image-text dataset, specifically including multilingual image-text data in fields such as multilingual payment, e-commerce platforms, multilingual trade, and logistics, covering multilingual product images and texts, trade material contract documents, product images, help manuals, logistics information attachments, screenshots and photographed information of the store review background, product labels, currencies of different countries, etc.

[0048] Specifically, when constructing a multilingual image-text pre-training dataset for the preset business domain, it includes sample image-text pair data and sample text data in the preset business domain, and the two are mixed in a certain training ratio. This training ratio can be exemplified as 3:7, etc. The training ratio refers to the data ratio input in each iteration during the training process; at the same time, multiple open-source image-text data are obtained and sampled according to the data source to obtain an open-source image-text dataset, which can include only image-text pair data or can also include image-text pair data and pure text data at the same time, and are mixed with the pre-training dataset in the preset business domain in a certain data ratio according to the task requirements to obtain a multimodal training dataset.

[0049] Specifically, image enhancement processing can also be performed on the images in the sample image-text pairs, such as using StyleGAN3 for image enhancement.

[0050] In possible implementation manners, the types of sample graphic-text pair data in a preset business field are determined based on the amount of text in the image, including a first type, a second type, a third type, and a fourth type with an increasing amount of text. The image corresponding to the first type is a textless image, and the image corresponding to the fourth type is a dense document image, such as a text document image. Exemplarily, the textless image can be a multilingual product picture on a multilingual e-commerce platform, etc., and the corresponding product description text with the attached product picture. The image of the second type can be a sparse text image, such as a screenshot of a store audit background and photographed information. The image of the third type can be a multi-text image, such as a multilingual trade material contract document, a logistics information attachment, etc. The image of the fourth type can be a dense long-text image, such as a multilingual help manual in the payment field, etc., which can be a multi-page PDF document. The original data types of the image sources in the sample graphic-text pair data can be PDF, picture, Word, PowerPoint, Excel, TXT, etc., and the corresponding image data is obtained through image conversion.

[0051] It can be understood that after converting different types of data into initial images in a preset field through image conversion, preprocessing is performed on the initial images, including but not limited to image data filtering and deduplication, etc., to eliminate low-quality, non-standard, and unreasonable images, such as file format zip, rar, 7z compressed files, etc.; it can also include image cleaning and enhancement, such as denoising the image, enhancing the contrast and clarity, to improve the recognizability of digital areas. The cleaning and enhancement methods include but not limited to adjusting color difference, randomly rotating, randomly horizontally or vertically flipping, scaling, cropping, translating, etc. For example, for a low-quality photographed photo scene in a logistics scenario, deblurring and enhancement are performed to correct the tilt and distortion in the image and ensure that the numbers are presented at the correct angle.

[0052] Correspondingly, after obtaining the images in the preset business field, image data annotation and text generation are performed on them. The acquisition methods of the sample graphic-text pair data in the preset business field include S301 - S307:

[0053] S301: For the images of the first type, obtain the description text in multiple language versions of the image, and combine the description text in each language version with the image respectively to obtain the sample graphic-text pair data.

[0054] Specifically, for the textless images of the first type, the description text attached to the image can be directly extracted from the database to generate a graphic-text pair. For example, for the multilingual product pictures on a multilingual e-commerce platform, obtain their description text to generate a product picture - product description pair, and at the same time translate the description text into multiple language texts to obtain multilingual graphic-text pairs. It can be understood that the description text can be used as the first sample annotation of the images of the first type.

[0055] S303: For images of the second type, perform content understanding and content extraction on the images based on a multi-modal model to obtain indication information and a first sample annotation corresponding to the indication information; combine the image, the indication information, and the first sample annotation to obtain sample image-text pair data.

[0056] Specifically, for sparse text images of the second type, a multi-modal model, such as a multi-modal large model, is used to set indication information (prompt) and corresponding first sample annotations in the corresponding language based on the language of the picture. The indication information can be, for example, a content question, and the first sample annotation can be the corresponding answer. For example, "question: What is the store name in this picture?", "answer: Star"; or, "question: Quel est le nom du magasin sur cette photo?", "answer: étoile".

[0057] S305: For images of the third type, perform layout area analysis on the images to obtain graphic elements and document elements in the images; obtain indication information and a first sample annotation corresponding to at least one of the graphic elements and the document elements; combine the image, the indication information, and the first sample annotation to obtain sample image-text pair data.

[0058] Specifically, for multi-text images of the third type, such as multi-language trade material contract documents, logistics information attachments, etc., the picture formats are diverse, and may be electronic versions, scanned versions, etc., and may contain irregular tables, pictures, text and other elements. By performing layout area analysis on the images, such as layout analysis, the graphic elements and document elements are marked with annotation boxes, and the coordinates and categories of the annotation boxes are output. The document elements such as text and tables, and elements such as pictures in the images can also be parsed and saved in a lightweight markup language format as the first sample annotation. In this way, the data in the preset business field is classified based on the text volume, and corresponding types of images, indication information, and first sample annotations are generated, so that the model can not only learn multi-language domain knowledge but also acquire the recognition and understanding abilities of different text images in combination with the characteristics of the business field, and further better master the multi-modal data understanding ability of the preset business field for subsequent model fine-tuning of specific tasks.

[0059] Specifically, the indication information of the figure element and the document element can be used to instruct the model to perform position marking, text content recognition, or text content understanding of the corresponding elements. Correspondingly, the types of the indication information of the third type of image include content question indication for the document element, position detection indication for the picture element or the document element, and content question indication and position detection indication for at least one of the picture element and the document element. The content question indication includes document content recognition or content understanding Q&A, etc. The position detection indication is used to instruct the model to recognize the position of the corresponding element in the image. The first annotation information is the reply information corresponding to the indication information, such as the text content in the document corresponding to the content question indication or the answer to the content understanding question, or the annotation box coordinates of the document element or the figure element corresponding to the position detection indication. In this way, fine-grained and multi-type indication information and annotation information are generated for the third type of multi-text image, further improving the model's understanding ability of multi-text images in the field to achieve accurate understanding of multi-modal fine-grained content.

[0060] Exemplarily, the indication information and the first sample annotation of the third type of image can only include text content recognition in the content question indication, such as annotating text recognition data, such as question: " \nPlease give the content of the text in the figure", answer: "***********"; or, the indication information and the first sample annotation can only include the position detection indication, such as indicating annotation detection data, such as question: " struct with bbox", answer: " <h2>< / h2> <box> 190,841,315,859< / box> <lb> <lb> <h2>< / h2> <box> 199,154,478,174< / box> <lb> <lb> <box> 148,719,853,823< / box> <lb> <lb> <box> 149,89,850,136< / box> <lb> <lb> <h2>< / h2> <box> 198,525,498,545< / box> <lb> <lb> <h2>< / h2> <box> 198,682,437,701< / box> <lb> <lb> <footer>< / footer> <box> 491,928,509,939< / box> <lb> <lb> <footer>< / footer> <box> 900,957,997,989< / box> <lb> <lb> <box> 190,878,514,895< / box> <lb> <lb>"; alternatively, the indication information and the first sample annotation may simultaneously include a content question indication and a location detection indication, such as annotation detection + text recognition data, e.g., question: " \n I want to know the content of the text in the figure and their respective positions", answer: " <ref>Test <box> 190,841,315,859< / box> < / ref> <ref>Result <box> 199,154,478,174< / box> < / ref> "; alternatively, the indication information and the first sample annotation may include content understanding Q&A, such as multi-turn Q&A data, e.g., question: ' \n what is the "seller"? Answer the question using a single word or phrase.'}, answer: "*****", question: "What is the account number?", answer: "*****".

[0061] S307: For the fourth type of image, input the target document corresponding to the image into a large language model for content understanding and content extraction to obtain the indication information corresponding to the image and the first sample annotation corresponding to the indication information; combine the image, the indication information, and the first sample annotation to obtain sample image-text pair data; the target document is obtained by splitting the original text document, and the fourth type of image is obtained by converting the target document into a picture.

[0062] Specifically, for the fourth type of dense document image, such as the image of a multi-page PDF document like a help manual in the cross-border payment field, parse and split the original long document into a size that the model can handle to obtain multiple target documents, and then convert them into picture format to construct sample image-text pair data of the multi-image-text type. At the same time, the target document, the indication information corresponding to the document content, and the first sample annotation can be combined as sample text data. Among them, a large language model (LLM) can be used to perform content understanding and content extraction on the split target document to set indication information for the fourth type of image and obtain an image annotation as the first sample annotation, such as "question: Which sites does XX e-commerce platform support for receiving payments?", "answer: "*****""; "question: How much is the cross-border payment collection handling fee?", "answer: "*****"" etc.

[0063] Specifically, during the process of obtaining sample text-image pair data in the general domain or a preset business domain, image enhancement methods can also be used to synthesize diverse data. For example, rendering the text in a question-answer dataset into an image, rendering tabular data into a composite image of a table and a chart, or generating multi-language domain-specific composite images such as multi-language product labels, images of currencies in different countries, etc. Exemplarily, the image enhancement method can be StyleGAN3, etc.

[0064] By constructing a multi-scenario, multi-language, and multi-modal professional dataset that includes the general domain and the preset business domain, the model can be deeply adapted to the multi-language business domain. Moreover, through diverse training data, the generality of the model in the multi-language scenario is enhanced, and its ability to process different types of inputs is improved.

[0065] Specifically, the initial model includes a vision encoder, a projection module, and a decoding module connected in sequence. The decoding module is constructed based on large language models, such as multi-language large models like internlm2-chat and Qwen2, to build a multi-modal model architecture. The vision encoder is used to process images to generate visual features, the projection module is used to map visual features to the text feature space and process text to generate text features, and the decoding module is used to receive the output of the projection module and perform content understanding to obtain the output result corresponding to the indication information. Exemplarily, the projection module can adopt a multi-layer perceptron model (MLP).

[0066] S_{203}: Based on multiple sample text-image pair data, perform contrastive learning training for aligning visual features and text features on the initial model. During the training process, freeze the model parameters of the decoding module and adjust the model parameters of the vision encoder and the projection module until the first end condition is met.

[0067] S_{205}: Based on multiple sample text-image pair data and multiple sample text data, perform constraint training for content understanding on the initial model that meets the first end condition. During the training process, adjust the model parameters of the vision encoder, the projection module, and the decoding module until the second end condition is met.

[0068] Specifically, the pre-training method of this application is divided into two stages: a contrastive learning training stage and a joint training stage for mixed data. The first stage is contrastive learning training based on image-text pairs, which is used to align the visual features (visual embeddings) of the visual encoder and the projection module with the input feature space of the large language model of the decoding module. By freezing the large language model of the decoding module, the visual encoder and the projection module are trained on a large scale of image-text pair data to perform visual-text feature alignment. During this period, a contrastive learning strategy is used to optimize the visual encoder and the projection module so that they learn the mapping relationship between visual and multilingual text representations. The second stage is the joint training of the visual encoder, the projection module, and the decoding module based on sample image-text pair data and sample text data. In the second stage, all modules are unfrozen, and the large language models of the visual encoder, the projection module, and the decoding module are jointly trained using a large scale of sample image-text pair data and pure text sample text data. During the training, a progressive unfreezing strategy is adopted to gradually unfreeze and fine-tune the parameters of each layer in the large language model of the decoding module to improve the training effect and reduce the consumption of training resources.

[0069] Cross-modal pre-training that combines images and text enables the model to better understand cross-language and cross-domain data, improving the transfer effect of the pre-trained model in specific domains. By using image-text pairs and pure text data for multi-modal fusion training, the problem of the decline in the language model's ability during the alignment training of different modalities is greatly reduced, further enhancing the multi-modal understanding ability across domains and languages.

[0070] S207: Determine the initial model that meets the second end condition as the target model.

[0071] Specifically, for the text in the sample image-text pair data and the sample text data, tokenization needs to be performed. During the training process, the text in the sample image-text pair data and the sample text data in the preset business domain are tokenized in combination with the target word library corresponding to the preset business domain as the input of the projection module. Exemplarily, the target word library can include professional terms in multiple language versions in the cross-border field, and the BPE (Byte Pair Encoding) method can be used for text tokenization. Based on the custom target word library as the knowledge base for tokenization, the model can be deeply adapted to the preset business domain, enhancing the model's understanding ability of specific languages in the preset business domain.

[0072] In some embodiments, S203 may include S401 - S415:

[0073] S401: Sample multiple sample image-text pair data to obtain positive sample image-text pairs and negative sample image-text pairs;

[0074] S403: Tokenize and perform feature embedding on the text of the positive sample image-text pairs and the text of the negative sample image-text pairs to obtain the first initial text features;

[0075] S405: Input the images of the positive sample image-text pairs and the images of the negative sample image-text pairs into a visual encoder for feature encoding to obtain the first visual features;

[0076] S407: Input the first visual features and the first initial text features into a projection module for text space mapping of the visual features and text feature extraction to obtain the first mapped features corresponding to the first visual features and the first text features corresponding to the first initial text features;

[0077] S409: Concatenate the first mapped features and the first text features to obtain the first combined feature;

[0078] S411: Input the first combined feature into the decoding module for content understanding of the combined indication information to obtain the first output result;

[0079] S413: Determine the first model loss based on the first output result and the first sample annotation of the positive sample image-text pairs, and the first output result and the first sample annotation of the negative sample image-text pairs;

[0080] S415: Train the initial model based on the first model loss to adjust the model parameters of the visual encoder and the projection module until the first end condition is met.

[0081] Specifically, positive sample image-text pairs refer to image-text pairs where the image and the text match, and negative sample image-text pairs refer to image-text pairs where the image and the text do not match. Some original sample image-text pair data are obtained as positive sample image-text pairs through sampling, and multiple negative sample image-text pairs are obtained by combining images with non-matching texts.

[0082] Specifically, tokenize the multi-language texts in the positive sample image-text pairs and the negative sample image-text pairs to obtain text tokens, and perform feature embedding through word embedding conversion to generate vector representations of the text tokens. Concatenate the vector representations of the text tokens to obtain the first initial text features. During the tokenization process, use the target vocabulary as the tokenization knowledge base to tokenize the text corresponding to the preset business domain.

[0083] Specifically, the images of each positive sample image-text pair and each negative sample image-text pair obtained by sampling are respectively input into the visual encoder, and feature encoding and feature extraction are performed on the images, and the first visual features corresponding to each image are output. Then, the first visual features and the first initial text features of each positive sample text pair are input into the projection module, and the first visual features and the first initial text features of each negative sample text pair are input into the projection module, so that the input first visual features are aligned in the text space through the projection module, mapped into the first mapping features, and the input first initial text features are feature extracted to obtain the first text features mapped to the text space. It can be understood that the text space here refers to the text feature space of the large language model of the decoding module.

[0084] Furthermore, the first mapping feature and the first text feature of the positive sample image-text pair, as well as the first mapping feature and the first text feature of the negative sample image-text pair are spliced together to obtain their respective first joint features, which are used as the input of the decoding module. Combined with the content understanding guidance of the indication information, a first output result is obtained. The first output result is the reply information output by the decoding module for the indication information, and the first sample annotation is the annotation true value corresponding to the indication information.

[0085] Combined with the training method of contrastive learning, when the parameters of the decoding module are frozen, the distance between the first output result of the positive sample image-text pair and the first sample annotation is shortened, that is, the distance between the first mapping feature and the first text feature is shortened accordingly, and the distance between the first output result of the negative sample image-text pair and the first sample annotation is increased, that is, the distance between the first mapping feature and the first text feature is increased, thereby determining the first model loss, which at least includes the contrast loss. Based on the first model loss, the model parameters of the visual encoder and the projection module are adjusted to obtain an updated initial model, and the aforementioned S401-S411 is repeatedly executed with the updated initial model, and so on, until the first end condition is met. Meeting the first end condition here means that the number of training iterations meets the preset number, or the first model loss is less than the preset loss, etc. It can also be other state conditions that meet the end of training, which are not limited here.

[0086] In this way, by constructing professional domain datasets and target lexicons covering multiple scenarios, multiple languages, and multiple modalities such as multilingual payment, e-commerce, and trade, combined with contrastive learning, the visual encoder's feature extraction capabilities for general domain images and preset business domain images are improved, as well as the projection module's ability to align domain visual features and text features, thereby providing high-quality input for the subsequent decoding module's comprehension training, enabling it to acquire excellent domain visual text comprehension capabilities.

[0087] In some embodiments, S205 may include S501-S513:

[0088] S501: Tokenize the text of the sample text-image pair data and the sample text data, and perform feature embedding to obtain the second initial text feature corresponding to the sample text-image pair data and the third initial text feature corresponding to the sample text data;

[0089] S503: Input the image of the sample text-image pair data into a visual encoder for feature encoding to obtain a second visual feature;

[0090] S505: Input the second visual feature and the second initial text feature of the sample text-image pair data, and the third initial text feature into a projection module respectively to perform text space mapping of the visual feature and text feature extraction, obtaining the second mapped feature and the second text feature corresponding to the sample text-image pair data, and the third text feature corresponding to the sample text data;

[0091] S507: Concatenate the second mapped feature and the second text feature corresponding to the sample text-image pair data to obtain a second combined feature;

[0092] S509: Input the second combined feature and the third text feature into a decoding module respectively to perform content understanding of the combined indication information. The indication information is used to provide the guiding information required for the decoding module to perform content understanding. The indication information includes prompt texts in multiple language versions;

[0093] S511: Determine a second model loss based on the second output result and the first sample annotation of the sample text-image pair data, and the third output result and the second sample annotation of the sample text data;

[0094] S513: Train an initial model that meets the first end condition based on the second model loss to adjust the model parameters of the visual encoder, the projection module, and the decoding module until the second end condition is met.

[0095] Specifically, in the tokenization process, the target word library is used as the tokenization knowledge library to tokenize the text corresponding to the preset business domain. The tokenization process in S501 is the same as the tokenization process involved previously and will not be elaborated here.

[0096] In the second stage, using the sample image-text pair data and sample text data in the general domain and the preset business domain in the multi-modal training dataset as training inputs, image feature extraction is performed through a data encoder to obtain corresponding second visual features, and the second visual features are used as inputs to a projection module for text space alignment to obtain corresponding second mapping features. At the same time, the second initial text features corresponding to the sample image-text pair data and the third initial text features corresponding to the sample text data are respectively input into the projection module for feature extraction to obtain the second text features corresponding to the second initial text features and the third text features corresponding to the third initial text features. Then, the second mapping features and the second text features corresponding to the sample image-text pair data are concatenated, and the resulting second joint features are used as inputs to a decoding module for feature-based content understanding, and the third text features are used as inputs to the decoding module for content understanding to respectively output a second output result and a third output result. It can be understood that the input to the decoding module also includes indication information, such as the indication information carried in the aforementioned sample image-text pair data. Correspondingly, the labeled answer of the indication information in the text of the image-text pair is the corresponding first sample label, such as the position coordinates of a graphic element or the answer to a question about the content of a graphic element, etc. The indication information can also be a question about the text content carried in the sample text data, etc. The labeled answer of the indication information carried in the sample text data is the corresponding second sample label. For example, the indication information can be "What is the liquidated damages proposed in the contract document?", and the second sample label is "The liquidated damages are XXX US dollars". It can be understood that the indication information is used to provide the guiding information required for the decoding module to perform content understanding. The indication information includes prompt texts in multiple language versions. Correspondingly, the first sample label or the second sample label also includes the corresponding labeled answers in multiple language versions.

[0097] Specifically, at least based on the differences between the second output result and the corresponding first sample label, and between the third output result and the corresponding second sample label, calculate the second model loss to adjust the model parameters of each module in the initial model that meets the first end condition until the second end condition is met. The second end condition is similar to the limitation of the aforementioned first end condition and will not be elaborated here.

[0098] In this way, in the second training stage, the sample image-text pair data and the sample text data can be alternately input into the initial model that meets the first end condition, allowing the model to effectively integrate and utilize visual information while maintaining language understanding ability, avoiding problems such as the decline in language model ability, knowledge update, and forgetting caused by alignment training, improving the transfer effect and generalization ability of the pre-trained model in a specific domain, and enhancing the performance of the model in multi-modal tasks.

[0099] In a specific embodiment, a large number of image-text interaction pairs and pure text corpora are used in the second training stage to train the visual encoder, projection module, and decoding module simultaneously. In this stage, the visual encoder processes the input image and generates visual features, which are aligned with text tokens through the projection module. The projection module maps text and visual embeddings into the same semantic space based on the self-attention mechanism, and then the processed mapping features are added to the front of the text features to form a joint representation. This joint representation is passed to the large language model in the decoding module for processing. For each layer of the large language model (including the decoder layer), the query (Q), key (K), value (V) matrices in its attention mechanism and the feed-forward neural network (FFN) are trained separately to further ensure the model's multi-language understanding ability, visual information understanding ability, and the integration ability of the two.

[0100] Based on some or all of the above embodiments, in some embodiments, the visual encoder includes a feature extraction module, a feature fusion module, and a feature extraction module. Among them, the feature extraction module is used to extract features from the sub-images corresponding to the input image to obtain feature maps. The feature fusion module takes the respective feature maps corresponding to the image as inputs for feature fusion to obtain fused features, which are then used as the inputs to the feature extraction module to obtain corresponding visual features. Correspondingly, the feature encoding process of the visual encoder for the input image includes S601-S605:

[0101] S601: Input the image into the feature extraction module, and perform local feature extraction on the image based on the self-attention mechanism of the sliding window to obtain multi-scale feature maps;

[0102] S603: Input the multi-scale feature maps into the feature fusion module for feature fusion to obtain fused features;

[0103] S605: Input the fused features into the feature extraction module for feature extraction to obtain the first visual feature or the second visual feature corresponding to the image.

[0104] Specifically, in the first training stage and the second training stage, the feature extraction module performs regional feature extraction on the image based on the sliding window, and based on the self-attention mechanism, uses the features of other regions corresponding to the current region as context for feature cross to obtain multi-scale feature maps. The multi-scale feature maps are fused by the feature fusion module to obtain fused features, which are input into the feature extraction module for further feature extraction to obtain the first visual feature in the first training stage or the second visual feature in the second training stage.

[0105] In this way, the granularity of image feature extraction is improved through multi-feature map extraction and the self-attention mechanism to obtain fine-grained visual features, and the full expression of visual information is realized through the method of fusion and then extraction, thereby improving the comprehensiveness and accuracy of the information input to the decoding module.

[0106] In some embodiments, S601 may include S6011 - S6012:

[0107] S6011: Perform adaptive image segmentation on the image to obtain multiple sub - images corresponding to the image;

[0108] S6012: Perform local feature extraction on each sub - image based on the self - attention mechanism of the sliding window to obtain multi - scale feature maps corresponding to the sub - images.

[0109] Through the adaptive segmentation of the image and the extraction of multi - scale feature maps, it is possible to avoid the problem of overly large patch features caused by too large image input, adapt to image inputs of different sizes and contents, and at the same time improve the accuracy and comprehensiveness of image feature extraction.

[0110] Specifically, an adaptive image segmentation algorithm is used to perform adaptive segmentation on the image. The patch size is dynamically adjusted according to the aspect ratio and resolution of each image, and it is cropped into a certain number of sub - images. The resolution of each sub - image can be fixed at a preset resolution, such as fixed at 448×448 pixels. The number of sub - images for image segmentation can be exemplified as 1 to 12.

[0111] Specifically, the feature extraction module serves as the backbone network of the video encoder. Based on the sliding window, the input sub - image is segmented into multiple visual regions of a preset size, so as to divide the image into multiple patches of a preset size according to the image resolution as image tokens (e.g., segmented into 1024 small patches of 14*14). Then, a convolutional connection structure such as a convolutional kernel is used to process each small patch to obtain image tokens. For example, a convolutional kernel with a size and stride of 1x4 processes 4 adjacent patches horizontally each time and merges them into a new feature, so that the number of image tokens for each sub - image is 256, reducing the length of the sequence to adapt to the subsequent model processing requirements. During the sub - image processing of the feature extraction module, multi - scale feature maps are output for each sub - image. Each attention layer in the feature extraction module obtains multi - size information, and each attention layer can perform selective merging of image patches based on different receptive fields to achieve the merging and information extraction of relevant image regions, and retain some image patches to obtain fine - grained features and better process targets of different scales. Exemplarily, the feature extraction module of the visual encoder can be constructed based on SwinTransformer, and at the same time, a Shunted Self - Attention module is adopted in each attention layer inside the Swin Transformer to focus on different parts of the feature map.

[0112] In some embodiments, the feature fusion module takes the multi-scale feature maps output by the feature extraction module as the input of the feature fusion module, constructs a feature pyramid through upsampling and lateral connections; performs object detection or semantic segmentation on different layers of the feature pyramid, uses multi-scale features to improve the performance of the model, and takes the output fused features as the input of the feature extraction module. Exemplarily, the feature fusion module can be, for example, FPN (Feature Pyramid Networks).

[0113] In some embodiments, the sample text-image pair data in each preset business domain includes multiple types of text-image pairs. The feature extraction module includes a gating network and multiple expert networks matching multiple types. S605 may include: estimating the type of the image based on the gating network, and inputting the fused features into the expert network matching the estimated type for feature extraction to obtain the first visual feature or the second visual feature.

[0114] Specifically, the gating network is used to determine the image type of the text-image pair corresponding to the input fused features and route it to the expert network corresponding to the type. The expert network is used to process the fused features to obtain the features for inputting into the decoding module. The text-image pair type here refers to the type determined based on the text volume as described above, such as the first type, the second type, the third type, and the fourth type, etc. It can be understood that the network structures of each expert network can be the same or specifically set. Each expert network processes images of different text volume types to expand the scale of the visual encoder, improve the training efficiency, reduce the computational resource occupancy, and enhance the processing and understanding capabilities of different images. Exemplarily, the feature extraction module can be constructed based on MoE (Mixed Expert Models).

[0115] Through the data processing solution of the above video encoder, it is possible to solve the problems of poor representation effect and bad model performance for text information-intensive images in the prior art through the setting of expert networks, achieve efficient and flexible visual processing, improve the model's ability to process complex visual information, and be able to perform fine-grained feature extraction on images.

[0116] Based on the above partial or all embodiments, in some embodiments, the method for obtaining the first model loss includes S701 - S709:

[0117] S701: Generate a contrastive loss based on the difference between the first output result of the positive sample text-image pair and the first sample annotation, and the difference between the first output result of the negative sample text-image pair and the first sample annotation.

[0118] S703: Calculate the auxiliary loss of the feature extraction module based on the selection probabilities of each expert network for the images input within a training batch and the number of expert networks, to obtain the load balancing loss;

[0119] S705: Calculate the auxiliary loss of the feature extraction module based on the probability index data obtained by the gating network for estimating the type of the fused features, to obtain the router z loss. The probability index data includes the probabilities that the images corresponding to the fused features belong to each expert network respectively;

[0120] S707: For the sample graphic-text pair data of the multi-language text type, calculate the cross-language consistency loss based on the output results in the multi-language version in the first output result, to obtain the cross-language consistency loss;

[0121] S709: Fuse the contrast loss, load balancing loss, router z loss, and cross-language consistency loss to obtain the first model loss.

[0122] Specifically, the contrast loss can be calculated based on the existing contrast loss calculation method. Using the text or position markers in the first output result and the differences between the text ground truth and position ground truth of the first sample annotation as the input of the contrast loss function, with the goal of bringing closer the first output result of the positive sample graphic-text pair and the first sample annotation, and pulling farther apart the first output result and the first sample annotation in the negative sample graphic-text pair, to obtain the contrast loss.

[0123] Specifically, the load balancing loss and the router z loss are the auxiliary losses for the gating network and expert networks in the feature extraction module, used to maintain load balance among each expert network in the visual encoder and projection module. Different types of images tend to be processed by different expert networks, and there may be a data imbalance problem. By statistically calculating the selection probabilities of the images input within a training batch for each expert network, that is, calculating the auxiliary loss by combining the probabilities of the images processed by each expert network, the load balancing loss is obtained. The input for calculating the router z loss includes the probability index data obtained by the gating network for estimating the type of the fused features. After the fused features are input into the gating network, the gating network estimates which expert network the image token corresponding to the fused features belongs to, to obtain the probabilities that each image token belongs to each expert network, and then generate the probability index data. The router z loss encourages the model to generate more appropriate probability values, thereby preventing the model from generating overly extreme outputs. Exemplarily, the router z loss function and load balancing loss function of the MoE model can be adopted.

[0124] Specifically, the cross - language consistency loss is a loss function introduced to ensure that the model performs consistently across different languages. The outputs of multiple language versions in the first output result of the same image - text pair data are used as the input of the loss function, and the loss calculation is carried out with the goal of minimizing the semantic differences between the output results of different languages of the same image - text pair data, so as to improve the cross - language generalization ability of the model.

[0125] By fusing the contrastive loss, the load - balancing loss, the router z loss, and the cross - language consistency loss, the alignment ability of each module for visual features and text features can be comprehensively optimized, thereby improving the training effect. Specifically, the second model loss L1 = Σ(w i * L i ), where w i is the adaptive weight, and L i represents the loss to be fused.

[0126] Based on some or all of the above - mentioned embodiments, in some embodiments, the method for obtaining the second model loss includes S801 - S809:

[0127] S801: Generate a cross - entropy loss based on the differences between the second output result and the first sample annotation, and between the third output result and the second sample annotation;

[0128] S803: Based on the selection probabilities of each expert network for the input images in the training batch and the number of expert networks, calculate the auxiliary loss of the feature extraction module to obtain the load - balancing loss;

[0129] S805: Based on the probability index data obtained by the gating network for estimating the type of the fused features, calculate the auxiliary loss of the feature extraction module to obtain the router z loss. The probability index data includes the probabilities that the images corresponding to the fused features belong to each expert network;

[0130] S807: For the sample image - text pair data or text data of the multi - language text type, based on the output results of multiple language versions in the second output result and the output results of multiple language versions in the third output result, calculate the cross - language consistency loss to obtain the cross - language consistency loss;

[0131] S809: Fuse the cross - entropy loss, the load - balancing loss, the router z loss, and the cross - language consistency loss to obtain the second model loss.

[0132] Specifically, the cross - entropy loss can be based on the existing cross - entropy loss calculation method. Taking the text or position markers in the second output result, the differences between the text ground truth and position ground truth of the first sample annotation, and the differences between the text in the third output result and the text ground truth of the second sample annotation as the input of the cross - entropy loss function, and then obtaining the cross - entropy loss.

[0133] Specifically, the load balancing loss and the router z loss here are similar to the aforementioned calculation methods and will not be elaborated here. The calculation input of the cross - language consistency loss includes the output results of multiple language versions in the second output result and the output results of multiple language versions in the third output result. It is a loss function introduced to ensure the consistent performance of the model in different languages. By minimizing the semantic differences between the output results in different languages, the cross - language generalization ability of the model is improved.

[0134] By fusing the cross - entropy loss, the load balancing loss, the router z loss, and the cross - language consistency loss, it is possible to comprehensively optimize the semantic expression and understanding of each module for different languages and different types of images, thereby improving the training effect. Specifically, the second model loss L2 = Σ(w j * L j ), where w j is the adaptive weight, and L j represents the loss to be fused.

[0135] Specifically, during the training process, set the model hyperparameters to train with the goal of minimizing the model loss until the model converges, and evaluate the performance of the model that meets the second end condition based on the test set in the preset business domain (such as the cross - border business domain). Comprehensive evaluation metrics (such as image - text matching, text generation, image understanding, etc.) can be designed to test the model's performance. According to the evaluation results, iterate and optimize the model architecture and training strategy. Also, the process data of the pre - training can be detected, including recording and visualizing training metrics, detecting GPU utilization and memory usage, saving model checkpoints, etc.

[0136] In addition, distributed training can be used to implement the above - mentioned model pre - training process. For example, use DeepSpeed ZeRO - 3 for distributed training, apply gradient accumulation and model parallelism strategies, simulate large - batch training, and achieve pipeline parallelism to improve GPU utilization.

[0137] The embodiment of this application also provides a model pre - training device 800 for multi - language tasks, as Figure 3 shown, Figure 3 shows the structural schematic diagram of a model pre - training device for multi - language tasks provided by the embodiment of this application. The device may include the following modules.

[0138] Acquisition Module 10: It is used to acquire a multi-modal training dataset and an initial model. The training dataset includes multiple sample text data and multiple sample text-image pair data. The multiple sample text-image pair data and multiple sample text data include multiple language contents, and the multiple sample text-image pair data include sample text-image pair data in the general domain and sample text-image pair data in the preset business domain in the target scenario. The multiple sample text data include text data in the preset business domain. The initial model includes a visual encoder, a projection module, and a decoding module connected in sequence. The decoding module is constructed based on a large language model;

[0139] First Training Module 20: It is used to perform contrastive learning training on the initial model for visual feature and text feature alignment based on multiple sample text-image pair data. During the training process, freeze the model parameters of the decoding module and adjust the model parameters of the visual encoder and the projection module until the first end condition is met;

[0140] Second Training Module 30: It is used to perform constraint training on the content understanding of the initial model that meets the first end condition based on multiple sample text-image pair data and multiple sample text data. During the training process, adjust the model parameters of the visual encoder, the projection module, and the decoding module until the second end condition is met;

[0141] Model Generation Module 40: It is used to determine the initial model that meets the second end condition as the target model; during the training process, perform word segmentation on the text in the sample text-image pair data and the sample text data in the preset business domain in combination with the target word library corresponding to the preset business domain as the input of the projection module.

[0142] In some embodiments, the First Training Module 20 is specifically used for:

[0143] Sample the multiple sample text-image pair data to obtain positive sample text-image pairs and negative sample text-image pairs;

[0144] Perform word segmentation and feature embedding on the text of the positive sample text-image pairs and the text of the negative sample text-image pairs to obtain first initial text features. During the word segmentation process, use the target word library as the word segmentation knowledge base to perform word segmentation on the text corresponding to the preset business domain;

[0145] Input the images of the positive sample text-image pairs and the images of the negative sample text-image pairs into the visual encoder for feature encoding to obtain first visual features;

[0146] Input the first visual features and the first initial text features into the projection module for text space mapping of visual features and text feature extraction to obtain first mapping features corresponding to the first visual features and first text features corresponding to the first initial text features;

[0147] Concatenate the first mapping features and the first text features to obtain first joint features;

[0148] Input the first combined feature into the decoding module for content understanding of the combined indication information to obtain a first output result;

[0149] Determine a first model loss based on the first output result and the first sample annotation of the positive sample text-image pair, as well as the first output result and the first sample annotation of the negative sample text-image pair;

[0150] Train the initial model based on the first model loss to adjust the model parameters of the visual encoder and the projection module until the first end condition is met.

[0151] In some embodiments, the second training module 30 may be specifically configured to:

[0152] Perform word segmentation processing and feature embedding on the text of the sample text-image pair data and the sample text data to obtain a second initial text feature corresponding to the sample text-image pair data and a third initial text feature corresponding to the sample text data. During the word segmentation process, use the target word library as the word segmentation knowledge base to segment the text corresponding to the preset business domain;

[0153] Input the image of the sample text-image pair data into the visual encoder for feature encoding to obtain a second visual feature;

[0154] Input the second visual feature, the second initial text feature, and the third initial text feature of the sample text-image pair data into the projection module respectively for text space mapping of the visual feature and text feature extraction to obtain a second mapped feature and a second text feature corresponding to the sample text-image pair data, and a third text feature corresponding to the sample text data;

[0155] Concatenate the second mapped feature and the second text feature corresponding to the sample text-image pair data to obtain a second combined feature;

[0156] Input the second combined feature and the third text feature into the decoding module respectively for content understanding of the combined indication information to obtain a second output result corresponding to the second combined feature and a third output result corresponding to the third text feature;

[0157] Determine a second model loss based on the second output result and the first sample annotation of the sample text-image pair data, and the third output result and the second sample annotation of the sample text data;

[0158] Train the initial model that meets the first end condition based on the second model loss to adjust the model parameters of the visual encoder, the projection module, and the decoding module until the second end condition is met.

[0159] In some embodiments, the visual encoder includes a feature extraction module, a feature fusion module, and a feature extraction module; the device further includes a visual encoding module for:

[0160] Input the image into the feature extraction module, and perform local feature extraction on the image based on the self-attention mechanism of the sliding window to obtain multi-scale feature maps;

[0161] Input the multi-scale feature maps into the feature fusion module for feature fusion to obtain fused features;

[0162] Input the fused features into the feature extraction module for feature extraction to obtain the first visual feature or the second visual feature corresponding to the image.

[0163] In some embodiments, the sample image-text pair data in each preset business domain includes multiple types of image-text pairs. The feature extraction module includes a gating network and multiple expert networks matching multiple types. The visual encoding module is further specifically configured to:

[0164] Estimate the type of the image based on the gating network, and input the fused features into the expert network matching the estimated type for feature extraction to obtain the first visual feature or the second visual feature.

[0165] In some embodiments, the visual encoding module is further specifically configured to:

[0166] Perform adaptive image segmentation on the image to obtain multiple sub-images corresponding to the image;

[0167] Perform local feature extraction on each sub-image based on the self-attention mechanism of the sliding window to obtain multi-scale feature maps corresponding to the sub-images.

[0168] In some embodiments, the type of the sample image-text pair data in the preset business domain is determined based on the amount of text in the image, including a first type, a second type, a third type, and a fourth type with increasing text amounts. The first type of image corresponds to an image without text, and the fourth type of image corresponds to a dense document image. The acquisition module 10 is specifically configured to:

[0169] For the first type of image, obtain the descriptive text in multiple language versions of the image, and combine each language version of the descriptive text with the image respectively to obtain the sample image-text pair data;

[0170] For the second type of image, perform content understanding and content extraction on the image based on the multi-modal model to obtain the indication information and the first sample annotation corresponding to the indication information; combine the image, the indication information, and the first sample annotation to obtain the sample image-text pair data;

[0171] For the third type of image, perform layout area analysis on the image to obtain the graphic elements and document elements in the image; obtain the indication information of at least one of the graphic elements and document elements and the first sample annotation corresponding to the indication information; combine the image, the indication information, and the first sample annotation to obtain the sample image-text pair data;

[0172] For the fourth type of image, the target document corresponding to the image is input into a large language model for content understanding and content extraction to obtain the indication information corresponding to the image and the first sample annotation corresponding to the indication information; the image, the indication information, and the first sample annotation are combined to obtain the sample text-image pair data; the target document is obtained by splitting the original text document, and the fourth type of image is obtained by converting the target document into a picture.

[0173] In some embodiments, the types of the indication information of the third type of image include content question indication for document elements, position detection indication for picture elements or document elements, and content question indication and position detection indication for at least one of picture elements and document elements.

[0174] In some embodiments, the first training module 20 is further specifically configured to:

[0175] Generate a contrastive loss based on the difference between the first output result of the positive sample text-image pair and the first sample annotation, and the difference between the first output result of the negative sample text-image pair and the first sample annotation;

[0176] Perform auxiliary loss calculation of the feature extraction module based on the selection probability of each expert network for the images input in the training batch and the number of expert networks to obtain a load balancing loss;

[0177] Perform auxiliary loss calculation of the feature extraction module based on the probability index data obtained by the gating network for type estimation of the fused features, where the probability index data includes the probabilities that the images corresponding to the fused features belong to each expert network respectively, to obtain a router z loss;

[0178] For the sample text-image pair data of the multilingual text type, perform cross-lingual consistency loss calculation based on the output results of the multilingual versions in the first output result to obtain a cross-lingual consistency loss;

[0179] Fuse the contrastive loss, the load balancing loss, the router z loss, and the cross-lingual consistency loss to obtain a first model loss.

[0180] In some embodiments, the second training module 30 is further specifically configured to:

[0181] Generate a cross-entropy loss based on the difference between the second output result and the first sample annotation, and the difference between the third output result and the second sample annotation;

[0182] Perform auxiliary loss calculation of the feature extraction module based on the selection probability of each expert network for the images input in the training batch and the number of expert networks to obtain a load balancing loss;

[0183] Based on the probability index data obtained by estimating the type of the fused features through a gating network, the auxiliary loss of the feature extraction module is calculated to obtain the router z loss. The probability index data includes the probabilities that the images corresponding to the fused features belong to each expert network respectively.

[0184] For the sample image-text pairs or text data of the multi-language text type, based on the output results of the multi-language versions in the second output result and the output results of the multi-language versions in the third output result, the cross-language consistency loss is calculated to obtain the cross-language consistency loss.

[0185] The cross-entropy loss, the load balancing loss, the router z loss, and the cross-language consistency loss are fused to obtain the second model loss.

[0186] It should be noted that the above device embodiment and the method embodiment are based on the same implementation manner.

[0187] An embodiment of the present application provides a device, which can be a terminal or a server, including a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the model pre-training method or the neural network training method for multi-language tasks provided by the above method embodiment.

[0188] The memory can be used to store software programs and modules. The processor executes various functional applications and anomaly detections by running the software programs and modules stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory can also include a memory controller to provide the processor with access to the memory.

[0189] The method embodiment provided by the embodiment of the present application can be executed in electronic devices such as mobile terminals, computer terminals, servers, or similar computing devices. Figure 4 It is a hardware structure block diagram of an electronic device for a model pre-training method for multi-language tasks provided by an embodiment of the present application. As Figure 4 As shown, the electronic device 900 can vary significantly due to different configurations or performances. It may include one or more central processing units (CPUs) 910 (the processor 910 may include, but is not limited to, processing devices such as a microcontroller unit (MCU) or a field-programmable gate array (FPGA)), a memory 930 for storing data, and one or more storage media 920 for storing application programs 923 or data 922 (such as one or more mass storage devices). Among them, the memory 930 and the storage media 920 can be transient storage or persistent storage. The program stored in the storage media 920 may include one or more modules, and each module may include a series of instruction operations on the electronic device. Further, the central processor 910 can be configured to communicate with the storage media 920 and execute a series of instruction operations in the storage media 920 on the electronic device 900. The electronic device 900 may also include one or more power supplies 960, one or more wired or wireless network interfaces 950, one or more input / output interfaces 940, and / or one or more operating systems 921, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0190] The input / output interface 940 can be used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by the communication provider of the electronic device 900. In one example, the input / output interface 940 includes a network interface controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the input / output interface 940 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0191] Those of ordinary skill in the art can understand that Figure 4 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the electronic device 900 may also include more or fewer components than those Figure 4 shown, or have a different configuration from that Figure 4 shown.

[0192] An embodiment of the present application also provides a computer-readable storage medium. The storage medium can be disposed in the electronic device to store at least one instruction or at least one segment of a program related to an anomaly detection method in the method embodiment. The at least one instruction or the at least one segment of the program is loaded and executed by the processor to implement the anomaly detection method provided in the above method embodiment.

[0193] Optionally, in this embodiment, the above storage medium may be located in at least one of multiple network servers in a computer network. Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0194] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above various optional implementation manners.

[0195] For the model pre-training method, device, equipment, storage medium, server, terminal, and program product for multilingual tasks provided by the present application, the multi-modal training data set adopted by the technical solution of the present application includes multiple sample text data and multiple sample text-image pair data. The multiple sample text-image pair data and the multiple sample text data include multiple language contents, and the multiple sample text-image pair data include sample text-image pair data in the general field and sample text-image pair data in a preset business field in the target scenario. The multiple sample text data include text data in the preset business field, so as to provide multi-language graphic and text data and pure text data covering the preset business field, improve the adaptability of the multi-modal model to multilingual tasks in a specific business field, and have the understanding ability of both multi-modal data and text data. And during the pre-training process, a two-stage training method is adopted, and visual feature and text feature alignment training and content understanding constraint training are respectively performed, so as to realize the transfer of multi-modal capabilities from high-resource languages to low-resource languages through multi-language and multi-modal pre-training, reduce data requirements, shorten the training time, and improve the model performance and generality at the same time.

[0196] It should be noted that: the above sequence of embodiments of the present application is only for description and does not represent the advantages or disadvantages of the embodiments. And the above specific embodiments of the present application have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be executed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multi-task processing and parallel processing are also possible or may be advantageous.

[0197] The embodiments in the present application are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.

[0198] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing the relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a disk, or an optical disc, etc.

[0199] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.< / lb> < / lb> < / lb> < / lb> < / lb> < / lb> < / lb> < / lb> < / lb> < / lb> < / lb> < / lb> < / lb> < / lb> < / lb> < / lb> < / lb> < / lb>

Claims

1. A model pre-training method for multilingual tasks, characterized in that, The method includes: Obtaining a multi-modal training data set and an initial model, where the training data set includes multiple sample text data and multiple sample text-image pair data. The multiple sample text-image pair data and the multiple sample text data include content in multiple languages, and the multiple sample text-image pair data include sample text-image pair data in the general domain and sample text-image pair data in a preset business domain in the target scenario. The multiple sample text data include text data in the preset business domain. The initial model includes a visual encoder, a projection module, and a decoding module connected in sequence, and the decoding module is constructed based on a large language model; the target scenario is a cross-border business scenario, and the preset business domain is a multi-language business domain; Based on the multiple sample text-image pair data, performing contrastive learning training on the initial model for aligning visual features and text features. During the training process, freeze the model parameters of the decoding module and adjust the model parameters of the visual encoder and the projection module until the first end condition is met; Based on the multiple sample text-image pair data and the multiple sample text data, performing constraint training on the initial model that meets the first end condition for content understanding. During the training process, adjust the model parameters of the visual encoder, the projection module, and the decoding module until the second end condition is met; the type of the sample text-image pair data in the preset business domain is determined based on the amount of text in the image, and includes a first type, a second type, a third type, and a fourth type with increasing text amount. The image of the first type is an image without text, the image of the second type is a sparse text image, the image of the third type is a multi-text image, and the image of the fourth type is a text document image; the visual encoder includes a feature extraction module, a feature fusion module, and a feature extraction module. The feature extraction module is used to perform adaptive image segmentation on the image to obtain multiple sub-images corresponding to the image; perform local feature extraction on each sub-image based on the self-attention mechanism of the sliding window to obtain multi-scale feature maps corresponding to the sub-images; the feature fusion module is used to fuse the multi-scale feature maps to obtain a fused feature; the feature extraction module includes a gating network and multiple expert networks matching multiple types. The gating network is used to determine the image type of the sample text-image pair data, and each expert network is used to process the fused features of images of different text amount types; Determine the initial model that meets the second end condition as the target model; during the training process, perform word segmentation on the text in the sample text-image pair data and the sample text data in the preset business domain in combination with the target vocabulary corresponding to the preset business domain as the input of the projection module.

2. The method according to claim 1, wherein The performing contrastive learning training on the initial model for aligning visual features and text features based on the multiple sample text-image pair data includes: Sampling the multiple sample text-image pair data to obtain positive sample text-image pairs and negative sample text-image pairs; Tokenize the text of the positive sample text-image pair and the text of the negative sample text-image pair, and perform feature embedding to obtain the first initial text feature. During the tokenization process, use the target vocabulary as the tokenization knowledge base to tokenize the text corresponding to the preset business domain. Input the images of the positive sample text-image pair and the negative sample text-image pair into the visual encoder for feature encoding to obtain the first visual feature. Input the first visual feature and the first initial text feature into the projection module for text space mapping of the visual feature and text feature extraction to obtain the first mapping feature corresponding to the first visual feature and the first text feature corresponding to the first initial text feature. Concatenate the first mapping feature and the first text feature to obtain the first combined feature. Input the first combined feature into the decoding module for content understanding of the combined indication information to obtain the first output result. Determine the first model loss based on the first output result and the first sample annotation of the positive sample text-image pair, and the first output result and the first sample annotation of the negative sample text-image pair. Train the initial model based on the first model loss to adjust the model parameters of the visual encoder and the projection module until the first end condition is met.

3. The method according to claim 1, characterized in that, The constrained training for content understanding of the initial model that meets the first end condition based on the multiple sample text-image pair data and the multiple sample text data includes: Tokenize the text of the sample text-image pair data and the sample text data, and perform feature embedding to obtain the second initial text feature corresponding to the sample text-image pair data and the third initial text feature corresponding to the sample text data. During the tokenization process, use the target vocabulary as the tokenization knowledge base to tokenize the text corresponding to the preset business domain. Input the images of the sample text-image pair data into the visual encoder for feature encoding to obtain the second visual feature. Input the second visual feature and the second initial text feature of the sample text-image pair data, and the third initial text feature into the projection module respectively for text space mapping of the visual feature and text feature extraction to obtain the second mapping feature and the second text feature corresponding to the sample text-image pair data, and the third text feature corresponding to the sample text data. Concatenate the second mapping feature and the second text feature corresponding to the sample text-image pair data to obtain the second combined feature. Input the second combined feature and the third text feature into the decoding module respectively for content understanding of the combined indication information to obtain the second output result corresponding to the second combined feature and the third output result corresponding to the third text feature. Determine the second model loss based on the second output result and the first sample annotation of the sample text-image pair data, and the third output result and the second sample annotation of the sample text data. Train the initial model that meets the first end condition based on the second model loss to adjust the model parameters of the visual encoder, the projection module, and the decoding module until the second end condition is met.

4. The method according to claim 2 or 3, characterized in that, The visual encoder includes a feature extraction module, a feature fusion module, and a feature extraction module; the feature encoding process of the visual encoder for the input image includes: Input the image into the feature extraction module, and perform local feature extraction on the image based on the self-attention mechanism of the sliding window to obtain multi-scale feature maps; Input the multi-scale feature maps into the feature fusion module for feature fusion to obtain fused features; Input the fused features into the feature extraction module for feature extraction to obtain the first visual feature or the second visual feature corresponding to the image.

5. The method according to claim 4, wherein The sample text-image pair data of each of the preset business fields includes multiple types of text-image pairs. The step of inputting the fused features into the feature extraction module for feature extraction to obtain the first visual feature or the second visual feature corresponding to the image includes: Estimate the type of the image based on the gating network, and input the fused features into the expert network matching the estimated type for feature extraction to obtain the first visual feature or the second visual feature.

6. The method according to claim 4, characterized in that The acquisition method of the sample text-image pair data of the preset business field includes: For the images of the first type, obtain the description texts in multiple languages of the images, and combine each description text with the image to obtain the sample text-image pair data; For the images of the second type, perform content understanding and content extraction on the images based on a multi-modal model to obtain indication information and the first sample annotation corresponding to the indication information; combine the image, the indication information, and the first sample annotation to obtain the sample text-image pair data; For the images of the third type, perform layout area analysis on the images to obtain graphic elements and document elements in the images; obtain the indication information of at least one of the graphic elements and the document elements and the first sample annotation corresponding to the indication information; combine the image, the indication information, and the first sample annotation to obtain the sample text-image pair data; For the images of the fourth type, input the target document corresponding to the image into a large language model for content understanding and content extraction to obtain the indication information corresponding to the image and the first sample annotation corresponding to the indication information; combine the image, the indication information, and the first sample annotation to obtain the sample text-image pair data; the target document is obtained by splitting the original text document, and the images of the fourth type are obtained by converting the target document into pictures.

7. The method according to claim 6, wherein The types of the indication information of the images of the third type include content question indications for the document elements, position detection indications for the picture elements or the document elements, and content question indications and position detection indications for at least one of the picture elements and the document elements.

8. The method according to claim 5, characterized in that, The method for obtaining the first model loss includes: Generate a contrastive loss based on the difference between the first output result of the positive sample text-image pair and the first sample annotation, and the difference between the first output result of the negative sample text-image pair and the first sample annotation. Based on the selection probabilities of each of the expert networks for the images input within a training batch and the number of the expert networks, calculate the auxiliary loss of the feature extraction module to obtain the load balancing loss; Based on the probability metric data obtained by the gating network for estimating the type of the fused feature, calculate the auxiliary loss of the feature extraction module to obtain the router z loss, where the probability metric data includes the probabilities that the images corresponding to the fused feature belong to each of the expert networks; For the sample text-image pair data of the multi-language text type, based on the output results of the multi-language versions in the first output result, calculate the cross-language consistency loss to obtain the cross-language consistency loss; Fuse the contrast loss, the load balancing loss, the router z loss, and the cross-language consistency loss to obtain the first model loss.

9. The method according to claim 5, characterized in that The method for obtaining the second model loss includes: Based on the differences between the second output result and the first sample annotation, and the differences between the third output result and the second sample annotation, generate the cross-entropy loss; Based on the selection probabilities of each of the expert networks for the images input within a training batch and the number of the expert networks, calculate the auxiliary loss of the feature extraction module to obtain the load balancing loss; Based on the probability metric data obtained by the gating network for estimating the type of the fused feature, calculate the auxiliary loss of the feature extraction module to obtain the router z loss, where the probability metric data includes the probabilities that the images corresponding to the fused feature belong to each of the expert networks; For the sample text-image pair data or text data of the multi-language text type, based on the output results of the multi-language versions in the second output result and the output results of the multi-language versions in the third output result, calculate the cross-language consistency loss to obtain the cross-language consistency loss; Fuse the cross-entropy loss, the load balancing loss, the router z loss, and the cross-language consistency loss to obtain the second model loss.

10. A model pre-training device for multilingual tasks, characterized in that, The device includes: An acquisition module: used to acquire a multi-modal training data set and an initial model, where the training data set includes a plurality of sample text data and a plurality of sample text-image pair data, the plurality of sample text-image pair data and the plurality of sample text data include multiple language contents, and the plurality of sample text-image pair data include sample text-image pair data in the general domain and sample text-image pair data in a preset business domain in the target scenario, the plurality of sample text data include the text data in the preset business domain, the initial model includes a visual encoder, a projection module, and a decoding module connected in sequence, and the decoding module is constructed based on a large language model; the target scenario is a cross-border business scenario, and the preset business domain is a multi-language business domain; A first training module: used to perform contrastive learning training on the initial model for aligning visual features and text features based on the plurality of sample text-image pair data, and freeze the model parameters of the decoding module and adjust the model parameters of the visual encoder and the projection module during the training process until the first end condition is met; Second training module: It is used to perform constrained training on the content understanding of the initial model that meets the first end condition based on the multiple sample image-text pair data and the multiple sample text data, and adjust the model parameters of the visual encoder, the projection module, and the decoding module during the training process until the second end condition is met; the type of the sample image-text pair data in the preset business domain is determined based on the amount of text in the image, and includes a first type, a second type, a third type, and a fourth type with increasing text amounts. The image of the first type is an image without text, the image of the second type is a sparse text image, the image of the third type is a multi-text image, and the image of the fourth type is a text document image; the visual encoder includes a feature extraction module, a feature fusion module, and a feature extraction module. The feature extraction module is used to perform adaptive image segmentation on the image to obtain multiple sub-images corresponding to the image; perform local feature extraction on each sub-image based on the self-attention mechanism of the sliding window to obtain multi-scale feature maps corresponding to the sub-images; the feature fusion module is used to fuse the multi-scale feature maps to obtain a fused feature; the feature extraction module includes a gating network and multiple expert networks matching multiple types. The gating network is used to determine the image type of the sample image-text pair data, and each expert network is used to process the fused features of images of different text amount types; Model generation module: It is used to determine the initial model that meets the second end condition as the target model; during the training process, perform word segmentation on the text in the sample image-text pair data and the sample text data in the preset business domain in combination with the target word library corresponding to the preset business domain as the input of the projection module.

11. The device according to claim 10, wherein The contrastive learning training for aligning the visual features and text features of the initial model based on the multiple sample image-text pair data includes: Sampling the multiple sample image-text pair data to obtain positive sample image-text pairs and negative sample image-text pairs; Performing word segmentation and feature embedding on the text of the positive sample image-text pairs and the text of the negative sample image-text pairs to obtain first initial text features. During the word segmentation process, use the target word library as the word segmentation knowledge base to perform word segmentation on the text corresponding to the preset business domain; Inputting the images of the positive sample image-text pairs and the negative sample image-text pairs into the visual encoder for feature encoding to obtain first visual features; Inputting the first visual features and the first initial text features into the projection module for text space mapping of the visual features and text feature extraction to obtain a first mapping feature corresponding to the first visual feature and a first text feature corresponding to the first initial text feature; Concatenating the first mapping feature and the first text feature to obtain a first combined feature; Inputting the first combined feature into the decoding module for content understanding combined with the indication information to obtain a first output result; Determine a first model loss based on the first output result and the first sample annotation of the positive sample image-text pair, and the first output result and the first sample annotation of the negative sample image-text pair. Train the initial model based on the first model loss to adjust the model parameters of the visual encoder and the projection module until the first termination condition is met.

12. The device according to claim 10, characterized in that, The constrained training for content understanding of the initial model that meets the first termination condition based on the multiple sample image-text pair data and the multiple sample text data includes: Perform word segmentation processing and feature embedding on the text of the sample image-text pair data and the sample text data to obtain a second initial text feature corresponding to the sample image-text pair data and a third initial text feature corresponding to the sample text data. During the word segmentation process, use the target vocabulary as the word segmentation knowledge base to segment the text corresponding to the preset business domain. Input the image of the sample image-text pair data into the visual encoder for feature encoding to obtain a second visual feature. Input the second visual feature and the second initial text feature of the sample image-text pair data, and the third initial text feature into the projection module respectively to perform text space mapping of the visual feature and text feature extraction, to obtain a second mapping feature and a second text feature corresponding to the sample image-text pair data, and a third text feature corresponding to the sample text data. Concatenate the second mapping feature and the second text feature corresponding to the sample image-text pair data to obtain a second combined feature. Input the second combined feature and the third text feature into the decoding module respectively for content understanding of the combined indication information to obtain a second output result corresponding to the second combined feature and a third output result corresponding to the third text feature. Determine a second model loss based on the second output result and the first sample annotation of the sample image-text pair data, and the third output result and the second sample annotation of the sample text data. Train the initial model that meets the first termination condition based on the second model loss to adjust the model parameters of the visual encoder, the projection module and the decoding module until the second termination condition is met.

13. The device according to claim 11 or 12, characterized in that, The visual encoder includes a feature extraction module, a feature fusion module and a feature extraction module; the process of feature encoding of the input image by the visual encoder includes: Input the image into the feature extraction module, and perform local feature extraction on the image based on the self-attention mechanism of the sliding window to obtain a multi-scale feature map. Input the multi-scale feature map into the feature fusion module for feature fusion to obtain a fused feature. Input the fused feature into the feature extraction module for feature extraction to obtain the first visual feature or the second visual feature corresponding to the image.

14. The device according to claim 13, wherein The sample image-text pair data of each preset business domain includes multiple types of image-text pairs. The process of inputting the fused feature into the feature extraction module for feature extraction to obtain the first visual feature or the second visual feature corresponding to the image includes: Estimate the type of the image based on the gating network, and input the fused feature into the expert network matching the estimated type for feature extraction to obtain the first visual feature or the second visual feature.

15. The device according to claim 13, characterized in that, The acquisition methods of the sample text-image pair data in the preset business field include: For the images of the first type, obtain the descriptive texts of the multi-language versions of the images, and combine the descriptive texts of each language version with the image respectively to obtain the sample text-image pair data; For the images of the second type, perform content understanding and content extraction on the images based on a multi-modal model to obtain indication information and a first sample annotation corresponding to the indication information; combine the image, the indication information, and the first sample annotation to obtain the sample text-image pair data; For the images of the third type, perform layout area analysis on the images to obtain graphic elements and document elements in the images; obtain indication information and a first sample annotation corresponding to at least one of the graphic elements and the document elements; combine the image, the indication information, and the first sample annotation to obtain the sample text-image pair data; For the images of the fourth type, input the target document corresponding to the image into a large language model for content understanding and content extraction to obtain indication information corresponding to the image and a first sample annotation corresponding to the indication information; combine the image, the indication information, and the first sample annotation to obtain the sample text-image pair data; the target document is obtained by splitting the original text document, and the images of the fourth type are obtained by converting the target document into pictures.

16. The device according to claim 15, characterized in that, The types of the indication information of the images of the third type include content question indications for the document elements, position detection indications for the picture elements or the document elements, and content question indications and position detection indications for at least one of the picture elements and the document elements.

17. The device according to claim 14, characterized in that The method for obtaining the first model loss includes: Generate a contrast loss based on the differences between the first output results of the positive sample text-image pairs and the first sample annotations, and the differences between the first output results of the negative sample text-image pairs and the first sample annotations; Perform auxiliary loss calculation on the feature extraction module based on the selection probabilities of each expert network for the images input in the training batch and the number of expert networks to obtain a load balancing loss; Perform auxiliary loss calculation on the feature extraction module based on the probability index data obtained by estimating the type of the fused feature by the gating network, and obtain a router z loss, where the probability index data includes the probabilities that the images corresponding to the fused features belong to each expert network; For the sample text-image pair data of the multi-language text type, perform cross-language consistency loss calculation based on the output results of the multi-language versions in the first output results to obtain a cross-language consistency loss; Fuse the contrast loss, the load balancing loss, the router z loss, and the cross-language consistency loss to obtain the first model loss.

18. The device according to claim 14, wherein, The method for obtaining the second model loss includes: Generate a cross-entropy loss based on the difference between the second output result and the first sample annotation, and the difference between the third output result and the second sample annotation. Based on the selection probabilities of each of the expert networks for the images input within the training batch and the number of the expert networks, calculate the auxiliary loss of the feature extraction module to obtain a load balancing loss. Based on the probability metric data obtained by the gating network for estimating the type of the fused feature, calculate the auxiliary loss of the feature extraction module to obtain a router z loss, where the probability metric data includes the probabilities that the images corresponding to the fused feature belong to each of the expert networks. For the sample image-text pairs or text data of the multi-language text type, based on the output results in the multi-language versions in the second output result and the output results in the multi-language versions in the third output result, calculate a cross-language consistency loss to obtain a cross-language consistency loss. Fuse the cross-entropy loss, the load balancing loss, the router z loss, and the cross-language consistency loss to obtain the second model loss.

19. A computer device, characterized in that, The device includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the model pre-training method for multi-language tasks as described in any one of claims 1-9.

20. A computer-readable storage medium, characterized in that, At least one instruction or at least one program segment is stored in the storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the model pre-training method for multi-language tasks as described in any one of claims 1-9.

21. A server, characterized in that, The server includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the model pre-training method for multi-language tasks as described in any one of claims 1-9.

22. A terminal, characterized in that, The terminal includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the model pre-training method for multi-language tasks as described in any one of claims 1-9.

23. A computer program product, characterized in that, The computer program product includes computer instructions, and when the computer instructions are executed by a processor, the model pre-training method for multi-language tasks as described in any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Image content analysis method and device, equipment and medium

    CN116824278A

  • Model training method and device, computer equipment and storage medium

    CN118194963A