Word list expansion method of large visual model

By building an automatic regression model to generate new visual vocabulary and integrate it with the original vocabulary, the problem of insufficient vocabulary coverage in the existing technology is solved, and the fine-grained perception ability of large-scale vision-language models is significantly improved, and it is suitable for multilingual and complex visual tasks.

CN120031074APending Publication Date: 2025-05-23LINKER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411850429.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing large-scale vision-language models have problems with insufficient vocabulary coverage when dealing with non-English scenes or fine-grained visual tasks requiring higher resolution, resulting in the inability to effectively process charts or high-definition documents.

Method used

By building an automatic regression model based on convolutional neural networks and small decoders, a new visual vocabulary is generated and integrated with the original visual vocabulary is expanded to improve its comprehension and perception.

Benefits of technology

It significantly improves the fine-grained perception ability of large-scale vision-language models in specific visual tasks, can effectively handle multilingual and complex visual input, and is suitable for tasks such as chart comprehension and document analysis.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses a vocabulary expansion method for a large visual model, which comprises the following steps of: 1, forming a vocabulary network for generating a new vocabulary based on a convolutional neural network and an automatic regression model of a small decoder, and generating a more efficient vocabulary by gradually predicting a next word in a visual vocabulary; and 2, integrating the new word list generated in the step 1 with a visual word list of an original large visual model to complete word list expansion of the large visual model. According to the vocabulary expansion method of the large visual model, the new vocabulary network can be effectively generated by setting the step 1 to the step 2, and then the new vocabulary is fused into the original vocabulary to realize the vocabulary expansion of the large visual model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a large-scale visual-language model, and more particularly to a vocabulary expansion method for a large-scale visual model. Background Art

[0002] Most existing large-scale visual-language models use general visual vocabulary (such as the CLIP vocabulary). However, when dealing with non-English scenarios (such as Chinese) or fine-grained visual tasks that require higher resolution (charts or high-definition documents), there is a problem of insufficient vocabulary coverage, which makes existing large-scale visual-language models unable to process charts or high-definition documents. Summary of the invention

[0003] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a vocabulary expansion method for a large-scale visual model, which can effectively expand the visual vocabulary used by the large-scale visual-language model, improve the understanding and perception capabilities of the large-scale visual model in specific visual tasks, such as text recognition, document parsing, and chart analysis, etc., thereby enabling the large-scale visual language model to process charts or high-definition documents.

[0004] To achieve the above object, the present invention provides the following technical solution: a method for expanding the vocabulary of a large visual model, characterized in that it comprises the following steps:

[0005] Step 1: A vocabulary network for generating new vocabulary is constructed based on an auto-regressive model based on a convolutional neural network and a small decoder, which is used to generate a more efficient vocabulary by gradually predicting the next word in the visual vocabulary;

[0006] Step 2: Integrate the new vocabulary generated in step 1 with the visual vocabulary of the original large-scale visual model to complete the vocabulary expansion of the large-scale visual model.

[0007] As a further improvement of the present invention, the vocabulary network in step 1 is obtained by training through the following steps: step 1, using a high-resolution image data set and a high-resolution image encoder for training, and encoding positive sample data and negative sample data during the training process;

[0008] In step 1 and 2, a lightweight small language model is used as a decoder to decode the data encoded in step 1, and an autoregressive generation method is used to generate a new vocabulary to obtain a vocabulary network.

[0009] As a further improvement of the present invention, the vocabulary network in step one generates a more efficient vocabulary in the following specific manner: first, a step-by-step prediction is performed, and then the step-by-step prediction process is continuously iterated, the entire vocabulary sequence is gradually generated through the model, and the iteration is stopped after reaching a preset sequence length or encountering an end marker, and finally a new vocabulary is generated.

[0010] As a further improvement of the present invention, the specific method of the step-by-step prediction in step 1 is: after the model generates a tag, the tag is added to the sequence, and then used to predict the next tag.

[0011] As a further improvement of the present invention, the specific method of integrating the new vocabulary with the visual vocabulary of the original large-scale visual model in step 2 is to process the input layers of the new and old vocabulary independently, and then fuse their features before entering the large-scale language model to ensure that the knowledge of the new vocabulary does not cover the original vocabulary.

[0012] As a further improvement of the present invention, in step 2, two convolutional layers are added after the new and old vocabulary lists are merged to adjust the feature shape. Specifically, the first convolutional layer converts the feature shape to 32×32×512, and the second convolutional layer further converts it to 16×16×1024, and finally the feature is flattened into a shape of 256×1024 to align with the image markers of CLIP.

[0013] The beneficial effect of the present invention is that a vocabulary network for generating new vocabulary is constructed by using a convolutional neural network and an autoregressive model of a small decoder, and then a new vocabulary is generated by using the vocabulary network, and then the new vocabulary is integrated with the original visual vocabulary, so that the vocabulary expansion can be realized simply and effectively. By expanding the visual vocabulary, the perception and understanding ability of large-scale visual-language models is significantly improved, especially in non-English tasks such as Chinese document parsing and chart understanding. The extended model performs well in fine-grained perception tasks, can effectively process multilingual and complex visual inputs, provides a wider application prospect, and effectively improves the performance of existing large visual models. DETAILED DESCRIPTION

[0014] The present invention will be further described in detail with reference to the given embodiments below.

[0015] The vocabulary expansion method of a large visual model in this embodiment mainly includes the following two stages:

[0016] In the first stage, a vocabulary network is designed to generate new vocabulary. The network consists of an auto-regressive model based on a convolutional neural network (CNN) and a small decoder, which generates more efficient vocabulary by gradually predicting the next word in the visual vocabulary.

[0017] In this embodiment, in order to ensure that the new vocabulary can effectively encode complex visual information, a large number of high-resolution image (1024×1024) data sets and high-resolution image encoders (such as ViTDet) are used in the training process to adapt to high-resolution visual information input (such as document-level OCR and chart parsing tasks). Furthermore, the new vocabulary generation network encodes positive sample data (high-resolution document images and chart data) and negative sample data (natural images) during training. The purpose of this is to hope that the new visual vocabulary network can perform well in processing artificial images (i.e., documents and charts) to make up for the shortcomings of the original network model. At the same time, it is also hoped that it will not become noise of the original network model when marking natural images.

[0018] In this embodiment, the new vocabulary generation network finally uses a lightweight small language model (OPT-125M) as a decoder and uses an autoregressive generation method to generate a new vocabulary. The small language model is used as a decoder, and its task is to help generate a new visual vocabulary. Since the model is small, only less computing resources are required to perform autoregressive vocabulary generation operations, and training can be completed faster, thereby achieving more efficient vocabulary expansion. The method of autoregressive vocabulary generation is to generate new vocabulary by gradually predicting the next word or token. The method first encodes the input image into an initial set of image tags through a vocabulary network. These tags are used as the beginning of the sequence to generate subsequent vocabulary. Then a step-by-step prediction is performed, and at each step, the model predicts the next tag based on the current existing vocabulary sequence. This part is the key to the autoregressive method, that is, after the model generates a tag, the tag is added to the sequence and then used to predict the next tag. By continuously iterating this process, the model gradually generates the entire vocabulary sequence until the specified sequence length is reached or the end tag (such as EOS, End of Sequence) is encountered, and finally a new vocabulary is generated. The advantage of this autoregressive generation method is that it allows the model to continuously adjust its understanding of visual information during the generation process, thereby generating an efficient vocabulary sequence, which is particularly suitable for tasks that require detailed perception, such as document parsing and complex scenarios such as OCR. In this way, the vocabulary network can fully learn the detailed features in these tasks, thereby generating a vocabulary suitable for high-density visual perception.

[0019] In the second stage, the generated new vocabulary is integrated with the original visual vocabulary of the large visual-language model (LVLM) (usually using the CLIP vocabulary) to retain the original performance of the large visual-language model in general tasks while significantly improving its fine-grained perception ability in specific tasks.

[0020] The specific approach of this embodiment is to process the input layers of the new and old vocabulary independently, and then fuse their features before entering the large-scale language model to ensure that the knowledge of the new vocabulary does not overwrite the original vocabulary. The final integration process does not require retraining the entire model, but only freezes the parameters of the new and old vocabulary, which simplifies the expansion process and improves training efficiency.

[0021] It is worth mentioning that since the feature resolution (64×64×256) of the final output of the high-resolution image encoder used in the new vocabulary generation network is inconsistent with the output resolution of the CLIP encoding (256×1024), two convolutional layers are added at the end of the vocabulary network to adjust the feature shape. Specifically, the first convolutional layer converts the feature shape to 32×32×512, and the second convolutional layer further converts it to 16×16×1024, and finally flattens the feature into a shape of 256×1024 to align with the image tag of CLIP. Since the feature resolution output by the SAM model is inconsistent with the output of CLIP (SAM output shape is 64×64×256, while CLIP is 256×1024), two convolutional layers are added at the end of the vocabulary network to adjust the feature shape. Specifically, the first convolutional layer transforms the feature shape to 32×32×512, the second convolutional layer further transforms it to 16×16×1024, and finally the feature is flattened to a shape of 256×1024 to align with the image landmarks using the CLIP model.

[0022] In summary, this embodiment provides an efficient visual vocabulary expansion method, which enables the large visual-language model (LVLM) to have advantages in fine-grained and multilingual tasks while retaining general performance, providing strong support for the application of large-scale visual-language models in complex scenarios.

[0023] Based on the above extension method, this embodiment provides the following examples:

[0024] The training process of generating new vocabulary

[0025] In this embodiment, a method for generating a new visual vocabulary by training a high-resolution document image and chart image dataset is proposed. This method is mainly used for fine-grained visual tasks, especially for processing chart understanding tasks containing complex structural information. In order to improve the performance of the model, we first prepared a high-resolution image dataset containing Chinese and English chart data with rich chart elements (such as axes, legends, titles, labels, etc.). The dataset includes common statistical charts (such as bar charts, line charts, pie charts, etc.) as well as professional charts in the fields of finance and science.

[0026] Dataset construction:

[0027] Sources of chart data: We collected a large number of Chinese and English chart images from public chart datasets (such as the Chart Understanding Dataset and the Chinese Chart Dataset) as well as commercial and academic publications. These charts cover a variety of topics, including economics, medicine, science and technology. Each chart image is accompanied by detailed annotation information, including the location, category, relationship of chart elements, and related text information (such as chart title, axis label, value label, etc.).

[0028] Data preprocessing: All chart images are uniformly resized to 1024×1024 resolution to ensure that the model can handle high-definition image data. To enhance data diversity, we also perform data augmentation such as random cropping, rotation, and color transformation on the chart images. In addition, in order to adapt to high-resolution image encoders (such as ViTDet), we also segment the images to avoid information loss when inputting the network.

[0029] Training process:

[0030] Visual vocabulary generation: A specialized visual vocabulary generation network is trained using a high-resolution image dataset (1024×1024). The network is based on an auto-regressive model of a convolutional neural network (CNN) and a small decoder to gradually generate new visual vocabulary. These vocabulary can effectively encode elements in a chart, such as axes, data points, lines, text labels, etc.

[0031] Autoregressive generation: During the training process, the visual vocabulary generation network encodes the chart image to obtain preliminary image features (such as the position information and relationship of each element in the chart), and gradually predicts the corresponding vocabulary for each chart element through autoregression. This method ensures that the network maintains a detailed perception of the chart structure when generating new vocabulary.

[0032] Model evaluation and optimization: By evaluating the performance of the generated vocabulary in the chart understanding task, the model training process is optimized by combining indicators such as precision, recall, and F1 value. Experimental results show that the generated new vocabulary performs well in the recognition and understanding of chart elements, especially when dealing with charts containing multiple languages ​​(such as Chinese and English) and complex graphics, which can effectively improve the fine-grained perception ability of the model.

[0033] Integration of new vocabulary with existing vocabulary

[0034] In this example, we integrate the generated new vocabulary with the visual vocabulary in the original large-scale visual-language model (such as CLIP) to improve the performance of large-scale visual-language models in graph understanding tasks. The specific operations are as follows:

[0035] Vocabulary integration process: We first process the input layers of the generated new vocabulary and CLIP vocabulary independently to ensure that the two do not interfere with each other. For the input chart image, it is first encoded through the generated new vocabulary to extract the high-resolution features of the chart elements; then, these features are fused with the features of the CLIP vocabulary to ensure that the model can simultaneously utilize the advantages of the new vocabulary in fine-grained visual tasks and the powerful capabilities of the CLIP vocabulary in general visual tasks.

[0036] Feature Fusion and Adjustment: Since the feature resolution of the new vocabulary is different from that of the CLIP vocabulary, we add two convolutional layers at the end of the vocabulary network to adjust the feature shape. Specifically, the first convolutional layer adjusts the feature shape from 64×64×256 to 32×32×512, and the second convolutional layer further adjusts it to 16×16×1024, and finally flattens the feature to 256×1024 to ensure that the feature is aligned with the image tag of CLIP.

[0037] No retraining of the entire model: After the vocabulary is integrated, we do not need to retrain the entire large vision-language model. Instead, we freeze the parameters of the new and old vocabulary, which simplifies the expansion process and significantly improves training efficiency.

[0038] Application in graph understanding tasks:

[0039] In the graph understanding task, the model showed excellent performance by training with graph data containing Chinese and English. Specific application scenarios include:

[0040] Financial data analysis: The model can identify and understand data trends, growth rates, market share and other information in Chinese financial charts, providing users with accurate data interpretation and reports.

[0041] Scientific research reports: The model can effectively process scientific charts in Chinese or English, such as experimental results charts, research data charts, etc., extract key data points and provide accurate interpretation, helping researchers quickly extract the key points of the data.

[0042] Enterprise management decision-making: The model can automatically analyze and interpret various statistical charts in enterprise management, helping decision makers understand operational conditions, financial conditions and market trends.

[0043] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A method for expanding the vocabulary of a large visual model, characterized by: The steps include: Step 1: A vocabulary network for generating new vocabulary is constructed based on an auto-regressive model based on a convolutional neural network and a small decoder, which is used to generate a more efficient vocabulary by gradually predicting the next word in the visual vocabulary; Step 2: Integrate the new vocabulary generated in step 1 with the visual vocabulary of the original large-scale visual model to complete the vocabulary expansion of the large-scale visual model.

2. The method for expanding the vocabulary of a large visual model according to claim 1, characterized in that: The vocabulary network in step 1 is obtained by training through the following steps: Step 1, using a high-resolution image dataset and a high-resolution image encoder for training, and encoding positive sample data and negative sample data during the training process; In step 1 and 2, a lightweight small language model is used as a decoder to decode the data encoded in step 1, and an autoregressive generation method is used to generate a new vocabulary to obtain a vocabulary network.

3. The method for expanding the vocabulary of a large visual model according to claim 2, characterized in that: The specific way in which the vocabulary network in step one generates a more efficient vocabulary is as follows: first, a step-by-step prediction is performed, and then the step-by-step prediction process is continuously iterated, the entire vocabulary sequence is gradually generated through the model, and the iteration is stopped after reaching a preset sequence length or encountering an end marker, and finally a new vocabulary is generated.

4. The method for expanding the vocabulary of a large visual model according to claim 3, characterized in that: The specific method of step-by-step prediction in step 1 is: after the model generates a tag, it adds the tag to the sequence and then uses it to predict the next tag.

5. The method for expanding the vocabulary of a large visual model according to any one of claims 1 to 4, characterized in that: The specific method of integrating the new vocabulary with the visual vocabulary of the original large-scale visual model in step 2 is to process the input layers of the new and old vocabulary independently, and then fuse their features before entering the large-scale language model to ensure that the knowledge of the new vocabulary does not cover the original vocabulary.

6. The method for expanding the vocabulary of a large visual model according to claim 5, characterized in that: In step 2, two convolutional layers are added after the new and old vocabulary lists are merged to adjust the feature shape. Specifically, the first convolutional layer converts the feature shape to 32×32×512, and the second convolutional layer further converts it to 16×16×1024. Finally, the feature is flattened into a shape of 256×1024 to align with the image markers of CLIP.