Model training method, method for associating text, and related device

By freezing non-critical parameters through the Bezier curve denoising training method and combining it with the pre-training and fine-tuning of the Transformer architecture, the problem of neglected relevance of text instances is solved, and efficient text detection and association processing is achieved.

CN119399748BActive Publication Date: 2025-09-30SHANGHAI BILIBILI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411537550.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-30
Publication Date
2025-09-30
Estimated Expiration
2044-10-30

AI Technical Summary

Technical Problem

In the existing technology, scene text detection and recognition do not focus on the correlation between text instances, resulting in low text reading efficiency. In addition, traditional model training methods are inefficient and it is difficult to complete large-scale text detection in a short period of time.

Method used

A Bezier curve-based denoising training method is used to freeze non-critical parameters in the model, and only the vocabulary and task-related structures are trained. The model efficiency is improved through pre-training and fine-tuning, and the Transformer architecture is used for text instance association processing.

Benefits of technology

It can detect hundreds of objects in text images in a short period of time, provide correlation information between text instances, and improve model training efficiency and text information processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399748B_ABST
    Figure CN119399748B_ABST
Patent Text Reader

Abstract

The present application provides a model training method, a method for associating text, an apparatus, an electronic device, a computer-readable medium, and a computer program product. The method of the present application includes: obtaining a training image set corresponding to a target task, wherein the target task is to detect text instances in an image and associate multiple target text instances with the same attributes; using a Bezier curve-based denoising training method to train a target model, wherein the method freezes target parameters to train the vocabulary mapping layer of the target model and the portion corresponding to the target task; after unfreezing the frozen parameters, fine-tuning the target model based on the training image set. The present application trains a model by adopting a Bezier curve-based denoising method, and then uses the model to detect text instances contained in an input image and link related text instances.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a model training method, a method for associating text, a device, an electronic device, a computer-readable medium, and a computer program product. Background Art

[0002] Scene text-related tasks are an important research direction in computer vision. They involve multiple subtasks such as detecting, recognizing, and linking text in natural scenes. These tasks are of great significance to image search, instant translation, robot navigation, industrial automation, and other fields.

[0003] Existing scene text solutions primarily focus on the detection and recognition of text instances, but fail to consider the interdependencies between them. In other words, each text instance is treated as a separate entity, without any dependencies between them. In fact, grouping related words together often facilitates text reading. Summary of the Invention

[0004] Various aspects of the present application provide a model training method, a method for associating text, an apparatus, an electronic device, a computer-readable medium, and a computer program product.

[0005] In one aspect of the present application, a training method is provided, wherein the method comprises:

[0006] Obtaining a training image set corresponding to a target task, wherein the target task is to detect text instances in an image and associate multiple target text instances with the same attributes;

[0007] Training the target model using a Bezier curve-based denoising training method, wherein the method freezes parameters other than the target parameters in the target model to train the portion corresponding to the target parameters;

[0008] After unfreezing the frozen parameters, fine-tune the target model based on the training image set.

[0009] In one aspect of the present application, a method for associating text is provided, wherein the method comprises:

[0010] Acquire the target scene image to be processed;

[0011] Use the trained target model to perform text detection on the text instances contained in the target scene image;

[0012] Determining the group of the detected text instance based on preset group label information;

[0013] Associating multiple text instances belonging to the same group;

[0014] The target model is obtained by training using the model training method of the embodiment of the present application.

[0015] In one aspect of the present application, a model training device is provided, wherein the device comprises:

[0016] A device for obtaining a training image set corresponding to a target task, wherein the target task is to detect text instances in an image and associate multiple target text instances with the same attributes;

[0017] A device for training a target model using a Bezier curve-based noise reduction training method, wherein the device performs training on a portion corresponding to the target parameter by freezing parameters other than the target parameter in the target model;

[0018] A device for fine-tuning a target model based on a target training image set after unfreezing the frozen parameters.

[0019] In one aspect of the present application, a device for associating texts is provided, wherein the device comprises:

[0020] Device for acquiring a target scene image to be processed;

[0021] a device for performing text detection on text instances contained in a target scene image using a trained target model;

[0022] means for determining the group of the detected text instance based on preset group label information;

[0023] means for associating a plurality of text instances belonging to the same group;

[0024] The target model is obtained by training using the model training method of the embodiment of the present application.

[0025] In another aspect of the present application, a computer-readable storage medium is provided, on which computer program instructions are stored. The computer program instructions can be executed by a processor to implement the method of the embodiment of the present application.

[0026] In another aspect of the present application, a computer program product is provided, including a computer program, which implements the method of the embodiment of the present application when executed by a processor.

[0027] In the solution provided in the embodiment of the present application, a denoising method based on Bezier curves is adopted to train the model, and then the trained model is used to detect text instances contained in the input image and associate the related text instances. Compared with the traditional text detection and recognition scheme, it provides information on whether the text instances are associated with each other, which is helpful for subsequent text information processing; in the training stage, parameters other than the vocabulary and the parameters related to the task to be executed are frozen, so that only the vocabulary mapping layer and the task-related structure are trained, and in the fine-tuning stage, all learnable parameters are unfrozen to fine-tune the model based on the data set corresponding to the text scene task. This training method reduces the cost of training and fine-tuning on the target data set, and can complete the detection of hundreds of targets in text images within a few hours. At the same time, the vocabulary can be easily expanded, further improving the efficiency of model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, a brief introduction will be given below to the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0029] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0030] Figure 1 A flow chart of a model training method provided in an embodiment of the present application is shown;

[0031] Figure 2 A flowchart of a method for associating texts according to an embodiment of the present application is shown.

[0032] Figure 3 shows an exemplary model structure schematic diagram according to an embodiment of the present application;

[0033] Figure 4 A schematic diagram of the structure of the model training device provided in an embodiment of the present application is shown;

[0034] Figure 5 A schematic structural diagram of a device for associating texts according to an embodiment of the present application is shown;

[0035] Figure 6 A structural diagram of a device suitable for implementing the solution in the embodiments of the present application is shown.

[0036] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION

[0037] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0038] In a typical configuration of the present application, the terminal and the equipment of the service network each include one or more processors (CPUs), input / output interfaces, network interfaces and memories.

[0039] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0040] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer program instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc-read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0041] Figure 1 A flow chart of a model training method provided in an embodiment of the present application is shown, wherein the method comprises at least step S101, step S102 and step S103.

[0042] In practical scenarios, the execution subject of this method can be a network device or an application running on a network device. The network device includes, but is not limited to, a network host, a single network server, a set of multiple network servers, or a collection of computers based on cloud computing, and can be used to implement some of the processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a group of loosely coupled computers forming a virtual computer.

[0043] The concepts involved in the embodiments of this application are explained below.

[0044] Transformer architecture: The model based on the Transformer architecture is a deep learning model.

[0045] Bezier curve: The Bezier curve method is a technology used for natural scene text detection and recognition (STR), which uses Bezier curves to represent and fit the shape of text areas instead of traditional rectangles or polygons.

[0046] In some embodiments, the model of the embodiment of the present application is an end-to-end model based on the Transformer architecture, which is used to detect text instances in images and associate multiple target text instances with the same attributes.

[0047] In some embodiments, when training the model of the present application, all parameters other than the vocabulary and the parameters related to the target task to be performed are frozen, so that only the vocabulary mapping layer and task-related structures are trained. During the fine-tuning phase, all learnable parameters are unfrozen, so that the model is fine-tuned based on the dataset corresponding to the target task.

[0048] exist Figure 1 Before step S101 shown, the method further includes step S104.

[0049] In step S104, the target model is pre-trained so that the target model can learn a universal feature representation.

[0050] Specifically, pre-training on a large public dataset allows the target model to learn universal feature representations that can capture the basic patterns in images and general properties of objects.

[0051] Refer to the following Figure 1 To illustrate, in step S101 , a training image set corresponding to a target task is obtained.

[0052] The target task is to detect and recognize text instances in images and associate multiple text instances with the same attributes.

[0053] The text instances with the same attributes refer to the association between these text instances, and the association includes but is not limited to:

[0054] 1) Relevance based on geographic location; for example, related to a specific region or place.

[0055] 2) Relevance based on the semantic content of the text; for example, relevance to a specific event, period, or context.

[0056] Those skilled in the art may define the relevance of text instances based on actual needs, and treat text instances with relevance as text instances with the same attributes by setting group labels corresponding to different attributes.

[0057] According to one embodiment, in a scenario of processing a map dataset, the association process links text instances (such as place names, landmarks, boundaries, etc.) in an image with their geographical locations on the map or other related information in the map.

[0058] In the context of historical maps, this linking might involve identifying text on the map and linking it to the map's semantic content, geographic data, or historical context.

[0059] The links include but are not limited to:

[0060] Linking text to geographic information: On historical maps, text often represents geographic locations, place names, or other geographic features. Linking problems may involve matching these text instances with actual locations on the map or with related geographic information.

[0061] Linking text to historical context: The text on the map may be associated with a specific historical event, period, or context. Linking questions may involve placing the text within a broader historical and cultural context to understand its meaning and significance.

[0062] Linking between text instances: In some cases, different text instances on a map may be related to each other. For example, the name of a region may be associated with a specific feature or event in that region. Linking problems may involve identifying and understanding the relationships between these text instances.

[0063] The training image set includes annotated data. The annotated information includes text and its corresponding entity or concept link information. This data can be obtained from an annotated dataset, for example, by parsing an annotation file in JSONL format, which contains images, text, and corresponding label information.

[0064] According to one embodiment, the method sets the group label through the following steps S105 to S106.

[0065] In step S105 , a corresponding group label is assigned to each text instance detected in the training image based on preset label definition information.

[0066] The tag definition information is used to determine which entities or concepts belong to the same group.

[0067] The group label is used to indicate which texts belong to the same group of entities or concepts.

[0068] The group label includes various information that can uniquely represent the corresponding group, such as a combination of numbers or letters.

[0069] In step S106 , based on the obtained group labels, annotation data corresponding to the training image is generated, where the annotation data includes the text instances contained in the training image and their group labels.

[0070] By using labeled data to train the model, the model learns how to identify the associations between text segments based on the labels.

[0071] Continue to refer to Figure 1 To illustrate, in step S102, a Bezier curve-based denoising training method is used to train the target model.

[0072] The method freezes parameters other than the target parameters in the target model to perform training on the part corresponding to the target parameters.

[0073] According to one embodiment, the target parameters include parameters related to the vocabulary and parameters related to the target task.

[0074] The vocabulary-related parameters include parameters related to the model's embedding layer or other parameters directly related to the vocabulary. Optionally, the vocabulary-related parameters include parameters of the word embedding layer of the expanded new vocabulary. This approach, after the vocabulary is expanded, can achieve excellent detection results for special characters with a small sample size.

[0075] According to one embodiment, the parameters related to the target task include parameters related to the model's detection structure, recognition structure, and linking structure. The detection structure is used to identify the location and shape of text regions in an image, the recognition structure is used to identify the text content corresponding to the text region, and the linking structure assigns a group label to the detected text instance.

[0076] The detection and recognition structures of the model are two parallel branches, and both structures output results simultaneously. The model optimization parameters corresponding to the recognition structure are shared with the detection structure, allowing the recognition structure to effectively collaborate with the detection structure to optimize the model parameters and effectively improve the model's detection performance.

[0077] By freezing all parameters in the target model except the target parameters, the model can focus on training the vocabulary and the parts related to the target task to be performed. Based on this training method, the model does not need to learn all parameters at the same time, thus improving the model training efficiency.

[0078] According to one embodiment, the method may determine which parameters need to be frozen by obtaining pre-stored freezing parameter indication information. Alternatively, the device may determine the parameters need to be frozen in real time based on user input.

[0079] According to one embodiment, the method includes steps S1021 to S1025 in which the method uses a Bezier curve-based denoising training method to train the target model in S102.

[0080] The Bezier curve-based denoising training method is targeted at text recognition tasks based on the Transformer architecture. By improving traditional denoising training, the method makes the model more stable and efficient when processing text of arbitrary shapes.

[0081] In step S1021 , the query of the noise reduction part is decomposed into a noise position query and a noise content query.

[0082] In step S1022 , a noise position query is generated using the four control points of the Bezier curve, and after adding noise to the Bezier control points, a new set of Bezier control points is obtained.

[0083] In step S1023 , the coordinates of multiple points are uniformly sampled from the generated Bezier curve using the Bezier control points to which noise is added.

[0084] In step S1024, the obtained coordinates are position-encoded and processed by a two-layer multilayer perceptron (MLP) to obtain a corresponding noise position query.

[0085] In step S1025 , a noise content query is initialized by a mask character sliding method to assist in the alignment of text content and position.

[0086] The Mask Character Sliding (MCS) method is used to assist in aligning text content and position.

[0087] Specifically, we first determine the number of valid characters t, which refers to the number of characters in the input sequence that actually make sense. Then, we perform an integer division operation To calculate the number of times each valid character should be cloned so that the total length T of the sequence is evenly distributed to each valid character. In addition, since T may not be completely divisible by t, there will be a remainder , indicating that there are k additional spaces to be allocated. In order to distribute these additional spaces fairly, an additional clone will be allocated to each of the first k valid characters in the sequence to ensure that the distribution of each character is as even as possible.

[0088] After obtaining the sliding characters, a masking operation is used to control the number of consecutive characters. This means that a portion of the consecutive characters are flipped to background labels with a certain probability. After processing the positive part, noise is added to the characters in the positive and negative parts, causing these characters to flip to other characters with a probability of λ. These characters are then embedded to obtain the noise content query, where all the characters initialized in the negative part are background.

[0089] According to one embodiment, when performing a task using a trained model, the portion of applying the Bezier curve-based denoising training method is removed during model inference.

[0090] Among them, the part that applies the Bezier curve-based denoising training method can be removed during model inference by setting the denoising training part to an inactive state in the model configuration file or code, or directly disabling the layer in the model that applies denoising training.

[0091] By removing the part of the model that applies the Bezier curve-based denoising training method during the model inference phase, the efficiency of the inference process is ensured without adding additional computational burden.

[0092] According to one embodiment, the target model is a model based on the Transformer architecture, the structure of which includes a backbone network, an encoder, and a decoder. This embodiment trains the model by performing the following steps:

[0093] The backbone network extracts global features from the input image to obtain corresponding global and local features. Optionally, the input image is preprocessed before feature extraction. This may include resizing, normalization, contrast enhancement, and other operations to ensure that the model receives input in a consistent format. The global features obtained by the backbone network help the model understand the global context of the image.

[0094] Next, the encoder further extracts global features of the input image. The global features obtained by the backbone network help the model understand the global context of the image.

[0095] Next, the decoder further extracts local features from the input image based on the global features to obtain the corresponding local features. To gain a more detailed understanding of the image content, the model needs to further extract local features. This is usually achieved by using deeper network structures that can identify detailed information such as edges, corners, and textures in the image. Local features are particularly important for text detection because they help the model recognize the outline and shape of text.

[0096] Next, a Bezier curve-based denoising training method is applied to the decoder to train the model to effectively handle noise and interference in the image. During model inference, this Bezier curve-based denoising training is removed to ensure efficient inference without adding additional computational burden.

[0097] The local features output by the decoder layer are used to locate the text regions in the image and identify the corresponding text content.

[0098] The Transformer-based model in this embodiment also includes a detection head, a recognition head, and a linking head. The method further includes: detecting the location of text instances in the image using the detection head; identifying the text content of the text instances using the recognition head; and assigning a group label to the detected text instances using the linking head.

[0099] The detection and recognition heads of the model are two parallel branches that output results simultaneously. The model optimization parameters corresponding to the recognition head are shared with the detection head. The recognition head can effectively collaborate with the detection structure to optimize the model parameters and effectively improve the model's detection performance.

[0100] Continue to refer to Figure 1 To illustrate, in step S103 , after the frozen parameters are unfrozen, the target model is fine-tuned based on the target training image set.

[0101] Fine-tuning is a process that preserves the useful features learned by the model during pre-training while adapting it to the new task with a small amount of adjustments. This process can significantly improve the performance of the model on a specific task while reducing training cost and time.

[0102] Optionally, during the fine-tuning process, the pre-trained target model is structurally adjusted according to the requirements of the target task to be performed.

[0103] According to the method of this embodiment, a denoising method based on Bezier curves is adopted to train the model, and then the trained model is used to detect text instances contained in the input image and associate the associated text instances. Compared with traditional text detection and recognition schemes, information on whether the text instances are associated with each other is provided, which is helpful for subsequent text information processing; in the training phase, parameters other than the vocabulary and parameters related to the task to be executed are frozen, so that only the vocabulary mapping layer and task-related structures are trained, and in the fine-tuning phase, all learnable parameters are unfrozen to fine-tune the model based on the data set corresponding to the text scene task. This training method reduces the cost of training and fine-tuning on the target data set, and can complete the detection of hundreds of targets in text images within a few hours. At the same time, the vocabulary can be easily expanded, further improving the efficiency of model training.

[0104] Figure 2 A flowchart of a method for associating texts according to an embodiment of the present application is shown. Figure 2 The steps shown include step S201 to step S204.

[0105] Reference Figure 2 , in step S201, the target scene image to be processed is obtained.

[0106] The target scene images may include photos or video frames of various types of scenes. For example, for urban scenes, the target scene images may include images showing urban buildings, streets, traffic, road signs, etc. For indoor scenes, the target scene images may include images showing indoor environments such as homes, offices, and stores.

[0107] In step S202, text detection is performed on text instances contained in the target scene image using the trained target model.

[0108] The trained target model is obtained by training and fine-tuning the above steps S101 to S103 based on the data set corresponding to the task.

[0109] Specifically, when performing text location detection, the trained target model is used to perform boundary detection to predict the vertices or boundary points of the text region. The detected boundary points are then converted into the location information of one or more text regions. For example, the detected boundary points can be converted into adaptive text region representations, such as polygons or curves.

[0110] In step S203 , the group to which the detected text instance belongs is determined based on preset group label information.

[0111] In step S204 , multiple text instances belonging to the same group are associated.

[0112] The association processing includes but is not limited to linking multiple text instances belonging to the same group.

[0113] According to the method of this embodiment, the trained model is used to detect text instances contained in the input image and the associated text instances are associated.

[0114] Refer to the following Figure 3 The exemplary model structure shown is used to illustrate the method of the embodiment of the present application. The model of this example is a model based on the Transformer structure.

[0115] Reference Figure 3 The model in this example mainly consists of three parts: the backbone network, the encoder, and the decoder. The encoder contains multiple encoder layers, and the decoder contains multiple decoder layers. These three parts are introduced below.

[0116] 1) Backbone network: The backbone network in this example is the ViTAE-v2 backbone network, which processes the input image and extracts global features. These features contain the basic visual information of the image and provide the basis for subsequent text detection and recognition.

[0117] 2) Encoder: The encoder further extracts global features of the input image.

[0118] 3) Decoder: The decoder extracts local features of the input image. The local features output by the decoder layer are used to locate the text area in the input image and recognize the text content.

[0119] The Bezier curve-based denoising training method is applied at the decoder layer to train the model to learn how to effectively handle noise and interference in images, thereby further optimizing the model's performance. Furthermore, the denoising training method helps the model better understand and handle changes in text shape, such as curved or tilted text, thereby providing more reliable detection or recognition results in practical applications. When performing model inference, the application of the Bezier curve-based denoising training method is removed to ensure the efficiency of the inference process without adding additional computational burden.

[0120] The transformer-based model in this example also includes a detection head, a recognition head, and a linking head. The detection head is used to detect the location of text instances in the image, while the recognition head is used to identify the text content of the detected text instances. The linking head is used to assign group labels to the detected text instances, thereby linking detected text instances with the same attributes.

[0121] This example uses a hybrid training method to train Figure 3The model shown. The process of this hybrid training method includes:

[0122] 1) Freeze the parameters other than the vocabulary and the parameters related to the target task to be executed to train the parameters of the word embedding layer of the expanded new vocabulary and the parameters of the detection head, recognition head, and linking head related to the target task;

[0123] 2) Obtain the target data set for the task to be performed;

[0124] 3) Unfreeze all learnable parameters and fine-tune the model based on the target dataset.

[0125] After unfreezing all learnable parameters, the model will rescale all parameters, including the weights of the embedding layers, detection heads, recognition heads, and linking heads. This process allows the model to better adapt to the specific target dataset. Fine-tuning typically requires several thousand steps of training to ensure good performance on the new dataset. Through fine-tuning, the model not only leverages the knowledge learned during pre-training but also further optimizes based on the characteristics of the target dataset, thereby improving classification accuracy.

[0126] The detection and recognition heads of the model are two parallel branches that output results simultaneously. The model optimization parameters corresponding to the recognition head are shared with the detection head. The recognition head can effectively collaborate with the detection structure to optimize the model parameters and effectively improve the model's detection performance.

[0127] When performing the text detection linking task, the input image is input into the trained model of this example. The detection head of the model is used to predict the text examples and their positions contained in the input image based on the local features output by the decoder. The linking head is used to assign the group labels to the detected text instances, and the detection results of one or more text instances in the input image and the group labels corresponding to each text instance are used as the model output results.

[0128] Figure 4 A schematic diagram of the structure of a model training device provided in an embodiment of the present application is shown. The device includes: a device for acquiring a training image set corresponding to a target task (hereinafter referred to as "training image acquisition device 101"), a device for training a target model using a Bezier curve-based denoising training method (hereinafter referred to as "denoising training device 102"), and a device for thawing frozen parameters and fine-tuning the target model based on the target training image set (hereinafter referred to as "model fine-tuning device 103").

[0129] According to one embodiment, the device further comprises a pre-training device.

[0130] The pre-training device pre-trains the target model so that the target model can learn a universal feature representation.

[0131] Specifically, the pre-training device pre-trains the target model on a large public dataset, allowing the target model to learn universal feature representations that can capture the basic patterns in images and general properties of objects.

[0132] Refer to the following Figure 4 To explain, the training image acquisition device 101 acquires a training image set corresponding to a target task.

[0133] The target task is to detect and recognize text instances in images and associate multiple text instances with the same attributes.

[0134] The text instances with the same attributes refer to the association between these text instances, and the association includes but is not limited to:

[0135] 1) Relevance based on geographic location; for example, related to a specific region or place.

[0136] 2) Relevance based on the semantic content of the text; for example, relevance to a specific event, period, or context.

[0137] Those skilled in the art may define the relevance of text instances based on actual needs, and treat text instances with relevance as text instances with the same attributes by setting group labels corresponding to different attributes.

[0138] According to one embodiment, in a scenario of processing a map dataset, the association process links text instances (such as place names, landmarks, boundaries, etc.) in an image with their geographical locations on the map or other related information in the map.

[0139] In the context of historical maps, this linking might involve identifying text on the map and linking it to the map's semantic content, geographic data, or historical context.

[0140] The links include but are not limited to:

[0141] Linking text to geographic information: On historical maps, text often represents geographic locations, place names, or other geographic features. Linking problems may involve matching these text instances with actual locations on the map or with related geographic information.

[0142] Linking text to historical context: The text on the map may be associated with a specific historical event, period, or context. Linking questions may involve placing the text within a broader historical and cultural context to understand its meaning and significance.

[0143] Linking between text instances: In some cases, different text instances on a map may be related to each other. For example, the name of a region may be associated with a specific feature or event in that region. Linking problems may involve identifying and understanding the relationships between these text instances.

[0144] The training image set includes annotated data. The annotated information includes text and its corresponding entity or concept link information. This data can be obtained from an annotated dataset, for example, by parsing an annotation file in JSONL format, which contains images, text, and corresponding label information.

[0145] According to one embodiment, the apparatus sets the group label by operating a label assigning device and a annotation generating device.

[0146] The label assigning device assigns a corresponding group label to each text instance detected in the training image based on preset label definition information.

[0147] The tag definition information is used to determine which entities or concepts belong to the same group.

[0148] The group label is used to indicate which texts belong to the same group of entities or concepts.

[0149] The group label includes various information that can uniquely represent the corresponding group, such as a combination of numbers or letters.

[0150] The annotation generating device generates annotation data corresponding to the training image based on the obtained group labels, and the annotation data includes the text instances included in the training image and their group labels.

[0151] By using labeled data to train the model, the model learns how to identify the associations between text segments based on the labels.

[0152] The noise reduction training device 102 uses a noise reduction training method based on Bezier curves to train the target model.

[0153] The noise reduction training device 102 performs training on the part corresponding to the target parameter by freezing the parameters other than the target parameter in the target model.

[0154] According to one embodiment, the target parameters include parameters related to the vocabulary and parameters related to the target task.

[0155] The vocabulary-related parameters include parameters related to the model's embedding layer or other parameters directly related to the vocabulary. Optionally, the vocabulary-related parameters include parameters of the word embedding layer of the expanded new vocabulary. This approach, after the vocabulary is expanded, can achieve excellent detection results for special characters with a small sample size.

[0156] According to one embodiment, parameters related to the target task include parameters related to the model's detection structure, recognition structure, and linking structure. The detection structure is used to identify the location and shape of text regions in an image; the recognition structure is used to identify the text content corresponding to the text region; and the linking structure assigns group labels to detected text instances.

[0157] The detection and recognition structures of the model are two parallel branches, and both structures output results simultaneously. The model optimization parameters corresponding to the recognition structure are shared with the detection structure, allowing the recognition structure to effectively collaborate with the detection structure to optimize the model parameters and effectively improve the model's detection performance.

[0158] By freezing all parameters in the target model except the target parameters, the model can focus on training the vocabulary and the parts related to the target task to be performed. Based on this training method, the model does not need to learn all parameters at the same time, thus improving the model training efficiency.

[0159] According to one embodiment, the device may determine which parameters need to be frozen by acquiring pre-stored freezing parameter indication information. Alternatively, the device may determine the parameters need to be frozen in real time based on user input.

[0160] According to one embodiment, the noise reduction training device 102 further includes a query decomposition device, a control point generation device, a coordinate point sampling device, a position query acquisition device, and a noise initialization device.

[0161] The Bezier curve-based denoising training method is targeted at text recognition tasks based on the Transformer architecture. By improving traditional denoising training, the method makes the model more stable and efficient when processing text of arbitrary shapes.

[0162] The query decomposition device decomposes the query of the noise reduction part into a noise position query and a noise content query.

[0163] The control point generating device generates a noise position query using four control points of the Bezier curve, and obtains a set of new Bezier control points after adding noise to the Bezier control points.

[0164] The coordinate point sampling device uses Bezier control points with added noise to uniformly sample multiple point coordinates from the generated Bezier curve.

[0165] The location query acquisition device obtains the corresponding noise location query by performing location encoding and two-layer multilayer perceptron (MLP) processing on the obtained coordinates.

[0166] The noise initialization device initializes the noise content query through the mask character sliding method.

[0167] The noise device is initialized by using the Mask Character Sliding (MCS) method to assist in the alignment of text content and position.

[0168] According to one embodiment, in the stage of performing a task using a trained model, when performing model inference, the apparatus removes the portion of applying a Bezier curve-based noise reduction training method.

[0169] Among them, the part that applies the Bezier curve-based denoising training method can be removed during model inference by setting the denoising training part to an inactive state in the model's configuration file or code or directly disabling the layer in the model that applies denoising training.

[0170] By removing the part of the model that applies the Bezier curve-based denoising training method during the model inference phase, the efficiency of the inference process is ensured and additional computational burden is avoided.

[0171] Continue to refer to Figure 4 To explain, after unfreezing the frozen parameters, the model fine-tuning device 103 fine-tunes the target model based on the target training image set.

[0172] Fine-tuning is a process that preserves the useful features learned by the model during pre-training while adapting it to the new task with a small amount of adjustments. This process can significantly improve the performance of the model on a specific task while reducing training cost and time.

[0173] Optionally, during the fine-tuning process, the model fine-tuning device 103 performs structural adjustment processing on the pre-trained target model according to the requirements of the target task to be performed.

[0174] According to the device of this embodiment, a denoising method based on Bezier curves is adopted to train the model, and then the trained model is used to detect text instances contained in the input image and associate the associated text instances. Compared with the traditional text detection and recognition scheme, it provides information on whether each text instance is associated with each other, which is helpful for subsequent text information processing; in the training stage, parameters other than the vocabulary and the parameters related to the task to be executed are frozen, so that only the vocabulary mapping layer and the task-related structure are trained, and in the fine-tuning stage, all learnable parameters are unfrozen to fine-tune the model based on the data set corresponding to the text scene task. This training method reduces the cost of training and fine-tuning on the target data set, and can complete the detection of hundreds of targets in text images within a few hours. At the same time, the vocabulary can be easily expanded, further improving the efficiency of model training.

[0175] Figure 5 A schematic structural diagram of an apparatus for associating texts according to an embodiment of the present application is shown.

[0176] The device includes: a device for acquiring a target scene image to be processed (hereinafter referred to as "input image acquisition device 201"), a device for performing text detection on text instances contained in the target scene image using a trained target model (hereinafter referred to as "text detection device 202"), a device for determining the group of the detected text instance based on preset group label information (hereinafter referred to as "group determination device 203"), and a device for associating multiple text instances belonging to the same group (hereinafter referred to as "association processing device 204").

[0177] The input image acquisition device 201 acquires the target scene image to be processed.

[0178] The target scene images may include photos or video frames of various types of scenes. For example, for urban scenes, the target scene images may include images showing urban buildings, streets, traffic, road signs, etc. For indoor scenes, the target scene images may include images showing indoor environments such as homes, offices, and stores.

[0179] The text detection device 202 uses the trained target model to perform text detection on the text instances contained in the target scene image.

[0180] The trained target model is obtained by training and fine-tuning the training image acquisition device 101, the noise reduction training device 102 and the model fine-tuning device 103 based on the data set corresponding to the task.

[0181] Specifically, when the text detection device 202 detects text positions, it uses the trained target model to perform boundary detection to predict vertices or boundary points of the text region, and then converts the detected boundary points into position information of one or more text regions.

[0182] The group determining means 203 determines the group to which the detected text instance belongs based on preset group label information.

[0183] The association processing means 204 performs association processing on multiple text instances belonging to the same group.

[0184] The association processing includes but is not limited to linking multiple text instances belonging to the same group.

[0185] According to the apparatus of this embodiment, the trained model is used to detect text instances contained in the input image and to perform association processing on the text instances with association.

[0186] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present application. The method corresponding to the electronic device can be the model training method or the method for associating text in the aforementioned embodiment, and its principle of solving the problem is similar to that of the method. The electronic device provided in an embodiment of the present application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the methods and / or technical solutions of the aforementioned multiple embodiments of the present application.

[0187] The electronic device may be a user device, or a device formed by integrating a user device and a network device via a network, or an application running on the above device. The user device includes but is not limited to various terminal devices such as computers, mobile phones, tablets, smart watches, and bracelets. The network device includes but is not limited to network hosts, single network servers, multiple network server sets, or cloud computing-based computer collections, and can be used to implement some of the processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing (Cloud Computing), where cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computers.

[0188] Figure 6The structure of a device suitable for implementing the method and / or technical solution in the embodiment of the present application is shown. The device 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage part 1208 into the random access memory (RAM) 1203. Various programs and data required for system operation are also stored in RAM 1203. CPU 1201, ROM 1202 and RAM 1203 are connected to each other through a bus 1204. Input / output (I / O) interface 1205 is also connected to bus 1204.

[0189] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, a touch screen, a microphone, an infrared sensor, and the like; an output section 1207 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), an LED display, an OLED display, and a speaker; a storage section 1208 including one or more computer-readable media such as a hard disk, an optical disk, a magnetic disk, and a semiconductor memory; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication section 1209 performs communication processing via a network such as the Internet.

[0190] In particular, the methods and / or embodiments of the present application can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 1201, the above-mentioned functions defined in the method of the present application are performed.

[0191] Another embodiment of the present application further provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of the present application.

[0192] Specifically, the present embodiment can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.

[0193] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0194] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0195] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0196] The flow chart or block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the equipment, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code include one or more executable instructions for realizing the logical function of the specification. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs the function or operation of the specification, or can be implemented with a combination of dedicated hardware and computer instructions.

[0197] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0198] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or page components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0199] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0200] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0201] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute some steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code.

[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

[0203] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device through software or hardware. Terms such as "first" and "second" are used to indicate names and do not imply any particular order.

Claims

1. A model training method, comprising: Obtaining a training image set corresponding to a target task, wherein the target task is to detect text instances in an image and associate multiple target text instances with the same attributes, wherein the association includes linking multiple text instances belonging to the same group; Training the target model using a Bezier curve-based denoising training method, wherein the method freezes parameters other than the target parameters in the target model to train the portion corresponding to the target parameters; After unfreezing the frozen parameters, fine-tune the target model based on the training image set; The method further comprises: Based on the preset label definition information, a corresponding group label is assigned to each text instance detected in the training image; Based on the obtained group labels, annotation data corresponding to the training image is generated, where the annotation data includes the text instances contained in the training image and their group labels.

2. The method according to claim 1, wherein The method of using the Bezier curve-based denoising training method to train the target model includes: Decompose the noise reduction part query into noise location query and noise content query; Generate noise position query using the four control points of the Bezier curve, and obtain a new set of Bezier control points after adding noise to the Bezier control points; Using the Bezier control points with added noise, uniformly sample multiple point coordinates from the generated Bezier curve; By performing position encoding and two-layer MLP processing on the obtained coordinates, the corresponding noise position query is obtained; Noisy content queries are initialized through a masked character sliding method to assist in the alignment of text content and location.

3. The method according to claim 1 or 2, wherein: The target model is a model based on the Transformer architecture, the structure of which includes a backbone network, an encoder, and a decoder. The method trains the model by the following steps: Perform global feature extraction on the input image in the backbone network to obtain corresponding global features and local features; Further extract global features of the input image in the encoder; In the decoder, based on the global features, the input image is further subjected to local feature extraction processing to obtain corresponding local features; A Bezier curve-based denoising training method is applied in the decoder to train the model.

4. The method according to claim 3, wherein: The transformer-based model further includes a detection head, a recognition head, and a linking head, and the method further includes: Detecting locations of text instances in the image using the detection head; Identifying the text content of the text instance by the identification head; The detected text instance is assigned a group label through the link header.

5. A method for detecting and associating text, wherein: The method comprises: Acquire the target scene image to be processed; Use the trained target model to perform text detection on the text instances contained in the target scene image; Based on the preset group label information, determine the group to which the detected text instance belongs; Associating multiple text instances belonging to the same group; Wherein, the target model is obtained by training using the model training method described in any one of claims 1 to 4.

6. The method according to claim 5, wherein: The method further comprises: When performing model inference, remove the part that applies the Bezier curve-based denoising training method.

7. A model training device, wherein: The device comprises: A device for obtaining a training image set corresponding to a target task, wherein the target task is to detect text instances in an image and associate multiple target text instances with the same attributes, wherein the association includes linking multiple text instances belonging to the same group; A device for training a target model using a Bezier curve-based noise reduction training method, wherein the device performs training on a portion corresponding to the target parameter by freezing parameters other than the target parameter in the target model; a device for fine-tuning a target model based on a target training image set after unfreezing the frozen parameters; Wherein, the device further includes: a label assigning device for assigning a corresponding group label to each text instance detected in the training image based on preset label definition information; The annotation generating device is used to generate annotation data corresponding to the training image based on the obtained group label, and the annotation data includes the text instance contained in the training image and its group label.

8. An apparatus for detecting and associating text, wherein: The device comprises: Device for acquiring a target scene image to be processed; a device for performing text detection on text instances contained in a target scene image using a trained target model; means for determining the group to which the detected text instance belongs based on preset group label information; means for associating a plurality of text instances belonging to the same group; Wherein, the target model is obtained by training using the model training method described in any one of claims 1 to 4.

9. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4, or the method according to claim 5 or 6.

10. A computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 4, or to perform the method according to claim 5 or 6.

11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 4, or performs the method according to claim 5 or 6.

Citation Information

Patent Citations

  • Model training method and device for determining text relevancy, electronic equipment and readable storage medium

    CN111832290A

  • Policy text correlation analysis method and system

    CN112580348A