Model error acquisition method and device, electronic equipment, readable medium and computer program

By performing data enhancement and analysis of text recognition results on images of multimodal generative large models, the problem of model generation illusion is solved, and the recognition of model errors and accuracy is improved.

CN120299058APending Publication Date: 2025-07-11BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510450851.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing multimodal generative models are prone to hallucinations when generating content, and cannot effectively avoid errors or inaccurate generation. The complexity makes behavior difficult to predict and understand, resulting in the generated images, text or audio content that may contain non-compliant with real-world logic or false information.

Method used

By augmenting the original image data, multiple enhancement images are generated, and these images are input into the trained multimodal generation model to obtain the text recognition results corresponding to each enhancement image, and the error of the model is obtained based on the differences between these results.

Benefits of technology

It can identify errors in multimodal generative large models, avoid using inaccurate models, and improve the accuracy and reliability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299058A_ABST
    Figure CN120299058A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model error acquisition method and device, electronic equipment, a readable medium and a computer program, and the method comprises the steps: carrying out the data enhancement of an original image, and obtaining a plurality of enhanced images; the multiple enhanced images are input into a trained multi-modal generation type large model, text recognition results corresponding to the enhanced images are obtained, and the multi-modal generation type large model is used for carrying out text recognition on the input images; and based on the difference between the text recognition results corresponding to the enhanced images, obtaining the error of the multi-modal generative large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method, device, electronic device, readable medium, and computer program for obtaining model errors. Background Art

[0002] The method for a multi-modal generative (Transformer) large model to process and recognize images combines the capabilities of an image encoder (Vision Transformer, ViT) and a traditional Transformer architecture. However, there is a common problem similar to human hallucinations in multi-modal generative large models. This "hallucination" refers to errors or inaccuracies that occur when the model generates content, resulting in generated image, text, or audio content that may contain elements that do not conform to the logic or physical laws of the real world, false or misleading information, including racial, gender, or other forms of bias, etc.

[0003] Currently, it is impossible to avoid the generation of hallucinations in generative multi-modal generative large models. On the one hand, real-world data sets are almost always imperfect and may contain noise, biases, and outliers. These imperfections will be learned by the model and may be amplified. On the other hand, generative multi-modal generative large models may learn some overly generalized rules from the training data, and these rules may not apply in specific situations. If the distribution of the training data and test data (or data in the actual application scenario) of the multi-modal generative large model is inconsistent, the multi-modal generative large model may generate more serious hallucinations on new data.

[0004] In addition, multi-modal generative large models are usually very complex, which makes their behavior difficult to predict and it is difficult to fully understand the process of generating content by them. The optimization objectives of multi-modal generative large models may not be exactly the same as the results expected by humans, which may lead to unsatisfactory results in some cases. Therefore, how to obtain the errors of multi-modal generative large models is a technical problem that needs to be solved in related technologies. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a method, device, electronic device, readable medium, and computer program for obtaining model errors, which can solve the problem of how to obtain model errors.

[0006] To solve the above technical problems, the embodiments of this application are implemented through the following aspects.

[0007] In a first aspect, an embodiment of the present application provides a method for obtaining model error, including: performing data augmentation on an original image to obtain multiple augmented images; inputting the multiple augmented images into a trained multi-modal generative large model to obtain text recognition results corresponding to each of the augmented images, where the multi-modal generative large model is used to perform text recognition on the input image; and obtaining the error of the multi-modal generative large model based on the differences between the text recognition results corresponding to each of the augmented images.

[0008] In a second aspect, an embodiment of the present application provides a device for obtaining model error, including: an image preprocessing module for performing data augmentation on an original image to obtain multiple augmented images; a recognition module for inputting the multiple augmented images into a trained multi-modal generative large model to obtain text recognition results corresponding to each of the augmented images, where the multi-modal generative large model is used to perform text recognition on the input image; and an error control module for obtaining the error of the multi-modal generative large model based on the differences between the text recognition results corresponding to each of the augmented images.

[0009] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor, and computer executable instructions stored on the memory and executable on the processor, where the computer executable instructions, when executed by the processor, implement the method for obtaining model error described in the first aspect above.

[0010] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium for storing computer executable instructions, where the computer executable instructions, when executed by a processor, implement the method for obtaining model error described in the first aspect above.

[0011] In a fifth aspect, an embodiment of the present application provides a computer program product, where the computer program product includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the method for obtaining model error described in the first aspect above.

[0012] In the embodiment of the present application, by performing data augmentation on the original image to obtain multiple augmented images, and respectively inputting the multiple augmented images into a trained multi-modal generative large model to obtain text recognition results corresponding to each of the augmented images, based on the differences between the text recognition results corresponding to each of the augmented images, the error of the multi-modal generative large model can be obtained, so as to know whether the multi-modal generative large model is accurate and avoid inaccurate text recognition results caused by using a multi-modal generative large model with a large error. Description of the Drawings

[0013] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0014] Figure 1 Fig. 4 shows a schematic flowchart of a method for obtaining model error provided by an embodiment of the present application;

[0015] Figure 2 Fig. 8 shows a schematic flowchart of a method for obtaining model error provided by another embodiment of the present application;

[0016] Figure 3 Fig. 12 shows a schematic structural diagram of a device for obtaining model error provided by an embodiment of the present application;

[0017] Figure 4 Fig. 16 is a schematic hardware structure diagram of an electronic device for executing a method for obtaining model error provided by an embodiment of the present application. Detailed implementation manners

[0018] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0019] Image document recognition and optical character recognition (OCR) is a technology that can recognize and extract text content from images. OCR technology is usually used to convert the text in scanned paper documents, photos, or PDF files into a machine-editable format, such as a plain text file or a Word document.

[0020] The multi-modal generative large model in the embodiments of the present application is a multi-modal generative large model, and the function of image document recognition is realized through the multi-modal generative large model. This multi-modal generative large model is a deep learning model that can process and generate data containing multiple modalities (such as text, images, audio, etc.). The model internally maps data of different modalities to a common embedding space to capture the mutual relationships between them. These models usually contain hundreds of millions or even billions of parameters and are trained on a large amount of multi-modal data.

[0021] Document recognition using a multi-modal generative large model generally includes the following steps:

[0022] 1) Document scanning.

[0023] In this step, a scanner or a high-resolution camera can be used to capture an image of the document. To ensure accurate subsequent recognition, it is necessary to ensure that the image is clear, has a moderate contrast, and has no obvious tilt or distortion.

[0024] 2) Image preprocessing.

[0025] In this step, the scanned image is preprocessed, such as denoising, binarization, skew correction, etc., to improve the accuracy of OCR.

[0026] 3) Image segmentation.

[0027] In this step, the preprocessed image can be segmented into multiple image patches, and each patch may contain a line of text or a few words. In this step, the image is segmented into blocks of a fixed size (e.g., blocks of 16x16 or 32x32 pixels). These image patches are regarded as elements in a sequence.

[0028] 4) Input to the multi-modal generative large model.

[0029] In this step, each image patch can be input into the multi-modal generative large model. The multi-modal generative large model processes these image patches and extracts features. In this step, the multi-modal generative large model can utilize the visual and text features learned during its training to perform character recognition on the image patches. The multi-modal generative large model may use a self-attention mechanism to understand the character order and context in the image patches.

[0030] In addition, since the Transformer model itself does not have the ability to handle sequence order, position encoding can be added to each image patch to preserve spatial information.

[0031] In the multi-modal generative large model, each image patch is flattened and converted into a vector of a fixed dimension through a linear layer, and this process is called linear embedding.

[0032] These embedding vectors are then fed into the Transformer image encoder. The image encoder consists of multiple self-attention layers and feed-forward networks and can capture the relationships between image patches.

[0033] After being processed by their respective encoders, the images and text are mapped into a common embedding space. In this space, the features of the images and text can be compared and combined with each other. The Transformer language decoder uses the cross-attention mechanism, enabling the model to consider both image and text information simultaneously. For example, text tokens can focus on the most relevant parts of the image, and vice versa.

[0034] 5) Text output.

[0035] The multimodal generative large model can output the text sequences corresponding to each image patch.

[0036] In some embodiments, these text sequences may be post-processed to correct recognition errors and improve text coherence.

[0037] 6) Post-processing.

[0038] In this step, the text output by the model can be post-processed, such as proofreading, formatting, removing extra spaces and punctuation marks, etc.

[0039] In addition, if needed, layout analysis can be performed on the text to restore the structure of the original document.

[0040] The image encoder (Vision Transformer, ViT) is a model that applies the Transformer architecture to image recognition tasks. The key features of ViT are as follows:

[0041] 1) No convolution: Different from traditional convolutional neural networks, ViT does not use convolutional layers and instead relies entirely on the self-attention mechanism to process images.

[0042] 2) Self-attention: ViT learns the relationships between image patches through the self-attention mechanism, enabling the model to effectively capture global context information.

[0043] 3) Pre-training and transfer learning: ViT is usually pre-trained on large-scale image datasets and can then be transferred to different visual tasks.

[0044] Through these methods, the multimodal generative large model can effectively process and recognize images, and align the image content with the text description, thus demonstrating powerful performance in various tasks, such as image caption generation, visual question answering, and image-text retrieval, etc.

[0045] However, current generative large models cannot avoid generating hallucinations. Therefore, multimodal generative large models also cannot avoid generating hallucinations. How to identify the errors existing in multimodal generative large models is a technical problem to be solved in related technologies. In response to this problem, the embodiments of the present application provide a scheme for obtaining errors of a multimodal generative large model. The following describes the scheme for obtaining errors of the multimodal generative large model provided by the embodiments of the present application with reference to the accompanying drawings.

[0046] Figure 1 FIG. 1 shows a flowchart of a method for obtaining model errors provided by an embodiment of the present application. This method can be executed by an electronic device, which can be a terminal or a server. In other words, the method can be executed by software installed in the electronic device or the hardware of the electronic device. As Figure 1 shown, the method may include the following steps:

[0047] S110: Perform data augmentation on the original image to obtain multiple enhanced images.

[0048] In the embodiments of the present application, an image preprocessing module can be used to perform data augmentation on the original image to obtain enhanced images. Data augmentation refers to augmenting data by applying a series of transformations to the original data. For example, operations such as cropping, rotation, and resizing can be used to transform the original image to obtain enhanced images. These transformations can be implemented through interpolation functions, which are used to smoothly interpolate pixel values during the transformation process.

[0049] In the embodiments of the present application, the original image can be an image input by the user.

[0050] In some embodiments, enhanced images can be obtained by randomly cropping the original image. In these embodiments, a new sub-image can be randomly cropped from the original image. When cropping, nearest neighbor interpolation or bilinear interpolation can be used to calculate the pixel values in the new image. Nearest neighbor interpolation directly takes the pixel value in the original image that is closest to the new position, while bilinear interpolation considers the surrounding four pixel values to calculate the pixel value at the new position.

[0051] In some other embodiments, an enhanced image can be obtained by rotating the original image around a center point by a predetermined angle. Herein, the predetermined angle can be any angle, and the specific value can be determined according to the actual application. In these embodiments, when rotating, bilinear interpolation or bicubic interpolation can be used to calculate the pixel values in the new image. Bicubic interpolation calculates the pixel value at the new position considering the surrounding 16 pixel values, so it is smoother than bilinear interpolation but has a higher computational cost.

[0052] In still some other embodiments, the original image can be enlarged or reduced to obtain an enhanced image of the original image. In these embodiments, image enhancement is performed by changing the size of the original image, usually by enlarging or reducing. In these embodiments, when scaling, bilinear interpolation or bicubic interpolation can be used. Bilinear interpolation considers the surrounding four pixel values when calculating the pixel values in the new image, while bicubic interpolation considers the surrounding 16 pixel values.

[0053] In the PIL (Pillow) library of Python, the interpolation functions in the above various embodiments can be implemented through the Image.resize() and Image.rotate() methods.

[0054] In the embodiments of the present application, the above multiple enhanced images can be obtained by enhancing in the above one way, or can be obtained by enhancing in the above multiple ways.

[0055] For example, within the range of the length and width of the original image, 5 sets of cropping parameters are randomly selected, which are the displacements of the upper left corner and the lower right corner of the cropping frame relative to the original image. For example, [20 pixels downward and 50 pixels to the right at the upper left corner, 0 pixels upward and 10 pixels to the left at the lower right corner]. Within the range of plus or minus 30 degrees, 5 sets of rotation angles are randomly selected. For example, [rotate 5 degrees counterclockwise][rotate 25 degrees clockwise]. Within the range of the ratio [0.3, 3.0], 5 sets of rotation angles are randomly selected. For example, [scale to 0.5 times the length and width of the original image][scale to 1.5 times the length and width of the original image]. The original image is called z_1, and the enhanced images obtained by the above 15 sets of data augmentation strategies can be called z_2,...., z_16.

[0056] S120, input the multiple enhanced images into the trained multi-modal generative large model to obtain text recognition results corresponding to the respective enhanced images, wherein the multi-modal generative large model is used for text recognition of the input image.

[0057] In the embodiments of the present application, the multimodal generative large model can be a multimodal language model. The multimodal language model is a language model based on an autoregressive model of multiple data forms. Among them, the probability of a sequence x appearing in natural data is the probability jointly calculated based on the autoregressive model and the non-autoregressive model:

[0058]

[0059] The multimodal generative large model uses the trained VIT to act on the image z to obtain supplementary information of the language information x.

[0060] The prefix of x is a question string used to guide the multimodal large model to answer questions related to the text in the image. For example, x = "What is the text in the figure? Please answer line by line from top to bottom."

[0061] The image encoder of the multimodal generative large model maps each fixed-size block of the image z (for example, a block of 16x16 or 32x32 pixels) into a vector with the same dimension as the language decoder. For example, for a language decoder with 6 billion parameters, the vector dimension of each fixed-size image block and each word is 4096.

[0062] The output layer of the multimodal generative large model is usually responsible for converting the vector representation inside the model into the final prediction result. For example, in the multimodal generative large model, it is converted into a semantically meaningful output (such as the probability distribution of the next word), or into the class probability in a classification task.

[0063] The beginning of the output layer of the multimodal generative large model is a linear layer, also known as a single-layer version of the fully connected layer or the multi-layer perceptron (MLP). This linear layer maps the output vector of the last Transformer layer to a vector with a dimension equal to the vocabulary size of the multimodal generative large model (for the language model) and the key information vocabulary (for the classification task). The output of the linear layer is usually the unnormalized log probability, which is then converted into a normalized probability distribution through the Softmax function.

[0064] One or more augmented images obtained by randomly selecting data augmentation strategies can be input into the multimodal generative large model. Let z_n represent such an image. The process of the multimodal generative large model generating the language sequence x from the image information z is as follows:

[0065] Assume that for the sequence of the first t words, the output vector of the last Transformer layer is h_t. Due to the autoregressive nature of the language model, h is the result of transforming the first t words and all key information using the function F, that is:

[0066] ht = F([x1, x2, …, x t-1 , {z1, z2, …, z n}).

[0067] Among them, the function represented by F is a neural network with a multi-layer Transformer structure.

[0068] After that, the multi-modal generative large model executes the following steps through the output layer:

[0069] 1) Linear transformation.

[0070] Assume that the output vector of the last Transformer layer is h_t. The linear layer performs transformation through the weight matrix w and the bias vector b:

[0071] y t = wh t + b

[0072] 2) Apply the Softmax function.

[0073] Pass the output z of the linear transformation through the Softmax function to obtain the probability (prob) of each category:

[0074] prob(t) = softmax(y t )

[0075] Through this step, it can be ensured that the sum of all probabilities is 1, and each probability is between 0 and 1.

[0076] 3) Select the category with the highest probability.

[0077] In the classification task, usually select the category with the highest probability as the prediction result of the model:

[0078] x t+1 = argmax(prob(t))

[0079] This probability distribution can be directly used to generate text. For example, select the next word through sampling or greedy decoding.

[0080] In this way, the Transformer model can convert the internal vector representation into useful outputs, such as the probability distribution of the next word in text generation or the category probability in the classification task.

[0081] Use the first T words to represent the question string, which is used to guide the multimodal large model to answer questions related to the text in the image. For example, if x = "What is the text in the document in the figure? Please answer line by line from top to bottom", then the language decoder of the multimodal generative large model generates all text sequences with subscripts from T to infinity by gradually calling the above process. Infinity here is a mathematical definition. In reality, the user defines the stop symbols for the text, such as carriage return and period, etc., and the maximum decoding text length to automatically stop the model from generating more words.

[0082] x T+1 :(z) = argmax(prob(x

[0083] = What is the text in the document in the figure? Please answer line by line from top to bottom, z))

[0084] Since multiple enhanced images z_n are generated in S110, the multimodal generative large model can calculate the corresponding text answer x(z) for any one or a combination of several of them, that is, the text recognition result of the enhanced image.

[0085] In practical applications, the multimodal language model can be trained through two techniques: pre-training and fine-tuning of the multimodal generative large model of related technologies.

[0086] S130, based on the differences between the text recognition results corresponding to each enhanced image, obtain the error of the multimodal generative large model.

[0087] For example, in the above example, the error of the multimodal generative large model can be predicted by calculating the degree of difference of x(z) when z takes different images.

[0088] In some embodiments, such as Figure 2 shown, S130 may include the following steps:

[0089] S131, obtain the similarity degree between the text recognition results corresponding to any two of the enhanced images;

[0090] S132, based on the average value of multiple similarity degrees, obtain the error of the multimodal generative large model.

[0091] In the above embodiment, the difference between the text recognition results corresponding to each enhanced image can be the average value of the similarity degrees of the text recognition results obtained from any two enhanced images z_1 and z_2.

[0092] In some embodiments, the accuracy criterion can be used to obtain the accuracy rate corresponding to the text recognition results of any two enhanced images, and the accuracy rate is used as the similarity degree.

[0093] Among them, accuracy refers to the proportion of samples actually belonging to the positive class among those predicted as the positive class by the model. It measures the accuracy of the model's prediction.

[0094] Mathematically, accuracy can be expressed as:

[0095]

[0096] Among them, TP (True Positive) is the number of samples correctly predicted as the positive class, and FP (False Positive) is the number of samples wrongly predicted as the positive class (i.e., actually belonging to the negative class).

[0097] In some other embodiments, the recall criterion can be used to obtain the recall rate corresponding to the text recognition results of any two enhanced images, and the recall rate is used as the similarity degree.

[0098] Recall rate refers to the proportion of samples actually belonging to the positive class that are correctly predicted as the positive class by the model. It measures the ability of the model to find all positive-class samples.

[0099] Mathematically, recall rate can be expressed as:

[0100]

[0101] Among them, FN (False Negative) is the number of samples wrongly predicted as the negative class (i.e., actually belonging to the positive class).

[0102] In still some other embodiments, the harmonic mean of the accuracy rate and the recall rate corresponding to the text recognition results of any two enhanced images is used as the similarity degree.

[0103] There is a certain trade-off relationship between accuracy and recall rate. Generally, increasing the accuracy may reduce the recall rate, and vice versa. For example, in the application of spam filtering, if the filtering criteria are tightened, it may reduce misjudgments (i.e., reduce FP), thus increasing the accuracy, but at the same time, it may also miss some real spam emails (i.e., increase FN), thus reducing the recall rate.

[0104] Therefore, in the embodiments of this application, in order to consider both accuracy and recall rate, the text similarity scoring device also uses the F1 score as a comprehensive evaluation index, which is the harmonic mean of accuracy and recall rate:

[0105]

[0106] By calculating the average value of the text similarity in terms of accuracy, recall rate, and F1 score, and taking the average value as the error value of the multi-modal generative large model.

[0107] In the above embodiment, the recall rate and the precision rate can be based on the recall rate and precision rate of n-gram. That is, the similarity between any two text sequences can be calculated based on the two indicators of the recall rate and precision rate of n-gram. N-gram is to divide the text sequence into n consecutive units (tokens), and these units can be characters, words, etc. For example, for the sentence "I like natural language processing", when n=2, the 2-grams are "I like", "like", "happy from", "natural", etc. The recall rate (Recall) refers to the ratio of the number of similar n-grams predicted in the comparison of two sentences to the number of similar n-grams actually existing in the two sentences. The precision rate (Precision) refers to the ratio of the number of similar n-grams that are truly correct to the number of similar n-grams predicted among the predicted similar n-grams.

[0108] For example, sentence A: Today's weather is very good, suitable for going out for a walk. Sentence B: Today's weather is good, very suitable for going out for a walk. When n=4, the 4-grams of sentence A include "Today's weather", "Every day the weather is very good", "The weather is very good", etc., and the 4-grams of sentence B include "Today's weather", "Every day the weather is not good", "The weather is good", etc. Similar 4-grams include "Today's weather", a total of 1. The total number of 4-grams of sentences A and B is 8, assuming that the actual similarity is this 1. If 2 similar 4-grams are predicted, but only "Today's weather" is correct, then the recall rate is 1 / 8 and the precision rate is 1 / 2.

[0109] In some embodiments, the average of the above multiple similarities can be used as the confidence of the multimodal generative large model.

[0110] In some embodiments, when the confidence of the multimodal generative large model is less than a preset threshold, the multimodal generative large model is prohibited from returning any text recognition results. For example, when the average value of the F1 score is less than a preset threshold, such as 0.9, the multimodal generative large model is forced to not return any text recognition results, that is, the multimodal generative large model is prohibited from performing text recognition on the image.

[0111] In some embodiments, when the average value of the above multiple similarities is less than a preset threshold, the average value of the multiple similarities can be output as the confidence, for example, the average value of the F1 score can be output as the confidence.

[0112] Through the technical solution provided by the embodiments of the present application, by performing data augmentation on the original image, multiple enhanced images are obtained. The multiple enhanced images are respectively input into a trained multi-modal generative large model to obtain the text recognition results corresponding to each enhanced image. Based on the differences between the text recognition results corresponding to each enhanced image, the error of the multi-modal generative large model can be obtained, and it can be known whether the multi-modal generative large model is accurate, avoiding inaccurate text recognition results caused by using a multi-modal generative large model with a large error.

[0113] Figure 3 FIG. 4 shows a schematic structural diagram of a model error acquisition device provided by an embodiment of the present application. The device 300 includes: an image preprocessing module 310, a recognition module 320, and an error control module 330.

[0114] In the embodiments of the present application, the image preprocessing module 310 is configured to perform data augmentation on the original image to obtain multiple enhanced images; the recognition module 320 is configured to input the multiple enhanced images into a trained multi-modal generative large model to obtain text recognition results corresponding to each of the enhanced images, where the multi-modal generative large model is used to perform text recognition on the input image; the error control module 330 is configured to obtain the error of the multi-modal generative large model based on the differences between the text recognition results corresponding to each of the enhanced images.

[0115] In a possible implementation, the image preprocessing module 310 performs data augmentation on the original image to obtain multiple enhanced images, including at least one of the following:

[0116] Randomly cropping the original image to obtain the enhanced image;

[0117] Rotating the original image around the center point by a predetermined angle to obtain the enhanced image;

[0118] Enlarging or reducing the original image to obtain the enhanced image.

[0119] In a possible implementation, the error control module 330 obtains the error of the multi-modal generative large model based on the differences between the text recognition results corresponding to each of the enhanced images, including:

[0120] Obtaining the similarity degree between the text recognition results corresponding to any two of the enhanced images;

[0121] Obtaining the error of the multi-modal generative large model based on the average value of the multiple similarity degrees.

[0122] In a possible implementation, the error control module 330 obtains the similarity degree between the text recognition results corresponding to any two of the enhanced images, including one of the following:

[0123] Using the accuracy criterion, obtain the accuracy rate corresponding to the text recognition results corresponding to any two of the enhanced images, and use the accuracy rate as the similarity degree;

[0124] Using the recall criterion, obtain the recall rate corresponding to the text recognition results corresponding to any two of the enhanced images, and use the recall rate as the similarity degree;

[0125] Use the harmonic mean of the accuracy rate and the recall rate corresponding to the text recognition results corresponding to any two of the enhanced images as the similarity degree.

[0126] In a possible implementation, the error control module 330 obtains the error of the multimodal generative large model based on the average value of multiple similarity degrees, including:

[0127] Use the average value of multiple similarity degrees as the confidence level of the multimodal generative large model.

[0128] In a possible implementation, the error control module 330 is further configured to prohibit the multimodal generative large model from returning any text recognition results when the confidence level of the multimodal generative large model is less than a preset threshold.

[0129] In a possible implementation, the multimodal generative large model is a multimodal language model.

[0130] The apparatus 300 provided in the embodiments of the present application can execute the various methods described in the foregoing method embodiments, and implement the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be elaborated herein.

[0131] Figure 4 The hardware structure diagram of the electronic device 400 that executes the model error acquisition method provided in the embodiments of the present application is shown. Referring to this figure, at the hardware level, the electronic device includes a processor 410. Optionally, it includes an internal bus 420, a network interface 430, and a memory. Among them, the memory may include a memory 440, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory 450, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other tasks.

[0132] The processor 410, network interface 430, and memory can be interconnected through an internal bus 420, which can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a bidirectional arrow is used to represent it in this figure, but it does not mean that there is only one bus or one type of bus.

[0133] A memory for storing programs. Specifically, the program can include program code, and the program code includes computer operation instructions. The memory can include a memory 440 and a non-volatile memory 450, and provide instructions and data to the processor 410.

[0134] The processor 410 reads the corresponding computer program from the non-volatile memory 450 into the memory 440 and then runs it, forming a device for locating target users at the logical level. The processor 410 executes the program stored in the memory and is specifically used to execute Figure 1-2 The method described in the embodiment and achieve the same or corresponding technical effects.

[0135] As described above in this application Figure 1-2The method disclosed in the embodiments shown can be applied to a processor or implemented by processor 410. Processor 410 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in hardware or instructions in software form in processor 410. The above-mentioned processor 410 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or implemented and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and processor 410 reads the information in the memory and combines its hardware to complete the steps of the above method.

[0136] The electronic device can also execute the various methods described in the foregoing method embodiments and achieve the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be elaborated here.

[0137] Of course, in addition to the software implementation, the electronic device of the present application does not exclude other implementation manners, such as a logic device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and may also be hardware or a logic device.

[0138] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable medium stores one or more programs. When the one or more programs are executed by an electronic device including multiple application programs, the electronic device is caused to perform the following operations: perform data augmentation on an original image to obtain multiple augmented images; input the multiple augmented images into a trained multi-modal generative large model to obtain text recognition results corresponding to each of the augmented images, where the multi-modal generative large model is used to perform text recognition on the input image; and obtain an error of the multi-modal generative large model based on differences between the text recognition results corresponding to each of the augmented images.

[0139] The electronic device can also execute the various methods described in the foregoing method embodiments and achieve the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be elaborated herein.

[0140] Among them, the computer-readable storage medium includes a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, an optical disc, or the like.

[0141] Further, an embodiment of the present application also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the following process is implemented: perform data augmentation on an original image to obtain multiple augmented images; input the multiple augmented images into a trained multi-modal generative large model to obtain text recognition results corresponding to each of the augmented images, where the multi-modal generative large model is used to perform text recognition on the input image; and obtain an error of the multi-modal generative large model based on differences between the text recognition results corresponding to each of the augmented images.

[0142] In summary, the foregoing are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

[0143] The system, apparatus, module, or unit illustrated in the above embodiments can be specifically implemented by a computer chip or an entity, or by a product with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0144] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0145] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0146] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.

Claims

1. A method for obtaining model error, comprising: Performing data augmentation on an original image to obtain multiple augmented images; Inputting the multiple augmented images into a trained multi-modal generative large model to obtain text recognition results corresponding to each of the augmented images, wherein the multi-modal generative large model is used for text recognition of the input image; Obtaining the error of the multi-modal generative large model based on the differences between the text recognition results corresponding to each of the augmented images.

2. The method according to claim 1, wherein The performing data augmentation on the original image to obtain multiple augmented images includes at least one of the following: Randomly cropping the original image to obtain the augmented image; Rotating the original image around a center point by a predetermined angle to obtain the augmented image; Enlarging or reducing the original image to obtain the augmented image.

3. The method according to claim 1 or 2, wherein The obtaining the error of the multi-modal generative large model based on the differences between the text recognition results corresponding to each of the augmented images includes: Obtaining the similarity degree between the text recognition results corresponding to any two of the augmented images; Obtaining the error of the multi-modal generative large model based on the average value of the multiple similarity degrees.

4. The method according to claim 3, wherein, The obtaining the similarity degree between the text recognition results corresponding to any two of the augmented images includes one of the following: Using the accuracy criterion, obtaining the accuracy corresponding to the text recognition results corresponding to any two of the augmented images, and taking the accuracy as the similarity degree; Using the recall criterion, obtaining the recall corresponding to the text recognition results corresponding to any two of the augmented images, and taking the recall as the similarity degree; Taking the harmonic mean of the accuracy and the recall corresponding to the text recognition results corresponding to any two of the augmented images as the similarity degree.

5. The method according to claim 3, wherein, The obtaining the error of the multi-modal generative large model based on the average value of the multiple similarity degrees includes: Taking the average value of the multiple similarity degrees as the confidence level of the multi-modal generative large model.

6. The method according to claim 5, wherein, The method further includes: In the case where the confidence level of the multi-modal generative large model is less than a preset threshold, prohibiting the multi-modal generative large model from returning any text recognition results.

7. The method according to claim 1 or 2, wherein The multi-modal generative large model is a multi-modal language model.

8. A device for obtaining model error, comprising: An image preprocessing module for performing data augmentation on an original image to obtain multiple augmented images; A recognition module for inputting the multiple augmented images into a trained multi-modal generative large model to obtain text recognition results corresponding to each of the augmented images, wherein the multi-modal generative large model is used for text recognition of the input image; An error control module for obtaining the error of the multi-modal generative large model based on the differences between the text recognition results corresponding to each of the augmented images.

9. An electronic device, comprising: A processor; And A memory arranged to store computer-executable instructions that, when executed, use the processor to execute the model error obtaining method according to any one of claims 1-7.

10. A computer-readable medium storing one or more programs, which when executed by an electronic device including a plurality of application programs, cause the electronic device to execute the model error acquisition method according to any one of claims 1-7.

11. A computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions that, when executed by a computer, cause the computer to execute the model error acquisition method according to any one of claims 1-7.