Text recognition method and device based on multi-modal large model, equipment and medium

CN122841928APending Publication Date: 2026-09-29BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610976648.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-29

AI Technical Summary

Benefits of technology

[0009]由此,可以通过统一的中间服务层实现多个多模态大模型的接入,以实现客户端对多个多模态大模型的集成,这样客户端可以通过统一且标准的文本识别请求进行文本识别,无需针对不同的后端模型维护不同的接口,有效降低客户端的集成和维护成本,同时也可以支持多模态大模型的快速接入,提高文本识别时可调用的多模态大模型的多样性,提升文本识别的响应效率。另外,可以基于返回类型对识别模型的识别结果进行类型转换,以为客户端返回符合其格式要求的识别结果,提升OCR结果的深度和广度,为下游的文档理解应用提供有效的数据支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841928A_ABST
    Figure CN122841928A_ABST
Patent Text Reader

Abstract

A text recognition method, device and equipment based on a multi-modal large model and a medium are provided, and belong to the technical field of computers.The method comprises the following steps: receiving a text recognition request sent by a client based on an intermediate service layer, wherein the text recognition request contains a first image to be recognized, and a plurality of multi-modal large models are accessed in the intermediate service layer; determining a recognition model for text recognition of the first image from the plurality of multi-modal large models; performing text recognition on the first image based on the recognition model to obtain a first recognition result output by the recognition model; performing type conversion on the first recognition result based on a return type corresponding to the text recognition request to obtain a second recognition result, and sending the second recognition result to the client. Thus, the client can perform text recognition through a unified and standard text recognition request, without the need to maintain different interfaces for different back-end models, thereby effectively reducing the integration and maintenance costs of the client.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This content relates to the field of computer technology, specifically to a text recognition method, apparatus, device, and medium based on a multimodal large model. Background Technology

[0002] With the development of technologies such as computers, networks, and communications, users can quickly perform text recognition using large models. In related technologies, multiple large models capable of text recognition can be used to provide text recognition services, with each large model deployed with a dedicated backend service. When a user needs text recognition, the client calls the interfaces of different backend services to invoke the corresponding large model. Summary of the Invention

[0003] This content section is provided to briefly introduce the concepts, which will be described in detail in the examples section later. This content section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] Firstly, a text recognition method based on a multimodal large model is provided, including: The intermediate service layer receives text recognition requests sent by the client, the text recognition requests contain a first image to be recognized, and the intermediate service layer is connected to multiple multimodal large models; From the plurality of multimodal large models, determine the recognition model for text recognition of the first image; Based on the recognition model, text recognition is performed on the first image to obtain the first recognition result output by the recognition model; Based on the return type corresponding to the text recognition request, the first recognition result is converted to obtain a second recognition result, and the second recognition result is sent to the client.

[0005] Secondly, a text recognition device based on a multimodal large model is provided, comprising: The receiving module is used to receive text recognition requests sent by clients based on the intermediate service layer. The text recognition requests contain a first image to be recognized, and the intermediate service layer is connected to multiple multimodal large models. The first determining module is used to determine the recognition model for text recognition of the first image from the plurality of multimodal large models; The recognition module is used to perform text recognition on the first image based on the recognition model, and obtain the first recognition result output by the recognition model; The processing module is used to perform type conversion on the first recognition result based on the return type corresponding to the text recognition request, obtain a second recognition result, and send the second recognition result to the client.

[0006] Thirdly, a computer-readable medium is provided having a computer program stored thereon, wherein the computer program, when executed by a processing device, implements the steps of the method described in the first aspect.

[0007] Fourthly, an electronic device is provided, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect.

[0008] Fifthly, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in the first aspect.

[0009] Therefore, a unified middleware service layer can be used to access multiple multimodal large models, enabling client-side integration of these models. Clients can then perform text recognition through a unified and standardized text recognition request, eliminating the need to maintain different interfaces for different backend models. This effectively reduces client integration and maintenance costs, while also supporting rapid access to multimodal large models, increasing the diversity of available models for text recognition, and improving response efficiency. Furthermore, the recognition results can be type-converted based on the return type to return formatted results to the client, enhancing the depth and breadth of OCR results and providing effective data support for downstream document understanding applications.

[0010] Other features and advantages of the technical solution will be described in detail in the following examples section. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the technical solution will become more apparent when taken in conjunction with the accompanying drawings and the following examples. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is an architecture diagram of how the client calls the backend model using related technologies.

[0012] Figure 2 This is a flowchart illustrating a text recognition method based on a multimodal large model, illustrating certain scenarios.

[0013] Figure 3This is an architecture diagram of a text recognition system based on a multimodal large model, illustrated according to certain scenarios.

[0014] Figure 4 This is a block diagram of a text recognition device based on a multimodal large model, shown according to some scenarios.

[0015] Figure 5 These are schematic diagrams of the structure of electronic devices shown in various scenarios. Detailed Implementation

[0016] The technical solution will now be described in more detail with reference to the accompanying drawings. Although certain scenarios are shown in the drawings, it should be understood that the technical solution can be implemented in various forms and should not be construed as limited to the scenarios described herein. Rather, these scenarios are provided to provide a more thorough and complete understanding of the technical solution. It should be understood that the accompanying drawings and the scenarios described are for illustrative purposes only and are not intended to limit the scope of protection of the technical solution.

[0017] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of the technical solution is not limited in this respect.

[0018] The term "comprising" and its variations as used herein can be open-ended, meaning "including but not limited to". The term "based on" can mean "at least partially based on". The term "one case" means "at least one case"; the term "another case" means "at least one additional case"; the term "some cases" means "at least some cases". Definitions of other terms will be given in the following description.

[0019] It should be noted that the concepts of "first" and "second" mentioned here are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.

[0020] It should be noted that the terms "one" and "more" used here are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0021] The names of messages or information exchanged between the multiple devices in the implementation are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0022] It is understandable that before using the technical solutions provided here, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in accordance with relevant laws and regulations, and their authorization should be obtained through appropriate means.

[0023] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations described herein.

[0024] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0025] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the technical solution. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the technical solution.

[0026] At the same time, it is understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws, regulations and related provisions.

[0027] As described in the background section, different large models in related technologies require deployment with their dedicated backend services, such as... Figure 1 As shown, when a user wants to perform text recognition, the client can call model M1 through the interface of service A1, model M2 through the interface of service A2, and model M3 through the interface of service A3. Models M1, M2, and M3 are used to represent the models used for text recognition. The interfaces of services A1, A2, and A3 are usually heterogeneous interfaces. Therefore, if the client wants to use multiple large models for data inference, it needs to adapt to multiple heterogeneous interfaces, which makes the client logic complex and the maintenance cost high.

[0028] To address the aforementioned issues, a text recognition method based on a multimodal large model is proposed, such as... Figure 2 As shown, the method may include: In step 21, a text recognition request sent by the client is received based on the intermediate service layer. The text recognition request contains a first image to be recognized. The intermediate service layer is connected to multiple multimodal large models.

[0029] A multimodal large model refers to a model trained by combining multimodal information such as text, images, videos, and audio. For example, in a text recognition scenario, this multimodal large model can include a Vision-Language Model (VLM), enabling integrated processing of images and text by combining a large language model with a visual encoder. For instance, a multimodal large model can be used to implement Optical Character Recognition (OCR) to recognize received images.

[0030] As an example, to shield clients from the differences in calling different backend models, a middleware service layer can be used to achieve unified abstraction and standardized access for different multimodal large models. For example... Figure 3 As shown, an intermediate service layer can be used to adapt the client to multiple multimodal large models, and multiple multimodal large models can access this intermediate service layer. Text recognition requests sent by the client will be uniformly sent to the intermediate service layer for processing.

[0031] In step 22, a recognition model for text recognition of the first image is determined from multiple multimodal large models.

[0032] As an example, when a user initiates a text recognition request in the client, they can select the model to apply based on their needs. For instance, the middleware layer can synchronize the names of the integrated multimodal models to the client, displaying them on the client's interface for easy user selection. For example, the interface could display interactive controls for each integrated multimodal model, allowing users to select a model by clicking on the controls. Alternatively, the interface could display multimodal models via dropdown menus. For instance, a default multimodal model could be displayed in the dropdown menu, and users could select the desired model from the selected dropdown menu.

[0033] In step 23, text recognition is performed on the first image based on the recognition model to obtain the first recognition result output by the recognition model.

[0034] The intermediate service layer can store the addresses and invocation methods of the accessed multimodal large models. Therefore, after determining the recognition model, an invocation request for the recognition model can be constructed based on its invocation method, and the recognition model can be invoked based on its address and the invocation request. The recognition model can perform model inference on the first image based on the invocation request and output a first recognition result. The recognition model can then return the first recognition result to the intermediate service layer.

[0035] In step 24, the first recognition result is converted based on the return type corresponding to the text recognition request to obtain the second recognition result, and the second recognition result is sent to the client.

[0036] The return type corresponding to the text recognition request is determined based on the user's selection. For example, candidate types can be displayed on the interface. These candidate types include plain text, structured data, and lightweight markup language. Plain text returns the text contained in the first image; structured data returns the structured text data corresponding to the first image, which may include element information of each element in the first image, including boundary information, type, and text; lightweight markup language can be Markdown, where a document is written in an easy-to-read and easy-to-write plain text format and then converted into a valid XHTML (eXtensible HyperText Markup Language) or HTML (HyperText Markup Language) document.

[0037] For example, if a user selects a type based on their needs, the type identifier of the selected candidate type can be added to the text recognition request. This allows the type identifier to be obtained by parsing the text recognition request, thus determining the return type. If the user does not select a type, a pre-configured default type can be used as the return type.

[0038] After receiving the first recognition result from the recognition model, the first recognition result can be converted based on the return type to obtain a second recognition result that meets the user's needs and returned to the client so that the text recognition result of the first image can be displayed on the client.

[0039] Therefore, a unified middleware service layer can be used to access multiple multimodal large models, enabling client-side integration of these models. Clients can then perform text recognition through a unified and standardized text recognition request, eliminating the need to maintain different interfaces for different backend models. This effectively reduces client integration and maintenance costs, while also supporting rapid access to multimodal large models, increasing the diversity of available models for text recognition, and improving response efficiency. Furthermore, the recognition results can be type-converted based on the return type to return formatted results to the client, enhancing the depth and breadth of OCR results and providing effective data support for downstream document understanding applications.

[0040] In some cases, the method may further include: Upon receiving a text recognition request, permission verification can be performed on the text recognition request to determine whether text recognition can be performed on the first image.

[0041] For example, the text recognition request can be authenticated by at least one of the following: If the number of bytes corresponding to the first image is less than the first threshold, the text recognition request is determined to have passed the permission verification.

[0042] In this approach, permission verification can be performed based on the size of the first image, thereby constraining the size of the input first image. For example, the first threshold can be set based on the actual application scenario, such as 5MB. If the number of bytes in the first image is less than 5MB, then the text recognition request is deemed to have passed the permission verification, meaning that the first image can be recognized. This avoids the first image being too large, which could lead to abnormal model inference and ensures the stable operation of the multimodal large model.

[0043] If the query rate per second of the interface called by the text recognition request is less than the second threshold, it is determined that the text recognition request has passed the permission verification.

[0044] The interface invoked by the text recognition request can be a text recognition interface, a plain text interface, or a layout detection interface. The query rate per second (MRS) of this interface can then be obtained. If the MRS is less than a second threshold, the interface is considered to be under low load. In this case, it can be determined that the text recognition request has passed the authorization verification, thus avoiding high-load operation of the interface and ensuring the stable operation of the multimodal large model.

[0045] If the text recognition request contains parameter values ​​of the parameters of the interface called by the text recognition request, it is determined that the text recognition request has passed the permission verification.

[0046] The text recognition request needs to include the parameter values ​​of the interface it calls. Therefore, the text recognition request can be parsed to determine whether it contains the required parameter values. If the text recognition request contains the parameter values ​​of all the parameters of the interface called by the text recognition request, the interface can be called based on the parameter values. This determines that the text recognition request has passed the permission verification, thereby ensuring the accuracy of the interface call.

[0047] In response to determining that the text recognition request has passed the permission verification, a recognition model for text recognition of the first image is determined from the plurality of multimodal large models.

[0048] Therefore, upon receiving a text recognition request, permission verification can be performed to ensure the stable operation of the multimodal large model.

[0049] In some cases, determining the recognition model for text recognition of the first image from the plurality of multimodal large models may include: If the text recognition request contains a large model identifier, then the multimodal large model indicated by the large model identifier is used as the recognition model.

[0050] As mentioned above, if a user selects a multimodal large model to be invoked when triggering a text recognition request, a large model identifier of the user-selected model can be added to the text recognition request. In this way, when the intermediate service layer receives the text recognition request, it parses it to obtain the large model identifier, and then uses the multimodal large model indicated by the large model identifier as the recognition model.

[0051] If the text recognition request does not contain a large model identifier, then a candidate model is determined from the multiple multimodal large models based on the service to which the text recognition request belongs, and a recognition model for text recognition of the first image is determined based on the candidate model.

[0052] If a user does not select a multimodal large model to invoke when triggering a text recognition request, the text recognition request will not contain a large model identifier. In some cases, a pre-configured default multimodal large model can be used as the recognition model. In other cases, multimodal large models that can be invoked by various services can be pre-configured, allowing different services to invoke the same multimodal large model. Accordingly, the intermediate service layer can parse the text recognition request to determine the service to which the text recognition request belongs based on the service identifier in the text recognition request. Then, the multimodal large model that can be invoked by the service indicated by the service identifier is used as a candidate model.

[0053] After identifying candidate models, as an example, one can be randomly selected from the candidate models as the recognition model. As another example, the load information of the candidate models can be obtained, and the candidate model with the lowest load can be selected as the recognition model.

[0054] Therefore, upon receiving a text recognition request from the client, the recognition model used to respond to the text recognition request can be determined from multiple multimodal large models accessed by the intermediate service layer. Based on this recognition model, text recognition of the first image can be achieved, thereby improving the response efficiency of the text recognition request and also improving the accuracy of the text recognition result to a certain extent.

[0055] In some cases, each of the multimodal large models is deployed based on an image file of the multimodal large model, the image file containing code files, dependency library files and configuration files corresponding to running the multimodal large model.

[0056] For example, each multimodal large model can be packaged into an independent image file during deployment, such as a Docker image file, and stored in an image repository. The code file corresponding to running the multimodal large model can represent the model, such as the weight file. Dependency library files can contain the libraries required to run the large model, such as vLLM, a library used for large model inference and services. Configuration files can contain parameters defining the model structure, hyperparameters, and model runtime environment.

[0057] This allows for the rapid and independent deployment of large multimodal models based on the image file, making the deployment of large multimodal models unrestricted by the environment. For example, the large multimodal model can be quickly deployed on virtual machines and servers based on the image file, and it also supports cross-platform and cross-architecture deployment. Therefore, when configuring the backend model on the client side, the image file of the large multimodal model can be used to quickly access the model, simplifying the process of accessing new large models and effectively improving the diversity of models that can be called in text recognition requests, so as to provide users with accurate text recognition results.

[0058] In some cases, the multimodal large model accessed in the intermediate service layer is updated in the following way: In response to a version update command, the system determines the multimodal main model to be updated and the updated version. Configuration personnel can select the multimodal main model to be updated via the configuration interface and indicate the updated version to be switched to. In response to this configuration operation, a version update command can be sent to the middle service layer, containing the multimodal main model to be updated and the updated version.

[0059] Pull the first image file of the multimodal large model to be updated in the updated version from the image file library, and execute the first image file to switch the multimodal large model to the updated version.

[0060] Each major model, after being updated, can be packaged to generate its corresponding image file, which is then stored in an image file library. Therefore, the first image file can be obtained by querying the image file library based on the updated version and the model identifier of the multimodal major model to be updated. This first image file is then executed to access the updated version of the multimodal major model at the intermediate service layer.

[0061] Therefore, version switching of multimodal large models can be achieved by updating the image file, so as to support the update and rollback of large model versions, improve the availability of accessed multimodal large models, and thus improve the accuracy of text recognition results.

[0062] In some cases, the intermediate service layer is deployed based on its image file. This allows for version updates of the intermediate service layer based on the image file, decoupling the client, the intermediate service layer, and the multimodal large model, and enabling independent updates of the client, the intermediate service layer, and the multimodal large model.

[0063] In some cases, the configuration file contains a resource adjustment strategy, which is used to update the configuration of computing resources for model inference of the multimodal large model according to the calling status of the multimodal large model. The computing resources include at least one of graphics processing unit (GPU) resources, neural processing unit (NPU) resources, and instances of the multimodal large model.

[0064] In configuring the multimodal large model, initial configuration of computing resources can be achieved through configuration parameters, and resource adjustment strategies can be configured to achieve dynamic updates of computing resources. Different multimodal large models can have their own independently configured resource adjustment strategies, which can be the same or different. Correspondingly, during the application of the multimodal large model, its usage can be monitored in real time. If the number of calls to the multimodal large model exceeds a preset threshold, it indicates a large usage volume. In this case, the amount of computing resources configured for model inference can be increased, thereby achieving dynamic adjustment of computing resources, ensuring the inference efficiency of model inference based on the multimodal large model, and thus improving the response efficiency of text recognition.

[0065] As an example, if the number of times the multimodal large model is called within a continuous period is less than a preset threshold, it indicates that the multimodal large model has a low call volume. In this case, the configuration of computing resources for model inference of the multimodal large model can be updated to the initial configuration.

[0066] Therefore, the computing resources for model inference of multimodal large models can be dynamically adjusted through the configuration file of the multimodal large model, ensuring the stable operation of the multimodal large model and improving the inference efficiency of the multimodal large model.

[0067] In some cases, different configuration files can be configured for each multimodal large model and the client application on which it is based. For example, after the configuration personnel configure a specific application, they can generate a configuration file corresponding to that application. The configuration file corresponding to that application can replace the original configuration file of the multimodal large model to generate an image file, so that when deploying the multimodal large model based on the image file, personalized configuration for the client application can be achieved.

[0068] In some cases, the text recognition request also includes a first type of the element to be recognized in the first image. For example, a user can extract text from elements of a specific type in an image. For instance, to extract text from a table in an image, the first type could be a table type. Accordingly, the method further includes: In response to the text recognition request, the layout detection interface in the intermediate service layer is invoked to obtain the layout information returned by the layout detection interface. The layout information includes the boundary information of each element in the first image and the type of the element.

[0069] For example, the middle service layer can be configured with multiple interfaces, such as plain text interfaces, text recognition interfaces, and layout detection interfaces. Each interface in the middle service layer can call multiple multimodal large models accessed within the middle service layer. Specifically, the plain text interface performs text recognition on the input image to obtain the plain text content. The text recognition interface performs text recognition on the input image and performs type conversion on the recognition results. The layout detection interface returns the position information and type of each element in the input image, excluding the text of the elements.

[0070] Here, in response to the text recognition request, the layout detection interface in the intermediate service layer can be called first to obtain the layout information returned by the layout detection interface. For example, the layout information can be obtained by calling the corresponding multimodal large model based on the layout detection interface.

[0071] Sub-images are determined from the first image based on the boundary information of the elements of the first type.

[0072] After obtaining the layout information, you can extract the elements of the first type, such as the elements of type table.

[0073] Specifically, determining a sub-image from the first image based on the boundary information of elements of the first type can be done by taking the image within the bounding box corresponding to the boundary information of that element in the first image as the sub-image for each element of the first type.

[0074] Accordingly, the step of performing text recognition on the first image based on the recognition model to obtain the first recognition result output by the recognition model includes: performing text recognition on the sub-images in the first image based on the recognition model to obtain the first recognition result output by the recognition model.

[0075] For example, a text recognition interface can be called based on a sub-image, thereby calling a recognition model based on the sub-image through the text recognition interface, so that the recognition model can recognize the sub-image and obtain a first recognition result.

[0076] Therefore, layout information can be obtained by performing layout detection on the first image. Based on the layout information, the parts of the first image that need to be recognized for text can be determined, thereby realizing text recognition in specific areas of the first image, improving the accuracy of text recognition, and avoiding the waste of resources caused by recognizing unnecessary parts of the image.

[0077] Figure 4 The diagram shown is a block diagram of a text recognition device based on a multimodal large model, illustrating some scenarios. Figure 4 As shown, the device 10 includes: The receiving module 101 is used to receive a text recognition request sent by the client based on the intermediate service layer. The text recognition request contains a first image to be recognized. The intermediate service layer is connected to multiple multimodal large models. The first determining module 102 is used to determine the recognition model for text recognition of the first image from the plurality of multimodal large models; The recognition module 103 is used to perform text recognition on the first image based on the recognition model, and obtain the first recognition result output by the recognition model; Processing module 104 is used to perform type conversion on the first recognition result based on the return type corresponding to the text recognition request, obtain a second recognition result, and send the second recognition result to the client.

[0078] Therefore, a unified middleware service layer can be used to access multiple multimodal large models, enabling client-side integration of these models. Clients can then perform text recognition through a unified and standardized text recognition request, eliminating the need to maintain different interfaces for different backend models. This effectively reduces client integration and maintenance costs, while also supporting rapid access to multimodal large models, increasing the diversity of available models for text recognition, and improving response efficiency. Furthermore, the recognition results can be type-converted based on the return type to return formatted results to the client, enhancing the depth and breadth of OCR results and providing effective data support for downstream document understanding applications.

[0079] Optionally, the first determining module 102 includes: The first determining submodule is used to, if the text recognition request contains a large model identifier, use the multimodal large model indicated by the large model identifier as the recognition model; The second determining submodule is used to determine a candidate model from the plurality of multimodal large models based on the service to which the text recognition request belongs if the text recognition request does not contain a large model identifier, and to determine the recognition model for text recognition of the first image based on the candidate model.

[0080] Optionally, each of the multimodal large models is deployed based on an image file of the multimodal large model, the image file containing code files, dependency library files and configuration files corresponding to running the multimodal large model.

[0081] Optionally, the configuration file includes a resource adjustment strategy, which is used to update the configuration of computing resources for model inference of the multimodal large model according to the calling status of the multimodal large model. The computing resources include at least one of graphics processor resources, neural network processing units, and multimodal large model instances.

[0082] Optionally, the text recognition request may further include a first type of the element to be recognized in the first image; The device 10 further includes: The calling module is used to call the layout detection interface in the intermediate service layer in response to the text recognition request, and obtain the layout information returned by the layout detection interface. The layout information includes the boundary information of each element in the first image and the type of the element. The second determining module determines a sub-image from the first image based on the boundary information of the elements of the first type; The identification module is used for: Based on the recognition model, text recognition is performed on the sub-image in the first image to obtain the first recognition result output by the recognition model.

[0083] Optionally, the device 10 further includes: The verification module is used to verify the permission of the text recognition request by at least one of the following: If the number of bytes corresponding to the first image is less than the first threshold, it is determined that the text recognition request has passed the permission verification; If the query rate per second of the interface called by the text recognition request is less than the second threshold, it is determined that the text recognition request has passed the permission verification. If the text recognition request contains parameter values ​​of the parameters of the interface called by the text recognition request, it is determined that the text recognition request has passed the permission verification; In response to determining that the text recognition request has passed the permission verification, a recognition model for text recognition of the first image is determined from the plurality of multimodal large models.

[0084] The following is for reference. Figure 5The diagram illustrates a structural schematic of an electronic device (e.g., a terminal device or a server) 600 suitable for implementing the above-described technical solutions. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs (Televisions), desktop computers, etc. Figure 5 The electronic device shown is merely an example and should not be construed as limiting its functionality or scope of use.

[0085] like Figure 5 As shown, electronic device 600 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. The random access memory 603 also stores various programs and data required for the operation of electronic device 600. The processing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. An input / output (I / O) interface 605 is also connected to bus 604.

[0086] Typically, the following devices can be connected to the input / output interface 605: input devices 606 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 608 including, for example, magnetic tape, hard disk, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0087] In particular, depending on certain circumstances, the processes described in the flowchart above can be implemented as computer software programs. For example, a computer program product is provided, comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. This computer program can be downloaded and installed from a network via communication device 609, or installed from storage device 608, or installed from read-only memory 602. When the computer program is executed by processing device 601, it performs the functions defined in the above-described methods.

[0088] It should be noted that the aforementioned computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM, or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In one case, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In another case, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.

[0089] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), the internet (e.g., the Internet), and peer-to-peer networks (e.g., ad-hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0090] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0091] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the following to occur: receive a text recognition request sent by a client based on an intermediate service layer. The text recognition request includes a first image to be recognized, and the intermediate service layer is connected to multiple multimodal large models. The electronic device then determines a recognition model from the multiple multimodal large models to perform text recognition on the first image. Based on the recognition model, the electronic device performs text recognition on the first image to obtain a first recognition result output by the recognition model. Based on the return type corresponding to the text recognition request, the electronic device performs type conversion on the first recognition result to obtain a second recognition result and sends the second recognition result to the client.

[0092] Computer program code for performing the above operations can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages, as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0093] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the figures. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0094] The modules mentioned above can be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the module itself; for example, a receiving module can also be described as "a module that receives text recognition requests sent by clients based on an intermediate service layer."

[0095] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field-Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Parts (ASSPs), Systems on Chips (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0096] In this context, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0097] The above description is merely illustrative and explains the technical principles employed. Those skilled in the art should understand that the scope of the technical solution is not limited to specific combinations of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features provided herein that have similar functions.

[0098] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limitations on the scope of the technical solution. Certain features described in the context of a single example can also be implemented in combination in a single example. Conversely, various features described in the context of a single example can also be implemented individually or in any suitable sub-combination in multiple examples.

[0099] Although the technical solution has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims. Regarding the aforementioned apparatus, the specific manner in which each module performs its operation has already been described in detail in the section concerning the method, and will not be elaborated upon here.

Claims

1. A text recognition method based on a multimodal large model, comprising: The intermediate service layer receives text recognition requests sent by the client, the text recognition requests contain a first image to be recognized, and the intermediate service layer is connected to multiple multimodal large models; From the plurality of multimodal large models, determine the recognition model for text recognition of the first image; Based on the recognition model, text recognition is performed on the first image to obtain the first recognition result output by the recognition model; Based on the return type corresponding to the text recognition request, the first recognition result is converted to obtain a second recognition result, and the second recognition result is sent to the client.

2. The method according to claim 1, wherein determining the recognition model for text recognition of the first image from the plurality of multimodal large models comprises: If the text recognition request contains a large model identifier, then the multimodal large model indicated by the large model identifier shall be used as the recognition model; If the text recognition request does not contain a large model identifier, then a candidate model is determined from the multiple multimodal large models based on the service to which the text recognition request belongs, and a recognition model for text recognition of the first image is determined based on the candidate model.

3. The method according to claim 1, wherein each of the multimodal large models is deployed based on an image file of the multimodal large model, the image file containing code files, dependency library files and configuration files corresponding to running the multimodal large model.

4. The method according to claim 3, wherein the configuration file includes a resource adjustment strategy, the resource adjustment strategy being used to update the configuration amount of computing resources for model inference of the multimodal large model according to the calling status of the multimodal large model, wherein the computing resources include at least one of graphics processor resources, neural network processing units, and multimodal large model instances.

5. The method according to claim 1, wherein the text recognition request further includes a first type of the element to be recognized in the first image; The method further includes: In response to the text recognition request, the layout detection interface in the intermediate service layer is invoked to obtain the layout information returned by the layout detection interface. The layout information includes the boundary information of each element in the first image and the type of the element. Sub-images are determined from the first image based on the boundary information of the elements of the first type; The step of performing text recognition on the first image based on the recognition model to obtain a first recognition result output by the recognition model includes: Based on the recognition model, text recognition is performed on the sub-image in the first image to obtain the first recognition result output by the recognition model.

6. The method according to claim 1, further comprising: The text recognition request is authorized by at least one of the following: If the number of bytes corresponding to the first image is less than the first threshold, it is determined that the text recognition request has passed the permission verification; If the query rate per second of the interface called by the text recognition request is less than the second threshold, it is determined that the text recognition request has passed the permission verification. If the text recognition request contains parameter values ​​of the parameters of the interface called by the text recognition request, it is determined that the text recognition request has passed the permission verification; In response to determining that the text recognition request has passed the permission verification, a recognition model for text recognition of the first image is determined from the plurality of multimodal large models.

7. A text recognition device based on a multimodal large model, comprising: The receiving module is used to receive text recognition requests sent by clients based on the intermediate service layer. The text recognition requests contain a first image to be recognized, and the intermediate service layer is connected to multiple multimodal large models. The first determining module is used to determine the recognition model for text recognition of the first image from the plurality of multimodal large models; The recognition module is used to perform text recognition on the first image based on the recognition model, and obtain the first recognition result output by the recognition model; The processing module is used to perform type conversion on the first recognition result based on the return type corresponding to the text recognition request, obtain a second recognition result, and send the second recognition result to the client.

8. A computer-readable medium having a computer program stored thereon, wherein, When executed by a processing device, the computer program performs the steps of the method according to any one of claims 1-6.

9. An electronic device, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-6.

10. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-6.