Multimodal self-certifying material identification method, apparatus, and device

By combining a multimodal large model and a cross-validation network, the problems of low efficiency and poor accuracy in self-certifying material identification are solved, achieving rapid adaptation and efficient and accurate material identification, and enhancing error detection capabilities.

CN122313495APending Publication Date: 2026-06-30ANT GALAXY (CHONGQING) INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANT GALAXY (CHONGQING) INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-04-02
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

The review of self-certification materials in the current technology requires a more efficient and accurate recognition solution. Traditional customized OCR recognition and manual review have problems such as low efficiency, high cost and poor generalization ability. Single large model recognition has model illusion and waste of resources, and multi-model recognition lacks cross-validation and is difficult to correct systematic errors.

Method used

A multimodal large model is used for material identification. By preprocessing image information, key fields are determined, and cross-validation tasks are generated. Multiple multimodal large models are used to execute the tasks independently, and the overall recognition results are obtained through fusion processing. By combining the task window length and the number of redundant validation coverage, a cross-validation network is constructed to improve accuracy and efficiency.

Benefits of technology

It enables rapid adaptation to new material formats, reduces deployment costs, improves recognition accuracy, enhances error detection capabilities, and improves the reliability and processing speed of overall recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122313495A_ABST
    Figure CN122313495A_ABST
Patent Text Reader

Abstract

This specification discloses a method, apparatus, and device for identifying multimodal self-certifying materials. The scheme includes: acquiring image information corresponding to the self-certifying material and preprocessing the image information; determining the required key fields based on the material type corresponding to the self-certifying material; generating a corresponding number of cross-validation tasks for each key field based on the corresponding task window length and redundant validation coverage; independently executing each corresponding cross-validation task using multiple multimodal large models, outputting local recognition results for some key fields included in the cross-validation task; and fusing the local recognition results to obtain the overall recognition result for the key fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to methods, apparatus and equipment for identifying multimodal self-certifying materials. Background Technology

[0002] In some business processes, users can improve their credit rating by uploading supporting documents, thereby achieving corresponding business objectives. For example, in lending scenarios, after uploading proof of identity, income, and assets, the platform can assess their credit limit based on this information; a higher credit rating typically allows for a higher loan amount. In shared services scenarios, users can upload proof of identity and credit score screenshots to use shared equipment without a deposit or enjoy better services. In e-commerce scenarios, the purchase of certain high-value or cross-border goods may require users to upload real-name authentication information and proof of address to ensure the compliance and security of the transaction.

[0003] Generally, different types of self-certification materials may come from multiple providers, such as different companies, schools, organizations, etc. Therefore, the formats of different types of self-certification materials usually differ, and even the same type of self-certification material may vary due to differences in the issuing party, the date of issuance, and other factors. Accurate identification of key fields in the materials can directly affect a user's credit rating.

[0004] Traditional solutions can employ customized Optical Character Recognition (OCR) methods. These methods train a dedicated OCR model for each material format, requiring retraining when the material format changes. Alternatively, manual review can be used, where human reviewers examine the materials and input key fields.

[0005] Therefore, a more efficient and accurate identification scheme is needed for the review of self-certification materials. Summary of the Invention

[0006] This specification provides one or more embodiments of a multimodal self-certification material identification method, apparatus, device, and storage medium to solve the following technical problem: the review of self-certification materials requires a more efficient and accurate identification scheme.

[0007] To solve the above-mentioned technical problems, one or more embodiments of this specification are implemented as follows: This specification provides a method for identifying multimodal self-certifying materials through one or more embodiments, including: Obtain the image information corresponding to the self-certification materials, and preprocess the image information; Based on the material type corresponding to the self-certification materials, determine the required key fields; For the key fields, a corresponding number of cross-validation tasks are generated based on the corresponding task window length and the number of redundant validation coverage times; Multiple multimodal large models are used to independently execute their respective cross-validation tasks, and output local recognition results for some key fields included in the cross-validation task. The results of each local identification are fused together to obtain the overall identification result for the key field.

[0008] This specification provides one or more embodiments of a multimodal self-certifying material identification device, comprising: The material acquisition module acquires image information corresponding to the self-certification materials and preprocesses the image information. The field determination module determines the required key fields based on the material type corresponding to the self-certification material; The task generation module generates a corresponding number of cross-validation tasks for the key fields, based on the corresponding task window length and the number of redundant validation coverage times. The task execution module independently executes its respective cross-validation task using multiple multimodal large models, and outputs local recognition results for some key fields included in the cross-validation task. The result fusion module merges the local recognition results to obtain the overall recognition result for the key field.

[0009] This specification provides one or more embodiments of a multimodal self-certifying material identification device, comprising: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: Obtain the image information corresponding to the self-certification materials, and preprocess the image information; Based on the material type corresponding to the self-certification materials, determine the required key fields; For the key fields, a corresponding number of cross-validation tasks are generated based on the corresponding task window length and the number of redundant validation coverage times; Multiple multimodal large models are used to independently execute their respective cross-validation tasks, and output local recognition results for some key fields included in the cross-validation task. The results of each local identification are fused together to obtain the overall identification result for the key field.

[0010] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured as follows: Obtain the image information corresponding to the self-certification materials, and preprocess the image information; Based on the material type corresponding to the self-certification materials, determine the required key fields; For the key fields, a corresponding number of cross-validation tasks are generated based on the corresponding task window length and the number of redundant validation coverage times; Multiple multimodal large models are used to independently execute their respective cross-validation tasks, and output local recognition results for some key fields included in the cross-validation task. The results of each local identification are fused together to obtain the overall identification result for the key field.

[0011] The above-described at least one technical solution adopted in one or more embodiments of this specification can achieve the following beneficial effects: 1. Rapid adaptation capability: No need to train a dedicated OCR model for each material format. Adaptation to new material formats can be completed within minutes using pre-labeled images in the template information, reducing deployment time and cost.

[0012] 2. Improve recognition accuracy: Through cross-validation and task decomposition, each model processes only a small number of fields, reducing the complexity of prompt words and reducing recognition errors caused by model illusions, thereby improving the recognition accuracy of key fields.

[0013] 3. Enhance error detection capability: By constructing a cross-validation network based on the logical relationships between fields, it can identify systematic errors or contradictory results that are difficult for a single model to detect, such as identifying obviously illogical combinations of fields, thereby improving the reliability of the overall identification results. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A flowchart illustrating a multimodal self-certifying material identification method provided for one or more embodiments of this specification; Figure 2A detailed flowchart illustrating a multimodal self-certifying material identification method for one or more embodiments of this specification. Figure 3 A schematic diagram of a pre-labeled image in an application scenario provided for one or more embodiments of this specification; Figure 4 This is a flowchart illustrating the fusion processing of local recognition results in an application scenario, provided for one or more embodiments of this specification. Figure 5 A schematic diagram of the structure of a multimodal self-certifying material identification device provided for one or more embodiments of this specification; Figure 6 This is a schematic diagram of the structure of a multimodal self-certifying material identification device provided for one or more embodiments of this specification. Detailed Implementation

[0016] This specification provides a method, apparatus, device, and storage medium for identifying multimodal self-certifying materials.

[0017] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0018] Traditional methods of manual review suffer from high costs, low efficiency, and low accuracy.

[0019] For customized OCR recognition, a large number of samples are collected and labeled for each material format, and customized OCR models and information extraction rules are trained. Since a dedicated OCR model needs to be trained for each format of self-certification material, the development cycle is long. Each new material format typically requires 2-4 weeks for sample collection, model training, and testing. Furthermore, when the material format is updated, the model needs to be retrained, resulting in high maintenance costs. In addition, the generalization ability is poor, with low recognition rates for format variants not covered by the training data, and it struggles to handle complex semantics, making it difficult to accurately determine semantic relationships within the material (e.g., monthly income and monthly salary referring to the same information).

[0020] Based on this, with the development of large model technology, and in response to the problems existing in traditional solutions, the embodiments of this specification can use multimodal large models for material identification, taking self-certifying materials as input, and using corresponding prompt words to output the required key fields through multimodal large models.

[0021] However, large models may face the problem of model illusion. Large models are prone to misidentification when processing image information. When faced with images with complex formats or poor quality, the error rate will increase, making it difficult to use them directly in the business execution process.

[0022] This specification provides two specific schemes for material identification using a multimodal large model.

[0023] In the first approach, a single multimodal large model is used for recognition. The self-evidence material is used as the input of the large model to recognize the entire image, and the output of the large model is obtained as the required key field.

[0024] However, this approach suffers from severe model illusion, with error rates reaching 15% to 20% on complex materials, failing to meet the error rate requirements of real-world applications (e.g., in credit operations, error rate requirements are typically below 1%). Furthermore, relying on a single large model lacks a validation mechanism, making it difficult to determine the reliability of the recognition results. It is also highly sensitive to the choice of prompt words, with result quality heavily dependent on prompt word design, resulting in poor stability. Moreover, it exhibits low processing efficiency when dealing with large amounts of information during word processing, leading to long inference times and slow response speeds.

[0025] In the second approach, multiple multimodal large models are used to identify the same certification material, and the same fields are determined by majority voting to determine the final result.

[0026] However, this approach still has some drawbacks. It struggles to identify systematic errors; if multiple models produce the same error for the same reason, the voting mechanism is insufficient to correct it; models typically lack informational correlation and cross-validation, making it difficult to utilize the logical consistency between identification results for verification. Furthermore, it often fails to differentiate between key fields, lacking different validation thresholds for different fields, making it difficult to meet the high accuracy requirements of business applications for key fields. Moreover, the process requires each large model to process all information, resulting in low computational resource utilization and potential resource waste.

[0027] Based on this, the following was proposed: Figure 1 The multimodal self-certifying material identification method is shown. Figure 1This diagram illustrates a multimodal self-certifying material identification method provided in one or more embodiments of this specification. This method can be applied to various business sectors, such as internet finance, e-commerce, instant messaging, gaming, and government services. The process can be executed by computing devices relevant to the specific sector (e.g., intelligent customer service servers or intelligent mobile terminals for payment services). Certain input parameters or intermediate results within the process can be manually adjusted to improve accuracy.

[0028] Figure 1 The process may include the following steps: S102: Obtain the image information corresponding to the self-certification material, and preprocess the image information.

[0029] Generally, self-certification materials can include various file formats, such as text and image formats. Text formats can include Word and PDF formats. For self-certification materials in these readable text formats, key fields can be directly obtained through text reading without further processing.

[0030] However, in practice, to ensure stable data transmission and prevent information tampering or formatting errors, user-uploaded supporting documents are typically presented in image format. Image formats can include JPG, PNG, etc.

[0031] Figure 2 This specification provides a detailed flowchart illustrating a multimodal self-certification material recognition method for one or more embodiments in an application scenario. Image information can be acquired through methods such as uploading via a user terminal, scanning with a scanning device, capturing images with a camera, or interacting with a cloud or platform storing the self-certification materials. When there are multiple self-certification materials, each material can be processed individually.

[0032] Image preprocessing primarily involves standardizing image information, such as resizing, rotation correction, and denoising. Resizing scales the image to a preset size, ensuring consistent input specifications for subsequent model processing and preventing size differences from affecting recognition accuracy. Rotation correction adjusts tilted images to a horizontal position by detecting text direction or edge features, ensuring proper reading of text lines. Denoising uses algorithms like Gaussian filtering and median filtering to remove noise interference such as spots and stripes generated during image capture or transmission, improving image clarity and contrast and laying the foundation for accurate recognition of key fields.

[0033] S104: Based on the material type corresponding to the self-certification material, determine the required key fields.

[0034] Self-certifying materials can include various types of materials, such as identity documents (e.g., ID card, passport), income documents (e.g., bank statements, pay slips), asset documents (e.g., property ownership certificate, vehicle registration certificate), and educational documents (e.g., graduation certificate, degree certificate).

[0035] Specifically, based on the material type corresponding to the self-certification materials, the corresponding template information is determined. The pre-marked key fields in the template information, and the location range of the pre-marked fields for each key field, are then determined.

[0036] Pre-set corresponding template information for self-certification materials of each material type. The template information contains the key fields that need to be identified under that material type.

[0037] The template information can be a pre-set labeled image. The pre-labeled image has the same size as the image information to facilitate the alignment of the pre-labeled image with the pre-processed self-certification material image information in the subsequent multimodal large model. No field values ​​are set for each key field in the pre-labeled image, making it a standard format image without specific user information. The location range of each key field is marked with a bounding box and an identifier.

[0038] Figure 3 This is a schematic diagram of a pre-labeled image in an application scenario provided by one or more embodiments of this specification. In the pre-labeled image corresponding to the ID card, six key fields are set, namely name, gender, ethnicity, birth, address, and ID number. The location range of these fields is marked by corresponding recognition boxes. The color of the recognition boxes can be a high-contrast color such as red or orange, and a corresponding serial number is set in the upper right corner of the recognition box as an identification mark.

[0039] Besides ID card, key fields required for bank statements may include account name, account number, transaction date, transaction amount, transaction type, and balance; key fields required for property ownership certificates may include property owner's name, property address, building area, property type, and property certificate number; key fields required for payslips may include employee name, employee number, department, month of payment, gross salary, and net salary.

[0040] Of course, in addition to pre-labeled images, the location range of key fields can also be defined by coordinate ranges. For example, by using pixel coordinates or relative coordinates, the approximate area of ​​each key field in the image can be determined so that information can be extracted in a targeted manner later.

[0041] S106: For the key fields, generate a corresponding number of cross-validation tasks based on the corresponding task window length and the number of redundant validation coverage times.

[0042] Cross-validation refers to a computer task designed to identify key fields using cross-validation. This task is executed through a multimodal large model to identify the key fields. Specifically, cross-validation involves performing independent identification on a single key field using multiple tasks, obtaining independent identification results from each task, and then cross-validating these results to determine the final identification result for that key field.

[0043] Based on this, the corresponding task window length and redundant verification coverage number are determined for the key fields.

[0044] The task window length determines the number of key fields in a single validation task; a longer task window results in a larger number of key fields. Typically, the task window length for each cross-validation task is kept consistent, as it is a preset static value. However, in some scenarios, different lengths can be set.

[0045] At this point, for each key field, the number of corresponding fields is determined, and the corresponding task window length is determined based on the number of fields. The task window length increases as the number of key fields increases, thereby ensuring that each task window contains an appropriate number of key fields to maintain the contextual relevance of the recognition.

[0046] Redundancy verification coverage refers to the number of cross-validation tasks required for the same key field. The number of redundancy verification coverages for different fields can be the same or different. For example, for a single key field, the number of redundancy verification coverages can be dynamically determined based on its importance. Key fields with higher importance (such as ID card numbers and bank account numbers) use a higher validation depth to increase their redundancy verification coverage, thus reducing the impact of single-shot errors through multiple independent validations. Conversely, key fields with lower importance can use a lower validation depth to reduce their redundancy verification coverage. By constructing a non-uniformly distributed redundancy verification coverage, important key fields are ensured to receive more validation, maintaining high accuracy of key information while reducing overall computational resource consumption.

[0047] The task window length is used as the number of key fields contained in a single cross-validation task, and through permutation and combination, a single key field can be covered by cross-validation tasks with a number of redundant validation coverage.

[0048] The process of permutation and combination can be achieved using a cyclic sliding window approach. Starting with the first key field, a consecutive number of key fields of the task window length are selected to form a cross-validation task. The window then slides forward one key field, and the same number of key fields are selected to form the next task, until all key fields are covered and the number of tasks containing each key field reaches its corresponding redundancy verification coverage count. For example, for an ID card, which contains 6 key fields to be identified: name, gender, ethnicity, date of birth, address, and ID number, the task window length and redundancy verification coverage count are both set to 3. This means that each cross-validation task contains 3 key fields, and each key field is covered by 3 cross-validation tasks. Through this process, a total of 6 cross-validation tasks are obtained: Task 1: (name, gender, ethnicity), Task 2: (gender, ethnicity, date of birth), Task 3: (ethnicity, date of birth, address), Task 4: (date of birth, address, ID number), Task 5: (address, ID number, name), and Task 6: (ID number, name, gender). In this way, each key field is covered by multiple task windows with different combinations, providing diverse identification criteria for subsequent cross-validation.

[0049] S108: Through multiple multimodal large models, each model independently executes its corresponding cross-validation task and outputs local recognition results for some key fields included in the cross-validation task.

[0050] Multimodal large models can handle data from various modalities, including text and image. Examples include General Multimodal Large Language Models (MLLMs) and Vision-Language Large Models (VLLMs). These models are pre-configured and can be hosted locally or in the cloud, accessed through appropriate interfaces. Furthermore, based on specific needs, open-source large models can be vertically trained using training samples from relevant domains to improve recognition accuracy in those domains.

[0051] Specifically, the relationship between the multimodal large model and the cross-validation task can be one-to-one or one-to-many, with one multimodal large model executing one or more cross-validation tasks. If the number of pre-set multimodal large models is sufficient and the computing power is sufficient, then a one-to-one approach is used to execute the tasks.

[0052] If the number of multimodal large models is insufficient, or the current available computing power is low, a one-to-many approach can be adopted, where one multimodal large model performs multiple cross-validation tasks. In this case, to increase the effectiveness of cross-validation, the number of times the same key field appears in the tasks performed by a single multimodal large model should be minimized to avoid the continuous impact of the model's own systematic bias on the multiple identifications of the same key field.

[0053] This explanation uses a one-to-one approach, establishing a one-to-one correspondence between a corresponding number of cross-validation tasks and multiple multimodal large models. When assigning these tasks, a random method can be used, or the assignment can be based on the historical performance of the multimodal large models. For example, cross-validation tasks containing important key fields can be assigned to models with higher past recognition accuracy to improve the reliability of key information recognition.

[0054] For a single cross-validation task, the input includes prompts, image information, and pre-labeled images corresponding to certain key fields. The corresponding multimodal large-scale model then outputs local recognition results for those key fields. The input includes pre-processed image information, pre-labeled images of the corresponding type, and the corresponding prompts. Using both image information and pre-labeled images as input allows the multimodal large-scale model to more accurately understand the material structure, reducing over-interpretation of the material content. Furthermore, location information aids recognition, minimizing model illusions. When the format of the self-validating material is updated, only the pre-labeled images need to be updated, eliminating the need for retraining and enabling rapid adaptation to new material formats.

[0055] Here is an example of a prompt: Please identify the information in the original image based on the location markers in the pre-labeled image; in the pre-labeled image, red box 1 represents [field A], red box 2 represents [field B], and red box 3 represents [field C]; please output only the following JSON format: {field A: “value”, field B: “value”, field C: “value”}; if you cannot determine a field, please output “N / A”.

[0056] For multimodal large models, the corresponding cross-validation task only includes the identification of some key fields. Therefore, the key fields identified are only a part of the key fields, and the corresponding identification results are called local identification results.

[0057] S110: The local recognition results are fused to obtain the overall recognition result for the key field.

[0058] Specifically, the fusion process mainly includes two core steps: conflict verification and result correction.

[0059] During conflict verification, for a single key field, its corresponding sub-identification result in each local identification result is determined. Taking a redundancy verification coverage of 3 times as an example, a single key field is covered by 3 cross-validation tasks, meaning that a sub-identification result of this key field exists in all 3 local identification results.

[0060] like Figure 2 As shown, if a preset number of sub-identification results are consistent, then that consistent sub-identification result will be used as the identification result corresponding to that key field in the fused overall identification result. A corresponding verification threshold is preset, and multiple verification modes can exist. In strict mode, the preset number corresponding to the verification threshold is usually set to require all sub-identification results to be consistent before accepting the sub-identification result. In lenient mode, the preset number corresponding to the verification threshold can be appropriately reduced. For example, if the redundant verification coverage is 3 times, the preset number can be set to 2. As long as two sub-identification results are consistent, the identification result is considered correct and used as the identification result in the final overall identification result.

[0061] For critical fields that fail conflict validation, a result correction step can be performed. In this environment, multiple correction methods can be tried. For example, retrying can be performed, reassigning the cross-validation task corresponding to the critical field to different multimodal large models, and setting an upper limit for the number of retries. If the validation still fails after reaching the upper limit, a corresponding error report can be generated. The error report includes the original image information, the local recognition results of each multimodal large model, and the pre-labeled image. The validation threshold can be downgraded from strict mode to lenient mode. Finally, the result can be submitted for manual review and confirmation.

[0062] Alternatively, an adaptive verification threshold mechanism can be implemented. The system automatically evaluates image quality (including dimensions such as sharpness and integrity) and records the historical recognition accuracy for each material, which serves as the basis for threshold adjustment. Based on material quality and historical recognition accuracy, the verification threshold is dynamically adjusted. For materials with high image quality, a stricter verification threshold is used (e.g., setting the preset number to all), while for materials with low image quality, a more lenient threshold is used (e.g., setting the preset number to 2 when the redundant verification coverage is 3). This reduces the number of retries and manual interventions while maintaining overall accuracy.

[0063] Alternatively, a progressive identification strategy can be adopted, first performing coarse-grained identification, and then performing fine identification for uncertain areas. In the first stage, a single large model is used to quickly identify all fields and identify potentially problematic fields. In the second stage, cross-validation is initiated only for problematic fields. In the third stage, more detailed prompts are used to identify fields that are still uncertain, thereby significantly reducing the consumption of computing resources and improving processing speed.

[0064] 1. Rapid adaptation capability: No need to train a dedicated OCR model for each material format. Adaptation to new material formats can be completed within minutes using pre-labeled images in the template information, reducing deployment time and cost.

[0065] 2. Improve recognition accuracy: Through cross-validation and task decomposition, each model processes only a small number of fields, reducing the complexity of prompt words and reducing recognition errors caused by model illusions, thereby improving the recognition accuracy of key fields.

[0066] 3. Enhance error detection capability: By constructing a cross-validation network based on the logical relationships between fields, it can identify systematic errors or contradictory results that are difficult for a single model to detect, such as identifying obviously illogical combinations of fields, thereby improving the reliability of the overall identification results.

[0067] In one or more embodiments of this specification, in the process of fusing the identification results of each locality, in addition to selecting the majority consensus method, other methods can also be adopted based on requirements.

[0068] Figure 4 This specification provides a flowchart illustrating the fusion processing of local recognition results in an application scenario, as provided in one or more embodiments. For each key field, the corresponding sub-recognition result within each local recognition result is determined, and the sub-confidence of each multimodal large model for each key field is determined. Here, a local recognition result refers to the recognition result for multiple key fields included in a single cross-validation task, while a sub-recognition result refers to the recognition result for a single key field within the local recognition results.

[0069] For large multimodal models, when outputting the results of each sub-identification, the corresponding sub-confidence score is usually also output simultaneously. Generally, the sub-confidence scores are not explicitly output; they can be obtained by parsing the internal output of the large multimodal model or through post-processing.

[0070] For one or more candidate identification values ​​appearing in the sub-identification results, the total confidence level corresponding to each candidate identification value is determined based on its sub-confidence level in each sub-identification result. A candidate identification value refers to different identification contents that may appear in a local identification result for the same key field. For example, if there are four multimodal large models, and for the key field of "user identifier," three multimodal large models output a sub-identification result of "U001," and one multimodal large model outputs a sub-identification result of "U007," then there are two candidate identification values: "U001" and "U007." When each multimodal large model outputs a sub-identification result and sub-confidence level, only the sub-identification result with the highest sub-confidence level needs to be output.

[0071] At this point, for a single candidate recognition value, the total confidence score is obtained by weighted summing of all its corresponding sub-confidence scores. Typically, each multimodal model can be assigned equal weight. For example, if there are four multimodal models, each model has a weight of 0.25. Assuming the sub-confidence scores of the three multimodal models with recognition results "U001" are 0.9, 0.8, and 0.9 respectively, the total confidence score of the candidate recognition value "U001" obtained by weighted summing is: 0.9*0.25 + 0.8*0.25 + 0.9*0.25 = 0.65. Similarly, assuming the sub-confidence score of the multimodal model with recognition result "U007" is 0.6, the total confidence score of the candidate recognition value "U007" obtained by weighted summing is: 0.6*0.25 = 0.15.

[0072] Based on the overall confidence score, the overall identification result for the key field is obtained. Generally, for each field, if there is only one candidate identification value, then as long as it is higher than the preset confidence score lower limit (e.g., it needs to be higher than 0.7, which can be dynamically adjusted based on requirements), the candidate identification value can be considered as the final sub-identification result and added to the overall identification result. However, if there are multiple candidate identification values, it is necessary to select the candidate identification value with the highest overall confidence score to determine whether it is higher than the preset confidence score lower limit.

[0073] Furthermore, considering that different multimodal large models may have different recognition capabilities for different types of key fields, such as Chinese, English, numbers, or printed and handwritten text, the recognition capabilities of different multimodal large models may differ. Simply assigning cross-validation tasks randomly may introduce noise during the fusion stage, affecting the final recognition results.

[0074] Based on this, for the target multimodal large model to be evaluated and corrected, the key field with the highest confidence in the single-field recognition task is determined among the key fields, and is used as the anchor field corresponding to the target multimodal large model. The key field with the lowest confidence in the single-field recognition task is determined as the stress field corresponding to the target multimodal large model.

[0075] Among them, the target multimodal large model to be evaluated and corrected refers to the multimodal large model whose recognition ability, stability, and performance on specific fields need to be evaluated in the current recognition process, and whose parameters need to be adjusted or task allocation optimized based on the evaluation results. Whether a multimodal large model is a target multimodal large model can be determined through the capability verification results of each multimodal large model, historical output results, manual selection, etc.

[0076] Single-field recognition tasks refer to tasks where the target multimodal large model only includes a single field during historical validation tasks, demonstrating the model's ideal recognition capabilities for each key field. However, as the number of fields in the validation task increases (i.e., the task window length increases), the model needs to process information from different fields simultaneously, which may lead to issues such as attention distraction and interference between fields, resulting in deviations in actual recognition performance compared to single-field recognition tasks.

[0077] At this point, for the anchor field and stress field, based on the corresponding task window length and redundant validation coverage, a corresponding number of cross-validation tasks are generated. The cross-validation tasks that simultaneously contain both the anchor field and stress field are executed by the target multimodal large model. Based on the sub-recognition results of the anchor field by the target multimodal large model, the capability of the target multimodal large model is evaluated and the results are corrected.

[0078] Based on the task window length and the number of redundant verification coverages, cross-validation tasks are generated in the manner described above. For example, still targeting the ID card, which contains 6 key fields to be identified, the task window length and the number of redundant verification coverages are both set to 3. A total of 6 cross-validation tasks are obtained: Task 1: (Name, Gender, Ethnicity), Task 2: (Gender, Ethnicity, Birth Date), Task 3: (Ethnicity, Birth Date, Address), Task 4: (Birth Date, Address, ID Card Number), Task 5: (Address, ID Card Number, Name), and Task 6: (ID Card Number, Name, Gender).

[0079] Assuming the anchor field is the name and the stress field is the ID number, the cross-validation tasks that include both the anchor field and the stress field are Task 5 and Task 6. If each multimodal large model executes a single cross-validation task, then one of these two tasks can be selected for execution by the target multimodal large model.

[0080] At this point, capability assessment and result correction can be carried out through multiple dimensions.

[0081] In the first dimension, the output consensus consistency of the target multimodal large model is obtained based on its first sub-identification result for the anchor field and the second sub-identification results of other multimodal large models for the anchor field. Output consensus consistency refers to the degree of consistency between the target multimodal large model's output and that of other multimodal large models in fields it excels at identifying. It reflects whether the target multimodal large model exhibits identification bias or instability during the current task execution. For example, if most other multimodal large models' second sub-identification results for the anchor field "name" are "first name," and the target multimodal large model's first sub-identification result for "name" is also "first name," it indicates a high degree of consensus consistency between the target multimodal large model and other models in the anchor field, suggesting relatively strong reliability of its identification results. Conversely, if the target multimodal large model's first sub-identification result is inconsistent with most second sub-identification results, identifying it as "second name," it indicates that the target multimodal large model may have some identification bias or instability in the current task. If the first sub-identification result of the target multimodal large model is consistent with the majority of the other second sub-identification results, the output consensus consistency level can be marked as 1. Otherwise, the output consensus consistency level is determined based on the number of values ​​in the minority of other second sub-identification results that are the same as the first sub-identification result.

[0082] In the second dimension, the confidence decay of the target multimodal large model is obtained based on the first sub-confidence of its first sub-identification result for the anchor field and the second sub-confidence of its identification result for the anchor field in a single-field identification task. The confidence decay refers to the decrease in the confidence of the target multimodal large model in identifying its preferred anchor field when handling cross-validation tasks involving multiple fields, compared to the confidence in identifying a single-field task containing only that anchor field. This metric reflects whether the multimodal large model's ability to identify its preferred fields in complex tasks (i.e., tasks involving multiple fields) is affected by interference from other field information or its own attention allocation. For example, in a single-field recognition task, the second sub-confidence of the target multimodal large model for the anchor field "name" is 0.95, while in a cross-validation task that includes "name" and "ID number" (stress field), its first sub-confidence for "name" is 0.85, and the confidence decay is 0.1.

[0083] In the third dimension, the confidence pressure impact level of the target multimodal large model is obtained based on the third sub-confidence of the third sub-identification result of the stress field. The confidence pressure impact level refers to the recognition confidence level of the target multimodal large model for the stress field (i.e., the field with the lowest confidence in its single-field recognition task) when processing cross-validation tasks that simultaneously contain anchor fields and stress fields. This indicator can reflect the impact of the existence of the stress field on the overall recognition status of the target multimodal large model. For example, in single-field recognition tasks, the sub-confidence of the target multimodal large model for the recognition result of the stress field "ID number" is usually low, such as 0.6. However, in cross-validation tasks that include "name" (anchor field) and "ID number" (stress field), if the third sub-confidence for "ID number" further decreases to 0.5, it indicates that the stress field has a significant negative impact on the overall recognition status of the model, which may exacerbate the recognition difficulty and uncertainty of the model. If the third sub-confidence remains at around 0.6 or slightly increases, it indicates that the model can still maintain a certain degree of stability when dealing with complex tasks that include fields that it is not good at.

[0084] Based on the degree of consensus in the output, the degree of confidence decay, and the degree of influence of confidence pressure, corresponding correction coefficients are obtained. These correction coefficients are then used to adjust the confidence of each local recognition result output by the target multimodal large model. The correction coefficients are obtained by standardizing the degree of consensus in the output, the degree of confidence decay, and the degree of influence of confidence pressure, and then performing a weighted summation.

[0085] For example, assuming that in the first dimension, the first sub-identification result of the target multimodal large model for the anchor field is consistent with the majority of the sub-identification results, then the output consensus consistency level is set to 1.

[0086] Assuming the confidence decay rate in the second dimension is 0.1, after standardization, this confidence decay rate is compared with the preset standard confidence decay rate value. If it is lower than the standard confidence decay rate value (within the allowed decay range, for example, set to 0.08), it indicates that the target multimodal large model has a small decay in its ability to identify anchor fields in complex tasks, and the standardized confidence decay rate value can be set to 1. If it is higher than the standard confidence decay rate value, it indicates that the decay of the target multimodal large model is large, and the standardized value can be obtained by dividing the two. In this case, assuming the standard confidence decay rate value is 0.08, the standardized confidence decay rate is 0.08 / 0.1 = 0.8.

[0087] Assuming the third sub-confidence score of the output third sub-identification result in the third dimension is 0.6, similar to the second dimension, it can be compared with the preset standard confidence score pressure influence value. When it is higher than the standard confidence score pressure influence value (the allowable pressure influence range, for example, set to 0.8), it indicates that its recognition ability of the pressure field is high, and the standardized value of the confidence score pressure influence can be set to 1. When it is lower than the standard confidence score pressure influence value, the two are divided to obtain the standardized value. At this time, assuming the standard confidence score pressure influence value is 0.8, the standardized confidence score pressure influence is 0.6 / 0.8 = 0.75.

[0088] At this point, the consensus level, confidence decay, and confidence pressure impact of the standardized output are 1, 0.8, and 0.75, respectively. Multiplying these three values ​​yields the final correction coefficient: 1 * 0.8 * 0.75 = 0.6. This correction coefficient represents the overall reliability level of the target multimodal large model in the current cross-validation task. The smaller the correction coefficient, the greater the negative impact on the model's recognition stability and accuracy under complex task environments. After obtaining the correction coefficient, the sub-confidence of all local recognition results output by the target multimodal large model is multiplied by this correction coefficient to achieve dynamic adjustment of its confidence.

[0089] Based on the same idea, one or more embodiments of this specification also provide apparatus and devices corresponding to the above methods, such as... Figure 5 , Figure 6 As shown.

[0090] Figure 5 A schematic diagram of a multimodal self-certifying material identification device provided for one or more embodiments of this specification, the device comprising: The material acquisition module 502 acquires image information corresponding to the self-certification material and preprocesses the image information. The field determination module 504 determines the required key fields based on the material type corresponding to the self-certification material; The task generation module 506 generates a corresponding number of cross-validation tasks for the key fields based on the corresponding task window length and the number of redundant validation coverage times. The task execution module 508 independently executes its corresponding cross-validation task through multiple multimodal large models, and outputs the local recognition results for some key fields contained in the cross-validation task. The result fusion module 510 fuses the local recognition results to obtain the overall recognition result for the key field.

[0091] Optionally, the field determination module 504 determines the corresponding template information based on the material type corresponding to the self-certification material; Determine each key field that is pre-marked in the template information, and the position range of the pre-marked field corresponding to each key field.

[0092] Optionally, the template information is a pre-labeled image; The pre-labeled image has the same size as the image information. In the pre-labeled image, no field values ​​are set for each key field, and the position range of each key field is marked with a box and an identifier.

[0093] Optionally, the task generation module 506 determines the corresponding task window length and the number of redundant verification coverage times for the key fields; The task window length is used as the number of key fields contained in a single cross-validation task, and through permutation and combination, a single key field can be covered by cross-validation tasks with the number of redundant validation coverage times.

[0094] Optionally, the task generation module 506 determines the number of fields corresponding to the key fields and determines the corresponding task window length based on the number of fields. For a single key field, the number of redundant validation coverage times is dynamically determined based on its importance.

[0095] Optionally, the task execution module 508 establishes a one-to-one correspondence between the corresponding number of cross-validation tasks and multiple multimodal large models; For a single cross-validation task, the prompt words corresponding to some of its key fields, the image information, and the pre-labeled image are taken as input, and the corresponding multimodal large model is used to output the local recognition result for that part of the key fields.

[0096] Optionally, the result fusion module 510 determines the sub-identification result corresponding to a single key field in each local identification result; If a preset number of sub-identification results are consistent, then the consistent sub-identification result will be used as the identification result corresponding to the key field in the fused overall identification result.

[0097] Optionally, the result fusion module 510 determines the sub-identification result corresponding to each key field in each local identification result, and determines the sub-confidence of each multimodal large model for each key field. For one or more candidate identification values ​​appearing in the sub-identification results, the total confidence level corresponding to each candidate identification value is determined based on its sub-confidence level in each sub-identification result. Based on the total confidence level, the overall identification result for the key field is obtained.

[0098] Optionally, the task generation module 506, for the target multimodal large model to be evaluated and corrected, determines the key field with the highest confidence in the single-field recognition task among the key fields, as the anchor field corresponding to the target multimodal large model, and determines the key field with the lowest confidence in the single-field recognition task, as the stress field corresponding to the target multimodal large model. For the anchor field and the stress field, based on the corresponding task window length and the number of redundant verification coverage, a corresponding number of cross-validation tasks are generated. The cross-validation tasks that simultaneously include the anchor field and the stress field are executed by the target multimodal large model to evaluate the capabilities of the target multimodal large model and correct the results based on the sub-identification results of the anchor field by the target multimodal large model.

[0099] Optionally, the task generation module 506 obtains the output consensus consistency degree corresponding to the target multimodal large model based on the first sub-identification result of the target multimodal large model on the anchor field and the second sub-identification result of other multimodal large models on the anchor field. Based on the first sub-confidence of the first sub-identification result of the target multimodal large model for the anchor field, and the second sub-confidence of the identification result of the target multimodal large model for the anchor field in the single-field identification task, the confidence decay degree corresponding to the target multimodal large model is obtained. Based on the third sub-confidence of the third sub-identification result of the target multimodal large model for the stress field, the confidence stress influence degree corresponding to the target multimodal large model is obtained; Based on the degree of consensus in the output, the degree of confidence decay, and the degree of influence of confidence pressure, corresponding correction coefficients are obtained, and the confidence of each local recognition result output by the target multimodal large model is corrected using the correction coefficients.

[0100] Figure 6 A schematic diagram of a multimodal self-certifying material identification device provided for one or more embodiments of this specification, the device comprising: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: Obtain the image information corresponding to the self-certification materials, and preprocess the image information; Based on the material type corresponding to the self-certification materials, determine the required key fields; For the key fields, a corresponding number of cross-validation tasks are generated based on the corresponding task window length and the number of redundant validation coverage times; Multiple multimodal large models are used to independently execute their respective cross-validation tasks, and output local recognition results for some key fields included in the cross-validation task. The results of each local identification are fused together to obtain the overall identification result for the key field.

[0101] Based on the same idea, one or more embodiments of this specification also provide a non-volatile computer storage medium corresponding to the above method, storing computer-executable instructions, wherein the computer-executable instructions are configured as follows: Obtain the image information corresponding to the self-certification materials, and preprocess the image information; Based on the material type corresponding to the self-certification materials, determine the required key fields; For the key fields, a corresponding number of cross-validation tasks are generated based on the corresponding task window length and the number of redundant validation coverage times; Multiple multimodal large models are used to independently execute their respective cross-validation tasks, and output local recognition results for some key fields included in the cross-validation task. The results of each local identification are fused together to obtain the overall identification result for the key field.

[0102] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog are commonly used. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0103] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0104] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0105] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.

[0106] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0107] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0108] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0109] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0110] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0111] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0112] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0113] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0114] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0115] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0116] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0117] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.

Claims

1. A method for identifying multimodal self-verifying materials, comprising: Obtain the image information corresponding to the self-certification materials, and preprocess the image information; Based on the material type corresponding to the self-certification materials, determine the required key fields; For the key fields, a corresponding number of cross-validation tasks are generated based on the corresponding task window length and the number of redundant validation coverage times; Multiple multimodal large models are used to independently execute their respective cross-validation tasks, and output local recognition results for some key fields included in the cross-validation task. The results of each local identification are fused together to obtain the overall identification result for the key field.

2. The method as described in claim 1, based on the material type corresponding to the self-certification material, determines the required key fields, specifically including: Based on the material type corresponding to the self-certification materials, determine the corresponding template information; Determine each key field that is pre-marked in the template information, and the position range of the pre-marked field corresponding to each key field.

3. The method as described in claim 2, wherein the template information is a pre-labeled image; The pre-labeled image has the same size as the image information. In the pre-labeled image, no field values ​​are set for each key field, and the position range of each key field is marked with a box and an identifier.

4. The method as described in claim 1, wherein for the key field, based on the corresponding task window length and the number of redundant verification coverage times, a corresponding number of cross-validation tasks are generated, specifically including: For the aforementioned key fields, determine the corresponding task window length and the number of redundant verification coverage times; The task window length is used as the number of key fields contained in a single cross-validation task, and through permutation and combination, a single key field can be covered by cross-validation tasks with the number of redundant validation coverage times.

5. The method as described in claim 4, wherein for the key field, determining the corresponding task window length and redundant verification coverage count specifically includes: For the key fields, determine the number of fields corresponding to them, and determine the corresponding task window length based on the number of fields; For a single key field, the number of redundant validation coverage times is dynamically determined based on its importance.

6. The method as described in claim 3, wherein multiple multimodal large models are used to independently execute their respective corresponding cross-validation tasks, and local recognition results for some key fields included in the cross-validation task are output, specifically including: Establish a one-to-one correspondence between the corresponding number of cross-validation tasks and multiple multimodal large models; For a single cross-validation task, the prompt words corresponding to some of its key fields, the image information, and the pre-labeled image are taken as input, and the corresponding multimodal large model is used to output the local recognition result for that part of the key fields.

7. The method as described in claim 1, wherein the local recognition results are fused to obtain the overall recognition result for the key field, specifically including: For a single key field, determine its corresponding sub-identification result in each local identification result; If a preset number of sub-identification results are consistent, then the consistent sub-identification result will be used as the identification result corresponding to the key field in the fused overall identification result.

8. The method as described in claim 1, wherein the local recognition results are fused to obtain the overall recognition result for the key field, specifically including: For each key field, determine its corresponding sub-identification result in each local identification result, and determine the sub-confidence of each multimodal large model for each key field; For one or more candidate identification values ​​appearing in the sub-identification results, the total confidence level corresponding to each candidate identification value is determined based on its sub-confidence level in each sub-identification result. Based on the total confidence level, the overall identification result for the key field is obtained.

9. The method as described in claim 8, wherein for the key field, based on the corresponding task window length and the number of redundant verification coverage times, a corresponding number of cross-validation tasks are generated, specifically including: For the target multimodal large model to be evaluated and corrected, among each key field, the key field with the highest confidence in the single field recognition task is determined as the anchor field corresponding to the target multimodal large model, and the key field with the lowest confidence in the single field recognition task is determined as the stress field corresponding to the target multimodal large model. For the anchor field and the stress field, based on the corresponding task window length and the number of redundant verification coverage, a corresponding number of cross-validation tasks are generated. The cross-validation tasks that simultaneously include the anchor field and the stress field are executed by the target multimodal large model to evaluate the capabilities of the target multimodal large model and correct the results based on the sub-recognition results of the anchor field by the target multimodal large model.

10. The method as described in claim 9, wherein the target multimodal large model is subjected to capability assessment and result correction based on the sub-identification results of the anchor field by the target multimodal large model, specifically including: Based on the first sub-identification result of the target multimodal large model for the anchor field, and the second sub-identification results of other multimodal large models for the anchor field, the output consensus consistency degree corresponding to the target multimodal large model is obtained; Based on the first sub-confidence of the first sub-identification result of the target multimodal large model for the anchor field, and the second sub-confidence of the identification result of the target multimodal large model for the anchor field in the single-field identification task, the confidence decay degree corresponding to the target multimodal large model is obtained. Based on the third sub-confidence of the third sub-identification result of the target multimodal large model for the stress field, the confidence stress influence degree corresponding to the target multimodal large model is obtained; Based on the degree of consensus in the output, the degree of confidence decay, and the degree of influence of confidence pressure, corresponding correction coefficients are obtained, and the confidence of each local recognition result output by the target multimodal large model is corrected using the correction coefficients.

11. A multimodal self-certifying material identification device, comprising: The material acquisition module acquires image information corresponding to the self-certification materials and preprocesses the image information. The field determination module determines the required key fields based on the material type corresponding to the self-certification material; The task generation module generates a corresponding number of cross-validation tasks for the key fields, based on the corresponding task window length and the number of redundant validation coverage times. The task execution module independently executes its respective cross-validation task using multiple multimodal large models, and outputs local recognition results for some key fields included in the cross-validation task. The result fusion module merges the local recognition results to obtain the overall recognition result for the key field.

12. The apparatus of claim 11, wherein the field determination module determines the corresponding template information based on the material type corresponding to the self-certification material; Determine each key field that is pre-marked in the template information, and the position range of the pre-marked field corresponding to each key field.

13. The apparatus of claim 12, wherein the template information is a pre-labeled image; The pre-labeled image has the same size as the image information. In the pre-labeled image, no field values ​​are set for each key field, and the position range of each key field is marked with a box and an identifier.

14. The apparatus of claim 11, wherein the task generation module determines the corresponding task window length and the number of redundant verification coverage times for the key field; The task window length is used as the number of key fields contained in a single cross-validation task, and through permutation and combination, a single key field can be covered by cross-validation tasks with the number of redundant validation coverage times.

15. The apparatus of claim 14, wherein the task generation module determines the number of fields corresponding to the key fields, and determines the corresponding task window length based on the number of fields; For a single key field, the number of redundant validation coverage times is dynamically determined based on its importance.

16. The apparatus of claim 13, wherein the task execution module establishes a one-to-one correspondence between the corresponding number of cross-validation tasks and multiple multimodal large models; For a single cross-validation task, the prompt words corresponding to some of its key fields, the image information, and the pre-labeled image are taken as input, and the corresponding multimodal large model is used to output the local recognition result for that part of the key fields.

17. The apparatus of claim 11, wherein the result fusion module determines, for a single key field, the corresponding sub-identification result in each local identification result; If a preset number of sub-identification results are consistent, then the consistent sub-identification result will be used as the identification result corresponding to the key field in the fused overall identification result.

18. The apparatus of claim 11, wherein the result fusion module determines, for each key field, the corresponding sub-identification result in each local identification result, and determines the sub-confidence of each multimodal large model for each key field; For one or more candidate identification values ​​appearing in the sub-identification results, the total confidence level corresponding to each candidate identification value is determined based on its sub-confidence level in each sub-identification result. Based on the total confidence level, the overall identification result for the key field is obtained.

19. The apparatus of claim 18, wherein the task generation module, for the target multimodal large model to be evaluated and corrected, determines, among the key fields, the key field with the highest confidence in the single-field recognition task as the anchor field corresponding to the target multimodal large model, and determines, among the key fields, the key field with the lowest confidence in the single-field recognition task as the stress field corresponding to the target multimodal large model; For the anchor field and the stress field, based on the corresponding task window length and the number of redundant verification coverage, a corresponding number of cross-validation tasks are generated. The cross-validation tasks that simultaneously include the anchor field and the stress field are executed by the target multimodal large model to evaluate the capabilities of the target multimodal large model and correct the results based on the sub-recognition results of the anchor field by the target multimodal large model.

20. The apparatus of claim 19, wherein the task generation module obtains the output consensus consistency degree corresponding to the target multimodal large model based on the first sub-identification result of the target multimodal large model on the anchor field and the second sub-identification results of other multimodal large models on the anchor field; Based on the first sub-confidence of the first sub-identification result of the target multimodal large model for the anchor field, and the second sub-confidence of the identification result of the target multimodal large model for the anchor field in the single-field identification task, the confidence decay degree corresponding to the target multimodal large model is obtained. Based on the third sub-confidence of the third sub-identification result of the target multimodal large model for the stress field, the confidence stress influence degree corresponding to the target multimodal large model is obtained; Based on the degree of consensus in the output, the degree of confidence decay, and the degree of influence of confidence pressure, corresponding correction coefficients are obtained, and the confidence of each local recognition result output by the target multimodal large model is corrected using the correction coefficients.

21. A multimodal self-certifying material identification device, comprising: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: Obtain the image information corresponding to the self-certification materials, and preprocess the image information; Based on the material type corresponding to the self-certification materials, determine the required key fields; For the key fields, a corresponding number of cross-validation tasks are generated based on the corresponding task window length and the number of redundant validation coverage times; Multiple multimodal large models are used to independently execute their respective cross-validation tasks, and output local recognition results for some key fields included in the cross-validation task. The results of each local identification are fused together to obtain the overall identification result for the key field.