Information inspection method, device, equipment and storage medium based on large model

By combining the correction of the initial recognition results with the target scene information using a large model, the problem of low accuracy in identifying key information is solved, and highly accurate and intelligent recognition results are achieved, which is suitable for document verification scenarios.

CN119296115BActive Publication Date: 2025-09-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411171895.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2025-09-23
Estimated Expiration
2044-08-23

AI Technical Summary

Technical Problem

In the existing technology, the accuracy of key information recognition is low, especially in the document verification scenario, it is difficult to effectively improve the accuracy of the recognition results.

Method used

The initial recognition result of the image to be detected is corrected by using the large model, and the target scene information is combined to correct and output the target recognition result through the target large model, including the key information of the image to be detected.

Benefits of technology

It improves the accuracy and intelligence of recognition results, reduces manual intervention, and enhances user experience, especially in the bidding document verification scenario, achieving efficient and automated authenticity identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119296115B_ABST
    Figure CN119296115B_ABST
Patent Text Reader

Abstract

The present disclosure provides a large-scale model-based information detection method, device, equipment and storage medium, which relates to the fields of data processing and image processing, and in particular to the technical fields of artificial intelligence, large-scale models, etc. The specific implementation scheme is: obtaining an image to be detected, and obtaining an initial recognition result of the image to be detected; wherein, the initial recognition result is the result obtained after performing text recognition on the image to be detected; at least the image to be detected is input into the target large-scale model to obtain the target scene information of the image to be detected; the target scene information and the initial recognition result are input into the target large-scale model to obtain an output target recognition result, wherein the target recognition result is at least a result obtained by correcting the initial recognition result and matching the target scene information, including key information containing the image to be detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields of data processing and image processing, and in particular to technical fields such as artificial intelligence and large models. Background Art

[0002] The identification of key information in documents often occurs in various scenarios, and its importance cannot be ignored. For example, in the verification scenario of bidding / tendering documents, the accuracy of the key information in the documents is particularly important; however, in existing key information recognition scenarios, the accuracy of the key information is often low. Therefore, how to effectively improve the accuracy of the recognition results has become a problem that needs to be solved urgently. Summary of the Invention

[0003] The present disclosure provides a large model-based information inspection method, device, equipment and storage medium.

[0004] According to one aspect of the present disclosure, a large model-based information detection method is provided, comprising:

[0005] Acquire an image to be detected, and obtain an initial recognition result of the image to be detected; wherein the initial recognition result is a result obtained after performing text recognition on the image to be detected;

[0006] At least inputting the image to be detected into a target large model to obtain target scene information of the image to be detected;

[0007] The target scene information and the initial recognition result are input into the target large model to obtain an output target recognition result, wherein the target recognition result is at least a result obtained by correcting the initial recognition result and matching the target scene information, including key information of the image to be detected.

[0008] According to another aspect of the present disclosure, there is provided an information detection device based on a large model, comprising:

[0009] A pre-processing unit, configured to obtain an image to be detected and obtain an initial recognition result of the image to be detected; wherein the initial recognition result is a result obtained after performing text recognition on the image to be detected;

[0010] A target processing unit is used to input at least the image to be detected into a target large model to obtain target scene information of the image to be detected; input the target scene information and the initial recognition result into the target large model to obtain an output target recognition result; wherein the target recognition result is at least a result obtained by correcting the initial recognition result and matching the target scene information, including key information of the image to be detected.

[0011] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0012] at least one processor; and

[0013] a memory communicatively connected to the at least one processor; wherein,

[0014] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.

[0015] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0016] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.

[0017] In this way, the disclosed solution effectively utilizes the large model to correct the initial recognition results of the image to be detected. Compared with the existing recognition methods, the target recognition results obtained by the disclosed solution are more accurate. Moreover, the above process does not require human intervention, and is intelligent, efficient and accurate, thereby providing strong support for improving user experience.

[0018] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0020] Figure 1 This is a flow diagram of the information detection method based on a large model according to an embodiment of the present disclosure. Figure 1 ;

[0021] Figure 2 This is a flow diagram of the information detection method based on a large model according to an embodiment of the present disclosure. Figure 2 ;

[0022] Figure 3 1 is a schematic diagram of a process for determining training data required for training a prompt word generator according to an embodiment of the present disclosure;

[0023] Figure 41 is a schematic diagram of the processing flow of a prompt word generator in the information detection method based on a large model according to an embodiment of the present disclosure;

[0024] Figure 5 is a schematic diagram of the connection relationship between modules in the preset recognition model according to an embodiment of the present disclosure;

[0025] Figure 6 is a flowchart of a specific example of a large model-based information detection method according to an embodiment of the present disclosure;

[0026] Figure 7 is a structural diagram of an information detection device based on a large model according to an embodiment of the present disclosure;

[0027] Figure 8 It is a block diagram of an electronic device used to implement the large model-based information detection method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0029] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to multiple similar technical terms and distinguish them, and do not mean to limit the order or to limit to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.

[0030] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0031] The following describes the related technologies of the embodiments of the present disclosure. The following related technologies are optional solutions that can be combined with the technical solutions of the embodiments of the present disclosure in any way, and all of them fall within the protection scope of the embodiments of the present disclosure.

[0032] Large models usually refer to machine learning models with large parameters and the ability to handle complex tasks, especially deep learning models.

[0033] Fine-tuning is a key concept in machine learning, particularly in fields like natural language processing (NLP) and computer vision. Fine-tuning involves further training a pre-trained model using labeled data for a specific task to optimize its performance on that task. Pre-trained models are typically trained on large datasets to learn general feature representations, while fine-tuning builds on these general features to adapt the model to specific application scenarios.

[0034] Figure 1 This is a schematic flow chart of an information detection method based on a large model according to an embodiment of the present application. Figure 1 The method may be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.

[0035] Furthermore, the method includes at least part of the following contents. Figure 1 As shown, including:

[0036] Step S101: obtaining an image to be detected and obtaining an initial recognition result of the image to be detected.

[0037] Here, the initial recognition result is a result obtained after performing text recognition on the image to be detected. For example, in one example, the initial recognition result includes at least part of the key information in the image to be detected.

[0038] Step S102: at least input the image to be detected into a target large model to obtain target scene information of the image to be detected.

[0039] Here, it should be noted that the target scene information can indicate the scene corresponding to the image to be detected, thus laying the foundation for subsequently improving the accuracy of the recognition result.

[0040] Furthermore, in one example, to further improve the accuracy of the inference results and ensure that the target scene information accurately represents the actual scene of the image to be detected, the initial recognition results and the image to be detected can be input into the target macro model to obtain the target scene information. This lays the foundation for further improving the accuracy of target recognition results.

[0041] Step S103: inputting the target scene information and the initial recognition result into the target large model to obtain an output target recognition result.

[0042] The target recognition result is at least a result obtained by correcting the initial recognition result and matching the target scene information, including key information of the image to be detected.

[0043] In this way, the disclosed solution effectively utilizes the large model to correct the initial recognition results of the image to be detected. Compared with the existing recognition methods, the target recognition results obtained by the disclosed solution are more accurate. Moreover, the above process does not require human intervention, and is intelligent, efficient and accurate, thereby providing strong support for improving user experience.

[0044] In addition, since the disclosed solution fully considers the scene information (i.e., target scene information) of the image to be detected when using the large model to correct the initial recognition results, the large model is more targeted during the recognition process, laying the foundation for improving the accuracy of the recognition results.

[0045] In a specific example of the disclosed scheme, the disclosed scheme can also be applied to verification scenarios. For example, after obtaining the key information of the image to be detected, the key information of the image to be detected is compared with the target key information (for example, with the target key information obtained from other systems) to obtain the target detection result of the image to be detected. At this time, the target detection result can represent the authenticity of the image to be detected, or can be called the authenticity of the key information of the image to be detected.

[0046] For example, in a bidding scenario, the image to be detected can be a contract agreement image. In this case, the key information of the contract in the contract agreement image can be obtained using the disclosed solution. At this time, the key information of the contract in the obtained contract agreement image can be compared with the key information of the theoretical contract (i.e., the target key information) to determine the authenticity of the contract agreement image. The above process can automatically and intelligently distinguish the authenticity without human intervention. Compared with the existing manual authenticity identification, the disclosed solution can save labor costs and is more efficient.

[0047] Thus, the disclosed solution provides an intelligent authenticity detection solution that does not require human intervention, effectively saving labor costs compared to existing manual detection. Moreover, due to the high accuracy of the target recognition results obtained by the disclosed solution, the target detection results obtained by the disclosed solution are also more accurate than those obtained by manual detection, further effectively reducing the error rate. In other words, the authenticity detection solution provided by the disclosed solution is both efficient and intelligent, thus laying the foundation for further improving the user experience.

[0048] In a specific example of the present disclosure, the above-mentioned image to be detected can be obtained in the following manner. Specifically, before obtaining the image to be detected, or before obtaining the initial recognition result of the image to be detected, the method further includes:

[0049] Step S1001: Obtain the file to be detected.

[0050] Step S1002: performing keyword recognition on the file to be detected to obtain a keyword recognition result.

[0051] For example, in one example, the keyword recognition result includes at least one target keyword, so that it is easy to determine the image to be detected that contains the required key information based on the recognized target keyword, providing support for the subsequent effective recognition of the key information.

[0052] For example, in one example, the layout analysis of the file to be detected can be performed, and based on the layout analysis results, the area in the text to be detected that requires keyword recognition can be determined, and then keyword recognition can be performed on the area in the file to be detected that requires keyword recognition to obtain at least one target keyword.

[0053] Step S1003: Based on the keyword recognition result, determine the image to be detected that contains key information corresponding to the target keyword.

[0054] It should be noted that, in one example, the keyword recognition result may include a target keyword. In this case, at least one image to be detected that includes key information corresponding to the target keyword may be determined based on the target keyword.

[0055] It is understandable that in order to include all key information corresponding to the target keyword, the number of images to be detected corresponding to the target keyword may be one or more. In other words, there may not be a one-to-one correspondence between the images to be detected and the target keyword.

[0056] Alternatively, in another example, the keyword recognition result may include multiple target keywords. In this case, based on each target keyword, an image to be detected that includes key information corresponding to each target keyword may be determined.

[0057] It is understandable that when the keyword recognition results contain multiple target keywords, it is necessary to obtain an image to be detected that contains the key information corresponding to each target keyword. In other words, in this scenario, multiple images to be detected can be obtained. This can meet the different needs of different scenarios, laying the foundation for enriching the application scenarios of the disclosed solution, improving the adaptability of the disclosed solution, and further enhancing the user experience.

[0058] Thus, the disclosed solution provides a specific method for obtaining an image to be detected. This method can efficiently determine the image to be detected from a document requiring key information recognition. Furthermore, this process requires no human intervention, thus achieving both intelligence and efficiency. Furthermore, the disclosed solution places no restrictions on the document to be detected. In other words, the disclosed solution can identify any document requiring key information recognition, thus achieving universal applicability.

[0059] Specifically, Figure 2 This is a schematic flow chart of an information detection method based on a large model according to an embodiment of the present application. Figure 2 The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understood that the above Figure 1 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.

[0060] Furthermore, the method includes at least part of the following contents. Figure 2 As shown, including:

[0061] Step S201: Acquire an image to be detected, and obtain an initial recognition result of the image to be detected.

[0062] Here, the initial recognition result is a result obtained after performing text recognition on the image to be detected.

[0063] It should be noted that the relevant content about the initial recognition result can be found in the above description and will not be repeated here.

[0064] Step S202: at least input the image to be detected into a target large model to obtain target scene information of the image to be detected.

[0065] It should be noted that the relevant content about the target scene information can be found in the above description and will not be repeated here.

[0066] Here, in one example, to further improve the accuracy of the inference results and ensure that the target scene information accurately represents the actual scene of the image to be detected, the disclosed solution can also input the initial recognition results and the image to be detected into the target macro model to obtain the target scene information. This lays the foundation for further improving the accuracy of the target recognition results.

[0067] Step S203: Obtaining a target prompt word that matches the target scene information.

[0068] Here, the target prompt word is used to at least indicate the output of a target data structure that matches the target scene information. For example, in one example, the target scene information can be input into a prompt word generator, or the target scene information and the initial recognition result can be input into the prompt word generator to output the target prompt word. In other words, the prompt word generator can be used to quickly determine a data format that matches the target scene information.

[0069] It should be further explained that the target prompt word can indicate the data format required to be output for the target scene information, so that the data output by the subsequent model (that is, the target recognition result) meets the data format requirements, thereby laying the foundation for further improving the level of intelligence and enhancing user experience.

[0070] For example, in one example, the prompt generator can be trained in advance, for example, Figure 3 As shown, you can first determine multiple scene information, then determine the core fields (or core query fields, which can be recorded as scene queries) of each scene information in the multiple scene information. Finally, for the core fields of each scene information, determine the data format required for each scene information to accurately customize the structured output data format required for each scene. Furthermore, the multiple scene information and the data formats required for each scene information are unified as training data. Alternatively, the core fields of each scene information can also be used as training data. In this way, the prompt word generator is trained to obtain a prompt word generator that can accurately output the data format required by the scene information.

[0071] Furthermore, for example, Figure 4As shown, in one example, in the prompt word generator obtained after the training is completed, the initial recognition result (for example, in one example, the initial recognition result can be specifically an optical character recognition (OCR) result (which may include recognized characters and / or tables)) can be vector matched with the target scene information (for example, with the core field in the target scene information) to obtain the text information associated with the core field in the initial recognition result, and then determine the data format required for the text information associated with the core field to construct the target prompt word. In other words, in one example, the target prompt word can specifically indicate the data format required to output the text information associated with the core field of the target scene information in the initial recognition result.

[0072] It is understood that for the present disclosure, the target prompt word may include multiple prompt words, or may also include a prompt instance and / or prompt words associated with the prompt instance. The present disclosure does not impose specific restrictions on the target prompt word. For example, in the invoice scenario, the target prompt word may specifically be: Please output the invoice recognition results in JSON (JavaScript Object Notation) format, including the following keywords: invoice number (invoice_num), invoice type (invoice_type), invoice code (invoice_code), invoice amount (invoice_amount), etc.

[0073] Step S204: input the target prompt word and the initial recognition result into the target large model to obtain an output target recognition result having the target data structure.

[0074] Here, the target recognition result is at least a result obtained by correcting the initial recognition result and matching the target scene information, including key information of the image to be detected.

[0075] That is to say, in this example, the target prompt word and the initial recognition result can be input into the target large model together, so as to use the reasoning ability of the target large model to correct the initial recognition result. For example, after obtaining the input target prompt word and the initial recognition result, the target large model uses few-shot learning (multi-example prompt learning) and / or in-context learning (context learning) to adjust the content in the initial recognition result, such as text or the data format of the text, or regenerate new text, etc., to obtain a target recognition result containing the key information of the image to be detected. At this time, the output format of the key information in the target recognition result meets the requirements of the required data format, thereby providing strong support for subsequent processing.

[0076] In this way, the disclosed solution provides a refined solution for obtaining target recognition results. The solution first obtains a target prompt word that matches the scene information of the image to be detected (that is, the target scene information), and then uses the target prompt word to guide the reasoning of the large model, so that the large model can not only modify the initial recognition results and output target recognition results with higher accuracy, but also make the output target recognition results meet the data format requirements. In this way, on the basis of effectively ensuring accuracy, the degree of intelligence is further improved, and at the same time, it also lays the foundation for further improving the user experience.

[0077] In a specific example, the above-mentioned initial recognition result can be obtained in the following manner. That is, before obtaining the initial recognition result of the image to be detected, the method further includes:

[0078] The acquired image to be detected is input into a preset recognition model to obtain the output of the initial recognition result. In this way, the model is used to quickly obtain the initial recognition result, which provides strong support for the subsequent rapid acquisition of high-precision target recognition results.

[0079] Here, in one example, an OCR model can be used to perform optical character recognition to obtain an optical character recognition result (i.e., the initial recognition result described above), which may include, for example, recognized characters and / or tables. However, it should be noted that the accuracy of the initial recognition result obtained in this case is relatively low, for example, there may be missing characters, garbled characters, etc. Therefore, after obtaining the initial recognition result, the reasoning ability of the large model can be further utilized to correct the initial recognition result, thereby effectively improving the accuracy of the final recognition result.

[0080] Furthermore, in one example, Figure 5 As shown, the preset recognition model may mainly include a text detection module, a direction classifier (also called a text box correction module) and a text recognition module.

[0081] Furthermore, the text detection module is mainly used to detect the text area in the input image to be detected. Here, it should be noted that the image to be detected may contain multiple text areas. In this case, the text detection module can be used to detect all text areas of the image to be detected. For example, for the invoice scenario described above, the text detection module can be used to obtain text areas such as the invoice number (invoice_num), invoice type (invoice_type), invoice code (invoice_code), and invoice amount (invoice_amount). For example, the text box containing the above text can be specifically output.

[0082] Furthermore, the direction classifier is mainly used to determine whether the direction of the text in the text area meets the preset rules, thereby providing strong support for improving the accuracy of subsequent inspections.

[0083] Furthermore, the text recognition module is primarily configured to perform text recognition on text regions where the orientation of the text satisfies the preset rules, to obtain the initial recognition result. For example, the text recognition module is used to perform text recognition on text regions where the orientation of the text satisfies the preset rules, thereby obtaining text recognition results corresponding to each text region to obtain the initial recognition result.

[0084] In this way, the disclosed scheme provides a refined structure of the recognition model, which is simple and efficient and can quickly obtain the initial recognition results. Moreover, since the disclosed scheme fully considers the influence of the direction of the text on the recognition results, the disclosed scheme can also maximize the accuracy of the initial recognition results, thereby providing strong support for the subsequent rapid and efficient correction of the initial recognition results. At the same time, it also provides strong support for quickly obtaining target recognition results.

[0085] Furthermore, in one example, the direction classifier is also used to rotate the text area when it is determined that the direction of the text in the text area does not meet the preset rules, so that the direction of the text in the text area after the rotation operation meets the preset rules. For example, it can be determined whether the angle between the text area, such as the text in the text box, and the preset direction is a preset value (for example, whether it is zero degrees); for text at zero degrees, it can be considered that the text is in the horizontal direction. Otherwise, if it is not zero degrees, the row where the text is located can be rotated, for example, an angle rotation operation can be performed to make the row where the text is located at zero degrees, and then the text area after the rotation operation is sent to the text recognition module after affine transformation to recognize the text content.

[0086] In this way, the disclosed solution further refines the processing logic of the direction classifier. Moreover, the above processing logic is simple and efficient, which provides strong support for the subsequent rapid and efficient correction of the initial recognition results. At the same time, it also provides strong support for quickly obtaining target recognition results.

[0087] Furthermore, in a specific example, after obtaining the target recognition result described above, the target recognition result can also be used to fine-tune the preset recognition model, thereby further improving the recognition accuracy of the model. For example, after obtaining the output target recognition result, the method further includes:

[0088] The target recognition result is used to fine-tune at least some of the network parameters in the text detection module and at least some of the network parameters in the text recognition module of the preset recognition model to train the preset recognition model, thereby obtaining the preset recognition model after parameter fine-tuning. In this way, the similarity between the recognition result output by the fine-tuned preset recognition model and the target recognition result satisfies the preset rules. In this way, the target recognition result is used to further improve the recognition accuracy of the preset recognition model, thereby further providing strong support for improving the accuracy of the target recognition result.

[0089] The following is a specific application scenario of the disclosed solution, namely the verification scenario of bidding / tender documents:

[0090] During the project bidding process, the importance of verifying key information in bidding documents cannot be underestimated. For example, to ensure the authenticity of various materials in bidding documents, key information such as scanned business licenses, contracts, and relevant invoices must be verified. However, current verification methods present several challenges: The verification process is labor-intensive and requires high professional expertise; and accuracy is difficult to guarantee due to the sheer volume of information in bidding documents, making it difficult for staff to efficiently and accurately identify and verify various types of information within a limited timeframe.

[0091] Based on this, this example proposes a large-scale model-based intelligent file authentication assistant (hereinafter referred to as the intelligent authentication assistant). This intelligent assistant has the following advantages:

[0092] High Efficiency and Accuracy: The Intelligent Verification Assistant can rapidly process large volumes of bidding and tendering documents, automatically identifying and verifying the authenticity of invoices, contracts, and other documents, significantly improving verification efficiency. Compared to traditional manual verification methods, this significantly reduces processing time, making the verification process much faster. Furthermore, the Intelligent Verification Assistant accurately identifies and verifies various types of information, reducing the risk of misjudgments and omissions, thereby improving verification accuracy.

[0093] High degree of automation: The intelligent verification assistant can automatically execute the verification process without human intervention. It can automatically extract key information from the document and compare it with the information in the database to realize the automated verification process. This effectively reduces the cost and error rate of manual operation and improves work efficiency.

[0094] Strong scalability: The intelligent verification assistant based on the large model has strong scalability and can continuously adapt to new verification requirements and data changes. For example, new verification rules and algorithms can be added according to actual needs to meet more complex and diverse verification needs.

[0095] High security: Since the intelligent verification assistant uses a large model to perform the verification process, it can better protect privacy compared to manual operations. In addition, the large model also adopts strict data encryption and privacy protection measures when processing sensitive information, further ensuring the security of information.

[0096] Specifically, the following combination Figure 6 The following is a detailed description of the authentication process of the intelligent authentication assistant proposed in this example:

[0097] Step S601, file upload: upload the bidding / tendering documents in the smart management portal.

[0098] Here, you can also authenticate the identity of the currently logged-in account before uploading the bidding / tender documents, and upload the bidding / tender documents after the authentication is passed.

[0099] Step S602: Obtaining the image to be tested: Through keyword recognition and / or layout analysis, keyword recognition is performed on the areas of the bidding / tendering document (e.g., scanned documents or images) where keyword recognition is required. This automatically extracts key scanned information, obtains multiple target keywords, and then obtains the image to be tested that contains the key information corresponding to the target keywords. For example, an image of a business license scan, an invoice screenshot, or a contract agreement image is obtained as the image to be tested.

[0100] Step S603, initial text recognition: using the preset recognition model described above (including the text detection module, the direction classifier (also known as the text box correction module) and the text recognition module) to perform text detection and text recognition on the input image to be detected to obtain an initial recognition result.

[0101] Step S604, large model scene identification: The initial recognition results are fed into the target large model, which performs scene identification to obtain scene information corresponding to the image to be detected. For example, in one example, the scene information can be one of the following: invoice, contract, business license, etc.

[0102] Step S605, constructing a target prompt word: inputting the scene information of the image to be detected obtained above and the initial recognition result obtained above into a prompt word generator to obtain a target prompt word, which is used to indicate the data format required for the key information in the image to be detected to be output.

[0103] Step S606: the large model outputs the target recognition result: the obtained target prompt word and the initial recognition result are input into the target large model to obtain the target recognition result.

[0104] Step S607: authenticity comparison, compare and analyze the target recognition results obtained by the target large model with the real data one by one, and return the verification results.

[0105] It should be noted that, in actual applications, anomalies can also be marked based on the verification results to provide prompts.

[0106] The disclosed solution also provides an information detection device based on a large model, such as Figure 7 As shown, including:

[0107] The pre-processing unit 701 is configured to obtain an image to be detected and obtain an initial recognition result of the image to be detected; wherein the initial recognition result is a result obtained by performing text recognition on the image to be detected;

[0108] The target processing unit 702 is used to input at least the image to be detected into the target large model to obtain the target scene information of the image to be detected; input the target scene information and the initial recognition result into the target large model to obtain an output target recognition result; wherein, the target recognition result is at least a result obtained by correcting the initial recognition result and matching the target scene information, including key information of the image to be detected.

[0109] In a specific example of the present disclosure,

[0110] The pre-processing unit is further configured to obtain a target prompt word that matches the target scene information; the target prompt word is at least configured to instruct output of a target data structure that matches the target scene information;

[0111] The target processing unit is specifically used to input the target prompt word and the initial recognition result into the target large model to obtain the output target recognition result having the target data structure.

[0112] In a specific example of the disclosed solution, the preprocessing unit is further configured to input the image to be detected into a preset recognition model to obtain the output initial recognition result.

[0113] In a specific example of the disclosed solution, the preset recognition model includes a text detection module, a direction classifier and a text recognition module; wherein,

[0114] The text detection module is used to detect the text area of ​​the input image to be detected;

[0115] The direction classifier is used to determine whether the direction of the text in the text area meets the preset rules;

[0116] The text recognition module is used to perform text recognition on a text area where the direction of the text meets the preset rule to obtain the initial recognition result.

[0117] In a specific example of the presently disclosed scheme, the direction classifier is also used to rotate the text area when it is determined that the direction of the text in the text area does not meet the preset rules, so that the direction of the text in the text area after the rotation operation meets the preset rules.

[0118] In a specific example of the disclosed solution, the pre-processing unit is further configured to:

[0119] Utilizing the target recognition result, fine-tune at least some of the network parameters in the text detection module, and fine-tune at least some of the network parameters in the text recognition module, so as to train the preset recognition model and obtain the preset recognition model after parameter fine-tuning.

[0120] In a specific example of the disclosed solution, the pre-processing unit is further configured to:

[0121] Get the file to be tested;

[0122] Performing keyword recognition on the file to be detected to obtain a keyword recognition result; the keyword recognition result includes at least one target keyword;

[0123] Based on the keyword recognition result, the image to be detected containing key information corresponding to the target keyword is determined.

[0124] In a specific example of the disclosed solution, the present invention further includes: a comparison unit;

[0125] The comparison unit is used to compare the key information of the image to be detected with the target key information to obtain the target detection result of the image to be detected, and the target detection result is used to indicate the authenticity of the image to be detected.

[0126] For the description of specific functions and examples of each unit of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0127] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0128] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0129] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0130] like Figure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0131] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0132] The computing unit 801 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the information detection method based on the large model. For example, in some embodiments, the information detection method based on the large model can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the information detection method based on the large model described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the information detection method based on the large model by any other appropriate means (e.g., by means of firmware).

[0133] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0134] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0135] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0136] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0137] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0138] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0139] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0140] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A large model-based information detection method, comprising: Acquire an image to be detected, and obtain an initial recognition result of the image to be detected; wherein the initial recognition result is a result obtained after performing text recognition on the image to be detected; At least inputting the image to be detected into a target large model to obtain target scene information of the image to be detected; Inputting the target scene information and the initial recognition result into the target macro model to obtain an output target recognition result, wherein the target recognition result is at least a result obtained by correcting the initial recognition result, matching the target scene information, and used to verify the authenticity of the image to be detected, including key information of the image to be detected; The step of obtaining the initial recognition result of the image to be detected includes: The image to be detected is input into a preset recognition model to obtain the output initial recognition result; the preset recognition model includes a text detection module, a direction classifier and a text recognition module; wherein the text detection module is used to detect the text area in the input image to be detected; the direction classifier is used to determine whether the direction of the text in the text area meets the preset rules; the text recognition module is used to perform text recognition on the text area where the direction of the text meets the preset rules to obtain the initial recognition result.

2. The method according to claim 1, further comprising: Obtaining a target prompt word that matches the target scene information; The target prompt word is at least used to instruct output of a target data structure that matches the target scene information; The step of inputting the target scene information and the initial recognition result into the target large model to obtain an output target recognition result includes: The target prompt word and the initial recognition result are input into the target large model to obtain an output target recognition result having the target data structure.

3. The method according to claim 1, wherein The direction classifier is further configured to rotate the text region if it is determined that the direction of the text in the text region does not satisfy a preset rule, so that the direction of the text in the text region after the rotation operation satisfies the preset rule.

4. The method according to claim 1, wherein The method further comprises: Utilizing the target recognition result, fine-tune at least some of the network parameters in the text detection module, and fine-tune at least some of the network parameters in the text recognition module, so as to train the preset recognition model and obtain the preset recognition model after parameter fine-tuning.

5. The method according to claim 1 or 2, further comprising: Get the file to be tested; Perform keyword recognition on the file to be detected to obtain a keyword recognition result; The keyword recognition result includes at least one target keyword; Based on the keyword recognition result, the image to be detected containing key information corresponding to the target keyword is determined.

6. The method according to claim 1 or 2, further comprising: The key information of the image to be detected is compared with the target key information to obtain a target detection result of the image to be detected, and the target detection result is used to indicate the authenticity of the image to be detected.

7. An information detection device based on a large model, comprising: A pre-processing unit, configured to obtain an image to be detected and obtain an initial recognition result of the image to be detected; wherein the initial recognition result is a result obtained after performing text recognition on the image to be detected; a target processing unit, configured to input at least the image to be detected into a target macromodel to obtain target scene information of the image to be detected; and input the target scene information and the initial recognition result into the target macromodel to obtain an output target recognition result; wherein the target recognition result is at least a result obtained by correcting the initial recognition result, matching the target scene information, and used to verify the authenticity of the image to be detected, including key information of the image to be detected; Wherein, the pre-processing unit is specifically used for: The image to be detected is input into a preset recognition model to obtain the output initial recognition result; the preset recognition model includes a text detection module, a direction classifier and a text recognition module; wherein the text detection module is used to detect the text area in the input image to be detected; the direction classifier is used to determine whether the direction of the text in the text area meets the preset rules; the text recognition module is used to perform text recognition on the text area where the direction of the text meets the preset rules to obtain the initial recognition result.

8. The device according to claim 7, wherein The pre-processing unit is further configured to obtain a target prompt word that matches the target scene information; the target prompt word is at least configured to instruct output of a target data structure that matches the target scene information; The target processing unit is specifically used to input the target prompt word and the initial recognition result into the target large model to obtain the output target recognition result having the target data structure.

9. The device according to claim 7, wherein The direction classifier is further configured to rotate the text region if it is determined that the direction of the text in the text region does not satisfy a preset rule, so that the direction of the text in the text region after the rotation operation satisfies the preset rule.

10. The device according to claim 7, wherein The pre-processing unit is further used for: Utilizing the target recognition result, fine-tune at least some of the network parameters in the text detection module, and fine-tune at least some of the network parameters in the text recognition module, so as to train the preset recognition model and obtain the preset recognition model after parameter fine-tuning.

11. The device according to claim 7 or 8, wherein The pre-processing unit is further used for: Get the file to be tested; Performing keyword recognition on the file to be detected to obtain a keyword recognition result; the keyword recognition result includes at least one target keyword; Based on the keyword recognition result, the image to be detected containing key information corresponding to the target keyword is determined.

12. The device according to claim 7 or 8, wherein Also includes: The comparison unit is used to compare the key information of the image to be detected with the target key information to obtain the target detection result of the image to be detected, and the target detection result is used to indicate the authenticity of the image to be detected.

13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.

15. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Invoice picture identification and verification method and system, device, and readable storage medium

    CN110675546A

  • Intelligent verification method and device, computer equipment and storage medium

    CN113420657A