Information detection method and device, equipment and storage medium
By performing optical character recognition and multi-model detection on financial reimbursement materials, the problem of low efficiency in traditional financial reimbursement review has been solved, and efficient and accurate sensitive information detection has been achieved, ensuring compliance and accuracy, and reducing labor costs.
Patent Information
- Application Number
- CN202510853465.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-16
AI Technical Summary
Traditional financial reimbursement review methods rely on a large amount of manpower, are inefficient and prone to omissions, and cannot meet the requirements of modern corporate management for efficiency, accuracy and compliance.
By obtaining the original material data of financial reimbursement materials for optical character recognition, combined with rule matching and multi-model detection, including natural language discrimination model, visual feature discrimination model, large language model and multimodal large model, sensitive information is identified from multiple dimensions to generate the final sensitive detection results.
It has improved the comprehensiveness, accuracy, efficiency and flexibility of sensitive information detection in financial reimbursement materials, reduced misjudgments and missed judgments, saved labor costs, and achieved early detection, early warning and early disposal of risks.
Smart Images

Figure CN120656198A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information detection technology, and in particular to an information detection method, apparatus, device and storage medium. Background Art
[0002] In today's digital age, financial reimbursement review, as a crucial component of internal corporate management, faces increasingly complex challenges in information security and privacy protection. With the expansion of enterprise scale and diversification of business operations, the workload of reviewing financial reimbursement materials has increased dramatically. Traditional review methods, which rely on extensive manpower, are not only inefficient but also prone to oversights, making them unable to meet the requirements of modern enterprise management for efficiency, accuracy, and compliance.
[0003] Therefore, introducing advanced technical means to optimize the financial reimbursement review process and improve review efficiency and accuracy has become an urgent issue to be addressed in corporate management. Summary of the Invention
[0004] The main purpose of this application is to provide an information detection method, device, equipment and storage medium, aiming to solve the problem that traditional auditing methods rely on a large amount of manpower input, which is not only inefficient but also prone to omissions.
[0005] To achieve the above objectives, the present application proposes an information detection method, which includes:
[0006] Obtaining original material data corresponding to the financial reimbursement materials, and performing optical character recognition on the original material data to obtain a material recognition result corresponding to the original material data;
[0007] Perform rule matching based on the original material data and the material identification result to obtain a first sensitive detection result;
[0008] Performing multi-model detection based on the original material data and the material identification result to obtain a second sensitive detection result;
[0009] A target sensitive detection result is obtained based on the first sensitive detection result and the second sensitive detection result.
[0010] In one embodiment, performing rule matching based on the original material data and the material identification result to obtain a first sensitive detection result includes:
[0011] Obtain sensitive information rule base;
[0012] According to the sensitive information rule library, the original material data and the material identification result are matched word by word to check whether sensitive information is contained;
[0013] If so, the specific content and location information of the sensitive information are recorded, and the sensitive information is marked to generate a first sensitive detection result.
[0014] In one embodiment, the original material data includes original image data; and performing multi-model detection based on the original material data and the material identification result to obtain a second sensitive detection result includes:
[0015] Inputting the original material data and the material recognition result into a natural language discrimination model for sensitivity detection, and obtaining a third sensitivity detection result output by the natural language discrimination model;
[0016] Inputting the original image data into a visual feature discrimination model for sensitivity detection, and obtaining a fourth sensitivity detection result output by the visual feature discrimination model;
[0017] If the third sensitive detection result and the fourth sensitive detection result are information-insensitive results, the third sensitive detection result and the fourth sensitive detection result are associated and combined to generate a second sensitive detection result.
[0018] In one embodiment, the method further comprises:
[0019] If the third sensitive detection result and the fourth sensitive detection result are information-sensitive results, inputting the first sensitive detection result and the third sensitive detection result into the large language model for sensitive information rechecking to obtain a fifth sensitive detection result;
[0020] Inputting the fourth sensitive detection result into the multimodal large model for sensitive information rechecking to obtain a sixth sensitive detection result;
[0021] The third sensitive detection result, the fourth sensitive detection result, the fifth sensitive detection result and the sixth sensitive detection result are associated and combined to generate a second sensitive detection result.
[0022] In one embodiment, inputting the original material data and the material recognition result into a natural language discrimination model for sensitivity detection to obtain a third sensitivity detection result output by the natural language discrimination model includes:
[0023] Performing semantic analysis on the original material data and the material recognition result by using the natural language discrimination model to obtain contextual semantic information;
[0024] Determining whether the original material data and the material recognition result involve sensitive information based on the contextual semantic information by the natural language discrimination model;
[0025] If so, the specific content and location information of the sensitive information are recorded, and the sensitive information is marked, and the third sensitive detection result is generated and output through the natural language discrimination model.
[0026] In one embodiment, inputting the original image data into a visual feature discrimination model for sensitivity detection to obtain a fourth sensitivity detection result output by the visual feature discrimination model includes:
[0027] Performing image analysis on the original image data using the visual feature discrimination model to identify entity information in the image;
[0028] Determining whether the entity information meets the financial reimbursement standards by using the visual feature discrimination model;
[0029] If not, the entity information is marked, and the fourth sensitive detection result is generated and output through the visual feature discrimination model.
[0030] In one embodiment, before performing multi-model detection based on the original material data and the material identification result to obtain the second sensitive detection result, the method further includes:
[0031] Obtain the target user's sensitive detection needs;
[0032] Conduct intent analysis on the sensitive detection requirements to generate model role settings, detection task background, detection task requirements, and model output requirements;
[0033] Based on the model role setting, the detection task background, the detection task requirements and the model output requirements, model prompt words are generated to perform multi-model detection.
[0034] In addition, to achieve the above-mentioned purpose, the present application also proposes an information detection device, which includes:
[0035] an acquisition module, configured to acquire original material data corresponding to financial reimbursement materials, and perform optical character recognition on the original material data to obtain a material recognition result corresponding to the original material data;
[0036] a matching module, configured to perform rule matching based on the original material data and the material identification result to obtain a first sensitive detection result;
[0037] a detection module, configured to perform multi-model detection based on the original material data and the material identification result to obtain a second sensitive detection result;
[0038] A generation module is used to obtain a target sensitive detection result based on the first sensitive detection result and the second sensitive detection result.
[0039] In addition, to achieve the above-mentioned purpose, the present application also proposes an information detection device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the information detection method described above.
[0040] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by the processor, the steps of the information detection method described above are implemented.
[0041] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the information detection method described above.
[0042] The present application provides an information detection method, apparatus, device and storage medium. The information detection method obtains original material data corresponding to financial reimbursement materials, performs optical character recognition on the original material data, obtains a material identification result corresponding to the original material data, and then performs rule matching based on the original material data and the material identification result to obtain a first sensitive detection result, thereby performing multi-model detection based on the original material data and the material identification result to obtain a second sensitive detection result, and then obtains a target sensitive detection result based on the first sensitive detection result and the second sensitive detection result. Therefore, all information in the financial reimbursement materials is more comprehensively covered according to a multi-dimensional data processing method to avoid misjudgment caused by information omission. At the same time, rule matching and multi-model detection (including natural language processing model, visual recognition model, large language model and multimodal large model) are combined to judge sensitive information from different angles. Among them, rule matching can quickly filter out obviously insensitive content and reduce the computational burden of subsequent complex models, while multi-model detection can process complex semantic and image information and improve the accuracy of detection, thereby effectively improving the comprehensiveness, accuracy, efficiency and flexibility of sensitive information detection of financial reimbursement materials. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0045] Figure 1 A flowchart of the first embodiment of the information detection method of this application is provided;
[0046] Figure 2 A flow chart of the second embodiment of the information detection method of this application is provided;
[0047] Figure 3 A brief flowchart of the information detection method provided for this application;
[0048] Figure 4 A system architecture diagram provided for the information detection method of this application;
[0049] Figure 5 This is a schematic diagram of the module structure of the information detection device according to an embodiment of the present application;
[0050] Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the information detection method in the embodiment of the present application.
[0051] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0052] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0053] Currently, sensitive information detection technology plays an important role in information security, privacy protection, and content review. Especially in the field of financial reimbursement compliance risk review, every financial reimbursement requires strict review of the reimbursement materials, thereby strengthening source compliance control, process compliance management, and result compliance monitoring, and achieving "early detection, early warning, and early disposal" of risks. However, this process often requires huge manpower costs. Therefore, introducing an efficient and accurate sensitive information detection method can effectively save manpower and reduce costs.
[0054] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0055] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of implementing the above functions, a big data service platform, an information detection system, etc. The following uses the information detection system as an example to illustrate this embodiment and the following embodiments.
[0056] Based on this, the embodiment of the present application provides an information detection method, referring to Figure 1 , Figure 1 A flowchart illustrating the first embodiment of the information detection method of this application.
[0057] In this embodiment, the information detection method includes steps S11 to S14:
[0058] Step S11, obtaining original material data corresponding to the financial reimbursement materials, and performing optical character recognition on the original material data to obtain a material recognition result corresponding to the original material data;
[0059] It should be noted that the original material data refers to the original form of the financial reimbursement materials submitted by the user, including electronic documents (such as PDF, Word files) and image files (such as invoice photos, consumption list photos, etc.), and there is no restriction here.
[0060] It should be further explained that optical character recognition (OCR) technology is a technology that converts text content in an image into editable text. Through OCR technology, text information in an image can be extracted for subsequent processing.
[0061] Specifically, in this embodiment, the system first obtains the financial reimbursement materials submitted by the user. These materials exist in various forms, including electronic documents and image files, to obtain the original material data. For image files, the system processes them using OCR technology, extracts the text content, and generates material recognition results. The system supports a variety of OCR tools, such as the open-source Tesseract OCR or Baidu OCR, to meet different needs and scenarios. This allows the system to recognize text in multiple languages and fonts, ensuring the accuracy and completeness of the extracted text information. This converts the text information in the image into a processable text format for subsequent sensitive information detection.
[0062] For example, a user uploads a photo of an invoice as financial reimbursement material. The system calls the OCR model to recognize the text content in the image and obtains the OCR result of the image as the material recognition result. For example, the OCR model is called to recognize the invoice number, amount, date, etc. in the image and convert it into text format.
[0063] Step S12, performing rule matching based on the original material data and the material identification result to obtain a first sensitive detection result;
[0064] It should be noted that the first sensitive detection result refers to the sensitive detection result generated by the system through rule matching, which quickly identifies content that obviously violates regulations and records the specific content and location information of sensitive information. It utilizes the efficiency of rule matching to quickly screen out a large amount of non-sensitive content and reduce the computational burden of subsequent complex models.
[0065] Specifically, a sensitive information rule base is obtained, and then the original material data and the material identification result are matched word by word based on the sensitive information rule base to check whether sensitive information is contained. If so, the specific content and location information of the sensitive information is recorded, and the sensitive information is marked to generate a first sensitive detection result.
[0066] Step S13, performing multi-model detection based on the original material data and the material identification result to obtain a second sensitive detection result;
[0067] It should be noted that the second sensitive detection result refers to the sensitive detection information obtained after a comprehensive analysis of the original material data and material identification results by combining multiple different models. The models include natural language discrimination models, visual feature discrimination models, large language models and multimodal large models, etc. Among them, the natural language discrimination model is used to analyze the semantics of the text content and determine whether it involves sensitive information; the visual feature discrimination model is used to analyze the image content, identify the entity information in the image and determine whether it meets the financial reimbursement standards. Therefore, through multi-model detection, the system can judge sensitive information from different angles and improve the accuracy and comprehensiveness of the detection.
[0068] It should be further explained that multi-model detection can select different model combinations according to different detection requirements. For example, when processing reimbursement materials that are mainly text-based, the natural language discrimination model can be used as the focus; when processing reimbursement materials containing a large number of images, the visual feature discrimination model and the multimodal large model can be combined for detection, etc. There is no restriction here.
[0069] Specifically, the original material data and the material identification result are input into the natural language discrimination model for sensitive detection to obtain a third sensitive detection result output by the natural language discrimination model, and then the original image data is input into the visual feature discrimination model for sensitive detection to obtain a fourth sensitive detection result output by the visual feature discrimination model. If the third sensitive detection result and the fourth sensitive detection result are information-insensitive results, the third sensitive detection result and the fourth sensitive detection result are associated and combined to generate a second sensitive detection result.
[0070] Furthermore, if the third sensitive detection result and the fourth sensitive detection result are information sensitive results, the first sensitive detection result and the third sensitive detection result are input into the large language model for sensitive information recheck to obtain a fifth sensitive detection result, and then the fourth sensitive detection result is input into the multimodal large model for sensitive information recheck to obtain a sixth sensitive detection result, thereby associating and combining the third sensitive detection result, the fourth sensitive detection result, the fifth sensitive detection result and the sixth sensitive detection result to generate a second sensitive detection result.
[0071] Step S14: obtaining a target sensitive detection result based on the first sensitive detection result and the second sensitive detection result.
[0072] It should be noted that the target sensitive detection result refers to the final judgment result obtained by integrating the first sensitive detection result and the second sensitive detection result. By integrating the advantages of the two detection methods, the accuracy and reliability of the detection results are ensured, the false positive rate and missed judgment rate are effectively reduced, and the compliance of financial reimbursement materials is ensured.
[0073] Specifically, the first sensitive detection result and the second sensitive detection result are integrated. If both are not sensitive, they are judged to be not sensitive; otherwise, they are sensitive. If they are sensitive, the specific content and specific location of the sensitive information are generated to generate the final result.
[0074] In addition, in a possible embodiment, the system can dynamically adjust the generation logic of the target sensitive detection result based on the confidence or weight of the first sensitive detection result and the second sensitive detection result. For example, when the two detection results are inconsistent, the system can give priority to the detection result with higher confidence, or determine the final result through a further re-examination mechanism. There is no limitation here.
[0075] For example, when the system is testing a reimbursement document, the first sensitive detection result shows that the material contains the sensitive word "Moutai", while the second sensitive detection result (through multi-model detection) shows that the overall content of the material is not sensitive. The system combines the two results and takes into account the importance of the sensitive word "Moutai". It finally determines that the reimbursement material involves sensitive information and generates a target sensitive detection result of "non-compliant", while providing detailed review reasons and handling suggestions.
[0076] This embodiment obtains original material data corresponding to the financial reimbursement materials, performs optical character recognition on the original material data, obtains a material identification result corresponding to the original material data, and then performs rule matching based on the original material data and the material identification result to obtain a first sensitive detection result, and then performs multi-model detection based on the original material data and the material identification result to obtain a second sensitive detection result, and then obtains a target sensitive detection result based on the first sensitive detection result and the second sensitive detection result. Therefore, all information in the financial reimbursement materials is more comprehensively covered according to the multi-dimensional data processing method, avoiding misjudgment caused by information omission, and combining rule matching and multi-model detection (including natural language processing model, visual recognition model, large language model and multimodal large model) to judge sensitive information from different angles. Among them, rule matching can quickly filter out obviously insensitive content and reduce the computational burden of subsequent complex models, while multi-model detection can process complex semantic and image information and improve the accuracy of detection, thereby effectively improving the comprehensiveness, accuracy, efficiency and flexibility of sensitive information detection of financial reimbursement materials.
[0077] Based on this, the embodiment of the present application provides an information detection method, referring to Figure 2 , Figure 2 A flow chart illustrating the second embodiment of the information detection method of this application.
[0078] In a feasible implementation manner, performing rule matching based on the original material data and the material identification result to obtain a first sensitive detection result includes:
[0079] Step S21, obtaining a sensitive information rule library;
[0080] It should be noted that the sensitive information rule library refers to a predefined database that stores various sensitive words, phrases, patterns, and related compliance rules that may appear in financial reimbursement materials. This rule library is developed based on national policies, internal company systems, and industry standards, and is used to help the system quickly identify and determine sensitive information in financial reimbursement materials. The sensitive information rule library is the foundation of rule matching, and the accuracy and completeness of its content directly affect the effectiveness of rule matching.
[0081] Specifically, in this embodiment, the system first loads a sensitive information rule base from a preset storage location. This rule base can be a local file, a database table, or a data set retrieved from a remote server over the network. The rules in the rule base can be simple keyword lists or complex regular expressions used to match specific text patterns. For example, the rule base might contain sensitive terms such as "Moutai," "Hetian jade," and "XX cigarettes," as well as specific monetary limit rules such as "per capita consumption exceeding 500 yuan" and "single reimbursement amount exceeding 100,000 yuan." By obtaining and loading these rule bases, the system can provide the necessary basis for subsequent rule matching operations.
[0082] Additionally, rule matching can support a dynamically updated sensitive information rule base to adapt to ever-changing financial policies and company systems. For example, when the company updates its reimbursement standards or the list of prohibited items, the system can promptly update the rule base to ensure the accuracy and timeliness of rule matching.
[0083] Step S22: performing word-by-word matching on the original material data and the material identification result according to the sensitive information rule library to check whether sensitive information is contained;
[0084] It should be noted that word-by-word matching refers to the system examining each word or phrase in the original material data and the material identification results against the rules in a predefined sensitive information rule library to determine whether it matches the sensitive information in the rule library, thereby quickly identifying content that may involve sensitive information. Sensitive information refers to information that may violate national policies, company regulations, or industry standards, such as illegally reimbursed items or excessive amounts of consumption.
[0085] Specifically, the system compares the original material data (such as the text content on an expense report) and the material recognition results (such as the text extracted from an image using optical character recognition technology) against the rules in the sensitive information rule library. For example, if the rule library contains the sensitive term "Moutai," the system checks whether the term "Moutai" appears in both the original material data and the material recognition results. If a match is found, the material is deemed to contain sensitive information and requires further processing.
[0086] In addition, word-by-word matching can be combined with contextual information for more accurate judgment. For example, regular expressions can be used to match specific amount ranges or date formats. There are no restrictions here.
[0087] Step S23: If yes, record the specific content and location information of the sensitive information, mark the sensitive information, and generate a first sensitive detection result.
[0088] It should be noted that the specific content refers to the specific text or other forms of sensitive information, and the location information can be the line number, column number, character offset, etc. in the text, or the coordinate position in the image. The marking refers to the system's annotation of the identified sensitive information to enable subsequent reviewers or systems to quickly locate and process such sensitive information.
[0089] Specifically, when the system discovers sensitive information through word-by-word matching, it records the content of the sensitive information and its specific location within the document. For example, if the sensitive word "Moutai" is found in the "Item Name" field of an expense report, the system will record the specific content of the word "Moutai" and its row and column number within the expense report. The sensitive information is also marked, for example, by highlighting, adding tags, or generating annotations, to facilitate quick identification by subsequent reviewers.
[0090] Furthermore, all identified sensitive information and its location information are integrated into the first sensitive detection result to provide a reference for subsequent multi-model detection, which not only enables the rapid identification of sensitive information, but also provides clear clues for subsequent review and processing, ensuring the compliance of financial reimbursement materials.
[0091] For example, in one specific implementation, the system loads a rule library containing sensitive terms such as "Moutai," "Hetian jade," and "XX cigarettes." When matching rules against a reimbursement claim, the system discovers the sensitive term "XX cigarettes." The system records the specific content of the term, "XX cigarettes," and its location in the claim (e.g., row 3, column 2). The system also highlights the sensitive information and generates a first sensitivity detection result. This allows subsequent reviewers or the system to quickly locate and process the sensitive information based on these records, ensuring the compliance of the reimbursement claim.
[0092] This embodiment obtains a sensitive information rule base, and then performs word-by-word matching on the original material data and the material identification results based on the sensitive information rule base to check whether they contain sensitive information. If so, the specific content and location information of the sensitive information is recorded, and the sensitive information is marked to generate a first sensitive detection result, thereby performing preliminary screening on a large amount of data in a short period of time, so that the system can quickly identify content that clearly violates regulations without the need for complex model calculations and analysis, greatly improving processing speed and reducing the computational burden of subsequent complex models. At the same time, it can also be flexibly adjusted according to specific financial policies and company systems. For example, when the company updates the reimbursement standards or the list of items prohibited from reimbursement, the system can update the rule base in a timely manner to ensure the accuracy and timeliness of rule matching. By recording the specific content and location information of the sensitive information, sensitive information can be identified more accurately, avoiding misjudgments caused by fuzzy matching or contextual misjudgments, and providing clear clues for manual review, reducing the workload of manual review and improving review efficiency.
[0093] In a feasible implementation, the original material data includes original image data; and performing multi-model detection based on the original material data and the material identification result to obtain a second sensitive detection result includes:
[0094] Step S31: inputting the original material data and the material recognition result into a natural language discrimination model for sensitivity detection, and obtaining a third sensitivity detection result output by the natural language discrimination model;
[0095] It should be noted that the natural language discrimination model is a model based on natural language processing (NLP) technology, which is used to analyze the semantics of text content and determine whether it involves sensitive information. Through contextual semantic analysis, the model can understand the keywords, phrases and their semantic relationships in the text, so as to accurately determine whether the text content meets the compliance requirements of financial reimbursement. The third sensitive detection result is the detection result generated by the natural language discrimination model based on the input original material data and material recognition results, which is used to identify whether the text content involves sensitive information.
[0096] Specifically, the natural language discrimination model is used to perform semantic analysis on the original material data and the material identification result to obtain contextual semantic information, and then the natural language discrimination model is used to determine whether the original material data and the material identification result involve sensitive information based on the contextual semantic information. If so, the specific content and location information of the sensitive information are recorded, and the sensitive information is marked, and the third sensitive detection result is generated and output through the natural language discrimination model.
[0097] Step S32: inputting the original image data into a visual feature discrimination model for sensitivity detection to obtain a fourth sensitivity detection result output by the visual feature discrimination model;
[0098] It should be noted that the visual feature discrimination model is a model based on computer vision technology, which is used to analyze image content, identify entity information (such as objects, scenes, text, etc.) in the image, and determine whether it meets the financial reimbursement standards. Through image recognition technology, the model can accurately identify key information in the image, thereby determining whether the image content involves sensitive information. The fourth sensitive detection result is a detection result generated by the visual feature discrimination model based on the input original image data, which is used to identify whether the image content involves sensitive information.
[0099] Specifically, the original image data is analyzed by the visual feature discrimination model to identify entity information in the image, and then the visual feature discrimination model is used to determine whether the entity information meets the financial reimbursement standards. If not, the entity information is marked, and the fourth sensitive detection result is generated and output by the visual feature discrimination model.
[0100] Step S33: If the third sensitive detection result and the fourth sensitive detection result are information-insensitive results, the third sensitive detection result and the fourth sensitive detection result are associated and combined to generate a second sensitive detection result.
[0101] It should be noted that the association combination refers to a comprehensive analysis of the detection results of the natural language discrimination model and the visual feature discrimination model to generate a final second sensitive detection result, thereby ensuring the accuracy and reliability of the detection result by combining the detection results of the two models.
[0102] Specifically, the system checks the third sensitive detection result output by the natural language discrimination model and the fourth sensitive detection result output by the visual feature discrimination model. If the detection results of both models are "insensitive", the system will associate and combine these two results to generate a second sensitive detection result of "insensitive", thereby ensuring the reliability of the detection results through a double verification mechanism. For example, if the natural language discrimination model determines that the text content in the reimbursement form does not involve sensitive information, and the visual feature discrimination model determines that the items in the invoice photo meet the reimbursement standards, the system will combine these two results to generate a second sensitive detection result of "insensitive", so as to comprehensively and accurately process the text and image information in the financial reimbursement materials and ensure the reliability of the detection results.
[0103] For example, the system performs multi-model detection on a reimbursement material that includes photos of invoices and consumption lists. First, the text content of the invoice and consumption list is input into the natural language discrimination model, and the model determines whether it involves sensitive information through semantic analysis. For example, the model recognizes that the reimbursement reason contains the phrase "customer appreciation dinner", but no obvious violations are found, so it outputs a third sensitive detection result of "insensitive". Next, the consumption list photo is input into the visual feature discrimination model. The model identifies the items in the photo (such as high-end red wine, seafood platter, etc.) and determines whether these items meet the reimbursement standards. The model finds that although these items are high-end, they do not exceed the company's entertainment standards, so it outputs a fourth sensitive detection result of "insensitive". Finally, the system associates and combines these two "insensitive" results to generate a second sensitive detection result of "insensitive", and judges that the reimbursement material as a whole is "insensitive". You can refer to Figure 3 .
[0104] This embodiment inputs the original material data and the material identification result into a natural language discrimination model for sensitive detection to obtain a third sensitive detection result output by the natural language discrimination model, and then inputs the original image data into a visual feature discrimination model for sensitive detection to obtain a fourth sensitive detection result output by the visual feature discrimination model. If the third sensitive detection result and the fourth sensitive detection result are information-insensitive results, the third sensitive detection result and the fourth sensitive detection result are associated and combined to generate a second sensitive detection result. The natural language discrimination model and the visual feature discrimination model are then combined to achieve multimodal data processing capabilities, that is, they can process text content and image content at the same time, ensure comprehensive detection of all types of data, thereby effectively reducing the misjudgment rate through a double verification mechanism, and achieving multi-angle analysis, thereby more accurately identifying sensitive information and ensuring the reliability of the detection results. In addition, hierarchical detection can also be achieved. That is, by first performing detection on the natural language discrimination model and the visual feature discrimination model, the system can quickly screen out a large amount of non-sensitive content, reducing the computational burden of subsequent complex models (such as large language models and multimodal large models), thereby effectively optimizing resource utilization and improving overall processing efficiency.
[0105] In a feasible embodiment, the method further includes:
[0106] Step S41: If the third sensitive detection result and the fourth sensitive detection result are information-sensitive results, the first sensitive detection result and the third sensitive detection result are input into a large language model for sensitive information rechecking to obtain a fifth sensitive detection result;
[0107] It should be noted that the large language model is a natural language processing model based on deep learning, with powerful language understanding and generation capabilities. It can conduct in-depth analysis of text content, understand contextual semantic relationships, and review sensitive information based on predefined rules and prompt words. The fifth sensitive detection result is a re-check result generated by the large language model based on the input first and third sensitive detection results, and is used to further confirm whether the text content involves sensitive information.
[0108] Specifically, when the third sensitive detection result output by the natural language discrimination model and the fourth sensitive detection result output by the visual feature discrimination model are both "sensitive", the system inputs the first sensitive detection result (the result of rule matching) and the third sensitive detection result (the result of the natural language discrimination model) into the large language model for re-examination, wherein the large language model further analyzes these results through deep semantic analysis combined with predefined prompt word templates. For example, if the third sensitive detection result identifies that the reimbursement reason contains the phrase "customer appreciation dinner" and the amount is high, the large language model can further determine whether the reimbursement reason meets the company's entertainment standards based on the contextual semantics. If the large language model confirms that the content involves sensitive information, it outputs a fifth sensitive detection result of "sensitive"; otherwise, it outputs an "insensitive" result.
[0109] Step S42: inputting the fourth sensitive detection result into the multimodal large model for sensitive information rechecking to obtain a sixth sensitive detection result;
[0110] It should be noted that the multimodal large model is a model that can process multiple modal data (such as text and images) and has the capabilities of image recognition, natural language understanding, and multimodal fusion. It can identify entity information in the image and make a comprehensive judgment based on the text content. The sixth sensitive detection result is a re-examination result generated by the multimodal large model based on the input fourth sensitive detection result, which is used to further confirm whether the image content involves sensitive information.
[0111] Specifically, the system inputs the fourth sensitive detection result (the result of the visual feature discrimination model) into the multimodal large model for re-inspection, wherein the multimodal large model uses image recognition technology to identify the entity information in the image, and makes a comprehensive judgment based on the text content. For example, if the fourth sensitive detection result identifies that the items in the invoice photo are high-end red wine and seafood platters, the multimodal large model can further analyze whether these items meet the reimbursement standards. If the multimodal large model confirms that the image content involves sensitive information, it outputs the sixth sensitive detection result of "sensitive"; otherwise, it outputs the result of "insensitive", which can be referred to Figure 3 .
[0112] Step S43: Associating and combining the third sensitive detection result, the fourth sensitive detection result, the fifth sensitive detection result, and the sixth sensitive detection result to generate a second sensitive detection result.
[0113] It should be noted that the association combination refers to the comprehensive analysis of the results of multiple detection models to generate a final second sensitive detection result. This ensures the accuracy and reliability of the detection results by combining the advantages of multiple models. The second sensitive detection result is the final result generated by combining the results of multiple detection models and is used to identify whether financial reimbursement materials contain sensitive information.
[0114] Specifically, the system associates and combines the third sensitive detection result (the result of the natural language discrimination model), the fourth sensitive detection result (the result of the visual feature discrimination model), the fifth sensitive detection result (the result of the large language model) and the sixth sensitive detection result (the result of the multimodal large model), and generates the final second sensitive detection result by comprehensively analyzing these results.
[0115] For example, if the third and fourth sensitivity detection results are both "sensitive," while the fifth and sixth sensitivity detection results are both "insensitive," the system can further analyze the confidence and specific evidence for these results. If the confidence is high and the evidence is sufficient, the system can generate a second sensitivity detection result of "insensitive"; otherwise, it generates a "sensitive" result. This allows for comprehensive and accurate processing of text and image information in financial reimbursement documents, ensuring the reliability of the detection results.
[0116] In one embodiment, the system performed multi-model detection on a reimbursement document containing a photo of an invoice and a purchase order. Preliminary test results showed that the natural language recognition model outputted a "sensitive" third-sensitivity detection result, identifying the phrase "client appreciation dinner" and the high amount in the reimbursement reason. The visual feature recognition model also outputted a "sensitive" fourth-sensitivity detection result, identifying the items in the invoice photo as high-end red wine and a seafood platter.
[0117] The system then fed these preliminary test results into a large language model for re-examination. The large language model's fifth sensitive test confirmed that the "client appreciation dinner" in the reimbursement claim met the company's entertainment standards, but the amount was slightly higher. The multimodal large model then re-examined the invoice photo, outputting a sixth sensitive test that confirmed the items in the image met the reimbursement standards, but there was suspicion of excessive spending. Finally, the system correlated these results to generate a second sensitive test, determining that the entire reimbursement claim was "sensitive," and providing detailed review reasons and handling recommendations.
[0118] In this embodiment, if the third sensitive detection result and the fourth sensitive detection result are information sensitive results, the first sensitive detection result and the third sensitive detection result are input into the large language model for sensitive information recheck to obtain a fifth sensitive detection result, and then the fourth sensitive detection result is input into the multimodal large model for sensitive information recheck to obtain a sixth sensitive detection result, thereby associating and combining the third sensitive detection result, the fourth sensitive detection result, the fifth sensitive detection result and the sixth sensitive detection result to generate a second sensitive detection result, and then further verifying and refining the preliminary detection result by introducing a large language model and a multimodal large model for rechecking, wherein the large language model can use its powerful language comprehension ability to conduct in-depth analysis of text content; the multimodal large model can combine image and text information to provide a more comprehensive judgment, thereby The multi-level re-inspection mechanism effectively reduces the false positive rate and ensures the accuracy of the detection results. At the same time, it combines the advantages of rule matching, natural language discrimination model, visual feature discrimination model, large language model and multimodal large model, and can judge sensitive information from different angles. Among them, rule matching quickly screens out content that is obviously in violation of regulations, the natural language discrimination model and the visual feature discrimination model conduct preliminary detection, and the large language model and the multimodal large model conduct in-depth re-inspection to ensure the comprehensiveness and reliability of the detection results, enhance the system's ability to handle complex scenarios, and enable the system to process large amounts of data in a short time, thereby improving overall detection efficiency. Furthermore, through the layered detection mechanism, it can effectively optimize resource utilization, improve overall processing efficiency, and enhance the flexibility and scalability of the system, thereby achieving accurate review of financial reimbursement risks, effectively saving the labor cost of financial reimbursement review, and achieving "early detection, early warning, and early disposal" of financial compliance risks.
[0119] In a feasible implementation manner, inputting the original material data and the material identification result into a natural language discrimination model for sensitivity detection to obtain a third sensitivity detection result output by the natural language discrimination model includes:
[0120] Step S51, performing semantic analysis on the original material data and the material recognition result through the natural language discrimination model to obtain contextual semantic information;
[0121] It should be noted that the contextual semantic information refers to the meaning of words and phrases in the text content in a specific context and their mutual relationship, which is used to accurately determine whether the text content involves sensitive information and identify the implicit meaning in the text.
[0122] Specifically, the natural language discrimination model receives original material data (such as the text content in the reimbursement form) and material recognition results (such as the text content extracted from the image through OCR technology) as input. The model uses a pre-trained deep learning algorithm to analyze these text contents sentence by sentence and extract contextual semantic information.
[0123] Step S52: determining whether the original material data and the material recognition result involve sensitive information based on the contextual semantic information using the natural language discrimination model;
[0124] It should be noted that after extracting contextual semantic information, the natural language recognition model uses this information to determine whether the text content contains sensitive information. Sensitive information refers to information that may violate national policies, company regulations, or industry standards, such as illegally reimbursed items or excessive spending.
[0125] Specifically, in this embodiment, the natural language recognition model determines whether the textual content in the original material data and material recognition results involves sensitive information based on contextual semantic information. For example, if the model recognizes the phrase "customer appreciation dinner" and, based on the context, determines that the per capita expenditure at the event was high, exceeding the company's hospitality standards, the model will deem the textual content to involve sensitive information. By leveraging predefined rules and training data, the model can identify implicit sensitive information, rather than relying solely on keyword matching. This allows for deep semantic analysis to effectively reduce false positives and improve detection accuracy.
[0126] Step S53: If yes, record the specific content and location information of the sensitive information, mark the sensitive information, and generate and output the third sensitive detection result through the natural language discrimination model.
[0127] It should be noted that the specific content refers to the identified sensitive words or phrases.
[0128] Specifically, when the natural language discrimination model determines that the text content involves sensitive information, it will record the specific content and location information of the sensitive information in detail. For example, when the model recognizes the phrase "customer appreciation dinner" and determines that it involves sensitive information, the model will record the specific content of the phrase "customer appreciation dinner" and its location information in the reimbursement form (such as the third row and second column). At the same time, the model will mark the sensitive information, such as by highlighting, adding labels, or generating annotations, so that subsequent reviewers can quickly identify it. Finally, the model generates and outputs a third sensitive detection result of "sensitive", providing clear clues for subsequent review and processing.
[0129] In this embodiment, the natural language discrimination model performs semantic analysis on the raw material data and the material identification results to obtain contextual semantic information. The natural language discrimination model then determines whether the raw material data and the material identification results involve sensitive information based on the contextual semantic information. If so, the specific content and location information of the sensitive information are recorded and the sensitive information is marked. The natural language discrimination model generates and outputs the third sensitive detection result, thereby achieving deep semantic analysis. This not only identifies keywords but also understands the specific meaning of keywords in context, thereby more accurately determining whether the text content involves sensitive information. For example, the model can distinguish between the different meanings of "Moutai" as a sensitive word and "Moutai Town" as a common place name, and simultaneously identify implicit sensitive information. Even if sensitive words are not directly mentioned in the text, the model can process implicit information through the semantic association of the context and determine whether there is any violation. For example, "A customer appreciation event was reimbursed at a high-end hotel with high per capita consumption" may imply a violation, thereby improving the accuracy and reliability of detection, reducing the false positive rate, and providing explainable detection results.
[0130] In a feasible implementation manner, inputting the original image data into a visual feature discrimination model for sensitivity detection to obtain a fourth sensitivity detection result output by the visual feature discrimination model includes:
[0131] Step S61, performing image analysis on the original image data using the visual feature discrimination model to identify entity information in the image;
[0132] It should be noted that the entity information refers to objects, scenes, text and other elements in the image that have practical meaning, which are used to determine whether the image content meets the financial reimbursement standards.
[0133] Specifically, in one embodiment, the visual feature recognition model receives raw image data (such as a photo of an invoice or a purchase order) as input. Using deep learning algorithms such as convolutional neural networks (CNNs), the model analyzes the image pixel by pixel to extract key features. For example, the model can identify textual information such as item names, amounts, and dates on an invoice, as well as specific items on a purchase order (such as a high-end red wine or a seafood platter). This allows the model to convert the entity information in the image into processable data, providing a basis for further judgment.
[0134] Step S62, determining whether the entity information meets the financial reimbursement standards through the visual feature discrimination model;
[0135] It should be noted that the financial reimbursement standards refer to a series of regulations and requirements regarding financial reimbursement formulated by a company or organization, including the scope of items allowed for reimbursement, amount limits, scene limits, etc., so that the visual feature discrimination model can judge the identified entity information based on the predefined financial reimbursement standards to determine whether the image content complies with the regulations.
[0136] Specifically, the visual feature discrimination model makes a judgment based on the identified entity information combined with predefined financial reimbursement standards. For example, if the model identifies the item on the invoice as "Moutai," and company regulations prohibit reimbursement of high-end liquor, the model will determine that the image content does not meet the financial reimbursement standards. Similarly, if the model identifies the scene in the consumption list as a "high-end hotel" and the per capita consumption exceeds the company's hospitality standards, the model will also determine that the image content does not meet the financial reimbursement standards. This enables the model to accurately identify and judge sensitive information in images.
[0137] Step S63: If not, mark the entity information, generate and output the fourth sensitive detection result through the visual feature discrimination model.
[0138] Specifically, in one embodiment, when the visual feature discrimination model determines that the entity information in the image does not meet the financial reimbursement standards, the specific content and location information of the entity information will be recorded in detail. For example, the model recognizes that the item in the invoice is "Moutai" and determines that it does not meet the financial reimbursement standards. The model will record the entity information of "Moutai" and its position coordinates in the image (such as the upper left corner coordinates are (100,100), and the lower right corner coordinates are (300,300)). At the same time, the model will mark the entity information, such as by highlighting, adding labels or generating annotations, so that subsequent auditors can quickly identify it. Finally, the model generates and outputs a fourth sensitive detection result of "sensitive", providing clear clues for subsequent review and processing.
[0139] This embodiment performs image analysis on the original image data through the visual feature discrimination model to identify entity information in the image, and then determines whether the entity information meets the financial reimbursement standards through the visual feature discrimination model. If not, the entity information is marked, and the fourth sensitive detection result is generated and output through the visual feature discrimination model, and then the image content is fully processed by identifying the entity information (such as objects, scenes, text, etc.) in the image, which not only makes up for the shortcomings of relying solely on text information, but also can process image content that cannot be completely converted into text through OCR technology, thereby more accurately identifying sensitive information. For example, even if the text content extracted by OCR does not clearly identify a violation, the entity information in the image (such as high-end gifts) may imply a violation, thereby effectively reducing the false positive rate, ensuring the reliability of the detection results, and further improving the accuracy of the detection through a multimodal combination of image and text.
[0140] In a feasible implementation manner, before performing multi-model detection based on the original material data and the material identification result to obtain the second sensitive detection result, the method further includes:
[0141] Step S71, obtaining the target user's sensitive detection requirements;
[0142] It should be noted that the sensitive detection requirements refer to the specific requirements for sensitive information detection of financial reimbursement materials proposed by users based on their own business scenarios, industry standards, company policies, etc., which may include specific sensitive information types (such as high-end gifts, excessive consumption, etc.), detection scope (such as specific reimbursement items or amount ranges), and requirements for the format and content of detection results, etc. There are no restrictions here and they can be set according to actual conditions.
[0143] Specifically, in one embodiment, the system receives user-submitted sensitive information detection requests through a user interface or API. The target user can specify the type of sensitive information to be detected, the scope of detection, and the expected results by filling out a form, uploading a configuration file, or directly entering instructions. For example, a user can request that the system detect the presence of high-end gifts (such as Moutai, Hermès, etc.) or excessive entertainment expenses (such as per capita consumption exceeding 500 yuan) in reimbursement materials.
[0144] Step S72: performing intent analysis on the sensitive detection requirements to generate model role settings, detection task background, detection task requirements, and model output requirements;
[0145] It should be noted that the intent analysis refers to parsing the sensitive detection requirements submitted by the user to understand the user's true intention and generate a specific detection task description accordingly. Among them, the model role setting refers to defining the role played by the detection model in the detection task, such as "financial reimbursement review assistant". The detection task background refers to the background information describing the detection task, such as "conducting compliance audits according to the company's latest financial policies". The detection task requirements refer to clarifying the specific requirements of the detection task, such as "checking whether the entertainment expenses in the reimbursement materials exceed the standards". The model output requirements refer to specifying the format and content of the output results of the detection model, such as "outputting the detection results, including the specific content, location and audit suggestions of sensitive information".
[0146] Specifically, the intent analysis of sensitive detection requirements submitted by users is performed to extract key information and generate the following content: Model role setting: define the role of the detection model, such as "financial reimbursement review assistant", to clarify the model's responsibilities in the detection task; Detection task background: describe the background information of the detection task, such as "conduct compliance review according to the company's latest financial policy" to provide context for the detection task; Detection task requirements: clarify the specific requirements of the detection task, such as "check whether the entertainment expenses in the reimbursement materials exceed the standard, and identify whether there are high-end gifts", to ensure the pertinence and accuracy of the detection task; Model output requirements: specify the format and content of the detection model output results, such as "output detection results, including the specific content, location and audit suggestions of sensitive information", to ensure the readability and usability of the detection results.
[0147] Step S73 : generating model prompt words based on the model role setting, the detection task background, the detection task requirements, and the model output requirements to perform multi-model detection.
[0148] It should be noted that the model prompts refer to specific instructions generated based on user needs to guide the detection model in detecting sensitive information. This helps the detection model better understand the user's intent and thus detect sensitive information more accurately. Multi-model detection refers to the use of multiple detection models (such as natural language recognition models, visual feature recognition models, and large language models) for comprehensive detection to improve detection accuracy and comprehensiveness.
[0149] Specifically, the prompt word project generates specific model prompt words based on the generated model role settings, detection task context, detection task requirements, and model output requirements. These prompt words will serve as input instructions to guide the detection model to detect sensitive information.
[0150] In one embodiment, the large language model prompt word template may be:
[0151] 1. Role Setting
[0152] As a financial reimbursement compliance reviewer, you are responsible for conducting compliance reviews of reimbursement materials submitted by employees to ensure they comply with relevant national and company financial systems, legal regulations, and other requirements, thereby achieving systematic control and efficient review of financial risks.
[0153] II. Mission Background
[0154] In order to implement the spirit of the regulations, strengthen source compliance control, process compliance management and result compliance monitoring, and improve financial audit efficiency and risk prevention and control capabilities, it is now necessary to conduct compliance review of all reimbursement materials to ensure that their content is true, procedures are compliant, bills are complete, and expenses are reasonable, and to eliminate false reimbursements, excessive expenditures, illegal receptions and other behaviors, and achieve the financial risk prevention and control goals of "early detection, early warning, and early disposal."
[0155] III. Task Requirements
[0156] Compliance review: Strictly check the regulations, supplier access and risk management checklists, etc., and review the reimbursement content item by item to see if it is compliant.
[0157] Authenticity verification: verify whether the reimbursement items actually occurred, and whether there were any false reports, duplicate reimbursements, misappropriation of funds, etc.
[0158] Invoice integrity: Check whether the specific contents of invoices, contracts, approval forms, and payment vouchers are complete and compliant.
[0159] Expense standard control: Check whether there are overspending or illegal use by comparing with the company's travel, business entertainment, meetings, training and other expense standards.
[0160] Systemic risk identification: Through audits, potential financial risk points are discovered, a risk early warning mechanism is established, and source control is promoted.
[0161] Review of sensitive words prohibited from reimbursement: Review whether the reimbursement content involves sensitive words prohibited from reimbursement (such as Moutai, Hetian jade, deer antler, Hermes, XX cigarettes, etc.) to ensure that no illegal items or services are reimbursed.
[0162] 4. Output requirements
[0163] Audit conclusion: clearly marked as “compliant” or “non-compliant”, with reasons for the audit attached.
[0164] Risk warning: If any non-compliance is found, the specific illegal clauses must be pointed out (such as the article number of the regulations, the article number of the company system, etc.).
[0165] Suggested handling opinions: Make suggestions for handling non-compliant matters, such as return, correction, reporting, suspension of payment, etc.
[0166] Format specification: The output must be clearly structured, concise, and well-organized to facilitate automatic recognition and archiving by the system.
[0167] 5. Task Examples
[0168] Reimbursement materials content:
[0169] Reimbursement person: Zhang San
[0170] Reimbursement item: Business entertainment expenses in a certain place in May 2024
[0171] Amount: 8,000 yuan
[0172] Attachments: Catering invoice (amount 8000 yuan), consumption list (XX cigarettes), number of guests entertained: 10, approval form
[0173] Audit results:
[0174] Compliance judgment: Non-compliant
[0175] Reasons for review: The expenses for this business entertainment were relatively high, reaching 800 yuan per person, which may pose risks; no clear explanation of the reasons for the business entertainment was provided, posing a financial compliance risk; the consumption list contained the sensitive word "XX cigarettes" which is prohibited from reimbursement, which exceeded the reimbursement scope and did not comply with the requirement of "practicing diligence and thrift" in the regulations.
[0176] Handling suggestion: Return the reimbursement application and focus on verifying whether there is any illegal consumption behavior.
[0177] Risk warning: This reimbursement involves reimbursement of excessive entertainment and illegal items, which may violate Article 8 of the regulations and relevant company regulations and requires special attention.
[0178] Text content: {context} Note: {context} is a placeholder for text content;
[0179] Output:
[0180] When using the large model for judgment, the obtained text content is used for replacement. When the multimodal large model is rechecked, the image file is directly uploaded without specific text content. Its prompt word template is slightly different from the large language model template. The multimodal large model prompt word template is as follows:
[0181] 1. Role Setting
[0182] You are an AI assistant for compliance review of financial reimbursement images. With image recognition and natural language understanding capabilities, you are responsible for conducting compliance reviews of the image content in reimbursement materials to ensure that they comply with relevant national and company financial systems, legal regulations, and other requirements, thereby achieving systematic management and efficient review of financial risks.
[0183] II. Mission Background
[0184] In order to implement the spirit of legal regulations, strengthen source compliance control, process compliance management and result compliance monitoring, and improve financial audit efficiency and risk prevention and control capabilities, it is now necessary to conduct compliance reviews on the picture content in reimbursement materials (such as marketing gifts, meeting scenes, reception scenes, transportation tickets, etc.) to ensure that the content is true, the procedures are compliant, the tickets are complete, and the fees are reasonable, to eliminate false reimbursements, excessive expenditures, illegal receptions, etc., and achieve the financial risk prevention and control goals of "early detection, early warning, and early disposal."
[0185] III. Task Requirements
[0186] Image content recognition: Identify objects, scenes, people, text, and other information in images to determine whether they meet financial reimbursement standards and whether they involve prohibited reimbursement content (such as marketing gifts such as Moutai, Hetian jade, deer antlers, Hermès, XX cigarettes, and bird's nests).
[0187] Compliance review: Check legal regulations to determine whether there are any illegal receptions, excessive consumption, extravagance, etc. Check company financial regulations to determine whether expenses such as travel, meetings, business entertainment, and marketing gifts are in compliance.
[0188] Authenticity verification: Determine whether the image is a real scene, whether there are any problems such as forgery, splicing, blurring, etc.
[0189] Risk warning: If any non-compliant content is found, the specific illegal clauses must be pointed out (such as the legal regulations, company regulations, etc.) and suggestions for handling must be made.
[0190] 4. Output requirements
[0191] Audit conclusion: clearly marked as “compliant” or “non-compliant”, with reasons for the audit attached.
[0192] Risk warning: If any non-compliance is found, the specific illegal clauses must be pointed out (such as the legal regulations, company regulations, etc.).
[0193] Suggested handling opinions: Make suggestions for handling non-compliant matters, such as return, correction, reporting, suspension of payment, etc.
[0194] Format specification: The output must be clearly structured, concise, and well-organized to facilitate automatic recognition and archiving by the system.
[0195] Multimodal recognition description: If image recognition is involved, the recognized objects, scenes, text, and other information must be described.
[0196] 5. Task Examples
[0197] Image content description (identified by multimodal model):
[0198] The picture shows the banquet hall of a five-star hotel. High-end red wine, champagne, seafood platters, and customized gift boxes are placed on the table, with the company logo in the background.
[0199] There is text under the picture: "Customer Appreciation Dinner in May 2024".
[0200] Audit results:
[0201] Compliance judgment: Non-compliant
[0202] Reasons for review: The high-end red wine, champagne, seafood platter and other items shown in the picture clearly exceed the hospitality standards stipulated in the company's "Business Hospitality Management Measures" (such as no more than 300 yuan per person); customized gift boxes appear in the picture, and no relevant approval form and compliance instructions are provided, which does not comply with the company's marketing gift management regulations; the company logo appears in the background of the picture, but it does not indicate whether the reception activity has been approved, which poses a compliance risk; the scene is suspected of violating Article 8 of the regulations on "practicing diligence and thrift" and is wasteful.
[0203] Handling suggestion: Return the reimbursement application, require the completion of complete information and re-approval, and focus on verifying whether there is any illegal consumption behavior.
[0204] Risk warning: This reimbursement involves excessive reception and marketing gifts, which may violate Article 8 of the regulations and relevant company systems and requires special attention.
[0205] VI. Supplementary Notes
[0206] Multimodal models must have capabilities such as image recognition, OCR recognition, natural language understanding, and compliance rule matching. For images that cannot be recognized or are unclear, the model should prompt "Image content is unclear, manual review recommended."
[0207] This embodiment obtains the sensitive detection needs of the target user and then performs intent analysis on the sensitive detection needs to generate model role settings, detection task background, detection task requirements and model output requirements, and then generates model prompt words based on the model role settings, the detection task background, the detection task requirements and the model output requirements to perform multi-model detection, and then perform customized detection according to the user's specific requirements. Because different users may have different sensitive information definitions and detection standards, for example, some companies may be particularly sensitive to specific items (such as high-end gifts), while other companies may be more concerned about amount limits, etc., through intent analysis, detection models and prompt words that meet user needs are generated, while ensuring the timeliness and accuracy of the detection results, which helps to reduce misjudgments or missed judgments caused by misunderstanding user needs.
[0208] For example, to help understand the implementation process of the information detection method, please refer to Figure 4 , Figure 4 This is a system architecture diagram provided for the information detection method of this application.
[0209] Specifically, the overall architecture of the system mainly includes three modules. The first is the data preprocessing module, which mainly parses the input financial reimbursement materials into original text data, image data, and OCR results; the second is the traditional AI judgment module, which mainly uses rule matching, trained natural language discrimination models, and trained visual discrimination models for judgment; the third is the large model judgment module, which uses a large language model to re-check the original text and OCR results that are sensitive to traditional AI judgment, and at the same time uses a multimodal large model to re-check the image data that is sensitive to traditional AI judgment, to further ensure the accuracy of sensitive detection, and finally integrate the detection results to output the final conclusion.
[0210] It should be noted that the examples in the figure are only used to understand the present application and do not constitute a limitation on the information detection method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0211] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0212] This application also provides an information detection device, please refer to Figure 5 , the information detection device includes:
[0213] An acquisition module 51 is configured to acquire original material data corresponding to financial reimbursement materials, and perform optical character recognition on the original material data to obtain a material recognition result corresponding to the original material data;
[0214] A matching module 52 is configured to perform rule matching based on the original material data and the material identification result to obtain a first sensitive detection result;
[0215] A detection module 53 is configured to perform multi-model detection based on the original material data and the material identification result to obtain a second sensitive detection result;
[0216] The generating module 54 is configured to obtain a target sensitive detection result based on the first sensitive detection result and the second sensitive detection result.
[0217] The information detection device is also used for:
[0218] Obtain sensitive information rule base;
[0219] According to the sensitive information rule library, the original material data and the material identification result are matched word by word to check whether sensitive information is contained;
[0220] If so, the specific content and location information of the sensitive information are recorded, and the sensitive information is marked to generate a first sensitive detection result.
[0221] The information detection device is also used for:
[0222] Inputting the original material data and the material recognition result into a natural language discrimination model for sensitivity detection, and obtaining a third sensitivity detection result output by the natural language discrimination model;
[0223] Inputting the original image data into a visual feature discrimination model for sensitivity detection, and obtaining a fourth sensitivity detection result output by the visual feature discrimination model;
[0224] If the third sensitive detection result and the fourth sensitive detection result are information-insensitive results, the third sensitive detection result and the fourth sensitive detection result are associated and combined to generate a second sensitive detection result.
[0225] The information detection device is also used for:
[0226] If the third sensitive detection result and the fourth sensitive detection result are information-sensitive results, inputting the first sensitive detection result and the third sensitive detection result into the large language model for sensitive information rechecking to obtain a fifth sensitive detection result;
[0227] Inputting the fourth sensitive detection result into the multimodal large model for sensitive information rechecking to obtain a sixth sensitive detection result;
[0228] The third sensitive detection result, the fourth sensitive detection result, the fifth sensitive detection result and the sixth sensitive detection result are associated and combined to generate a second sensitive detection result.
[0229] The information detection device is also used for:
[0230] Performing semantic analysis on the original material data and the material recognition result by using the natural language discrimination model to obtain contextual semantic information;
[0231] Determining whether the original material data and the material recognition result involve sensitive information based on the contextual semantic information by the natural language discrimination model;
[0232] If so, the specific content and location information of the sensitive information are recorded, and the sensitive information is marked, and the third sensitive detection result is generated and output through the natural language discrimination model.
[0233] The information detection device is also used for:
[0234] Performing image analysis on the original image data using the visual feature discrimination model to identify entity information in the image;
[0235] Determining whether the entity information meets the financial reimbursement standards by using the visual feature discrimination model;
[0236] If not, the entity information is marked, and the fourth sensitive detection result is generated and output through the visual feature discrimination model.
[0237] The information detection device is also used for:
[0238] Obtain the target user's sensitive detection needs;
[0239] Conduct intent analysis on the sensitive detection requirements to generate model role settings, detection task background, detection task requirements, and model output requirements;
[0240] Based on the model role setting, the detection task background, the detection task requirements and the model output requirements, model prompt words are generated to perform multi-model detection.
[0241] The information detection device provided in this application, employing the information detection method of the above-described embodiment, can solve the technical problems described in the background art. Compared with the prior art, the beneficial effects of the information detection device provided in this application are the same as those of the information detection method provided in the above-described embodiment, and the other technical features of the information detection device are the same as those disclosed in the above-described embodiment and are not further described here.
[0242] The present application provides an information detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the information detection method in the above-mentioned embodiment one.
[0243] Reference below Figure 6, which shows a schematic structural diagram of an information detection device suitable for implementing an embodiment of the present application. The information detection device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The information detection device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0244] like Figure 6 As shown, the information detection device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the information detection device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the information detection device to communicate with other devices wirelessly or wired to exchange data. Although the figure shows an information detection device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or provided instead.
[0245] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0246] The information detection device provided in this application, employing the information detection method of the aforementioned embodiment, can resolve the technical problems described in the background art. Compared to the prior art, the beneficial effects of the information detection device provided in this application are the same as those of the information detection method provided in the aforementioned embodiment, and the other technical features of the information detection device are the same as those disclosed in the aforementioned embodiment, and are not further described here.
[0247] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0248] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0249] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer program) stored thereon, wherein the computer-readable program instructions are used to execute the information detection method in the above-mentioned embodiment.
[0250] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0251] The computer-readable storage medium may be included in the information detection device; or it may exist independently without being assembled into the information detection device.
[0252] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the information detection device, the information detection device:
[0253] Obtaining original material data corresponding to the financial reimbursement materials, and performing optical character recognition on the original material data to obtain a material recognition result corresponding to the original material data;
[0254] Perform rule matching based on the original material data and the material identification result to obtain a first sensitive detection result;
[0255] Performing multi-model detection based on the original material data and the material identification result to obtain a second sensitive detection result;
[0256] A target sensitive detection result is obtained based on the first sensitive detection result and the second sensitive detection result.
[0257] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0258] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0259] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0260] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned information detection method, and can solve the technical problems described in the background art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the information detection method provided in the above-mentioned embodiment, and are not further described here.
[0261] An embodiment of the present application provides a computer program product, including a computer program, which implements the steps of the above-mentioned information detection method when executed by a processor.
[0262] The computer program product provided in this application can solve the technical problems in the background technology. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiment of this application are the same as the beneficial effects of the information detection method provided in the above embodiment, which will not be repeated here.
[0263] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. An information detection method, characterized in that: include: Obtaining original material data corresponding to the financial reimbursement materials, and performing optical character recognition on the original material data to obtain a material recognition result corresponding to the original material data; Perform rule matching based on the original material data and the material identification result to obtain a first sensitive detection result; Performing multi-model detection based on the original material data and the material identification result to obtain a second sensitive detection result; A target sensitive detection result is obtained based on the first sensitive detection result and the second sensitive detection result.
2. The information detection method according to claim 1, wherein: The performing rule matching based on the original material data and the material identification result to obtain a first sensitive detection result includes: Obtain sensitive information rule base; According to the sensitive information rule library, the original material data and the material identification result are matched word by word to check whether sensitive information is contained; If so, the specific content and location information of the sensitive information are recorded, and the sensitive information is marked to generate a first sensitive detection result.
3. The information detection method according to claim 1, wherein: The original material data includes original image data; The performing multi-model detection based on the original material data and the material identification result to obtain a second sensitive detection result includes: Inputting the original material data and the material recognition result into a natural language discrimination model for sensitivity detection, and obtaining a third sensitivity detection result output by the natural language discrimination model; Inputting the original image data into a visual feature discrimination model for sensitivity detection, and obtaining a fourth sensitivity detection result output by the visual feature discrimination model; If the third sensitive detection result and the fourth sensitive detection result are information-insensitive results, the third sensitive detection result and the fourth sensitive detection result are associated and combined to generate a second sensitive detection result.
4. The information detection method according to claim 3, wherein: The method further comprises: If the third sensitive detection result and the fourth sensitive detection result are information-sensitive results, inputting the first sensitive detection result and the third sensitive detection result into the large language model for sensitive information rechecking to obtain a fifth sensitive detection result; Inputting the fourth sensitive detection result into the multimodal large model for sensitive information rechecking to obtain a sixth sensitive detection result; The third sensitive detection result, the fourth sensitive detection result, the fifth sensitive detection result and the sixth sensitive detection result are associated and combined to generate a second sensitive detection result.
5. The information detection method according to claim 3, wherein: The inputting the original material data and the material recognition result into the natural language discrimination model for sensitivity detection to obtain a third sensitivity detection result output by the natural language discrimination model includes: Performing semantic analysis on the original material data and the material recognition result by using the natural language discrimination model to obtain contextual semantic information; Determining whether the original material data and the material recognition result involve sensitive information based on the contextual semantic information by the natural language discrimination model; If so, the specific content and location information of the sensitive information are recorded, and the sensitive information is marked, and the third sensitive detection result is generated and output through the natural language discrimination model.
6. The information detection method according to claim 3, wherein: Inputting the original image data into the visual feature discrimination model for sensitivity detection to obtain a fourth sensitivity detection result output by the visual feature discrimination model includes: Performing image analysis on the original image data using the visual feature discrimination model to identify entity information in the image; Determining whether the entity information meets the financial reimbursement standards by using the visual feature discrimination model; If not, the entity information is marked, and the fourth sensitive detection result is generated and output through the visual feature discrimination model.
7. The information detection method according to claim 1, wherein: Before performing multi-model detection based on the original material data and the material identification result to obtain a second sensitive detection result, the method further includes: Obtain the target user's sensitive detection needs; Conduct intent analysis on the sensitive detection requirements to generate model role settings, detection task background, detection task requirements, and model output requirements; Based on the model role setting, the detection task background, the detection task requirements and the model output requirements, model prompt words are generated to perform multi-model detection.
8. An information detection device, characterized in that: include: an acquisition module, configured to acquire original material data corresponding to financial reimbursement materials, and perform optical character recognition on the original material data to obtain a material recognition result corresponding to the original material data; a matching module, configured to perform rule matching based on the original material data and the material identification result to obtain a first sensitive detection result; a detection module, configured to perform multi-model detection based on the original material data and the material identification result to obtain a second sensitive detection result; A generation module is used to obtain a target sensitive detection result based on the first sensitive detection result and the second sensitive detection result.
9. An information detection device, characterized in that: The information detection device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the information detection method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the information detection method according to any one of claims 1 to 7 are implemented.