Page checking method and device based on multi-modal AI, equipment and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN HEALTH INSURANCE CO LTD
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]传统页面校验方式多依赖正则表达式对“投保人姓名”“身份证号”等关键字段进行格式校验,或借助文本比对工具检测内容是否重复;但这类方法仅能实现表层格式校验与字面匹配,容易产生大量假阳性告警,同时无法理解页面内容的真实含义,从而导致页面校验的准确度不足
[0010]上述一种基于多模态AI的页面校验方法、装置、计算机设备及存储介质所实现的方案中,通过对页面截图数据进行光学字符识别,保证文本识别与定位准确;通过提取多模态元素,并构建三维映射表,将文本、位置、页面元素三者关联绑定,实现页面信息的结构化表达,为后续风险检测提供统一数据基础;通过对三维映射表中的页面文本内容语义解析,能够精准识别语义错误、逻辑矛盾和表述歧义,有效弥补传统关键词匹配方法的不足,大幅提升校验的准确性,通过对多模态元素进行多维条件校验,能够快速检测格式、范围、时间一致性等硬性违规问题,确保页面内容符合产品的规范要求,降低合规风险,通过将语义风险项与条件校验结果协同校准,双重校验交叉验证,排除误判、提升校验结果的准确率。
Smart Images

Figure CN122528880A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image detection technology, and in particular to page verification methods, apparatus, devices and media based on multimodal AI. Background Technology
[0002] With the rapid iteration of internet technology and digital products, various interactive pages such as web pages and app interfaces have become the core carriers for information display and business interaction, and are widely used in many fields such as healthcare, finance, government affairs, and media. For example, in the healthcare field, pages are mainly used to display diagnosis and treatment information, medication guidance, examination reports, and health education content; and in the fintech field, pages are often used to present key content such as account information, transaction data, business rules, and risk warnings.
[0003] Traditional page validation methods often rely on regular expressions to validate the format of key fields such as "insured's name" and "ID number," or use text comparison tools to detect duplicate content. However, these methods can only achieve surface format validation and literal matching, which can easily generate a large number of false positive alerts. At the same time, they cannot understand the true meaning of the page content, resulting in insufficient accuracy of page validation.
[0004] Therefore, in the face of the increasing demand for page validation, current page validation methods urgently need to be improved to solve the problem of insufficient accuracy of page validation results from existing methods. Summary of the Invention
[0005] This invention provides a page verification method, apparatus, device, and medium based on multimodal AI, which addresses how to improve the accuracy of page verification results.
[0006] Firstly, a page verification method based on multimodal AI is provided, including: Obtain the page screenshot data loaded by the browser engine and the page structure data corresponding to the page screenshot data, perform optical character recognition on the page screenshot data, and obtain the recognized page text content and the position information of the page text content; Multimodal elements are extracted from the page structure data, and a three-dimensional mapping table is constructed based on the page text content, the location information, and the multimodal elements. Semantic analysis is performed on the page text content in the three-dimensional mapping table to identify semantic risk items; Perform multidimensional condition verification on the multimodal elements in the three-dimensional mapping table to obtain the condition verification results; The semantic risk items and the conditional verification results are calibrated together to generate the verification results of the page screenshot data.
[0007] Secondly, a page verification device based on multimodal AI is provided, comprising: The text location recognition module is used to obtain page screenshot data loaded by the browser engine and page structure data corresponding to the page screenshot data, perform optical character recognition on the page screenshot data, and obtain the recognized page text content and the location information of the page text content. A multimodal element extraction module is used to extract multimodal elements from the page structure data; A three-dimensional mapping table construction module is used to construct a three-dimensional mapping table based on the page text content, the location information, and the multimodal elements; The semantic risk item determination module is used to perform semantic parsing on the page text content in the three-dimensional mapping table to determine semantic risk items. The condition verification result verification module is used to perform multi-dimensional condition verification on the multimodal elements in the three-dimensional mapping table to obtain the condition verification result; The verification result generation module is used to collaboratively calibrate the semantic risk item with the conditional verification result to generate the verification result of the page screenshot data.
[0008] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the aforementioned page verification method based on multimodal AI.
[0009] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the aforementioned page verification method based on multimodal AI.
[0010] The aforementioned solution, implemented using a multimodal AI-based page verification method, apparatus, computer equipment, and storage medium, ensures accurate text recognition and positioning by performing optical character recognition on page screenshot data. It extracts multimodal elements and constructs a three-dimensional mapping table, linking text, location, and page elements to achieve a structured representation of page information, providing a unified data foundation for subsequent risk detection. Semantic analysis of the page text content in the three-dimensional mapping table accurately identifies semantic errors, logical contradictions, and ambiguities, effectively compensating for the shortcomings of traditional keyword matching methods and significantly improving verification accuracy. Multi-dimensional conditional verification of multimodal elements quickly detects hard violations such as format, range, and time consistency, ensuring page content complies with product specifications and reducing compliance risks. Collaborative calibration of semantic risk items and conditional verification results, along with dual verification and cross-validation, eliminates misjudgments and improves the accuracy of verification results. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of an application environment for a page verification method based on multimodal AI according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a page verification method based on multimodal AI in one embodiment of the present invention; Figure 3 This is a schematic diagram of a page verification device based on multimodal AI in one embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to one embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] This invention provides a page verification method based on multimodal AI, which can be applied to applications such as... Figure 1In this application environment, the client communicates with the server via a network. The server can obtain page screenshot data loaded by the browser engine and the corresponding page structure data. It performs optical character recognition (OCR) on the page screenshot data to obtain the recognized page text content and its location information. Multimodal elements are extracted from the page structure data, and a three-dimensional mapping table is constructed based on the page text content, the location information, and the multimodal elements. Semantic parsing is performed on the page text content in the three-dimensional mapping table to identify semantic risk items. Multi-dimensional conditional verification is performed on the multimodal elements in the three-dimensional mapping table to obtain conditional verification results. The semantic risk items and the conditional verification results are then collaboratively calibrated to generate the verification result of the page screenshot data. The verification result is then fed back to the client. This invention provides a page verification device based on multimodal AI. For verification result processing, by collaboratively calibrating semantic risk items and conditional verification results, and performing dual verification and cross-validation, false judgments are eliminated and the accuracy of the verification results is improved. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The present invention will now be described in detail through specific embodiments.
[0015] Please see Figure 2 As shown, Figure 2 A flowchart illustrating a page verification method based on multimodal AI provided in an embodiment of the present invention includes the following steps: S1. Obtain the page screenshot data loaded by the browser engine and the page structure data corresponding to the page screenshot data, perform optical character recognition on the page screenshot data, and obtain the recognized page text content and the position information of the page text content.
[0016] In this embodiment of the invention, the page screenshot data refers to the image file generated by taking a screenshot of the content displayed on the current screen after the browser engine loads the H5 page, and the page structure data refers to the document object model structure data captured from the rendering engine after the browser engine loads the H5 page, which can also be called the document object model tree, including code structure, element composition and attribute information.
[0017] In this embodiment of the invention, the page text content refers to all text information that is visible to the user and extracted from the page screenshot using optical character recognition technology, including titles, body text, button labels, prompts, etc. The position information includes the coordinates of the top left corner, the coordinates of the bottom right corner, width, and height.
[0018] In the healthcare scenario, screenshots of medical service pages loaded by the browser engine and the corresponding page structure data are obtained. These pages include registration pages, consultation pages, prescription pages, test report pages, and health management pages. Optical character recognition is performed on these pages to obtain the text content of the medical-related pages and the location information of the text content within the medical pages. The text content includes patient information, symptom descriptions, medical orders, drug information, test indicators, and medical records.
[0019] In the financial health scenario, screenshots of financial business pages loaded by the browser engine and the corresponding page structure data are obtained. These pages include payment pages, transfer pages, wealth management pages, credit pages, identity verification pages, and transaction confirmation pages. Optical character recognition (OCR) is performed on the screenshots of financial business pages to obtain the recognized financial-related page text content and its location information within the financial pages. The text content includes account information, transaction amount, interest rate information, contract terms, risk warnings, and verification elements.
[0020] In this embodiment of the invention, the step of performing optical character recognition on the page screenshot data to obtain the recognized page text content and the location information of the page text content includes: The page screenshot data is preprocessed to generate a standard page image; Detect the text region in the standard page image; Character recognition is performed on the text area to obtain the page text content; The position of the page text content is parsed to determine the position information of the page text content in the page screenshot data.
[0021] In this embodiment of the invention, page screenshot data captured after the browser engine loads is obtained, and grayscale processing, noise removal, contrast enhancement, and tilt correction operations are sequentially performed on the page screenshot data to obtain standard page screenshot data that meets the recognition requirements.
[0022] Furthermore, the system identifies regions in the standard page screenshot data where pixel values change drastically. It then aggregates adjacent black pixels within the identified regions into independent connected regions, connects neighboring connected regions, merges them into text blocks, and further aggregates text blocks belonging to the same paragraph or with the same semantic meaning to form text regions.
[0023] Furthermore, the text regions are adjusted to a uniform size and pixel values are normalized. Visual features in the processed candidate text regions are extracted, and the contextual dependencies of the processed text regions are captured. Based on the visual features and contextual dependencies, the text regions are decoded to obtain the page text content corresponding to the text regions.
[0024] Furthermore, the relative position coordinates of each character in the text area are obtained, and the relative position coordinates are mapped back to the absolute coordinate system of the page screenshot data. The spacing and arrangement direction between characters in the text area are determined. Based on the spacing and arrangement direction, characters belonging to the same text line or the same semantic unit are merged into a complete text block. The overall bounding box coordinates of the complete text block are calculated and calibrated based on the mapped absolute coordinates. The overall bounding box coordinates are used as the position information of the page text content in the page screenshot data.
[0025] In this embodiment of the invention, image preprocessing of page screenshot data can eliminate interference such as image noise, tilt, and uneven lighting, thereby improving image quality. By detecting text regions in standard page screenshot data, the location of text can be accurately located, invalid background areas can be filtered out, the amount of data for subsequent character recognition can be reduced, and recognition efficiency and accuracy can be improved. By performing character recognition and position parsing on text regions, a one-to-one correspondence between text content and page position can be achieved, providing positional basis for subsequent processes such as risk positioning and collaborative verification.
[0026] S2. Extract multimodal elements from the page structure data, and construct a three-dimensional mapping table based on the page text content, the location information, and the multimodal elements.
[0027] In this invention, the multimodal elements refer to various types and forms of page components contained in the page structure data, including buttons, input boxes, labels, images, icons, links, drop-down lists, etc. The three-dimensional mapping table refers to a mapping relationship table that associates and stores text content, position coordinates, and page structure elements in a structured manner.
[0028] In the healthcare scenario, multimodal elements are extracted from the structural data of medical business pages. These multimodal elements include registration buttons, consultation input boxes, prescription labels, drug information controls, test report areas, and patient information modules. Based on the medical page text content, page text location information, and multimodal elements obtained by OCR recognition, a three-dimensional mapping table is constructed that associates medical text, location coordinates, and page controls. This table is used to achieve structured verification and risk detection of medical pages.
[0029] In fintech scenarios, multimodal elements are extracted from the structural data of financial business pages. These multimodal elements include transfer buttons, payment input boxes, contract terms areas, interest rate display labels, identity verification controls, and transaction confirmation modules. Based on the financial page text content, page text location information, and multimodal elements obtained through OCR recognition, a three-dimensional mapping table is constructed that associates financial text, location coordinates, and page controls. This table is used to achieve compliance verification and risk identification of financial pages.
[0030] In this embodiment of the invention, extracting multimodal elements from the page structure data includes: The page structure data is traversed to obtain a set of page nodes; The page nodes are classified according to preset categories to obtain category nodes with type labels; The attribute information of the classification nodes is parsed, and the classification nodes are spatially aligned according to the node position information in the attribute information to determine the multimodal elements.
[0031] In this embodiment of the invention, a depth-first traversal or a breadth-first traversal is performed on the page structure data to scan all hierarchical nodes in the page structure data in sequence, and all the scanned nodes are used as a set of page nodes. The basic type of each node in the set of page nodes is determined, and non-content nodes such as comment nodes and document type declaration nodes are filtered out according to the determined basic type. The retained nodes are classified into predefined categories according to the tag names of the retained nodes to obtain category nodes containing types.
[0032] For example, nodes with tags such as p, div, or span that contain text content are classified as text nodes; nodes with tags such as button or that have button role attributes are classified as button nodes.
[0033] Furthermore, the attribute information includes the display content, control type, and node position information of the corresponding node. In the process of parsing the attribute information of the category node, the category node is parsed, and the feature information carried by the category node itself is extracted according to the parsing result. Based on the feature information, the display text of the category node on the page, the corresponding control category, and the position information such as the coordinates and size of the category node on the page are obtained. The obtained relevant information is used as the attribute information of the category node.
[0034] For example, extract the rendered text content of text nodes or nodes containing text in category nodes, including the text value of the node itself and the additional text generated by CSS pseudo-elements. For button nodes in category nodes, extract not only the displayed text on the button, but also its associated jump link address, click event type, and button style status (such as enabled or disabled).
[0035] Furthermore, during the spatial alignment of the classification nodes, the absolute coordinate system of the page screenshot data is used as the reference. Based on the node position information in the attribute information, the classification nodes are matched with the corresponding areas in the page screenshot data and their spatial positions are calibrated to obtain multimodal elements with accurate positions.
[0036] In this embodiment of the invention, constructing a three-dimensional mapping table based on the page text content, the location information, and the multimodal elements includes: Establish a text coordinate mapping set based on the page text content and the location information; Construct an element region mapping set based on the attribute information of the multimodal elements; Spatial overlay analysis is performed between the text coordinate mapping set and the element region mapping set to determine the text element association set; A three-dimensional mapping set is generated based on the association set of the text elements.
[0037] In this embodiment of the invention, the page text content is used as a key value, and its corresponding position information is used as a position attribute, thereby combining them into a text coordinate mapping entry with a text coordinate mapping relationship. All identified page text content is traversed, and text coordinate mapping relationships are established one by one to generate a text coordinate mapping set.
[0038] Furthermore, extract information such as the type label, display content, control properties, and coordinates of the coverage area determined after spatial alignment for each modal element in the multimodal element. Combine the extracted information with the modal element to form an element region mapping entry, and then combine multiple element region mapping entries into an element region mapping set.
[0039] Furthermore, each text coordinate mapping entry in the text coordinate mapping set is traversed, and the spatial inclusion relationship between the position information in the text coordinate mapping entry and the coverage area coordinates of each modal element in the element area mapping set is determined. That is, if the position information in a certain text coordinate mapping entry is completely within the coverage area coordinates of a certain modal element, or has a clear spatial adjacency relationship with the coverage area coordinates of the modal element, then the association relationship between the page text content and the corresponding modal element is established. For page text content that has not established an association relationship, it is associated with the nearest modal element according to the spatial proximity rule to obtain the text element association set.
[0040] Furthermore, the text content and its corresponding location information, modal element type labels, display content, and control properties in the text element association set are combined into a three-dimensional mapping entry, and multiple three-dimensional mapping entries are combined into a three-dimensional mapping set.
[0041] In this embodiment of the invention, by constructing an element region mapping set, multimodal elements are associated with page regions, clearly recording the type, range, and position of page controls, which facilitates subsequent spatial matching with text. By performing spatial overlay analysis on the text coordinate mapping set and the element region mapping set, the control to which the text belongs is determined, and the correspondence between text and elements is established, transforming page information from discrete to structured association.
[0042] S3. Perform semantic parsing on the page text content in the three-dimensional mapping table to determine semantic risk items.
[0043] In this embodiment of the invention, semantic risk items refer to specific text fragments identified from the page text content that contain semantic errors, ambiguous expressions, or logical contradictions.
[0044] In the healthcare context, semantic analysis is performed on the text content of medical-related pages in the three-dimensional mapping table to identify statements involving illegal medical treatment promotions, false efficacy promises, inappropriate medication guidelines, privacy leaks, and non-compliance with medical standards, thereby determining semantic risk items on the page.
[0045] In fintech scenarios, semantic analysis is performed on the text content of financial-related pages in a 3D mapping table to identify statements involving false investment returns, illegal financial advertising, misleading loan terms, missing risk warnings, and violations of financial regulatory requirements, in order to determine semantic risk items on the page.
[0046] In this embodiment of the invention, the step of performing semantic parsing on the page text content in the three-dimensional mapping table to determine semantic risk items includes: The page text content in the three-dimensional mapping table is segmented into words to obtain a segmented text sequence. Identify the entity information of each segmented text in the segmented text sequence; Based on the entity information and preset domain semantic conditions, semantic relationship analysis is performed on adjacent texts of the segmented text to determine candidate risk items of semantic contradiction; A confidence analysis is performed on the candidate risk items, and the candidate risk items with a confidence level exceeding a preset confidence level threshold are aggregated into semantic risk items.
[0047] In this embodiment of the invention, the page text content in the three-dimensional mapping table is split into units such as words and phrases, redundant symbols and invalid characters are removed to form a word segmented text sequence. Based on a preset insurance domain dictionary, each word in the word segmented text sequence is labeled with an entity. Based on the labeled entities, key entity information such as business entities, institution names, amounts, clause keywords, and risk warnings are identified and used as entity information.
[0048] Furthermore, based on entity information and preset domain semantic conditions, the logical relationships, collocation relationships, and constraint relationships between adjacent texts in the segmented text are judged, and candidate risk items with problems such as expression contradictions, illegal publicity, and misleading descriptions are identified based on the judgment results.
[0049] For example, when "congenital diseases" are excluded in the "Liability Exclusion" module, while the "Product Manual" module contains a description that "covers congenital diseases", a logical contradiction is identified between the two and it is marked as a candidate risk item.
[0050] Furthermore, the semantic conflict intensity between the segmented text and adjacent text in the candidate risk items is analyzed; the entity matching degree between the pre-set compliance entity library and the entity information in the candidate risk items is calculated; the text features of the candidate risk items are extracted, and the violation feature matching degree between the text features and the pre-set domain violation feature template is calculated. The semantic conflict intensity, entity matching degree, and violation feature matching degree are weighted and summed to obtain the confidence level of each candidate risk item.
[0051] In this embodiment of the invention, by performing word segmentation on the page text content, long texts are split into standardized and ordered word sequences, reducing the difficulty of semantic parsing; by identifying the entity information of each word segmented text, key entities in the domain are accurately extracted, the core analysis object is locked, interference from invalid information is avoided, and the accuracy of subsequent semantic judgment is improved; by performing semantic relationship analysis on adjacent texts, suspicious content is quickly located, and candidates with violation risks are initially screened; by performing confidence analysis on candidate risk items, the reliability of the identification results is ensured, and real and effective semantic risk items are output.
[0052] S4. Perform multi-dimensional condition verification on the multimodal elements in the three-dimensional mapping table to obtain the condition verification results.
[0053] In this embodiment of the invention, the condition verification result refers to the data set that outputs the violating element, the violation type, the specific rule violated, and the corresponding element location information after verifying the multimodal element by preset rules (such as format, time, range, etc.).
[0054] In this embodiment of the invention, the step of performing multi-dimensional condition verification on the multimodal elements in the three-dimensional mapping table to obtain the condition verification result includes: Based on the type corresponding to each modal element in the multimodal element of the three-dimensional mapping table, the modal element is matched with a preset verification condition library to obtain the verification condition corresponding to the modal element; The modal elements are checked item by item according to the verification conditions to obtain the element verification results corresponding to the modal elements; The element verification results are aggregated into conditional verification results.
[0055] In this embodiment of the invention, the preset verification condition library includes format rules (text content must conform to a specific format), time rules (date expressions must be consistent), range rules (numerical values must be within the allowed range), and existence rules (required fields cannot be empty), etc.
[0056] In this embodiment of the invention, each modal element in the multimodal elements of the three-dimensional mapping table is traversed, and all verification conditions applicable to the type label are retrieved from the preset verification condition library according to the type label of the modal element. All the retrieved verification conditions are associated with the modal element to obtain the verification conditions corresponding to the modal element.
[0057] For example, for an element of type "form field" and subtype "age input box", the system will match the validation condition "age must be an integer and between 0 and 100".
[0058] Furthermore, based on the validation conditions, the text content, position layout, display style, interaction logic, and completeness of required fields of the modal element are checked item by item to determine whether they conform to the preset specifications. The element validation result is determined according to the judgment result, which includes the execution status of all validation conditions of the modal element. The element validation results are classified according to the type label of the modal element to obtain the condition validation result.
[0059] In this embodiment of the invention, by matching modal elements with a preset verification condition library, the verification content is ensured to be highly targeted and non-redundant, avoiding invalid verification and improving the rationality of verification. By verifying the modal elements item by item according to the verification conditions, the verification can be guaranteed to be comprehensive and detailed, and can accurately detect non-compliance issues of multimodal elements in terms of content, format and logic.
[0060] S5. Perform collaborative calibration of the semantic risk item and the conditional verification result to generate the verification result of the page screenshot data.
[0061] In this embodiment of the invention, the step of collaboratively calibrating the semantic risk item and the conditional verification result to generate the verification result of the page screenshot data includes: The semantic risk items are matched with the condition verification results to obtain matched risk items that match and non-matched risk items that do not match. The confidence level corresponding to the non-matching risk item is compared with a preset risk confidence threshold, and risk items to be confirmed are selected from the non-matching risk items based on the comparison result; The contents of the risk items to be confirmed are reviewed to obtain valid risk items that pass the review; The matching risk items and the valid risk items are merged and deduplicated to generate the verification result of the page screenshot.
[0062] Furthermore, based on the location information, control to which the semantic risk item belongs, and the content of the violation, it is compared one by one with the abnormal location and violation type of the modal element in the conditional validation result to determine whether the two point to the same risk area or the same violation issue. Those that corroborate each other and have the same location and type are recorded as matching risk items, and those that cannot be matched are recorded as non-matching risk items.
[0063] Furthermore, the confidence level corresponding to the non-matching risk item is compared with the pre-set risk confidence level threshold. Non-matching risk items with a confidence level greater than or equal to the threshold are retained and treated as risk items to be confirmed. Combining page screenshot data, text content in the 3D mapping table, and attribute information of multimodal elements, the text semantics, control logic, and display content of the risk items to be confirmed are manually or automatically reviewed to verify whether they actually violate regulations or contradictions. Risk items with violations or contradictions are deleted, and valid risk items that have been reviewed and confirmed to have risks are retained.
[0064] Furthermore, the matching risk items are merged with the valid risk items, and duplicate identification and labeling of the same risk item are deduplicated to form a risk list. Based on this risk list, the overall compliance judgment of the page screenshot data is made, and the risk location, risk type, risk level and confidence level of the risk list are marked to form the verification result.
[0065] In this embodiment of the invention, by performing confidence analysis on semantic risk items, a confidence score is assigned to each semantic risk item, thereby achieving quantitative differentiation of risks, facilitating subsequent accurate screening and verification, and improving the reliability of risk identification. By matching semantic risk items with conditional verification results, bidirectional verification enables risks to mutually corroborate each other, distinguishing between confirmed and questionable risks, and reducing misjudgments caused by single detection. By reviewing the content of risk items to be confirmed, system misjudgments are eliminated, ensuring that the finally identified risks are true and effective.
[0066] As can be seen, in the above solution, for the verification result business, the page screenshot data loaded by the browser engine and the page structure data corresponding to the page screenshot data are obtained. Optical character recognition is performed on the page screenshot data to obtain the recognized page text content and the location information of the page text content. Multimodal elements are extracted from the page structure data, and a three-dimensional mapping table is constructed based on the page text content, the location information, and the multimodal elements. Semantic parsing is performed on the page text content in the three-dimensional mapping table to determine semantic risk items. Multidimensional conditional verification is performed on the multimodal elements in the three-dimensional mapping table to obtain conditional verification results. The semantic risk items and the conditional verification results are collaboratively calibrated to generate the verification result of the page screenshot data. By collaboratively calibrating the semantic risk items and the conditional verification results, dual verification and cross-validation are performed to eliminate false judgments and improve the accuracy of the verification results.
[0067] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0068] In one embodiment, a page verification device based on multimodal AI is provided, which corresponds one-to-one with the page verification method based on multimodal AI in the above embodiments. For example... Figure 3 As shown, this page verification device based on multimodal AI includes a text location recognition module 101, a multimodal element extraction module 102, a three-dimensional mapping table construction module 103, a semantic risk item determination module 104, a conditional verification result verification module 105, and a verification result generation module 106. Detailed descriptions of each functional module are as follows: The text location recognition module 101 is used to obtain page screenshot data loaded by the browser engine and page structure data corresponding to the page screenshot data, perform optical character recognition on the page screenshot data, and obtain the recognized page text content and the location information of the page text content. The multimodal element extraction module 102 is used to extract multimodal elements from the page structure data; The three-dimensional mapping table construction module 103 is used to construct a three-dimensional mapping table based on the page text content, the location information, and the multimodal elements; The semantic risk item determination module 104 is used to perform semantic parsing on the page text content in the three-dimensional mapping table to determine semantic risk items. The condition verification result verification module 105 is used to perform multi-dimensional condition verification on the multimodal elements in the three-dimensional mapping table to obtain the condition verification result. The verification result generation module 106 is used to collaboratively calibrate the semantic risk item with the conditional verification result to generate the verification result of the page screenshot data.
[0069] In one embodiment, the text location recognition module 101, when performing optical character recognition on the page screenshot data to obtain the recognized page text content and the location information of the page text content, is used to: The page screenshot data is preprocessed to generate a standard page image; Detect the text region in the standard page image; Character recognition is performed on the text area to obtain the page text content; The position of the page text content is parsed to determine the position information of the page text content in the page screenshot data.
[0070] In one embodiment, the multimodal element extraction module 102, when extracting multimodal elements from the page structure data, is used to: The page structure data is traversed to obtain a set of page nodes; The page nodes are classified according to preset categories to obtain category nodes with type labels; The attribute information of the classification nodes is parsed, and the classification nodes are spatially aligned according to the node position information in the attribute information to determine the multimodal elements.
[0071] In one embodiment, when constructing a three-dimensional mapping table based on the page text content, the location information, and the multimodal elements, the three-dimensional mapping table construction module 103 is used to: Establish a text coordinate mapping set based on the page text content and the location information; Construct an element region mapping set based on the attribute information of the multimodal elements; Spatial overlay analysis is performed between the text coordinate mapping set and the element region mapping set to determine the text element association set; A three-dimensional mapping set is generated based on the association set of the text elements.
[0072] In one embodiment, the semantic risk item determination module 104, when performing semantic parsing on the page text content in the three-dimensional mapping table to determine semantic risk items, is used to: The page text content in the three-dimensional mapping table is segmented into words to obtain a segmented text sequence. Identify the entity information of each segmented text in the segmented text sequence; Based on the entity information and preset domain semantic conditions, semantic relationship analysis is performed on adjacent texts of the segmented text to determine candidate risk items of semantic contradiction; A confidence analysis is performed on the candidate risk items, and the candidate risk items with a confidence level exceeding a preset confidence level threshold are aggregated into semantic risk items.
[0073] In one embodiment, when the condition verification result verification module 105 performs multi-dimensional condition verification on the multimodal elements in the three-dimensional mapping table and obtains the condition verification result, it is used to: Based on the type corresponding to each modal element in the multimodal element of the three-dimensional mapping table, the modal element is matched with a preset verification condition library to obtain the verification condition corresponding to the modal element; The modal elements are checked item by item according to the verification conditions to obtain the element verification results corresponding to the modal elements; The element verification results are aggregated into conditional verification results.
[0074] In one embodiment, when the verification result generation module 106 performs collaborative calibration of the semantic risk item and the conditional verification result to generate the verification result of the page screenshot data, it is used to: The semantic risk items are matched with the condition verification results to obtain matched risk items that match and non-matched risk items that do not match. The confidence level corresponding to the non-matching risk item is compared with a preset risk confidence threshold, and risk items to be confirmed are selected from the non-matching risk items based on the comparison result; The contents of the risk items to be confirmed are reviewed to obtain valid risk items that pass the review; The matching risk items and the valid risk items are merged and deduplicated to generate the verification result of the page screenshot.
[0075] This invention provides a page verification device based on multimodal AI. For verification result processing, it acquires page screenshot data loaded by the browser engine and corresponding page structure data. Optical character recognition (OCR) is performed on the page screenshot data to obtain the recognized page text content and its location information. Multimodal elements are extracted from the page structure data, and a three-dimensional mapping table is constructed based on the page text content, the location information, and the multimodal elements. Semantic parsing is performed on the page text content in the three-dimensional mapping table to identify semantic risk items. Multidimensional conditional verification is performed on the multimodal elements in the three-dimensional mapping table to obtain conditional verification results. The semantic risk items and the conditional verification results are then collaboratively calibrated to generate the verification result for the page screenshot data. By collaboratively calibrating the semantic risk items and the conditional verification results, and performing dual verification and cross-validation, false judgments are eliminated, and the accuracy of the verification results is improved.
[0076] For specific limitations regarding a page verification device based on multimodal AI, please refer to the limitations of a page verification method based on multimodal AI mentioned above, which will not be repeated here. Each module in the aforementioned page verification device based on multimodal AI can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0077] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a multimodal AI-based page validation method on the server side.
[0078] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a page validation method based on multimodal AI.
[0079] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Obtain the page screenshot data loaded by the browser engine and the page structure data corresponding to the page screenshot data, perform optical character recognition on the page screenshot data, and obtain the recognized page text content and the position information of the page text content; Multimodal elements are extracted from the page structure data, and a three-dimensional mapping table is constructed based on the page text content, the location information, and the multimodal elements. Semantic analysis is performed on the page text content in the three-dimensional mapping table to identify semantic risk items; Perform multidimensional condition verification on the multimodal elements in the three-dimensional mapping table to obtain the condition verification results; The semantic risk items and the conditional verification results are calibrated together to generate the verification results of the page screenshot data.
[0080] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Obtain the page screenshot data loaded by the browser engine and the page structure data corresponding to the page screenshot data, perform optical character recognition on the page screenshot data, and obtain the recognized page text content and the position information of the page text content; Multimodal elements are extracted from the page structure data, and a three-dimensional mapping table is constructed based on the page text content, the location information, and the multimodal elements. Semantic analysis is performed on the page text content in the three-dimensional mapping table to identify semantic risk items; Perform multidimensional condition verification on the multimodal elements in the three-dimensional mapping table to obtain the condition verification results; The semantic risk items and the conditional verification results are calibrated together to generate the verification results of the page screenshot data.
[0081] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0082] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0083] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0084] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. It should be noted that if AI models, software tools, or components other than those of our company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The user personal information involved in the embodiments of this application is all authorized (knowing and agreeing) by the relevant parties or fully authorized by all parties, and the executing entity can obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with the provisions of relevant laws and regulations and do not violate public order and good morals. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A page verification method based on multimodal AI, characterized in that, include: Obtain the page screenshot data loaded by the browser engine and the page structure data corresponding to the page screenshot data, perform optical character recognition on the page screenshot data, and obtain the recognized page text content and the position information of the page text content; Multimodal elements are extracted from the page structure data, and a three-dimensional mapping table is constructed based on the page text content, the location information, and the multimodal elements. Semantic analysis is performed on the page text content in the three-dimensional mapping table to identify semantic risk items; Perform multidimensional condition verification on the multimodal elements in the three-dimensional mapping table to obtain the condition verification results; The semantic risk items and the conditional verification results are calibrated together to generate the verification results of the page screenshot data.
2. The page verification method based on multimodal AI as described in claim 1, characterized in that, The step of performing optical character recognition on the page screenshot data to obtain the recognized page text content and the location information of the page text content includes: The page screenshot data is preprocessed to generate a standard page image; Detect the text region in the standard page image; Character recognition is performed on the text area to obtain the page text content; The position of the page text content is parsed to determine the position information of the page text content in the page screenshot data.
3. The page verification method based on multimodal AI as described in claim 1, characterized in that, The extraction of multimodal elements from the page structure data includes: The page structure data is traversed to obtain a set of page nodes; The page nodes are classified according to preset categories to obtain category nodes with type labels; The attribute information of the classification nodes is parsed, and the classification nodes are spatially aligned according to the node position information in the attribute information to determine the multimodal elements.
4. The page verification method based on multimodal AI as described in claim 1, characterized in that, The step of constructing a three-dimensional mapping table based on the page text content, the location information, and the multimodal elements includes: Establish a text coordinate mapping set based on the page text content and the location information; Construct an element region mapping set based on the attribute information of the multimodal elements; Spatial overlay analysis is performed between the text coordinate mapping set and the element region mapping set to determine the text element association set; A three-dimensional mapping set is generated based on the association set of the text elements.
5. The page verification method based on multimodal AI as described in claim 1, characterized in that, The step of performing semantic parsing on the page text content in the three-dimensional mapping table to determine semantic risk items includes: The page text content in the three-dimensional mapping table is segmented into words to obtain a segmented text sequence. Identify the entity information of each segmented text in the segmented text sequence; Based on the entity information and preset domain semantic conditions, semantic relationship analysis is performed on adjacent texts of the segmented text to determine candidate risk items of semantic contradiction; A confidence analysis is performed on the candidate risk items, and the candidate risk items with a confidence level exceeding a preset confidence level threshold are aggregated into semantic risk items.
6. The page verification method based on multimodal AI as described in claim 1, characterized in that, Multidimensional condition validation is performed on the multimodal elements in the three-dimensional mapping table to obtain the condition validation results, including: Based on the type corresponding to each modal element in the multimodal element of the three-dimensional mapping table, the modal element is matched with a preset verification condition library to obtain the verification condition corresponding to the modal element; The modal elements are checked item by item according to the verification conditions to obtain the element verification results corresponding to the modal elements; The element verification results are aggregated into conditional verification results.
7. The page verification method based on multimodal AI as described in claim 1, characterized in that, The step of collaboratively calibrating the semantic risk item with the conditional verification result to generate the verification result of the page screenshot data includes: The semantic risk items are matched with the condition verification results to obtain matched risk items that match and non-matched risk items that do not match. The confidence level corresponding to the non-matching risk item is compared with a preset risk confidence threshold, and risk items to be confirmed are selected from the non-matching risk items based on the comparison result; The contents of the risk items to be confirmed are reviewed to obtain valid risk items that pass the review; The matching risk items and the valid risk items are merged and deduplicated to generate the verification result of the page screenshot.
8. A page verification device based on multimodal AI, characterized in that, include: The text location recognition module is used to obtain page screenshot data loaded by the browser engine and page structure data corresponding to the page screenshot data, perform optical character recognition on the page screenshot data, and obtain the recognized page text content and the location information of the page text content. A multimodal element extraction module is used to extract multimodal elements from the page structure data; A three-dimensional mapping table construction module is used to construct a three-dimensional mapping table based on the page text content, the location information, and the multimodal elements; The semantic risk item determination module is used to perform semantic parsing on the page text content in the three-dimensional mapping table to determine semantic risk items. The condition verification result verification module is used to perform multi-dimensional condition verification on the multimodal elements in the three-dimensional mapping table to obtain the condition verification result; The verification result generation module is used to collaboratively calibrate the semantic risk item with the conditional verification result to generate the verification result of the page screenshot data.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the page verification method based on multimodal AI as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the page verification method based on multimodal AI as described in any one of claims 1 to 7.