Commercial password application security assessment evidence identification method based on image identification
Through the commercial cryptographic application security assessment method based on image recognition, using the Qwen-VL model and graph database technology, the consistency and objectivity problems in traditional certificate evaluation methods are solved, multi-dimensional consistency checking and dynamic credibility adjustment of certificates and evidence are realized, and the accuracy and reliability of the assessment results are improved.
Patent Information
- Application Number
- CN202511220527.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Traditional certificate evaluation methods have difficulty processing massive amounts of data, cannot guarantee the consistency and objectivity of evaluation results, cannot effectively link certificates with evidence, lack a dynamic adjustment mechanism, and cannot automatically and carefully compare subtle differences, resulting in one-sided evaluation results.
The Qwen-VL visual language model is used to extract information from multimodal documents, generate high-dimensional vector representations and convert them into structured evidence nodes. An evidence network is established through a graph database to perform forward and reverse traceability, dynamically adjust the credibility of certificate entity groups, and screen out abnormal certificates.
It implements multi-dimensional consistency checks on certificates and evidence, dynamically adjusts certificate credibility, improves the objectivity and consistency of evaluation results, discovers and marks abnormal certificates, and generates visual reports.
Smart Images

Figure CN120746609A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a commercial cryptographic application security assessment evidence recognition method based on image recognition. Background Art
[0002] In many scenarios such as personnel qualification certification, contract review, and supply chain traceability, evaluating the authenticity, completeness, and consistency of submitted certificates and documents is a crucial step.
[0003] Traditional certificate evaluation methods often rely on manual review, making it difficult to process massive amounts of data and unable to guarantee the consistency and objectivity of evaluation results. Existing technologies usually treat certificates and evidence in isolation, making it impossible to effectively link them and form a complete chain of evidence, resulting in one-sided evaluation results. Traditional verification methods often focus on a single dimension, making it difficult to systematically check the comprehensive consistency or matching degree between the certificate and the associated evidence in multiple dimensions such as content, format, seal, source and context, and unable to automatically and carefully compare subtle differences. Moreover, the evaluation is often static and lacks a dynamic adjustment mechanism. Once certain key fields are matched or an anti-counterfeiting mark is identified, the certificate is easily judged to be valid. It is impossible to dynamically update the credibility assessment of the certificate based on the mutual support or contradiction between the evidence, as well as the quality and timeliness of the evidence itself. Summary of the Invention
[0004] The purpose of the present invention is to provide a commercial cryptographic application security assessment evidence identification method based on image recognition to solve the problems raised in the prior art.
[0005] To achieve the above object, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for identifying evidence for commercial cryptographic application security assessment based on image recognition, comprising: To evaluate the senior engineer certification submitted by an employee, the Qwen-VL visual language model was used to extract information from multimodal documents. A high-dimensional vector representation was generated for each extracted item, annotated with its source, type, and preliminary quality score, and converted into structured evidence nodes. Evidence nodes are stored in a graph database, and multiple edges are established based on physical location, semantic similarity, and reference relationships to form a preliminary evidence network, assigning initial credibility to the nodes. Certificate nodes are identified in the network, and their related nodes are clustered to form certificate entity clusters, which are assigned comprehensive initial credibility. Through forward tracing, starting from the certificate entity group, searching for supporting evidence along the reference association edge, and adjusting the credibility of the certificate entity group by checking the credibility, consistency, trusted list matching and context relevance of the supporting evidence; Through reverse tracing, check whether there is contradictory evidence pointing to the certificate, low-credibility strong associations, context conflicts, and outdated or replaced clues. If any are found, the credibility of the certificate entity will be reduced; Update the credibility of the certificate entity group based on the forward and reverse traceability results, propagate the credibility change along the reference edge, and affect the credibility of the relevant supporting and contradictory evidence nodes; Traverse all certificate entity groups, filter out those below the credibility threshold and mark them as abnormal, analyze the causes of their low credibility and contradictions; list abnormal certificate information, visualize the key evidence leading to the abnormality, contradictions and their associated paths, and generate text descriptions.
[0006] In conjunction with the first aspect, in a first implementation of the first aspect of this application, the step of evaluating a senior engineer certification certificate submitted by an employee and extracting information from a multimodal document using the Qwen-VL visual language model includes: Read the multimodal document of the submitted senior engineer certification certificate, including text, images, tables and files; use the Qwen-VL visual language model to identify and extract key information in the document as information items, including the name of the certificate holder, certificate number, issue date, name of the issuing agency, certificate level, certificate validity period, signature on the certificate, seal pattern and text and certificate photo.
[0007] In combination with the first aspect, in a second implementation of the first aspect of the present application, generating a high-dimensional vector representation for each extracted item, annotating its source, type, and preliminary quality score, and converting it into a structured evidence node includes: Generate a high-dimensional vector representation for each extracted information item, create a source annotation, record the specific information of its source document, and mark the type of the information item; give a preliminary quality score based on model confidence, content clarity and completeness; and structurally encapsulate the extracted information item content, high-dimensional vector representation, source annotation, type annotation and preliminary quality score and convert them into structured evidence nodes.
[0008] In conjunction with the first aspect, in a third implementation of the first aspect of the present application, the evidence nodes are stored in a graph database, multiple edges are established based on physical location, semantic similarity, and reference relationships to form a preliminary evidence network, and initial credibility is assigned to the nodes, including: Read and parse the information items and attached metadata contained in each encapsulated evidence node. In the graph database, based on the physical location relationship of the information items in the original document, establish edges representing spatial proximity for spatially adjacent or ordered information nodes. These edges are labeled as spatial proximity edges and are assigned labels of adjacent fields. Utilize the high-dimensional vector representation of the information item content and calculate the cosine similarity between vectors to identify information nodes that are highly semantically related or may have a mutually corroborating relationship. Establish edges representing semantic associations for these nodes as semantic similarity edges and assign labels of semantic association or mutual corroboration. Based on the reference or derivation relationship between information items, establish edges representing reference dependency as reference association edges and assign labels of reference relationship or dependency relationship. Connect nodes through analysis of physical location, semantic similarity, and reference relationship and creation of corresponding edges to form a multi-dimensional preliminary evidence network. When assigning initial credibility to a node, the preliminary quality score of the information item is extracted, and the importance of the information item, the authority of the issuing agency, and the connection pattern of the information item in the evidence network are analyzed; comprehensive factors are calculated and assigned a specific initial credibility value for each node in the graph database through a preset weighted calculation method, representing the probability that the information item is recognized as true in the initial stage of evaluation.
[0009] In combination with the first aspect, in a fourth implementation of the first aspect of the present application, identifying certificate nodes in the network, clustering related nodes to form a certificate entity group, and assigning a comprehensive initial credibility includes: In the initial evidence network, the Louvain algorithm is used to identify nodes representing the same certificate entity, and all related nodes directly or indirectly connected to the certificate node are found; these certificate nodes and all related nodes are regarded as a whole, and are aggregated into a certificate entity cluster through the Leiden algorithm; according to the initial credibility values of all nodes in the entity cluster, the connection strength between nodes and the overall structural characteristics of the entity cluster in the network, a random walk-based algorithm is used to calculate and assign a comprehensive initial credibility value for the certificate entity cluster.
[0010] In combination with the first aspect, in a fifth implementation of the first aspect of the present application, the forward tracing starts from the certificate entity group, searches for supporting evidence along the reference association edge, and adjusts the credibility of the certificate entity group by checking the credibility, consistency, trusted list matching, and context relevance of the supporting evidence, including: In the graph database, a Cypher query is used to locate the node representing the Senior Engineer Certification Certificate and its initial attributes. All evidence nodes claiming to support the certificate's validity are found along the referenced edges, and the initial quality scores of the evidence nodes are read. Key field values of the certificate entity group are extracted, and associated evidence nodes are found in the graph database. The corresponding field information is extracted, and the basic consistency between the certificate's internal information and each piece of evidence is evaluated. Check whether the evidence contains a seal and extract its key features; for image evidence, use image processing technology to locate the seal area based on shape, color, texture or position, and crop the main seal image and evidence seal image; for text or digital evidence, check whether the seal and its features are clearly mentioned; when the evidence does not have seal information, skip the comparison; for evidence with successfully extracted seal images, perform multi-dimensional feature comparison; calculate the clarity of the main seal and evidence seal by Laplace operator variance to determine whether they are consistent; perform color quantization and compare the RGB average values of the main pixels to determine whether the colors are consistent; apply the Canny operator for edge detection, analyze the shape features of the contour, and determine whether the shapes are consistent; use optical character recognition technology to extract the seal text, and compare the text content after standardization to determine whether it is exactly the same to determine the text consistency; summarize the comparison results of each evidence in terms of clarity, color, shape and text dimensions to form a seal state set for the evidence and add it to its overall state set; when the seal does not match in any key feature and the evidence presents a seal, a seal mismatch state is generated, and this state is not generated for a complete match or no seal; Compare the consistency of the issuing agency names, extract the full name of the main agency, directly search for text evidence, and use optical character recognition for image evidence; record all relevant agency names found as an evidence agency name set, and standardize these names and the full name of the main agency; traverse the evidence agency name set to determine whether each name is exactly the same as the full name of the main agency; when there is an exact match, determine that it is consistent; when there is an incomplete match, apply the rules to check whether it is an acceptable variant, and determine that it is consistent if any rule matches; when there is no match after traversal, determine that it is inconsistent; add the status description of the issuing agency to the overall status set of the evidence; for each associated evidence node, summarize all status descriptions in its overall status set, evaluate its overall consistency with the certificate entity group, and assign a credibility score accordingly; summarize the evaluation results of all associated evidence, analyze the evidence distribution, and obtain an overall consistency score; Check whether the key field value of the certificate exists in the pre-defined trusted list. If it exists, the trustworthiness of the field is increased. If it does not exist, the trustworthiness is reduced or not adjusted, depending on whether the field is required to be in the list. The trustworthiness check scores of all fields are weighted and summed to obtain the trusted list check score. Evaluate the contextual relevance of the evidence node, query its spatially adjacent nodes and semantically similar nodes to check whether they support the certificate validity or provide background information; evaluate the supportability of each adjacent node, taking into account the node weight; aggregate the supportability scores of all spatially adjacent and semantically similar nodes, and combine the spatial proximity weight and semantic similarity weight to calculate the contextual relevance score; Adjustment factors were calculated using a pre-defined weighting formula based on the initial quality score of the evidence, the consistency score, the credibility list check score, and the contextual relevance score.
[0011] In combination with the first aspect, in a sixth implementation of the first aspect of the present application, the reverse tracing is used to check whether there is contradictory evidence pointing to the certificate, low-credibility strong associations, context conflicts, and outdated and replacement clues, and if any are found, the credibility of the certificate entity group is reduced, including: Query all reference-related edges pointing to the certificate entity group nodes in the graph database, analyze the reverse-related nodes, and look for negative words and information conflicts. If a direct contradiction is found, mark it as high risk and calculate the direct contradiction score; Check the low-trust nodes pointing to the certificate and take their average impact score as the low-trust node impact score; The spatial proximity and semantically similar nodes of the certificate are queried. If there is more than a preset number of negative information or doubts, it is considered that there is a context conflict. The context relevance score obtained by forward tracing is used to calculate the context conflict score. Analyze reverse correlation nodes to find clues that the certificate has been updated, replaced, or invalidated. If found, mark it as high risk and calculate the certificate replacement risk score; The negative adjustment factor is calculated using a preset weighted formula by comprehensively combining the direct contradiction score, low credibility node impact score, context conflict score, and certificate replacement risk score.
[0012] In combination with the first aspect, in a seventh implementation of the first aspect of the present application, the credibility of the certificate entity group is updated according to the forward and reverse tracing results, and the credibility change is propagated along the reference edge to affect the credibility of the relevant supporting and contradictory evidence nodes, including: After forward and reverse tracing, an adjustment factor and a negative adjustment factor are obtained based on the comprehensive evaluation results. The adjustment factors are applied to the credibility of the certificate entity group. After traversing all supporting evidence, a total adjustment is made to obtain the adjusted credibility. The negative adjustment factor is applied to the credibility adjusted by forward tracing to obtain the final credibility. The final credibility is compared with the original credibility before adjustment to calculate the specific change rate of the credibility. Along all reference edges connected to the certificate entity group, the credibility change rate is transmitted to the relevant supporting evidence nodes and contradictory evidence nodes according to the preset weighted propagation rules; For nodes with supporting evidence, the improvement of certificate credibility will correspondingly improve the node's credibility; for nodes with contradictory evidence, the improvement of certificate credibility will reduce the node's credibility; The status of all relevant nodes is updated according to the credibility after propagation, triggering a new round of traceability and evaluation process.
[0013] In conjunction with the first aspect, in an eighth implementation of the first aspect of the present application, traversing all certificate entity groups, screening out those below a credibility threshold and marking them as abnormal, and analyzing the causes and contradictions of their low credibility include: Set a credibility threshold, traverse all identified certificate entity groups in the graph database, and check whether the final credibility of each certificate entity group is lower than the threshold; For certificate entity groups below the threshold, they are marked as abnormal, triggering the abnormality analysis process; During the anomaly analysis process, the forward and reverse traceability results of the certificate entity group are carefully reviewed to determine the specific cause of low credibility. Focus on the contradictions in the abnormal certificate entity group, that is, the conflicting information in the forward and reverse traceability. Record abnormal certificate entities and their analysis results, and trigger corresponding alarms or notification mechanisms to ensure that responsible personnel understand and handle these abnormal situations.
[0014] In conjunction with the first aspect, in a ninth implementation of the first aspect of the present application, the abnormal certificate information is listed, the key evidence, contradictions, and associated paths leading to the abnormality are visually displayed, and a text description is generated, including: Generate a list of abnormal certificates, listing the key information of all certificate entities marked as abnormal; for each abnormal certificate, automatically extract the key evidence and contradictions that lead to its low credibility, including low-credibility supporting evidence, high-credibility negation evidence, and node information with direct conflicts; Utilizing graph database visualization tools, the anomalous certificate, its associated key evidence, contradictions, and the paths between them are graphically displayed. Simultaneously with the visualization, a textual description is generated detailing the causes of the anomalous certificate's low credibility, the specific content of the key evidence, the manifestations of the contradictions, and the logic of their associations. This helps responsible personnel understand the anomaly and provides a reference for further investigation or decision-making. Combine abnormal certificate lists, visual displays, and text descriptions into a unified report for viewing and exporting.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention converts the extracted information into nodes and stores them in a graph database. Edges are established based on physical location, semantic similarity, and reference relationships to form a preliminary evidence association network. The relevant nodes are clustered into entity groups representing specific certificates, and their overall credibility is evaluated.
[0016] 2. The present invention performs an information consistency check between the certificate and the evidence, checking whether the seal in the evidence is consistent with the seal on the certificate in terms of clarity, color, shape and text content; checking whether the organization name in the evidence matches the organization name on the certificate, and making judgments based on different rules.
[0017] 3. The present invention dynamically adjusts the credibility of the certificate entity group by searching for supporting evidence through forward tracing and searching for contradictory evidence through reverse tracing. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A schematic diagram of the steps of the commercial cryptographic application security assessment evidence identification method based on image recognition according to the present invention; Figure 2 This is a schematic diagram of the seal image consistency detection steps of the commercial cryptographic application security assessment evidence recognition method based on image recognition of the present invention. DETAILED DESCRIPTION
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0020] Example: Figure 1-Figure 2 As shown, the present invention provides a technical solution. like Figure 1 As shown in the schematic diagram of the steps of the commercial encryption application security assessment evidence identification method based on image recognition, the present invention provides a commercial encryption application security assessment evidence identification method based on image recognition, comprising: Step S100: To evaluate the senior engineer certification certificate submitted by the employee, the Qwen-VL visual language model is used to extract information from the multimodal document; a high-dimensional vector representation is generated for each extracted item, and its source, type, and preliminary quality score are annotated, and it is converted into a structured evidence node; Specifically, it receives documents submitted by employees containing senior engineer certifications, including text, images, tables, and files. It uses optical character recognition technology to extract all visible text content, obtain all recognized text blocks and their position coordinates in the image, and directly extract the text content of native PDF or text files. It integrates text from different sources to form a collection containing all visible text. Analyze certificate images, including those of senior engineer certificates, communication protocols, communication passwords, product certificates, contracts, and information security filing certificates. Use image processing techniques, including edge detection, template matching, and color recognition, to locate key image areas, identify the specific location of signatures or stroke areas with handwriting characteristics; identify seal areas with specific shapes, colors, and internal patterns or text features; identify image areas with portrait features, usually located in the upper right corner of the certificate or a specified location; accurately crop the located image elements from the original image, save them as separate image files, and record their source document and location information; The prepared multimodal data is used as input, and the API or local interface of the Qwen-VL model is called to identify and extract key information items, including the name of the certificate holder, certificate number, issuance date, name of the issuing authority, certificate level, certificate validity period, signature information, seal pattern and text, and certificate photo. Each extracted information item is returned in the form of a key-value pair. Analyze input text and images, search for patterns or keywords matching names, numbers, dates, organization names, grades, and expiration dates in the text, and use contextual understanding to confirm the accuracy of the information. For signature images, recognize the signature text. For seal images, recognize the shape and color of the seal and use optical character recognition technology to attempt to read the text within the seal. For certificate photos, determine basic features such as gender and approximate age range. Combine text and image information for cross-validation. Output a structured information dictionary or list containing all identified key information items and their values. Perform simple format and logic checks on the output of Qwen-VL, including text information, bounding box coordinates, confidence scores, and image fragments, and organize each information item recognized by the model into a structured intermediate format along with its corresponding value, location information, and confidence score given by the model; Select the CLIP embedding model. For text information items, pass its fields as input to the embedding model to generate a fixed-length vector. For image information items, pass the extracted image region data to the embedding model to generate the corresponding image vector. Associate the generated vector with the corresponding information item and store it in the information item's data structure. Record the original document from which the information item came, its specific location in the document, the source type, and the type of information item extracted; Use the confidence scores given by Qwen-VL when extracting information to assess the image quality of information items from image sources; use image processing techniques to quantify clarity, and for text, check for garbled or blurred characters; check whether the extracted information item values are complete; and design a weighted formula to combine the above factors to calculate a comprehensive preliminary quality score. Design a standard evidence node data structure. For each extracted information item, use all relevant information to create a specific evidence node instance according to the designed node data structure. Store the generated structured evidence node in a temporary data structure for use in subsequent steps. To reduce the potential risks of Qwen-VL in this invention, the identified password-related fields are partially masked or transmitted encrypted to enhance data security; Qwen-VL is trained a second time based on commercial cryptographic terminology to simulate massive certificate and adversarial sample inputs to improve recognition accuracy; in the input preprocessing stage, digital watermark detection or noise analysis is used to identify PS traces and detect image tampering; at the same time, a manual intervention channel is retained for low-credibility results.
[0021] In a specific embodiment, an employee submitted a PDF file containing a senior engineer certification certificate. First, optical character recognition technology was used to process the PDF to extract text content, including name: Zhang Wei, certificate number: ENG2025068901, issue date: 2024-12-15, issuing agency: China Engineers Federation, certificate level: senior engineer, valid until: 2029-12-14, and other text blocks, and the page number and coordinate position of these texts in the PDF were recorded. Then, image processing technology was used to locate and crop the signature image of Zhang Wei in the upper right corner of the first page image of the PDF and save it as Signature_GW.png, record the source page number 1, coordinate range x:100-200, y:50-80, and the circular red seal image in the center was saved as Seal_2025.png, record the source page number 1, coordinate range x:300-500, y:400-600. The Qwen-VL model is called, taking the integrated text collection and the cropped image as input. Qwen-VL successfully extracts key information items, including the name Zhang Wei, certificate number ENG2025068901, issue date 2024-12-15, issuing authority China Engineers Federation, certificate level Senior Engineer, expiration date 2029-12-14, signature information, seal text, and certificate photo. It then returns a dictionary containing these key-value pairs, along with a model confidence score for each item: 0.98 for the name, 0.99 for the certificate number, and 0.85 for the seal text. The Qwen-VL output is formatted and confirmed to be correct. The CLIP model is then selected to generate a 768-dimensional vector for the text Zhang Wei, a 768-dimensional vector for the image Signature_GW.png, and a 768-dimensional vector for the image Seal_2025.png. These vectors are then associated with the corresponding information items. The source of each information item is recorded as Cert_2025-001.pdf, the source type is PDF, and the information item types are name, certificate number, and issue date. Based on the Qwen-VL confidence of 0.98, the signature image clarity of 0.95, the seal image clarity of 0.90, and the completeness of the information, a comprehensive preliminary quality score is calculated, with the name scoring 0.96, the certificate number 0.97, and the seal text 0.87. According to the designed evidence node data structure, a specific evidence node instance is created for each information item, such as Name: Zhang Wei, Certificate Number: ENG2025068901, containing the information item content, high-dimensional vector, source annotation, type annotation, and preliminary quality score. These structured evidence nodes are then stored in a temporary data structure.
[0022] Step S200: Evidence nodes are stored in a graph database. Multiple edges are established based on physical location, semantic similarity, and reference relationships to form a preliminary evidence network, and initial credibility is assigned to the nodes. Certificate nodes are identified in the network, and their related nodes are clustered to form certificate entity clusters, which are assigned comprehensive initial credibility. Specifically, all generated encapsulated evidence nodes are read, each of which contains extracted information items, high-dimensional vector representations, source annotations, type annotations, and preliminary quality scores; Analyze the physical layout of information items in the original multimodal document and create a spatial proximity edge in the graph database for those information nodes that are spatially adjacent or arranged in a logical order. This edge is given a proximity or similar label to indicate their proximity relationship in the document space. Utilizing the high-dimensional vector representation of each information item, the cosine similarity algorithm is used to calculate the similarity between these vectors. The cosine value of the angle between two vectors is calculated. If the two vectors are in exactly the same direction, the angle is 0 degrees and the cosine value is 1, indicating that they are most similar. If the directions are completely opposite, the angle is 180 degrees and the cosine value is -1, indicating that they are least similar. If the directions are perpendicular, the angle is 90 degrees and the cosine value is 0, indicating that they have no similarity in direction. Information nodes with highly relevant or mutually corroborating content are identified. When high-dimensional vectors show high similarity, semantic similarity edges are established for these nodes, and semantic association or mutual corroboration labels are assigned. Analyze whether there is a clear reference or derivative relationship between information items. Use natural language processing and knowledge graph technology to perform word segmentation, part-of-speech tagging, named entity recognition, etc. for information items in text form; use a pre-trained bidirectional encoder representation language model, input a sentence or paragraph containing two information items and their context, and the model learns and predicts whether there is a specific relationship between them; output the probability of the relationship existing, set a threshold, and determine that a relationship exists if the threshold is exceeded; for information items containing images, combine image recognition results with text information to determine the relationship; for example, identify the unit name on the seal, and then find whether there is a text information item that mentions the unit to determine whether there is a reference or attribution relationship; establish reference association edges for nodes with reference or dependency relationships, and assign reference relationship or dependency labels; Through the above analysis based on physical location, semantic similarity, and reference relationships, and the creation of corresponding edges, the originally isolated evidence nodes are connected, forming a preliminary multi-dimensional evidence network containing various edge types. Calculate an initial credibility for each node in the graph database. This is done by factoring in model confidence, content clarity and completeness, the fact that some information items are more critical to the authenticity of the certificate than others, identifying the issuing authority, and assigning higher initial credibility to items from authoritative organizations. The number and type of connections a node has in the network may also affect its credibility. Using a pre-set weighted calculation formula, we combine these factors to calculate a specific initial credibility value for each node in the graph database. This represents the probability that the information item will be considered authentic at the initial stage of the evaluation, based on its own quality and network position. In the preliminary evidence network, identify the node representing the core certificate information, and use the Louvain algorithm to identify all related nodes directly or indirectly connected to the certificate node. These related nodes include other information items on the certificate and information items in the associated evidence; The identified certificate node and all its directly or indirectly connected related nodes are considered as a whole, and the more precise Leiden algorithm is further used to aggregate these nodes into a certificate entity group; this entity group represents the collection of all evidence information related to the specific certificate; A comprehensive initial credibility value is calculated for this newly formed certificate entity group, including the distribution and average value of the initial credibility values of all nodes in the entity group, the type and number of connections between nodes within the entity group, and the overall structure of the entity group in the network, such as density, and connections with other external nodes. A random walk-based algorithm is used to simulate the transmission and mutual influence of information between nodes within the entity group. Each node in the graph database is regarded as a state in the random walk, and the edges in the graph are regarded as the possibility of state transfer. The initial credibility of each node is used as the initial value of the personalized vector, and random walk simulation and iterative calculation are performed; after multiple iterations, when the PageRank value converges, each node will obtain a final PageRank value, which comprehensively evaluates the credibility of the entire entity group, and finally assigns the certificate entity group a comprehensive initial credibility value, and quantitatively evaluates the overall preliminary authenticity of the certificate.
[0023] In one specific embodiment, all evidence nodes generated from the employee-submitted certificate Cert2025-001.pdf are read, including the name Zhang Wei (node ID: E1, vector similarity 0.98, quality score 0.96), the certificate number ENG2025068901 (node ID: E2, vector similarity 0.99, quality score 0.97), and the seal text "China Engineers Federation" (node ID: E3, vector similarity 0.85, quality score 0.87). Analysis of the physical layout reveals that the coordinates of E1's name Zhang Wei and E2's certificate number ENG2025068901 on the first page of the PDF are similar: E1: x: 50-150, y: 50-100; E2: x: 200-300, y: 50-100. Therefore, a spatially adjacent edge is created for E1 and E2 in the graph database, labeled "adjacent." Using CLIP vectors, the cosine similarity between E1's name, Zhang Wei, and E3's seal text, China Engineers Federation, was calculated to be 0.12. However, the cosine similarity between E2's certificate number and E3's seal institution name was 0.45, falling short of the pre-defined semantic similarity threshold of 0.6. Therefore, no semantic similarity edge was established. Analyzing citation relationships, a bidirectional encoder representation model analyzed the text context and found a high-probability citation relationship between E3, China Engineers Federation, and the text issuing institution, China Engineers Federation, with a quality score of 0.95. The probability was 0.92, exceeding the threshold of 0.8. Therefore, a citation edge was established between E3 and E4, labeled "citation relationship." After these steps, nodes E1, E2, E3, and E4 were connected, forming a preliminary evidence network. Initial credibility was calculated for each node: E1's name scored 0.96, E2's certificate number scored 0.97, E3's seal text scored 0.87, and E4's issuing institution text scored 0.95. Considering the critical importance of the certificate number to authenticity, E2 was given a higher weight, with a weight coefficient of 1.2. The final initial credibility of E2 was adjusted to 0.97 × 1.2 = 1.164, which is normalized to 0.98. Core certificate nodes, such as E2's certificate number, were identified, and the Louvain algorithm was used to find its directly connected nodes, such as E1, E3, and E4. These nodes were aggregated using the Leiden algorithm to form a certificate entity cluster C1. The comprehensive initial credibility of C1 was calculated, and the average initial credibility of the nodes within the cluster was 0.96 + 0.98 + 0.87 + 0.95) ÷ 4 = 0.935. There was one spatially adjacent edge and one citation-related edge between nodes, and the intra-cluster connectivity density was moderate. All nodes within the cluster were derived from the original certificate document. Using a random walk-based algorithm, which takes into account the mutual influence between nodes, the final comprehensive initial credibility value assigned to C1 was 0.92, indicating a high probability of authenticity during the initial evaluation phase.
[0024] Step S300: Starting from the certificate entity group, searching for supporting evidence along the reference association edge through forward tracing, and adjusting the credibility of the certificate entity group by checking the credibility, consistency, trusted list matching and context relevance of the supporting evidence; Specifically, in the graph database, a Cypher query is used to find the node representing the Senior Engineer Certification Certificate. Then, all evidence nodes claiming to support the validity of the certificate are traversed along the referenced associated edge types to find the initial quality score assigned to each evidence node during the initial extraction phase. In the graph database, find the evidence nodes directly or indirectly associated with the certificate and extract information related to key fields; evaluate the basic consistency between the certificate's internal information and the certificate and the single evidence; Check whether the evidence contains seal information and extract its key features; use image processing technology to locate the area on the certificate image that is suspected to be a seal by looking for a specific shape, specific color, specific texture or location; crop the located area to obtain a sub-image containing the seal image, which is recorded as the main seal image; for each associated evidence, if it exists in the form of an image, repeat the above steps to locate and crop the seal image on the evidence, and record it as the evidence seal image; when the associated evidence is a text description or digital record, check whether the seal information is explicitly mentioned. If so, record the key features of the description, including color, shape and the mentioned text; if there is no mention or the evidence itself does not contain an image, it is considered that the evidence does not provide seal information and the seal comparison step is skipped; For each piece of evidence whose evidence seal image is successfully extracted, feature dimensions are extracted and compared; clarity indexes are calculated for both the main seal image and the evidence seal image, and the Laplace operator variance is used to quantify the edge sharpness of the images. The calculated values are recorded as the main seal clarity value and the evidence seal clarity value, respectively; when the evidence seal clarity value is equal to the main seal clarity value, the status is described as clarity consistency; when they are not equal, the status is described as clarity inconsistency; Color quantization is performed on the main color areas of the main seal image and the evidence seal image, and the RGB average values of the main pixels in the image are calculated and recorded as the main seal color and the evidence seal color respectively. When the RGB value of the evidence seal color is exactly equal to the RGB value of the main seal color, the status is described as color consistency; when they are not equal, the status is described as color inconsistency. The Canny operator is used to perform edge detection on the main seal image and the evidence seal image. The detected edge contours are subjected to geometric shape analysis to calculate the minimum circumscribed rectangle, circularity, and ellipticity parameters of the contours. The main shape features obtained are recorded as the main seal shape feature and the evidence seal shape feature respectively. When the shape described by the evidence seal shape feature is exactly the same as the main seal shape feature, the state is described as shape consistency; when the shape feature descriptions are different, the state is described as shape inconsistency. Apply optical character recognition technology to the main seal image and the evidence seal image to extract the text content within the seal area; the extracted text content is recorded as the main seal text content and the evidence seal text content respectively; the extracted text is standardized by removing leading and trailing spaces, standardizing punctuation, and converting to uppercase or lowercase; check whether the standardized evidence seal text content is exactly the same as the standardized main seal text content; if they are exactly the same, the status description is generated as consistent text content; if they are not exactly the same, the status description is generated as inconsistent text content; For each piece of evidence, all generated status descriptions are summarized to form the seal status set of the evidence; all status descriptions in the seal status set are added to the overall status set generated by the evidence; ultimately, the status information of the evidence in all inspection dimensions is included; When the seal information in the evidence does not match the certificate seal in any of the above key features, and the seal information is presented in the evidence, the status description Seal Mismatch is generated; when the seal information is completely matched, or there is no seal information in the evidence, this status is not generated; Compare the name of the issuing authority displayed on the physical body of the certificate with the issuing authority information mentioned or reflected in the associated evidence to see if they are consistent, and handle any name discrepancies, abbreviations, and full names; Extract the full name of the issuing organization from the certificate image or text and record it as the full name of the principal organization. If the evidence is text, search the text for information related to the issuing organization, including the full name, common abbreviations, English name, and any name that varies in different documents. In the case of image evidence, optical character recognition technology is used to identify the text in the image and find the name associated with the issuing authority; Record all relevant institution names or logos found in the evidence to form a set, which is recorded as the evidence institution name set; Standardize the full name of the principal organization and each name in the set of evidence organization names, including unifying character case, removing spaces before and after the names, and removing unnecessary punctuation in the names; Traverse the evidence institution name set, and for each current evidence name in the set, determine whether it is exactly the same as the full name of the principal institution; if there is a current evidence name that is exactly the same as the full name of the principal institution, determine that the issuing institution information of the evidence is consistent with the certificate; Apply rule-based definitions to check whether it is an acceptable variant and perform rule matching based on known information; check whether it is a recognized abbreviation and determine whether the current evidence name is a widely recognized abbreviation of the full name of the principal organization that is used in formal occasions; if it matches, it is determined to be consistent; check whether it corresponds to the English name and determine whether the current evidence name is the official English name of the full name of the principal organization; determine whether the current evidence name has added or removed certain prefixes or suffixes based on the full name of the principal organization, and the core subject name has not changed, and define rules to determine whether such changes are reasonable and do not affect subject identification; if it is determined to be a reasonable variant through the rules, it can be determined to be consistent; check whether there is a known historical name and determine whether the current evidence name is the official name used by the institution at the time of certificate issuance or earlier, and perform rule matching based on historical information; After traversing the set of evidence institution names, if any of the above methods are used to determine consistency, the status is described as consistent issuing institutions; after traversing all names, if none of the consistency judgment rules are passed, the status is described as inconsistent issuing institutions; The generated status description of the issuing authority is added to the overall status set of the evidence. For each associated evidence node, all status descriptions in its overall status set are collected. Based on the summarized status descriptions, the overall consistency of the evidence node with the certificate entity group is evaluated. Based on the overall status of the evidence and the specific inconsistencies in the status set, a credibility score is assigned to each evidence node. The evaluation results of all associated evidence nodes are summarized. The distribution of evidence is analyzed, and a consistency score is obtained based on the aggregated information of all evidence. A trusted list is predefined for the key fields of the certificate entity group. Each key field value of the certificate entity group is queried to check whether the value exists in the corresponding trusted list. If it exists, the field is considered to have passed the verification and the preset credibility is increased. If it does not exist, it is determined whether the field is required to exist in the list. If it is required, the preset credibility is reduced. The credibility check scores of all key fields are weighted and summed according to their weights to obtain the trusted list check score. Query other nodes that have spatially adjacent edges to the evidence node to check whether these adjacent nodes support the validity of the certificate or contain relevant background information; query other nodes that have semantically similar edges to the evidence node to check whether these nodes point to supporting information; combine the information of spatially and semantically adjacent nodes to determine whether the evidence node is in a supportive context and calculate context relevance; Query nodes that have spatially adjacent edges to the current evidence node. These edges represent physical proximity or structural associations, and check whether these nodes contain information supporting the validity of the certificate or provide relevant background information. Query nodes that have semantically similar edges with the current evidence node. These edges indicate that the node content is similar in topic, concept, or referent, and check whether these nodes point to supporting information. For each identified neighboring node, evaluate whether it supports, is irrelevant to, or contradicts the certificate entity group; set a binary judgment, with support as 1 and irrelevant or contradictory as 0; consider the weight or credibility of the neighboring nodes; the support information provided by highly credible neighboring nodes is more important than that provided by low-credible nodes; Aggregate the support scores of all spatially adjacent nodes and semantically similar nodes, consider the node weights, and use a weighted average algorithm to derive the spatial proximity aggregation score and semantic similarity aggregation score. Combining the aggregation results of the two dimensions of spatial proximity and semantic similarity, the context relevance score CF is calculated. The formula is: ; Among them, CF is the context relevance score; SpS is the spatial proximity aggregation score; SeS is the semantic similarity aggregation score; w s is the weight of the spatial proximity score, w c is the weight of the semantic similarity score; Based on the above inspection results, the adjustment factor μ is calculated using the preset weighted formula, which is: ; Among them, μ is the adjustment factor, BF is the initial quality score of the extracted items, CS is the consistency score, LCF is the trust list check score, CF is the context relevance score, and w1, w2, w3, and w4 are the corresponding weights respectively.
[0025] In a specific embodiment, the entity group C1 representing the senior engineer certificate of employee Zhang Wei is found in the graph database through Cypher query, and its core node is the certificate number ENG2025068901, node ID: E2. Traversing along the referenced associated edge, a path is found to the associated evidence node E5, which is a scanned image of the same certificate submitted by Zhang Wei, and the initial quality score of E5 is read as 0.92. The consistency of the internal information of the certificate is evaluated, and it is found that the node information such as E1 name Zhang Wei, E2 certificate number, E3 seal text, and E4 issuing agency text are basically consistent. Checking the seal information, in the certificate entity group C1, the E3 node contains the seal text China Engineers Federation and the seal image. In the associated evidence E5, image processing technology is used to locate and crop the seal area to obtain the evidence seal image. The main seal image E3 and the evidence seal image were compared for feature clarity. The main seal clarity value was 120, while the evidence seal clarity value was 118. The status description was "Inconsistent clarity." The main color RGB average was calculated. The main seal color was (200, 0, 0), and the evidence seal color was (200, 0, 0). The status description was "Consistent color." The Canny operator was used for edge detection and shape analysis. The main seal shape feature was circular, and the evidence seal shape feature was also circular. The status description was "Consistent shape." Optical character recognition was used to extract text. The main seal text content was "China Engineers Federation" and the evidence seal text content was "China Engineers Federation." After standardization, the text content was identical. The status description was "Consistent text content." The seal status of E5 was summarized as inconsistent clarity, consistent color, consistent shape, and consistent text content. No seal mismatch status was generated. The issuing agency was compared. The full name of the main agency was "China Engineers Federation." The agency name extracted from evidence E5 was also "China Engineers Federation." After standardization, the name was identical. The status description was "Consistent issuing agency." The seal and institution status were added to the overall status set for E5. The overall consistency of E5 with C1 was assessed. While there were minor discrepancies in clarity, other key features were consistent, resulting in a calculated consistency score (CS) of 0.95. A query of spatially adjacent and semantically similar nodes related to E5 revealed that these nodes all supported certificate validity, with a spatial proximity aggregation score (SpS) of 0.9 and a semantic similarity aggregation score (SeS) of 0.85. A contextual relevance score (CF) was calculated as (0.9 × 0.85) ÷ 2 = 0.875. Key fields, such as the certificate number ENG2025068901, were checked to see if they were included in the pre-set trusted list. The number was not in the list, but the field's presence in the list was not mandatory, resulting in a trusted list check score (LCF) of 0.0. Based on E5's initial quality score (BF) of 0.92, consistency score (CS) of 0.95, trusted list score (LCF) of 0.0, and contextual relevance score (CF) of 0.875, the pre-set weighting formula yielded a calculated μ of 0.8265.
[0026] Step S400: Through reverse tracing, check whether there is contradictory evidence pointing to the certificate, low-credibility strong associations, context conflicts, and outdated and replaced clues. If any are found, the credibility of the certificate entity group is reduced; the credibility of the certificate entity group is updated based on the forward and reverse tracing results, and the credibility change is propagated along the reference edge to affect the credibility of related supporting and contradictory evidence nodes; Specifically, a certificate entity group node is obtained from the graph database as a starting point, and all reference-related edges pointing to the current certificate entity group node are queried; Analyze the content of the nodes found from the reverse reference edge, looking for words or sentences that negate the validity of the certificate, and check whether the information extracted from the node directly conflicts with the key information of the certificate entity group. If a direct contradiction is found, it is marked as high risk and the direct contradiction score DCS is recorded as 1, otherwise it is 0; Check the preliminary quality scores or credibility of all evidence nodes pointing to the certificate entity group; if multiple low-credibility nodes are found to be strongly associated with the certificate entity group through reference edges, take their average influence score as the low-credibility node influence score; if there are no low-credibility nodes, the low-credibility node influence score is 0; Query nodes that have spatially adjacent edges to the certificate entity group to check whether these nodes contain information that contradicts the validity of the certificate; query nodes that have semantically similar edges to the certificate entity group to check whether these nodes generally reflect negative situations or doubts; if the certificate entity group is in a context of negative information or doubts, it is considered that there is a context conflict, and the context relevance score CF obtained by forward tracing is calculated using the formula: ; Among them, CCS is the context conflict score, CF is the context relevance score, and ConflictRatio is the proportion of negative nodes; if there is no conflict, CCS is 0; Analyze the content of the nodes found by the reverse reference edge, looking for explicit or indirect clues that the certificate has been updated, replaced, or is no longer in use; check whether there is a new certificate entity group node with evidence that it has replaced the current one. When such clues are found, mark it as high risk and the certificate replacement risk score ORS is 1, otherwise it is 0; Based on the inspection results, determine whether there are factors that require a reduction in confidence; calculate the negative adjustment factor (NAF) using the preset weighting formula: ; Among them, DCS is the direct contradiction score; LNI is the low credibility node impact score; CCS is the context conflict score; ORS is the certificate replacement risk score; w5, w6, w7, and w8 are the corresponding weights respectively; Apply the adjustment factor to the current credibility of the certificate entity group, traverse all the supporting evidence found, and make a total adjustment to obtain the adjusted credibility T1; apply the negative adjustment factor to the adjusted credibility T1 of the certificate entity group to obtain the final credibility T2, and make a decision based on the final credibility.
[0027] In one specific example, a certificate entity cluster C1, representing Zhang Wei's Senior Engineer Certificate, was retrieved from a graph database. A query of all reference-related edges pointing to C1 revealed a link to the associated evidence node E6, with an initial quality score of 0.45. This node, representing another document, claimed that the certificate submitted by Zhang Wei was forged. The content of E6 was first analyzed for negative terms. It was found to contain explicit terms such as "forged" and "invalid," directly conflicting with the core information of C1. Therefore, a direct contradiction score (DCS) of 1 was assigned. An examination of the quality of the evidence nodes pointing to C1 revealed that only E6 had a low confidence score, with a score of 0.45. Therefore, a low confidence node influence score (LNI) of 0.45 was assigned. A query of spatially adjacent nodes to C1 revealed no negative information. A query of semantically similar nodes also revealed no widespread negative information, resulting in a context conflict score (CCS) of 0. E6's content contained no clues that the certificate had been updated, replaced, or outdated, resulting in an ORS of 0. Based on DCS = 1, LNI = 0.45, CCS = 0, and ORS = 0, the negative adjustment factor (NAF) is calculated using a pre-set weighting formula to be 0.635. Assume that the credibility of C1 after forward traceability is 0.76, which is used as the current credibility. Applying the negative adjustment factor of 0.635 to the current credibility of 0.76 yields a final credibility of approximately 0.28, a significant decrease. This indicates that the direct contradictory evidence E6 discovered through reverse traceability has cast serious doubt on the authenticity of certificate C1.
[0028] Step S500: Traverse all certificate entity groups, filter out those below the credibility threshold and mark them as abnormal, analyze the causes of their low credibility and contradictions; list abnormal certificate information, visualize the key evidence leading to the abnormality, contradictions and their associated paths, and generate text descriptions.
[0029] Specifically, a credibility threshold is set, representing the minimum acceptable credibility level. A traversal process is initiated to access all identified certificate entity groups stored in the graph database. For each certificate entity group, its adjusted final credibility value is read. If the final credibility value of a certificate entity group is lower than the preset threshold, the status of the entity group is marked as abnormal, triggering the subsequent abnormality analysis process. Once a certificate entity group is marked as abnormal, it will automatically enter the abnormality analysis process, and conduct a detailed inspection of the supporting evidence collected during the forward tracing process of the abnormal certificate entity group and its contribution to the credibility. It will also conduct a detailed inspection of the evidence found during the reverse tracing process that does not support or even negates the certificate and its weakening effect on the credibility. By comparing the results of the forward and reverse tracing processes, the specific reasons that led to the low credibility of the certificate entity group are analyzed, including the lack of sufficient high-credibility supporting evidence, the presence of multiple high-credibility negating or contradictory evidence, a broken chain of evidence, the lack of key information, and irreconcilable conflicts between evidence. Pay attention to and clearly identify direct conflicts or contradictions found in the forward and backward tracing results; for example, an evidence node is considered supportive in the forward tracing, but is found to contradict another high-confidence evidence in the backward tracing; All certificate entity group information marked as abnormal, along with detailed analysis results such as its final credibility value, analysis of the causes of low credibility, and identified contradictions, are uniformly recorded in a dedicated abnormality log or database table as a basis for audit tracking and processing. Based on preset rules or configurations, the corresponding alarm or notification mechanism is automatically triggered, including sending emails, push messages, and displaying alarm icons on the management interface to notify relevant personnel responsible for reviewing or handling abnormal situations, ensuring that they are promptly aware of problematic certificates that require attention. Generate a clear list of abnormal certificates. This list should include key information for each abnormal certificate entity group, including the certificate ID, associated employee information, final credibility value, and abnormality marking time. For each abnormal certificate in the list, extract the specific information that caused it to be marked as abnormal from the graph database, identify those evidence nodes with low credibility values that have a negative impact on overall credibility, clearly list the node information and conflicting content of direct conflicts or contradictions, extract the connection paths between these key evidence, contradictions, and the certificate entity group nodes, and display their specific connection relationships in the graph network. Using the visualization tools or integrated visualization engines provided by the graph database, the abnormal certificate entity clusters, extracted key evidence nodes, conflicting nodes, and the associated paths between them can be intuitively displayed graphically. For example, different colors or shapes can be used to distinguish normal nodes, abnormal nodes, supporting evidence, negating evidence, conflicting points, etc., and their relationships can be represented by connecting lines. While visualizing the results, a detailed text description is generated, summarizing the causes of low credibility, extracting key evidence, extracting conflicting information, and briefly explaining the association logic. This allows responsible personnel to quickly and accurately understand the core issues of the anomaly even without delving into the graph database. The generated abnormal certificate list, visual charts and corresponding text descriptions are integrated into a structured unified report, and an export function is provided for use by responsible personnel, management or archiving.
[0030] In one specific embodiment, the credibility threshold is set at 0.7. After initiating the traversal, the certificate entity group C1 is accessed, representing Zhang Wei's Senior Engineer Certificate. Its final credibility, adjusted through steps S300 and S400, is 0.28, below the threshold of 0.7. C1's status is immediately marked as abnormal, and the anomaly analysis process is triggered. The analysis reveals that while scanned image evidence E5, with a quality score of 0.92, was collected during the forward traceback and contributed to some credibility, the direct contradictory evidence E6 discovered during the reverse traceback, claiming the certificate was forged, with a quality score of 0.45 and a direct contradiction score of 1, and the resulting negative adjustment factor of 0.635, is the primary cause of the significant decrease in credibility. The cause of the low credibility is summarized as the presence of high-impact direct negation evidence. The contradiction is clearly identified: the forged character in E6 directly conflicts with the core information in C1. This information is recorded in the anomaly log and triggers an email alert to Li Ming, the head of the certification review team. The generated abnormal certificate list includes the following entries: Certificate ID: ENG2025068901, Associated Employee: Zhang Wei, Final Credibility: 0.28, and Anomaly Flagging Time: 2025-07-07 14:30. For C1, key evidence node E6 was extracted, with a quality of 0.45. Its content contains forgeries, and the contradiction between E6 and the core information of C1. The association path is that E6 points to C1 through a reference edge. In the visualization, the C1 node is marked red to indicate an anomaly, and the E6 node is marked yellow to indicate negative evidence. The reference edge between them is displayed in bold. The text description generated at the same time is: "Certificate ENG2025068901, Zhang Wei's final credibility is 0.28, which is lower than the threshold of 0.7 and is marked as an anomaly. The key evidence is node E6, the content of which claims that the certificate is forged, with a quality score of 0.45. The main contradiction is that the content of E6 directly conflicts with the core information of the certificate. The association path is that E6 points to the certificate entity group C1 through a reference association edge. The anomaly has been notified to Li Ming for processing." Finally, this unified report containing lists, charts and descriptions is generated and can be exported.
[0031] like Figure 2 As shown in the schematic diagram of the seal image consistency detection step of the commercial encryption application security assessment evidence recognition method based on image recognition, the present invention provides a commercial encryption application security assessment evidence recognition method based on image recognition, comprising: Check the consistency of information between the certificate and the evidence, and check whether the seal in the evidence is consistent with the seal on the certificate in terms of clarity, color, shape and text content. The specific steps are as follows: In the graph database, use Cypher queries to find the evidence nodes directly or indirectly associated with the certificate and extract the content related to the seal information. Check whether the evidence contains seal information and extract its key features; use image processing technology to locate the area on the certificate image that is suspected to be a seal by looking for a specific shape, specific color, specific texture or location; crop the located area to obtain a sub-image containing the seal image, which is recorded as the main seal image; for each associated evidence, if it exists in the form of an image, repeat the above steps to locate and crop the seal image on the evidence, and record it as the evidence seal image; when the associated evidence is a text description or digital record, check whether the seal information is explicitly mentioned. If so, record the key features of the description, including color, shape and the mentioned text; if there is no mention or the evidence itself does not contain an image, it is considered that the evidence does not provide seal information and the seal comparison step is skipped; For each piece of evidence whose evidence seal image is successfully extracted, feature dimensions are extracted and compared; clarity indexes are calculated for both the main seal image and the evidence seal image, and the Laplace operator variance is used to quantify the edge sharpness of the images. The calculated values are recorded as the main seal clarity value and the evidence seal clarity value, respectively; when the evidence seal clarity value is equal to the main seal clarity value, the status is described as clarity consistency; when they are not equal, the status is described as clarity inconsistency; Color quantization is performed on the main color areas of the main seal image and the evidence seal image, and the RGB average values of the main pixels in the image are calculated and recorded as the main seal color and the evidence seal color respectively. When the RGB value of the evidence seal color is exactly equal to the RGB value of the main seal color, the status is described as color consistency; when they are not equal, the status is described as color inconsistency. The Canny operator is used to perform edge detection on the main seal image and the evidence seal image. The detected edge contours are subjected to geometric shape analysis to calculate the minimum circumscribed rectangle, circularity, and ellipticity parameters of the contours. The main shape features obtained are recorded as the main seal shape feature and the evidence seal shape feature respectively. When the shape described by the evidence seal shape feature is exactly the same as the main seal shape feature, the state is described as shape consistency; when the shape feature descriptions are different, the state is described as shape inconsistency. Apply optical character recognition technology to the main seal image and the evidence seal image to extract the text content within the seal area; the extracted text content is recorded as the main seal text content and the evidence seal text content respectively; the extracted text is standardized by removing leading and trailing spaces, standardizing punctuation, and converting to uppercase or lowercase; check whether the standardized evidence seal text content is exactly the same as the standardized main seal text content; if they are exactly the same, the status description is generated as consistent text content; if they are not exactly the same, the status description is generated as inconsistent text content; For each piece of evidence, all generated status descriptions are summarized to form the seal status set of the evidence; all status descriptions in the seal status set are added to the overall status set generated by the evidence; ultimately, the status information of the evidence in all inspection dimensions is included; When the seal information in the evidence does not match the certificate seal in any of the above key features, and the seal information is presented in the evidence, the status description seal mismatch is generated; when the seal information matches completely, or there is no seal information in the evidence, this status is not generated.
[0032] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
Claims
1. A commercial cryptographic application security assessment evidence identification method based on image recognition, characterized in that: include: To evaluate the senior engineer certification submitted by employees, the Qwen-VL visual language model was used to extract information from multimodal documents. Generate a high-dimensional vector representation for each extracted item, annotate its source, type and preliminary quality score, and transform it into a structured evidence node; Evidence nodes are stored in a graph database, and multiple edges are established based on physical location, semantic similarity, and reference relationships to form a preliminary evidence network and assign initial credibility to the nodes. Identify certificate nodes in the network, cluster related nodes to form certificate entity groups, and assign comprehensive initial credibility; Through forward tracing, starting from the certificate entity group, searching for supporting evidence along the reference association edge, and adjusting the credibility of the certificate entity group by checking the credibility, consistency, trusted list matching and context relevance of the supporting evidence; Through reverse tracing, check whether there is contradictory evidence pointing to the certificate, low-credibility strong associations, context conflicts, and outdated or replaced clues. If any are found, the credibility of the certificate entity will be reduced; Update the credibility of the certificate entity group based on the forward and reverse traceability results, propagate the credibility change along the reference edge, and affect the credibility of the relevant supporting and contradictory evidence nodes; Traverse all certificate entity groups, filter out those below the credibility threshold and mark them as abnormal, analyze the causes of their low credibility and contradictions; list abnormal certificate information, visualize the key evidence leading to the abnormality, contradictions and their associated paths, and generate text descriptions.
2. The commercial cryptography application security assessment evidence identification method based on image recognition according to claim 1 is characterized in that: To evaluate the senior engineer certification submitted by an employee, the Qwen-VL visual language model was used to extract information from multimodal documents, including: Read the multimodal document of the submitted senior engineer certification certificate, including text, images, tables and files; use the Qwen-VL visual language model to identify and extract key information in the document as information items, including the name of the certificate holder, certificate number, issue date, name of the issuing agency, certificate level, certificate validity period, signature on the certificate, seal pattern and text and certificate photo.
3. The commercial cryptography application security assessment evidence identification method based on image recognition according to claim 1, characterized in that: The method generates a high-dimensional vector representation for each extracted item, annotates its source, type, and preliminary quality score, and converts it into a structured evidence node, including: Generate a high-dimensional vector representation for each extracted information item, create a source annotation, record the specific information of its source document, and mark the type of the information item; give a preliminary quality score based on model confidence, content clarity and completeness; and structurally encapsulate the extracted information item content, high-dimensional vector representation, source annotation, type annotation and preliminary quality score and convert them into structured evidence nodes.
4. The commercial cryptography application security assessment evidence identification method based on image recognition according to claim 1, characterized in that: The evidence nodes are stored in a graph database, and multiple edges are established based on physical location, semantic similarity, and reference relationships to form a preliminary evidence network and assign initial credibility to the nodes, including: Read and parse the information items and attached metadata contained in each encapsulated evidence node. In the graph database, based on the physical location relationship of the information items in the original document, establish edges representing spatial proximity for spatially adjacent or ordered information nodes. These edges are labeled as spatial proximity edges and are assigned labels of adjacent fields. Utilize the high-dimensional vector representation of the information item content and calculate the cosine similarity between vectors to identify information nodes that are highly semantically related or may have a mutually corroborating relationship. Establish edges representing semantic associations for these nodes as semantic similarity edges and assign labels of semantic association or mutual corroboration. Based on the reference or derivation relationship between information items, establish edges representing reference dependency as reference association edges and assign labels of reference relationship or dependency relationship. Connect nodes through analysis of physical location, semantic similarity, and reference relationship and creation of corresponding edges to form a multi-dimensional preliminary evidence network. When assigning initial credibility to a node, the preliminary quality score of the information item is extracted, and the importance of the information item, the authority of the issuing agency, and the connection pattern of the information item in the evidence network are analyzed; comprehensive factors are calculated and assigned a specific initial credibility value for each node in the graph database through a preset weighted calculation method, representing the probability that the information item is recognized as true in the initial stage of evaluation.
5. The commercial cryptography application security assessment evidence identification method based on image recognition according to claim 1, characterized in that: The identification of certificate nodes in the network, clustering of related nodes to form a certificate entity group, and assigning a comprehensive initial credibility includes: In the initial evidence network, the Louvain algorithm is used to identify nodes representing the same certificate entity, and all related nodes directly or indirectly connected to the certificate node are found; these certificate nodes and all related nodes are regarded as a whole, and are aggregated into a certificate entity cluster through the Leiden algorithm; according to the initial credibility values of all nodes in the entity cluster, the connection strength between nodes and the overall structural characteristics of the entity cluster in the network, a random walk-based algorithm is used to calculate and assign a comprehensive initial credibility value for the certificate entity cluster.
6. The commercial cryptography application security assessment evidence identification method based on image recognition according to claim 1, characterized in that: The forward tracing starts from the certificate entity group and searches for supporting evidence along the reference association edge. By checking the credibility, consistency, trusted list matching and context relevance of the supporting evidence, the credibility of the certificate entity group is adjusted, including: In the graph database, a Cypher query is used to locate the node representing the Senior Engineer Certification Certificate and its initial attributes. All evidence nodes claiming to support the certificate's validity are found along the referenced edges, and the initial quality scores of the evidence nodes are read. Key field values of the certificate entity group are extracted, and associated evidence nodes are found in the graph database. The corresponding field information is extracted, and the basic consistency between the certificate's internal information and each piece of evidence is evaluated. Check whether the evidence contains a seal and extract its key features; for image evidence, use image processing technology to locate the seal area based on shape, color, texture or position, and crop the main seal image and evidence seal image; for text or digital evidence, check whether the seal and its features are clearly mentioned; when the evidence does not have seal information, skip the comparison; for evidence with successfully extracted seal images, perform multi-dimensional feature comparison; calculate the clarity of the main seal and evidence seal by Laplace operator variance to determine whether they are consistent; perform color quantization and compare the RGB average values of the main pixels to determine whether the colors are consistent; apply the Canny operator for edge detection, analyze the shape features of the contour, and determine whether the shapes are consistent; use optical character recognition technology to extract the seal text, and compare the text content after standardization to determine whether it is exactly the same to determine the text consistency; summarize the comparison results of each evidence in terms of clarity, color, shape and text dimensions to form a seal state set for the evidence and add it to its overall state set; when the seal does not match in any key feature and the evidence presents a seal, a seal mismatch state is generated, and this state is not generated for a complete match or no seal; Compare the consistency of the issuing agency names, extract the full name of the main agency, directly search for text evidence, and use optical character recognition for image evidence; record all relevant agency names found as an evidence agency name set, and standardize these names and the full name of the main agency; traverse the evidence agency name set to determine whether each name is exactly the same as the full name of the main agency; when there is an exact match, determine that it is consistent; when there is an incomplete match, apply the rules to check whether it is an acceptable variant, and determine that it is consistent if any rule matches; when there is no match after traversal, determine that it is inconsistent; add the status description of the issuing agency to the overall status set of the evidence; for each associated evidence node, summarize all status descriptions in its overall status set, evaluate its overall consistency with the certificate entity group, and assign a credibility score accordingly; summarize the evaluation results of all associated evidence, analyze the evidence distribution, and obtain an overall consistency score; Check whether the key field value of the certificate exists in the pre-defined trusted list. If it exists, the trustworthiness of the field is increased. If it does not exist, the trustworthiness is reduced or not adjusted, depending on whether the field is required to be in the list. The trustworthiness check scores of all fields are weighted and summed to obtain the trusted list check score. Evaluate the contextual relevance of the evidence node, query its spatially adjacent nodes and semantically similar nodes to check whether they support the certificate validity or provide background information; evaluate the supportability of each adjacent node, taking into account the node weight; aggregate the supportability scores of all spatially adjacent and semantically similar nodes, and combine the spatial proximity weight and semantic similarity weight to calculate the contextual relevance score; Adjustment factors were calculated using a pre-defined weighting formula based on the initial quality score of the evidence, the consistency score, the credibility list check score, and the contextual relevance score.
7. The commercial cryptographic application security assessment evidence identification method based on image recognition according to claim 1, characterized in that: The reverse tracing process checks for contradictory evidence, low-credibility strong associations, contextual conflicts, and outdated or replaced clues pointing to the certificate. If any are found, the credibility of the certificate entity will be reduced, including: Query all reference-related edges pointing to the certificate entity group nodes in the graph database, analyze the reverse-related nodes, and look for negative words and information conflicts. If a direct contradiction is found, mark it as high risk and calculate the direct contradiction score; Check the low-trust nodes pointing to the certificate and take their average impact score as the low-trust node impact score; The spatial proximity and semantically similar nodes of the certificate are queried. If there is more than a preset number of negative information or doubts, it is considered that there is a context conflict. The context relevance score obtained by forward tracing is used to calculate the context conflict score. Analyze reverse correlation nodes to find clues that the certificate has been updated, replaced, or invalidated. If found, mark it as high risk and calculate the certificate replacement risk score; The negative adjustment factor is calculated using a preset weighted formula by comprehensively combining the direct contradiction score, low credibility node impact score, context conflict score, and certificate replacement risk score.
8. The commercial cryptography application security assessment evidence identification method based on image recognition according to claim 1, characterized in that: The credibility of the certificate entity group is updated based on the forward and reverse traceability results, and the credibility change is propagated along the reference edge, affecting the credibility of the relevant supporting and contradictory evidence nodes, including: After forward and reverse tracing, an adjustment factor and a negative adjustment factor are obtained based on the comprehensive evaluation results. The adjustment factors are applied to the credibility of the certificate entity group. After traversing all supporting evidence, a total adjustment is made to obtain the adjusted credibility. The negative adjustment factor is applied to the credibility adjusted by forward tracing to obtain the final credibility. The final credibility is compared with the original credibility before adjustment to calculate the specific change rate of the credibility. Along all reference edges connected to the certificate entity group, the credibility change rate is transmitted to the relevant supporting evidence nodes and contradictory evidence nodes according to the preset weighted propagation rules; For nodes with supporting evidence, the improvement of certificate credibility will correspondingly improve the node's credibility; for nodes with contradictory evidence, the improvement of certificate credibility will reduce the node's credibility; The status of all relevant nodes is updated according to the credibility after propagation, triggering a new round of traceability and evaluation process.
9. The commercial cryptography application security assessment evidence identification method based on image recognition according to claim 1, characterized in that: The method traverses all certificate entity groups, selects those below the credibility threshold and marks them as abnormal, and analyzes the causes and contradictions of their low credibility, including: Set a credibility threshold, traverse all identified certificate entity groups in the graph database, and check whether the final credibility of each certificate entity group is lower than the threshold; For certificate entity groups below the threshold, they are marked as abnormal, triggering the abnormality analysis process; During the anomaly analysis process, the forward and reverse traceability results of the certificate entity group are carefully reviewed to determine the specific cause of low credibility. Focus on the contradictions in the abnormal certificate entity group, that is, the conflicting information in the forward and reverse traceability. Record abnormal certificate entities and their analysis results, and trigger corresponding alarms or notification mechanisms to ensure that responsible personnel understand and handle these abnormal situations.
10. The commercial cryptography application security assessment evidence identification method based on image recognition according to claim 1, characterized in that: The abnormal certificate information is listed, and the key evidence, contradictions and related paths leading to the abnormality are visually displayed, and a text description is generated, including: Generate a list of abnormal certificates, listing the key information of all certificate entities marked as abnormal; for each abnormal certificate, automatically extract the key evidence and contradictions that lead to its low credibility, including low-credibility supporting evidence, high-credibility negation evidence, and node information with direct conflicts; Utilizing graph database visualization tools, the anomalous certificate, its associated key evidence, contradictions, and the paths between them are graphically displayed. Simultaneously with the visualization, a textual description is generated detailing the causes of the anomalous certificate's low credibility, the specific content of the key evidence, the manifestations of the contradictions, and the logic of their associations. This helps responsible personnel understand the anomaly and provides a reference for further investigation or decision-making. Combine abnormal certificate lists, visual displays, and text descriptions into a unified report for viewing and exporting.
Citation Information
Patent Citations
Agricultural product safe supply chain anti-fake management and control system
CN103325041A
Information interaction safety protection system and operation realization method thereof
CN104573547A
Grain and oil food full supply chain information security management system and method based on trusted identifier and IPFS
CN110879902A
Document consistency comparison method based on semantic analysis and keyword driving
CN119886103A
Cited By
Network fraud evidence consolidation method and system for multi-mode semantic consistency verification
CN122133674A
Network fraud evidence solidification method and system for multimodal semantic consistency verification
CN122133674B
Artificial intelligence-based password application risk intelligent assessment method and system
CN122372186A