Pre-press electronic file comparison method and device and storage medium

By combining a multi-scale feature pyramid network with a rule engine, the problems of low efficiency and poor accuracy in prepress document comparison are solved, and high-precision multi-language and multi-modal content comparison is achieved, meeting the high compliance requirements of pharmaceutical packaging and improving production efficiency and quality control.

CN120689885APending Publication Date: 2025-09-23SHENZHEN NINE STARS PRINTING & PACKAGING GRP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510805699.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing prepress document comparison technology in the field of pharmaceutical packaging and printing has problems such as low efficiency, high false detection rate, inability to recognize small-size text and complex chemical formulas, inability to meet multilingual content consistency requirements, and lack of pharmaceutical industry standard optimization, making it difficult to achieve high precision and high compliance requirements.

Method used

A multi-scale feature pyramid network is used for super-resolution reconstruction, conditional random fields and multi-language processing units are combined to identify text areas, chemical formula topology analysis is used to extract topological structure features, and multi-modal semantic feature matching detection is performed using a rule engine and pharmaceutical industry standard library to output a difference report.

Benefits of technology

It improves the accuracy and efficiency of pre-press document inspection, reduces quality control costs, improves production efficiency and quality levels, and significantly improves the compliance and accuracy of pharmaceutical packaging printing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689885A_ABST
    Figure CN120689885A_ABST
Patent Text Reader

Abstract

The invention discloses a pre-press electronic file comparison method and device and a storage medium, and relates to the field of image display. Carrying out super-resolution reconstruction on the input electronic file, and identifying and dividing a character area, a graph area and a chemical formula area; text semantic features are extracted, and an intermediate file with semantic annotations in a target text area is generated; the chemical formula is converted into fingerprint codes, and topological structure features are extracted through chemical formula topological analysis; performing format conversion on the graph, and extracting a graph feature matrix according to the pixel coordinate points and the image shape; and carrying out differential matching detection on the multi-modal semantic features, the topological structure features and the graphic feature matrix by utilizing a rule engine and a drug industry specification library, and outputting a difference report containing a confidence score. According to the scheme, multi-modal feature extraction, a dynamic rule engine and an industry knowledge base are fused, a new pre-press comparison normal form oriented to the field of medicine packaging is constructed, and the problems of small text missing detection, chemical formula misjudgment and multi-language ambiguity are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image display, and in particular to a pre-press electronic file comparison method, device, and storage medium. Background Art

[0002] In the field of pharmaceutical packaging printing, pharmaceutical packaging boxes and instructions are the core carriers of key pharmaceutical information, and they have extremely high requirements for content accuracy. Prepress electronic files are significantly complex and contain a variety of special elements: small fonts are common, such as 4-point fonts are often used to mark important information such as contraindications, usage and dosage of drugs; chemical molecular formulas are complex and diverse, such as C6H 12 Structures such as O6 contain specific atomic connections and spatial configurations; the accuracy of mathematical formulas, such as dosage calculation formulas, is directly related to medication safety; multilingual mixed typesetting is common, and the combination of Chinese, English, Japanese, Korean and other languages ​​poses challenges to cross-language semantic understanding in text comparison.

[0003] The traditional prepress file comparison methods in the current industry have many drawbacks. Manual visual inspection relies on manpower to compare line by line, which is extremely inefficient. For dozens of pages of instructions, it often takes several hours to complete. In addition, due to the limitations of human visual fatigue and concentration, the false detection rate exceeds tens of percentage points, making it difficult to ensure accurate recognition of small-size text and complex chemical formulas. General text comparison tools are not capable of processing professional typesetting files. They are unable to parse information such as layers, vector elements, and font metadata of prepress typesetting software files, resulting in poor format compatibility. In multilingual mixed scenarios, such tools can only perform simple character-level comparisons, and cannot compare changes in file formatting before and after editing. They are also unable to establish semantic associations between terms in different languages, making semantic matching ineffective. It is difficult to meet the strict requirements of pharmaceutical packaging for multilingual content consistency and to prompt designers to edit differences in historical versions.

[0004] Existing patented technologies, such as the "Automatic Prepress Graphics and Text Proofreading Device" (CN103336759A), attempt to achieve basic comparison through format conversion, but suffer from significant flaws. During the format conversion process, when converting PDFs to plain text, important information such as vector graphics, font metadata, and layers is inevitably lost, resulting in distorted file content and affecting the accuracy of subsequent comparisons. The comparison logic only compares differences at the character level, lacking semantic understanding and structural analysis of the content. This makes it impossible to identify differences between chemical formula isomers. For example, the different structures of ortho and para substituents can lead to significant changes in chemical properties, but existing technologies are unable to detect such structural risks. Furthermore, existing technologies are not optimized for the specific needs of the pharmaceutical industry and lack built-in pharmaceutical industry standards, such as dedicated rules for dosage unit consistency checks required by FDA regulations. Consequently, they lack dynamic adaptability and are unable to meet the high-precision and high-compliance requirements of pharmaceutical packaging printing. Summary of the Invention

[0005] The embodiments of the present application provide a prepress electronic file comparison method, device and storage medium for improving the accuracy and efficiency of prepress file inspection and reducing the mismatch between prepress files and printing process requirements.

[0006] In one aspect, the present application provides a pre-press electronic file comparison method, the method comprising: Perform super-resolution reconstruction of input electronic files based on a multi-scale feature pyramid network, and identify and segment text, graphics, and chemical formula areas from the reconstructed scanned images; Extract the semantic features of the text in the text area and generate an intermediate file with semantic annotations for the target text area; convert the chemical formula in the chemical formula area into a fingerprint code and extract the topological structure features through chemical formula topological analysis; convert the graphics in the graphic area into a format and extract the graphic feature matrix based on pixel coordinate points and image shape; The rule engine and pharmaceutical industry standard library are used to perform differential matching detection on multimodal semantic features, topological structure features, and graphic feature matrices, and output a difference report containing confidence scores.

[0007] Specifically, the super-resolution reconstruction of the input electronic file based on the multi-scale feature pyramid network and the identification and segmentation of text, graphics, and chemical formula areas from the reconstructed scanned image include: Scanning electronic documents based on a super-resolution reconstruction mechanism, performing four-fold upsampling on the scanned images that are lower than the target resolution to generate a sampled image; The sampled image and the original scanned image are fed into the residual network for fusion processing to generate a reconstructed scanned image; The reconstructed scanned images are extracted and fused layer by layer according to the resolution level. Combined with the semantic segmentation network, they are recognized at the pixel level, and the typesetting features such as character spacing, line height, and alignment are extracted. The text area and table area are delineated based on the typesetting features. The mathematical expression structure is parsed through the LaTeX syntax tree, and the formula area is delineated based on the structure. The path parameters of the vector graphics are extracted based on the contour detection algorithm of OpenCV, and the image area is divided according to the contour.

[0008] Specifically, extracting the semantic features of the text area and generating an intermediate file with semantic annotations for the target text area includes: Conditional random fields are used to post-process the recognition results and correct the misjudgment of isolated characters through contextual semantic constraints; Obtain a shared semantic space library, optimize cross-language term matching between entity relationships using a translation distance loss function, and determine the text language; Based on the font features and font size, the small font size of the text and the corresponding target text range are determined, and an intermediate file with semantic annotations is generated for semantic judgment.

[0009] Specifically, extracting the semantic features of the text area and generating an intermediate file with semantic annotations for the target text area includes: Conditional random fields are used to post-process the recognition results and correct the misjudgment of isolated characters through contextual semantic constraints; Obtain a shared semantic space library, optimize cross-language term matching between entity relationships using a translation distance loss function, and determine the text language; Based on the font features and font size, the small font size of the text and the corresponding target text range are determined, and an intermediate file with semantic annotations is generated for semantic judgment.

[0010] Specifically, the step of converting the chemical formula in the chemical formula area into a fingerprint code and extracting topological structure features through chemical formula topological analysis includes: The extended Morgan algorithm is used to generate atomic environment fingerprints. Centered on the target atom, the atom type, hybridization state, and substituent connection pattern within the three-layer bond length range are recursively extracted to generate a 1024-bit binary fingerprint code. Based on the molecular structure characterization model, the molecular formula is converted into an atom-bond graph. By aggregating the adjacent node features layer by layer, the fingerprint code is converted into a high-dimensional fingerprint feature vector containing atomic connectivity, hybridization state, and spatial configuration.

[0011] Specifically, the format conversion of the graphics in the graphics area and the extraction of the graphic feature matrix according to the pixel coordinate points and the image shape include: Use regular expressions to match basic primitives in vector path files and extract control point coordinates; A path descriptor containing several geometric parameters is constructed based on the coordinates of the control points to generate a graphic path and determine the graphic feature matrix.

[0012] Specifically, the use of the rule engine and the pharmaceutical industry specification library to perform differential matching detection on multimodal semantic features, topological structure features, and graphic feature matrices includes: A rule-based knowledge graph is constructed based on relevant regulations on pharmaceutical packaging. The core entities and semantic feature relationship keys in the graph are used for semantic verification, and semantic differences are determined based on the cosine similarity values ​​of cross-language terms. Measure the structural similarity of chemical formulas converted into fingerprint feature vectors, and determine the detection results based on the similarity value; Perform matrix decomposition on the graphic feature matrix, separate affine transformation and non-affine transformation, calculate the vector graphics transformation deviation, and determine the image distortion based on the deviation result.

[0013] Specifically, the process of determining semantic differences includes: Calculate cross-language semantic relevance based on embedded fusion terms, language encoding, and domain attributes , the formula is as follows:

[0014] in, and is the term node feature, is the set of adjacent nodes, represents the attention weight matrix, represents the activation function; The process of measuring structural similarity includes: Introducing the Tanimoto coefficient to calculate the structural similarity of chemical molecules , which is expressed as follows:

[0015] Among them, F A and F B They are respectively the fingerprint feature vectors of the molecules to be compared. When T<0.9, the auxiliary verification mechanism is triggered. By performing three-layer feature aggregation on the atom-bond graph, a structure vector containing atomic connectivity and spatial configuration is generated, and the joint discrimination output recognition result is constructed in combination with the fingerprint coding.

[0016] The process of calculating vector graphics transformation deviation includes: The transformation deviation measurement formula is defined as follows:

[0017] Among them and is the control point coordinate offset, M is the transformation matrix, is the determinant value; when When >0.5, it is determined that non-affine distortion exists.

[0018] On the other hand, the present application provides a computer device, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the prepress electronic file comparison method described in any of the above aspects.

[0019] On the other hand, the present application provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set or instruction set. The at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the prepress electronic file comparison method described in any of the above aspects.

[0020] The beneficial effects of the technical solutions provided in the embodiments of the present application include at least: This solution, integrating multimodal feature extraction, a dynamic rules engine, and an industry knowledge base, establishes a new paradigm for prepress comparison in the pharmaceutical packaging sector. It effectively addresses long-standing industry challenges such as missed small text, misjudgment of chemical formulas, and multilingual ambiguity. It can save companies over 50% in quality control costs, improve production efficiency and quality, and play a significant role in driving intelligent upgrades in the printing industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a structural diagram of a prepress electronic file comparison system provided in an embodiment of the present application; Figure 2 is a flow chart of a prepress electronic file comparison method provided in an embodiment of the present application; Figure 3 This is a flowchart of the differential matching detection provided by the embodiment of the present application; Figure 4 It is a structural diagram of a prepress electronic file comparison device provided in an embodiment of the present application; Figure 5 A structural block diagram of a computer device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION

[0022] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0023] In this document, "plurality" refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.

[0024] Figure 1 Figure 1 is a schematic diagram of the prepress electronic document comparison system provided in an embodiment of this application. This system utilizes a layered, decoupled technical architecture, achieving high-precision difference detection in complex scenarios through multimodal feature fusion and domain knowledge-driven rule reasoning. The system architecture, shown in Figure 1, comprises three core functional modules: a high-precision OCR enhancement module, a structured feature extraction module, and a dynamic adaptive rule engine module. Each module communicates and interacts with feature information via standardized data interfaces.

[0025] The high-precision OCR enhancement module builds a multi-stage image enhancement and semantic-aware character recognition framework optimized for the unique imaging characteristics of pharmaceutical documents. It employs deep learning-based super-resolution reconstruction models (such as ESRGAN) to perform a 4x upsampling of scanned images with resolutions below 300 dpi. This utilizes a residual dense network architecture to enhance edge clarity of small-size characters. During the feature extraction phase, a multi-scale Feature Pyramid Network (FPN) is deployed. Through a top-down feature fusion strategy, it integrates semantic information from different resolution levels (P2-P5), effectively capturing the detailed features of 4-8pt fonts.

[0026] The high-precision OCR enhancement module includes a chemical symbol recognition unit, a conditional random field (CRF) unit, and a multilingual processing unit. The chemical symbol recognition unit integrates a domain-specific character representation system, specifically a dedicated font library containing over 1,200 chemical glyphs. These include chemical symbols in the Unicode extension block (such as U+21CC to U+21D3, reaction arrows, U+2B06 to U+2B33, and geometric shapes), as well as custom molecular formula templates (supporting cyclic structures and stereo configuration annotations).

[0027] The conditional random field recognition unit can post-process the recognition results and correct the misjudgment of isolated characters through contextual semantic constraints, such as distinguishing the difference between the number "0" and the letter "O".

[0028] The multilingual processing unit is based on the cross-lingual pre-trained mBERT (multilingual BERT) model. Using a masked language model (MLM), it learns the shared semantic space of ten languages, including Chinese, English, Japanese, and Korean. Its purpose is to perform cross-lingual term alignment based on character-level recognition, i.e., multilingual mapping relationships. For example, a multilingual mapping relationship such as "contraindications - contraindications - contraindications - 금기증" can be established to address semantic discontinuity in mixed typesetting scenarios.

[0029] The structured feature extraction module is primarily designed to implement multi-level feature parsing from file formats to semantic structures, constructing a three-dimensional feature space encompassing visual features, logical structures, and domain semantics. It includes a vector element parsing unit specifically for vector graphics parsing, a semantic segmentation and region classification unit for distinguishing text, formulas, and images within images, and a chemical formula topology analysis unit specifically for analyzing chemical formulas and symbols.

[0030] Specifically, the vector element parsing unit has developed a parser-based metadata extraction engine for prepress software files (such as PDF / InDesign). This engine supports precise analysis of CMYK color gamut parameters (accuracy down to 0.1%), font features (weight, width, kerning parameters), and layer hierarchy (including lock status and transparency settings). This engine establishes a metadata description system encompassing 238 format features. During the actual feature extraction phase, vector element analysis and the construction of a graphic feature matrix are performed based on this metadata description system.

[0031] The semantic segmentation and region classification unit is primarily a prerequisite for the recognition phase, dividing the scanned image into regions and extracting feature information. In one possible implementation, a document layout analysis model can be constructed using the U-Net++ network architecture. This divides the document into multiple semantic regions, such as text, formula, image, and table, at the pixel level, and differentiated feature extraction strategies are designed for each region. For example, the recognition strategy is as follows: Text and table area: extract typesetting features such as character spacing, line height, and alignment; Formula area: parses mathematical expression structures through LaTeX syntax trees; for example, fractions, subscript and subscript hierarchical relationships, etc. Image area: Extract path parameters of vector graphics based on OpenCV's contour detection algorithm, such as using Bezier curves to control point coordinates with an accuracy of 0.01mm.

[0032] The chemical formula topology analysis unit is designed to identify chemical symbols and chemical formulas. This solution can build a molecular structure representation model based on graph neural networks, converting molecular formulas into atom-bond graphs (where nodes are atoms and edges are chemical bond types), and aggregate adjacent node features layer by layer through a graph convolutional network (GCN) to generate high-dimensional feature vectors containing atomic connectivity, hybridization state, and spatial configuration, providing semantic-level feature representation for subsequent structural comparisons.

[0033] The dynamic adaptive rule engine module matches and identifies multimodal feature information extracted in the early stages. By constructing a three-level architecture consisting of a "domain rule library - flexible matching mechanism - incremental learning module," it enables intelligent adjustment of matching strategies and achieves higher recognition rates. Specifically, it includes an industry rule knowledge graph unit and a multi-dimensional difference assessment unit.

[0034] The industry-specific knowledge graph unit builds an ontology model based on pharmaceutical packaging regulations (such as FDA 21 CFR Part 11 and ICH Q8). It includes 18 core entities (such as "unit of measurement," "contraindication term," and "approval document number") and 56 semantic relationships (such as "equivalent expression," "hierarchical inclusion," and "mandatory verification") to enable semantic verification of key terms. For example, it automatically identifies the equivalence of "three times a day" and "three times / day."

[0035] Multi-dimensional difference assessment model: This model establishes a comprehensive assessment system that includes visual differences (ΔE color difference, font size deviation Δpt), structural differences (molecular formula Tanimoto coefficient, vector path Hausdorff distance), and semantic differences (cross-language term cosine similarity). It supports users to customize the weight parameters of each dimension (weight range 0.1-0.9, step size 0.1) and acceptable thresholds (e.g., setting ΔE ≤ 3.0 as the color tolerance range) through a graphical interface.

[0036] In addition, this module can also include an incremental learning optimization unit, which adopts a false alarm filtering mechanism based on contrastive learning. By comparing logs, it constructs positive and negative sample pairs (positive samples: reasonable differences confirmed by humans; negative samples: system false alarm items), and uses the gradient boosting tree (GBRT) to dynamically adjust the algorithm parameters. It establishes exclusive feature templates for the layout habits of specific companies (such as a pharmaceutical company that always uses the Hanyi Shu Song font), reducing the repeated triggering rate of historical false alarms by 62%.

[0037] Figure 2 Flowchart of the prepress electronic file comparison method provided in an embodiment of the present application, comprising the following steps: S1. Super-resolution reconstruction of input electronic files based on a multi-scale feature pyramid network, and identification and segmentation of text, graphics, and chemical formula areas from the reconstructed scanned images; This step uses a super-resolution reconstruction model based on deep learning. When scanning electronic documents, images below the target resolution are upsampled by 4 times, and the edge clarity of small-size characters is improved through a residual dense network structure. The details can be summarized as follows: S11, scanning the electronic document based on the super-resolution reconstruction mechanism, performing a four-fold upsampling process on the scanned image that is lower than the target resolution to generate a sampled image; S12, the sampled image and the original scanned image are fed into a residual network for fusion processing to generate a reconstructed scanned image; S13 extracts semantic information from the reconstructed scanned image layer by layer according to the resolution level and fuses them. Combined with the semantic segmentation network, it performs pixel-level recognition, extracts typesetting features such as character spacing, line height, and alignment, and delineates text and table areas based on these typesetting features. It parses the structure of mathematical expressions through the LaTeX syntax tree and delineates the formula area based on the structure. It extracts the path parameters of vector graphics based on the contour detection algorithm of OpenCV and delineates the image area based on the contour.

[0038] S2. Extract semantic features of text in the text area and generate an intermediate file with semantic annotations for the target text area; convert chemical formulas in the chemical formula area into fingerprint codes and extract topological structural features through chemical formula topological analysis; convert the graphics in the graphic area into a format and extract a graphic feature matrix based on pixel coordinates and image shape; Next, we process and extract features of the data of each region by type for later difference evaluation.

[0039] 1. Text area: 1) Use conditional random fields to post-process the recognition results and correct the misjudgment of isolated characters through contextual semantic constraints; The purpose of this step is to continue to optimize and correct the text based on the recognition process to reduce the probability of errors, namely, correction of some special letters, numbers and symbols.

[0040] 2) Obtain a shared semantic space library, optimize cross-language term matching between entity relationships using a translation distance loss function, and determine the text language; Determining the language of a text involves two steps: identifying the text and establishing a language mapping. This mapping is determined through a shared semantic space, which helps resolve semantic gaps in mixed typesetting scenarios when extracting semantic information.

[0041] The specific implementation of cross-language term alignment can be achieved by using a multilingual semantic alignment method based on knowledge graphs and graph attention networks. The domain-specific multilingual knowledge graph (ML-KG) contains more than 32,000 cross-language term pairs and 12 types of semantic relations (equivalence, hyponymy, taboo association, etc.). The knowledge graph construction is divided into three stages: Ontology modeling: Based on the ISO 30042 terminology standard, define the "drug terminology" ontology class, including core entities such as "dosage unit", "contraindications", and "indications", and build cross-language attribute mapping (e.g., Chinese "Spec" → English "Specification" → Japanese "Spec"); Term extraction: We use the mBERT model to perform entity recognition on multilingual corpora such as FDA drug instructions and ICH guidelines, and use a rule engine to clean up noisy terms (for example, excluding non-medical meanings of "dose"). Graph completion: Using the TransR cross-language knowledge representation model, Chinese, English, Japanese, and Korean terms are embedded into a shared semantic space (dimension 200). The translation distance loss function is used to optimize the relationship between entities and solve the problem of data sparsity.

[0042] In the semantic alignment stage, a multi-layer graph attention network (GAT) is deployed for cross-language term matching.

[0043] 3) Based on the font features and font size, determine the small font size of the text and the corresponding target text range, and generate an intermediate file with semantic annotations for semantic judgment.

[0044] For small fonts, especially for notes, precautions, superscripts, and special text, the residual dense network structure improves edge clarity of small font characters. The small font and area positions are determined based on the font size, and an intermediate file with semantic annotations is generated, providing rich semantic information for subsequent processing. Semantic annotations can be generated based on actual conditions, such as text such as "Taboo, Focus on" for later manual verification and semantic judgment.

[0045] 2. Chemical formula area: 1) Generate atomic environment fingerprints using the extended Morgan algorithm. Centered on the target atom, recursively extract the atom type, hybridization state, and substituent connection pattern within the three-layer bond length range to generate a 1024-bit binary fingerprint code; This step introduces molecular fingerprinting-based chemical formula isomer recognition technology, particularly for isomers with similar chemical formulas (chemical symbols containing benzene ring structures), which are prone to identification errors. To address this issue, this application constructs a multi-level difference detection system, from symbolic representation to topological structure. Molecular fingerprinting technology converts two-dimensional chemical formulas into structured feature vectors.

[0046] An extended Morgan algorithm (with radius = 3) is used to generate atomic environment fingerprints. Centered on the target atom, the algorithm recursively extracts the atom type (e.g., C / H / N / O, etc.), hybridization state (sp² / sp³), and substituent connection pattern within a three-layer bond length range, generating a 1024-bit binary fingerprint code that can effectively capture ring structures (e.g., ortho-para substitutions on the benzene ring) and stereochemical differences (e.g., chiral carbon atom configurations).

[0047] 2) Based on the molecular structure representation model, the molecular formula is converted into an atom-bond graph. By aggregating the adjacent node features layer by layer, the fingerprint code is converted into a high-dimensional fingerprint feature vector containing atomic connectivity, hybridization state, and spatial configuration; This step utilizes the topological structure of chemical molecules and uses chemical principles to further analyze and encode the features of the atom-bond graph. This application can construct a molecular structure representation model based on graph neural networks, converting the molecular formula into an atom-bond graph (nodes are atoms, edges are chemical bond types), and then aggregating adjacent node features layer by layer through a graph convolutional network (GCN) to generate a high-dimensional feature vector containing atomic connectivity, hybridization state, and spatial configuration, providing semantic-level feature representation for subsequent structural comparison.

[0048] 3. Vector Graphics Area (Vector graphics are commonly used in medical printed electronic documents, including trademarks and images) 1) Use regular expressions to match basic primitives in the vector path file and extract the coordinates of control points; This step mainly involves the geometric feature characteristics of Liu Yongkai, and the core lies in the precise analysis of path parameters and transformation invariance analysis.

[0049] Specifically, you can use regular expressions to match basic primitives in SVG / PDF paths (including but not limited to M / m movement, L / l line, Q / q quadratic Bezier, and C / c cubic Bezier) and extract the coordinates of control points (with an accuracy of 0.01mm for device-independent pixels).

[0050] 2) Based on the coordinates of the control points, a path descriptor containing several geometric parameters is constructed to generate a graphic path and determine the graphic feature matrix.

[0051] Based on the control point coordinates, a path descriptor containing 168 geometric parameters is constructed. This path descriptor generates a graphic path and determines a feature matrix. This solution's path descriptor calculates and determines a point cloud (i.e., a descriptor) based on the image's contour curve. The greater the number of point clouds, the clearer the graphic path.

[0052] S3. Use the rule engine and pharmaceutical industry standard library to perform differential matching detection on multimodal semantic features, topological structure features, and graphic feature matrices, and output a difference report containing confidence scores.

[0053] This step is also divided into three types for differential evaluation, classification analysis output confidence score and output and difference report, see Figure 3 The flowchart shown.

[0054] S301. Analyze the semantic features of text by constructing a rule-based knowledge graph based on relevant regulations on pharmaceutical packaging, perform semantic level verification using the core entities and semantic feature relationship keys in the graph, and determine semantic differences based on cross-language term cosine similarity values; Based on the embedded fusion terms, language encoding, and domain attributes, the semantic relevance of cross-language term pairs is calculated through multi-head attention (8 heads) to calculate the semantic relevance of cross-language terms. , the formula is as follows:

[0055] in, and is the term node feature, is the set of adjacent nodes, represents the attention weight matrix, Represents the activation function. Attention calculation needs to interact with the input query (hi), key (hj) and other vectors. First perform a linear transformation on these vectors. This "guiding tool" is designed to make attention calculations more accurate and task-adaptive. It simultaneously adjusts the feature space and learns how to allocate attention. It is a key learnable component in multi-head attention that enables the model to "learn to pay attention." This mechanism enables the system to accurately identify the semantic equivalence of "medication interval" (Chinese), "DosingInterval" (English), and "Touyao jiejie" (Japanese). This resolves mismatching issues in mixed typesetting caused by word order differences (such as inverted modifiers in English), achieving a cross-language term alignment accuracy of 93.6%.

[0056] S302, measuring the structural similarity of the chemical formula converted into the fingerprint feature vector, and determining the detection result according to the similarity value; This process introduces the Tanimoto coefficient to calculate the structural similarity of chemical molecules , which is expressed as follows:

[0057] Among them, F A and F B The fingerprint feature vectors of the molecules to be compared are respectively. When T < 0.9, the graph neural network-assisted verification mechanism is triggered. A graph convolutional network (GCN) performs three-layer feature aggregation on the atom-bond graph, generating a 128-dimensional structure vector that includes atomic connectivity (such as the number of rings and branches) and spatial configuration (such as cis-trans isomerism). This is combined with fingerprint encoding to construct a joint discriminant model and identify output confidence information, improving the accuracy of isomer detection to 97.2% (a 42 percentage point increase over traditional character matching methods).

[0058] This technology breaks through the limitations of traditional methods that rely solely on molecular formula character matching. It can identify structural risks caused by substituent position isomerism (such as catechol vs. resorcinol) and differences in functional group connections (such as methyl ether vs. ethanol). It is particularly suitable for accurate comparison of chemical names and excipient ingredients in drug instructions.

[0059] S303: Perform matrix decomposition on the graphic feature matrix, separate affine transformation and non-affine transformation, calculate the vector graphic transformation deviation, and determine the image distortion according to the deviation result.

[0060] This process is mainly based on transformation matrix decomposition for evaluation. The geometric transformations (translation, rotation, scaling, shearing) applied to the path are decomposed into matrices, separating affine transformations (maintaining line parallelism) from non-affine transformations (such as local distortion). The process of calculating the vector graphics transformation deviation includes: The transformation deviation measurement formula is defined as follows:

[0061] Among them and is the control point coordinate offset, M is the transformation matrix, is the determinant value; when When >0.5, it is determined that non-affine distortion exists.

[0062] Furthermore, this solution can also incorporate Hausdorff distance to calculate shape differences in vector graphics, combined with the curvature change rate of Bezier curves (threshold 0.3 rad / mm) to detect local deformation. This method can identify path deviations as small as 0.1mm and angular deviations as small as 1°, improving detection accuracy tenfold compared to traditional pixel-level comparison (accuracy of 1mm).

[0063] During the implementation process, by retaining the original vector path data (non-rasterized processing), the loss of anchor points and curve fitting errors caused by format conversion are avoided, ensuring the complete transmission of CMYK color gamut parameters (accuracy 0.1%) and font outline features (font weight deviation ≤ 0.5pt), providing technical support for the compliance comparison of vector graphics such as registered trademarks and logos in pharmaceutical packaging.

[0064] The technical effects brought about by the technical solution of this application include the following: The first introduction of molecular fingerprint coding and multilingual knowledge graphs into the field of prepress comparison breaks the limitation of traditional comparison methods that rely solely on character-level comparison, and provides new technical ideas and methods for prepress document comparison. It breaks through the limitations of traditional character-level comparison and implements three-dimensional difference detection based on "semantic-structural-visual". It understands multilingual content at the semantic level, analyzes chemical formulas and vector graphics at the structural level, and detects format and image differences at the visual level, building a comprehensive and in-depth comparison system. This program has achieved good results in actual application. A pilot project at a pharmaceutical company showed that the program reduced the quality accident rate by 67% and increased the efficiency of passing FDA audits by 40%. It significantly improved the company's quality control level and work efficiency, and has strong practicality and promotion value.

[0065] Figure 4 : is a schematic structural diagram of a prepress electronic file comparison device provided in an embodiment of the present application, comprising: Recognition module 410, for performing super-resolution reconstruction of the input electronic file based on a multi-scale feature pyramid network, and identifying and dividing text, graphics, and chemical formula areas from the reconstructed scanned image; Feature extraction module 420 is used to extract semantic features of text in the text area and generate an intermediate file with semantic annotations for the target text area; convert chemical formulas in the chemical formula area into fingerprint codes and extract topological structural features through chemical formula topological analysis; convert the format of graphics in the graphic area and extract the graphic feature matrix based on pixel coordinate points and image shape; The output module 430 is used to use the rule engine and the pharmaceutical industry specification library to perform differential matching detection on the multimodal semantic features, topological structure features, and graphic feature matrix, and output a difference report including a confidence score.

[0066] The prepress electronic file comparison device provided in the embodiment of the present application can be applied to the prepress electronic file comparison method provided in the above embodiment. For relevant details, please refer to the above method embodiment. Its implementation principles and technical effects are similar and will not be repeated here.

[0067] It should be noted that the prepress electronic file comparison device provided in the embodiment of the present application is only illustrated by the division of the above-mentioned functional modules / functional units. In actual applications, the above-mentioned functions can be assigned to different functional modules / functional units as needed, that is, the internal structure of the prepress electronic file comparison device can be divided into different functional modules / functional units to complete all or part of the functions described above. In addition, the implementation method of the prepress electronic file comparison method provided in the above-mentioned method embodiment and the implementation method of the prepress electronic file comparison device provided in this embodiment are based on the same concept. The specific implementation process of the prepress electronic file comparison device provided in this embodiment is detailed in the above-mentioned method embodiment and will not be repeated here.

[0068] Figure 5 The following is a block diagram of the structure of a computer device provided by an exemplary embodiment of the present application. The computer device may be a desktop computer, a laptop computer, a PDA, a cloud server, or other computer device. The computer device may include, but is not limited to, a processor and a memory. The processor and the memory may be connected via a bus or other means. The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, graphics processing units (GPU), embedded neural network processors (NPU) or other dedicated deep learning coprocessors, discrete gate or transistor logic devices, discrete hardware components, or other chips, or a combination of the above chips.

[0069] The processor may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor may be implemented in at least one of the following hardware forms: a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is used to process data while awake, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data while in standby mode. In some embodiments, the processor may integrate a graphics processing unit (GPU), which is responsible for rendering and drawing content displayed on the display. In some embodiments, the processor may also include an artificial intelligence (AI) processor, which is used to handle computational operations related to machine learning.

[0070] As a non-transient computer-readable storage medium, the memory can be used to store non-transient software programs, non-transient computer executable programs and modules, such as program instructions / modules corresponding to the methods in the above-mentioned embodiments of the present application. The processor executes various functional applications and data processing of the processor by running the non-transient software programs, instructions and modules stored in the memory, that is, the method in the above-mentioned method embodiment is implemented. The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0071] The peripheral device interface can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor and memory. In some embodiments, the processor, memory, and peripheral device interface are integrated on the same chip or circuit board. In other embodiments, any one or two of the processor, memory, and peripheral device interface can be implemented on separate chips or circuit boards, although this embodiment is not limited to this.

[0072] The embodiment of the present application also discloses a computer-readable storage medium. Specifically, the computer-readable storage medium is used to store a computer program, and when the computer program is executed by the processor, the method in the above-mentioned method implementation is implemented. Those skilled in the art will understand that all or part of the processes in the above-mentioned method implementation of the present application can be completed by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the implementation of the above-mentioned methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (Flash Memory), a hard disk drive (HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memory.

[0073] This specific embodiment is merely an explanation of the present invention and is not intended to limit the present invention. After reading this specification, those skilled in the art may make non-creative modifications to this embodiment as needed. However, as long as such modifications are within the scope of the claims of the present invention, they are protected by patent law.

Claims

1. A pre-press electronic file comparison method, characterized in that: The method comprises: Perform super-resolution reconstruction of input electronic files based on a multi-scale feature pyramid network, and identify and segment text, graphics, and chemical formula areas from the reconstructed scanned images; Extract the semantic features of the text in the text area and generate an intermediate file with semantic annotations for the target text area; convert the chemical formula in the chemical formula area into a fingerprint code and extract the topological structure features through chemical formula topological analysis; convert the graphics in the graphic area into a format and extract the graphic feature matrix based on pixel coordinate points and image shape; The rule engine and pharmaceutical industry standard library are used to perform differential matching detection on multimodal semantic features, topological structure features, and graphic feature matrices, and output a difference report containing confidence scores.

2. The method according to claim 1, characterized in that The method of super-resolution reconstruction of the input electronic file based on the multi-scale feature pyramid network and identifying and dividing text, graphics, and chemical formula areas from the reconstructed scanned image includes: Scanning electronic documents based on a super-resolution reconstruction mechanism, performing four-fold upsampling on the scanned images that are lower than the target resolution to generate a sampled image; The sampled image and the original scanned image are fed into the residual network for fusion processing to generate a reconstructed scanned image; The reconstructed scanned images are extracted and fused layer by layer according to the resolution level. Combined with the semantic segmentation network, they are recognized at the pixel level, and the typesetting features such as character spacing, line height, and alignment are extracted. The text area and table area are delineated based on the typesetting features. The mathematical expression structure is parsed through the LaTeX syntax tree, and the formula area is delineated based on the structure. The path parameters of the vector graphics are extracted based on the contour detection algorithm of OpenCV, and the image area is divided according to the contour.

3. The method according to claim 2, characterized in that The step of extracting semantic features of text areas and generating an intermediate file with semantic annotations for the target text areas includes: Conditional random fields are used to post-process the recognition results and correct the misjudgment of isolated characters through contextual semantic constraints; Obtain a shared semantic space library, optimize cross-language term matching between entity relationships using a translation distance loss function, and determine the text language; Based on the font features and font size, the small font size of the text and the corresponding target text range are determined, and an intermediate file with semantic annotations is generated for semantic judgment.

4. The method according to claim 1, wherein The step of extracting semantic features of text areas and generating an intermediate file with semantic annotations for the target text areas includes: Conditional random fields are used to post-process the recognition results and correct the misjudgment of isolated characters through contextual semantic constraints; Obtain a shared semantic space library, optimize cross-language term matching between entity relationships using a translation distance loss function, and determine the text language; Based on the font features and font size, the small font size of the text and the corresponding target text range are determined, and an intermediate file with semantic annotations is generated for semantic judgment.

5. The method according to claim 4, characterized in that The step of converting the chemical formula in the chemical formula area into a fingerprint code and extracting topological structure features through chemical formula topological analysis includes: The extended Morgan algorithm is used to generate atomic environment fingerprints. Centered on the target atom, the atom type, hybridization state, and substituent connection pattern within the three-layer bond length range are recursively extracted to generate a 1024-bit binary fingerprint code. Based on the molecular structure characterization model, the molecular formula is converted into an atom-bond graph. By aggregating the adjacent node features layer by layer, the fingerprint code is converted into a high-dimensional fingerprint feature vector containing atomic connectivity, hybridization state, and spatial configuration.

6. The method according to claim 5, characterized in that The step of converting the format of the graphics in the graphics area and extracting the graphic feature matrix according to the pixel coordinate points and the image shape includes: Use regular expressions to match basic primitives in vector path files and extract control point coordinates; A path descriptor containing several geometric parameters is constructed based on the coordinates of the control points to generate a graphic path and determine the graphic feature matrix.

7. The method according to claim 6, characterized in that The use of the rule engine and the pharmaceutical industry specification library to perform differential matching detection on multimodal semantic features, topological structure features, and graphic feature matrices includes: A rule-based knowledge graph is constructed based on relevant regulations on pharmaceutical packaging. The core entities and semantic feature relationship keys in the graph are used for semantic verification, and semantic differences are determined based on the cosine similarity values ​​of cross-language terms. Measure the structural similarity of chemical formulas converted into fingerprint feature vectors, and determine the detection results based on the similarity value; Perform matrix decomposition on the graphic feature matrix, separate affine transformation and non-affine transformation, calculate the vector graphics transformation deviation, and determine the image distortion based on the deviation result.

8. The method according to claim 7, characterized in that The process of determining semantic differences includes: Calculate cross-language semantic relevance based on embedded fusion terms, language encoding, and domain attributes , the formula is as follows: in, and is the term node feature, is the set of adjacent nodes, represents the attention weight matrix, represents the activation function; The process of measuring structural similarity includes: Introducing the Tanimoto coefficient to calculate the structural similarity of chemical molecules , which is expressed as follows: Among them, F A and F B are the fingerprint feature vectors of the molecules to be compared. When T < 0.9, the auxiliary verification mechanism is triggered. By performing three-layer feature aggregation on the atom-bond graph, a structure vector containing atomic connectivity and spatial configuration is generated, and the joint discrimination output recognition result is constructed in combination with the fingerprint coding. The process of calculating vector graphics transformation deviation includes: The transformation deviation measurement formula is defined as follows: Among them and is the control point coordinate offset, M is the transformation matrix, is the determinant value; when When it is > 0.5, it is determined that non-affine distortion exists.

9. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the prepress electronic file comparison method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The readable storage medium stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the prepress electronic file comparison method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Device and method for automatically proofreading pre-printing image and text

    CN103336759A