Image-text association method and related equipment

Through deep learning and natural semantic processing technology, image density areas are extracted and feature fusion is performed. Combined with multimodal model analysis, the problems of insufficient universality and intelligence of graphic and text correlation systems in the existing technology are solved, and efficient and accurate automatic association between drawings and text descriptions are achieved.

CN119942567APending Publication Date: 2025-05-06NEW PRIME NUMBER (BEIJING) DATA CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510030335.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The automated graphic and text association systems in the prior art have low versatility and intelligence, which is difficult to adapt to complex graphic and text processing needs, resulting in insufficient accuracy and efficiency in the correlation between drawings and text descriptions.

Method used

Deep learning and natural semantic processing technology are used to extract the density areas of the image, use different feature extraction models to process image blocks, fuse feature information, and combine multimodal models to analyze the matching degree of image and text descriptions to generate accurate text descriptions.

Benefits of technology

It realizes the intelligent relationship between drawings and text descriptions, improves the matching accuracy of graphic and text content and the efficiency of automated processing, reduces manual intervention, and is suitable for technical document management and content review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942567A_ABST
    Figure CN119942567A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image-text association method and related equipment. The image-text association method comprises the following steps: extracting a first density region and a second density region in an image; the first density area is divided into first blocks, and the second density area is divided into second blocks; processing the first block through a first feature extraction model to obtain first feature information of the image; processing the second block through a second feature extraction model to obtain second feature information of the image; fusing the first feature information and the second feature information to obtain fused feature information; obtaining character description information of the image according to the fused feature information; respectively inputting the image description information and the character description information into a visual encoder and a language encoder of a multi-modal model for processing to obtain the matching degree of the image description information and the character description information; and obtaining description characters of the image according to the character description information of which the matching degree meets the condition, and associating the description characters with the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields of artificial intelligence, image processing and natural language processing, and in particular to a method, device, electronic device, computer-readable medium and computer program product for associating images and texts. Background Art

[0002] In various technical documents (such as design documents, patent application documents, academic papers, etc.), the association between drawings and text descriptions needs to be accurate and correct to ensure the integrity and correctness of the document content. The automatic association system in related technologies usually relies on template matching, which has low versatility and intelligence and is difficult to adapt to complex image and text processing needs. Summary of the invention

[0003] According to one aspect of the present disclosure, a method for associating an image with text is provided, including: extracting a first density area and a second density area in an image; dividing the first density area into a first block, and dividing the second density area into a second block; processing the first block through a first feature extraction model to obtain first feature information of the image; processing the second block through a second feature extraction model to obtain second feature information of the image; fusing the first feature information and the second feature information to obtain fused feature information; obtaining text description information of the image based on the fused feature information; inputting the image and the text description information into a visual encoder and a language encoder of a multimodal model for processing, respectively, to obtain a matching degree between the image and the text description information; obtaining a caption of the image based on the text description information whose matching degree meets a condition, and associating the caption with the image.

[0004] According to one aspect of the present disclosure, a device for associating images and texts is provided, including: an extraction unit for extracting a first density area and a second density area in an image; a segmentation unit for segmenting the first density area into a first block and the second density area into a second block; a processing unit for processing the first block through a first feature extraction model to obtain first feature information of the image; processing the second block through a second feature extraction model to obtain second feature information of the image; a fusion unit for fusing the first feature information and the second feature information to obtain fused feature information; the processing unit is further used to obtain text description information of the image based on the fused feature information; the processing unit is further used to input the image and the text description information into a visual encoder and a language encoder of a multimodal model for processing, respectively, to obtain a matching degree between the image and the text description information; the processing unit is further used to obtain a caption of the image based on the text description information whose matching degree meets a condition, and associate the caption with the image.

[0005] According to one aspect of an embodiment of the present disclosure, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the image-text association method as described in any embodiment of the present disclosure is implemented.

[0006] According to one aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: one or more processors; a storage device configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the image-text association method as described in any embodiment of the present disclosure.

[0007] According to one aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the image-text association method described in any embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 The flowchart of the image-text association method according to an embodiment of the present disclosure is schematically shown.

[0009] Figure 2 The flowchart of a method for associating images and texts according to another embodiment of the present disclosure is schematically shown.

[0010] Figure 3 The flowchart of a method for associating images and texts according to another embodiment of the present disclosure is schematically shown.

[0011] Figure 4 A schematic diagram of a user interface according to an embodiment of the present disclosure is schematically shown.

[0012] Figure 5 A block diagram of a device for associating images and texts according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0013] The disclosed embodiment provides an artificial intelligence-based image-text association system, which applies deep learning and natural semantic processing technology, combines automatic matching and analysis of images and texts, and realizes intelligent association between drawings (or images) and text descriptions (or explanatory texts), that is, realizes intelligent association of image and text content, and can automatically complete tasks such as image recognition, content matching, and description generation, providing an efficient and accurate solution for technical document management and content review, etc. The method steps of the disclosed embodiment can be executed by a terminal device or a server.

[0014] like Figure 1 As shown, the image-text association method provided by the embodiment of the present disclosure may include the following steps.

[0015] In S110 , a first density region and a second density region in the image are extracted.

[0016] The images in the embodiments of the present disclosure refer to images including drawings such as technical drawings, and therefore, in the present disclosure, "drawings" and "images" can be used interchangeably. The drawings or images in the embodiments of the present disclosure may include, but are not limited to, drawings in the specification drawings in the patent application documents, other forms of drawings or various technical documents used to present graphic content such as structures, functions and / or processes, such as design drawings, flow charts, assembly drawings (such as mechanical assembly drawings) and circuit diagrams, academic papers and other drawings. That is, the solutions provided by the embodiments of the present disclosure are universal.

[0017] In technical documents, manually annotating the relationship between drawings and text content involves a large amount of work and a high error rate. The method, device, and system provided in the embodiments of the present disclosure are intended to automatically achieve efficient association between drawings and text in technical documents, so as to improve annotation accuracy, reduce manual intervention, and improve overall work efficiency.

[0018] In the embodiment of the present disclosure, according to the distribution of pixel points such as lines, nodes, and texts of the drawings in the image, it can be divided into different density areas, such as a first density area and a second density area. The first density area is larger or smaller than the second density area.

[0019] In an exemplary embodiment, extracting a first density region and a second density region in an image includes: dividing the image into unit regions; calculating content density of the unit regions; and determining the first density region and the second density region based on the content density.

[0020] In the disclosed embodiment, the image can be divided into multiple (two or more) unit areas according to the actual scene, and then the content density of each unit area is calculated separately, that is, the pixel ratio of the lines, nodes, text and other pixel points to the total pixels of the entire unit area, so as to determine the first density area and the second density area in the image.

[0021] In an exemplary embodiment, calculating the content density of the unit area includes: obtaining a gradient value between different pixels in the unit area, and obtaining the content density of the unit area based on the gradient value between different pixels in the unit area and the total pixels in the unit area; or obtaining first edge information in the unit area, and obtaining the content density of the unit area based on the first edge information in the unit area and the total pixels in the unit area.

[0022] The embodiments of the present disclosure provide a variety of different ways to measure the content density of a unit area, for example, based on the gradient value between pixels within the unit area, or based on edge information within the unit area, which is referred to as the first edge information here to distinguish it from subsequent edge information.

[0023] In an exemplary embodiment, extracting a first density region and a second density region in an image includes: performing denoising and enhancement preprocessing on the image to obtain a preprocessed image; and extracting the first density region and the second density region in the preprocessed image.

[0024] In some embodiments, the image may be directly divided into different density regions. In other embodiments, the image may be preprocessed, such as denoising and / or enhancement, and the preprocessed image may be divided into a first density region and a second density region to increase the accuracy of density region division.

[0025] In an exemplary embodiment, the image is subjected to denoising and enhancement preprocessing to obtain a preprocessed image, including: performing denoising on the image using a Gaussian filter with a kernel size of 5×5 and a standard deviation of 1.0 to obtain a denoised image; performing contrast enhancement on the denoised image using a contrast limited adaptive histogram equalizer with a window size of 8×8 to obtain a contrast enhanced image; and inputting the contrast enhanced image into a U-net model for clarity enhancement to obtain the preprocessed image.

[0026] It can be understood that although the embodiments of the present disclosure use the Gaussian filter, contrast limited adaptive histogram equalizer and U-Net model with the above parameters for image preprocessing as examples, the present disclosure is not limited to this, and other suitable image preprocessing methods and / or other different parameters can also be used for image preprocessing.

[0027] In S120, the first density region is divided into first blocks, and the second density region is divided into second blocks.

[0028] In an exemplary embodiment, dividing the first density area into first blocks and dividing the second density area into second blocks includes: dividing the first density area into the first blocks according to a first size; dividing the second density area into the second blocks according to a second size, wherein the first size is smaller than the second size.

[0029] In some embodiments, the first density area and the second density area may be further divided into smaller first blocks and second blocks according to the first size and the second size, respectively. For example, the first size is adapted to the first density area, for example, the greater the content density of the first density area, the smaller the first size; otherwise, the greater the first size. The second size is adapted to the second density area.

[0030] In an exemplary embodiment, the first density area is divided into the first blocks according to a first size; and the second density area is divided into the second blocks according to a second size, including: extracting second edge information of the image; dynamically adjusting the size and shape of the first size and the second size based on the second edge information; dividing the first density area into the first blocks based on the adjusted first size; and dividing the second density area into the second blocks based on the adjusted second size.

[0031] In some embodiments, the first size of the first density region and the second size of the second density region are fixed. In other embodiments, the first size and the second size can be dynamically adjusted according to the second edge information of the image to enhance the flexibility of the first block and the second block, and the adaptability to the size and shape of the components, nodes, etc. in the drawing.

[0032] In S130, the first block is processed by a first feature extraction model to obtain first feature information of the image; and the second block is processed by a second feature extraction model to obtain second feature information of the image.

[0033] In the disclosed embodiment, the first feature extraction model and the second feature extraction model are both deep learning models capable of extracting features from images. The first feature extraction model and the second feature extraction model are different types of deep learning models or similar deep learning models with different model parameters. By respectively extracting the first feature information of each first block in the first density area and the second feature information of each second block in the second density area through the first feature extraction model and the second feature extraction model, parallel processing of images can be achieved to improve processing efficiency; and adaptive deep learning models can be used for feature extraction of different density areas to improve the accuracy of feature extraction while saving computing resources.

[0034] In an exemplary embodiment, the first feature extraction model includes an initial convolution layer, a maximum pooling layer, a first residual block and a second residual block; the second feature extraction model includes a first inverted residual block and a second inverted residual block, the first inverted residual block and the second inverted residual block respectively include 3 inverted residual modules, each inverted residual module includes a depth-separable convolution and an activation function. The specific structures of the two feature extraction models will be illustrated in the following embodiments.

[0035] In S140, the first feature information and the second feature information are fused to obtain fused feature information.

[0036] In an exemplary embodiment, fusing the first feature information and the second feature information to obtain fused feature information includes: determining a first weight of the first feature information and a second weight of the second feature information; obtaining the fused feature information based on the first feature information and its first weight, and the second feature information and its second weight; or, cascading the first feature information and the second feature information to obtain the fused feature information; or, fusing the first feature information and the second feature information into the fused feature information through a global average pooling operation.

[0037] The disclosed embodiments provide a variety of feature fusion methods, such as weighted sum fusion; cascade fusion, which can directly concatenate the first feature information and the second feature information; or perform a global average pooling operation on the first feature information and the second feature information.

[0038] In an exemplary embodiment, determining a first weight of the first feature information and a second weight of the second feature information includes: obtaining a first complexity of the first block and a second complexity of the second block; and determining the first weight and the second weight based on the first complexity and the second complexity.

[0039] In an exemplary embodiment, obtaining a first complexity of the first block includes: obtaining third edge information of the first block; obtaining the number of edge pixels in the first block based on the third edge information; obtaining an edge density of the first block according to the number of edge pixels in the first block and the total number of pixels of the first block; obtaining texture changes of the first block using a gradient amplitude histogram to obtain texture information density of the first block; detecting text information density of the first block; and obtaining the first complexity according to the edge density, texture information density and text information density of the first block.

[0040] In an exemplary embodiment, obtaining the second complexity of the second block includes: obtaining fourth edge information of the second block; obtaining the number of edge pixels in the second block based on the fourth edge information; obtaining the edge density of the second block according to the number of edge pixels in the second block and the total number of pixels of the second block; obtaining the texture change of the second block using a gradient amplitude histogram to obtain the texture information density of the second block; detecting the text information density of the second block; and obtaining the second complexity according to the edge density, texture information density and text information density of the second block.

[0041] The disclosed embodiment can calculate the complexity of the first block and the second block respectively, and determine the weights of the first feature information and the second feature information based on the complexity, so as to achieve weighted sum fusion of the first feature information and the second feature information, so that the obtained fused feature information is more reasonable and accurate.

[0042] In S150, text description information of the image is obtained according to the fused feature information.

[0043] The text description information in the embodiments of the present disclosure refers to identifying the drawing in the image and generating one or more text descriptions for the drawing based on the identification information, which is used to describe one or more of the function of the drawing, the nodes or components contained therein, the positions of the contained nodes or components in the image, the positional relationships and connection relationships between them, etc. The text description is text used to explain the content contained in the drawing.

[0044] In an exemplary embodiment, obtaining text description information of the image based on the fused feature information includes: inputting the fused feature information into a classification layer to obtain classification results, sequence information, and position information of nodes in the image; organizing the classification results and position information of the nodes into a node sequence based on the sequence information; and inputting the node sequence into a Transformer model to obtain the text description information.

[0045] In some embodiments, the fused feature information may be processed by a classification layer to obtain text description information of the image. In other embodiments, the result output by the classification layer may be further processed by a Transformer model to obtain text description information of the image.

[0046] In S160, the image and the text description information are respectively input into the visual encoder and the language encoder of the multimodal model for processing to obtain the matching degree between the image and the text description information.

[0047] In the embodiments of the present disclosure, the multimodal model refers to a deep learning model that can process different modalities, such as a model that processes text and images simultaneously. In the following embodiments, the multimodal model is illustrated as a CLIP (Contrastive Language-Image Pre-training) model, but the present disclosure is not limited to this. The CLIP model includes a visual encoder and a language encoder to process the image and the text description information respectively to obtain the matching degree between the image and the text description information. The matching degree here can also be referred to as similarity.

[0048] In an exemplary embodiment, the image and the text description information are respectively input into a visual encoder and a language encoder of a multimodal model for processing to obtain a matching degree between the image and the text description information, including: inputting the text description information into a pre-trained language model to obtain semantic information of the text description information; preprocessing the image to obtain a preprocessed image; and inputting the preprocessed image and the semantic information into a visual encoder and a language encoder of the multimodal model to obtain the matching degree.

[0049] The pre-trained language model in the embodiments of the present disclosure refers to a deep learning model that has been pre-trained and can perform natural language processing on text. In the following embodiments, BERT (Bidirectional Encoder Representations from Transformers) or BERT-base model is used as an example, but the present disclosure is not limited to this.

[0050] In some embodiments, the image and text description information may be input into the visual encoder and language encoder of the CLIP model, respectively. In other embodiments, the text description information may be first input into a pre-trained language model for processing to extract semantic information from the text description information, and then the image and the semantic information may be input into the visual encoder and language encoder of the CLIP model, respectively.

[0051] In S170, the caption text of the image is obtained according to the text description information whose matching degree satisfies the condition, and the caption text is associated with the image.

[0052] In an exemplary embodiment, the description text of the image is obtained based on the text description information whose matching degree satisfies the condition, and the description text is associated with the image, including: obtaining description templates for different types of images; determining the type of the image based on the text description information; selecting a corresponding description template according to the type of the image; and obtaining the description text of the image by combining the description template corresponding to the image and the text description information.

[0053] like Figure 2 As shown, the image-text association method provided by the embodiment of the present disclosure may include the following steps.

[0054] In S201, a drawing is input, for example, a technical drawing.

[0055] In S202, data is preprocessed.

[0056] In the disclosed embodiment, before the drawings are identified, image preprocessing or data preprocessing may be performed to obtain a preprocessed image or a preprocessed drawing. The data preprocessing in the disclosed embodiment may include image enhancement and denoising. A number of sample drawings were tested, and a Gaussian filter with a kernel size of 5x5 and a standard deviation of 1.0 was selected. This configuration takes into account both noise removal and retention of edge details. The CLAHE (Contrast Limited Adaptive Histogram Equalization) algorithm (window size 8×8) can be applied to enhance the image contrast. Setting the CLAHE window size to 8x8 can enhance the local detail contrast of lines and nodes in complex drawings (such as circuit diagrams) while avoiding the appearance of artifacts caused by over-enhancement.

[0057] In some embodiments, a U-Net model can be introduced. In response to the needs of technical drawing processing, the embodiment of the present disclosure adopts a shallow structure in U-Net, including 4 downsampling layers and 4 upsampling layers, which is more suitable for drawing processing of medium complexity than the general U-Net model. Two convolutional layers are used in each layer (the convolution kernel size is 3x3, and the activation function is ReLU (Linear rectification function, called rectified linear unit)), which can extract more edges and local features, and help process fine lines in the drawings. The U-Net model eliminates drawing noise and increases drawing clarity while maintaining drawing details. After the input drawing undergoes data preprocessing (such as denoising and enhancement) and then enters U-Net, the interference of input noise on model performance can be further reduced.

[0058] In some embodiments, the U-Net model can be used to segment drawings and extract regions of interest. A classic symmetrical U-shaped network can be used, including 4 layers of downsampling and upsampling. An accurate segmentation mask is generated to ensure that the shape and details of complex parts are preserved intact.

[0059] In some embodiments, data preprocessing also includes adaptive segmentation of drawings. For example, data preprocessing also includes edge detection. In the following embodiments, the Canny edge detection algorithm is used as an example. It can be understood that the edge information of the drawings is not limited to the Canny edge detection algorithm, and other algorithms or models, such as the U-Net model, can also be used. The edge information extracted in the preprocessing stage can be used to determine the first edge information of the unit area, the second edge information of the image, the third edge information of the first block, and the fourth edge information of the second block.

[0060] In the disclosed embodiment, the Canny edge detection algorithm (for example, setting the low threshold to 50 and the high threshold to 150) can be used to accurately obtain the edge information of the image. Assume that the edge information of the image is black pixels and the non-edge is white pixels. After extracting the edge information of the image, the content density of the unit area can be determined based on the ratio of the black pixels (i.e., the first edge information) in the unit area to the total pixels in the unit area. The drawing is divided into a first density area and a second density area based on the content density of the unit area to adjust the segmentation granularity, i.e., determine the first size and the second size. Canny edge detection is used to detect edge details and assist in subsequent feature extraction. Obtain boundary information with clear lines to facilitate segmentation and node detection.

[0061] In some embodiments, assuming that the first density area is a high-density area or a complex density area (such as a line-dense area), the first density area can be divided into 100×100 pixel blocks (i.e., the first size of the first block is 100×100 pixels). Assuming that the second density area is a low-density area, the second density area can be divided into 200×200 pixel blocks (i.e., the second size of the second block is 200×200 pixels). Experiments show that 100×100 pixel block segmentation improves the accuracy of image-text matching while retaining details, while 200×200 pixel block segmentation improves computational efficiency. The preprocessed image data is divided into blocks of different sizes (such as 100x100 or 200x200 pixels).

[0062] In other embodiments, the size of the small block can be adaptively adjusted according to the shape and size of the edge.

[0063] In the disclosed embodiment, adaptive edge segmentation can divide complex drawings into smaller, easier-to-process areas (small blocks), which can be used for local feature enhancement and local recognition in subsequent steps, and further enhance key structures (such as borders and lines) in small blocks to facilitate accurate feature extraction. Different areas in complex drawings (such as text areas and frame areas) may require different processing methods. The algorithm can be adjusted in a targeted manner through the regionalization operation after segmentation. For example, the first block uses the first feature extraction algorithm, and the second block uses the second feature extraction algorithm. In addition, segmentation can also be used to reduce the interference of the background noise of the entire drawing on the local feature extraction, that is, to achieve noise isolation. At the same time, after the drawing is divided into small blocks, different areas can be processed in parallel to improve processing efficiency.

[0064] In the disclosed embodiment, Canny edge detection can quickly extract clear edges and generate a preliminary edge map in a classic image processing manner. This method is very efficient when processing simple lines and boundaries, and is particularly suitable for technical drawings with clear structures. In addition, the edges generated by Canny can be used to guide regional segmentation, such as further determining the candidate area of ​​the step box, so as to achieve adaptive adjustment of the first size and / or second size of the small block. At the same time, Canny edge detection can be used as an auxiliary feature of the input Resnet model (which can be used as a first feature extraction model) to help the network focus on key areas more quickly and improve edge recognition accuracy. Deep learning models such as the first feature extraction model (such as ResNet) can extract edge features and shape information (contained in the first feature information and the second feature information) of the step box in the drawing. The goal of the ResNet model is deep feature extraction, focusing on global context and complex structures, but may not be as sensitive to fine edge features in the drawing (such as thin frame lines or broken lines) as the Canny algorithm.

[0065] The synergy between Canny edge detection and deep learning models such as ResNet can significantly improve the accuracy and robustness of drawing feature extraction, ensuring stable performance when processing complex technical drawings. Canny edge detection provides preliminary edge information for generating candidate regions or clarifying boundary features in the data preprocessing stage. Deep learning models such as ResNet further extract deep edge and shape information for complex feature analysis.

[0066] In S203, the region is divided into blocks.

[0067] After acquiring the edge information of the image, the granularity of the segmentation is dynamically adjusted according to the content density of the image. In some embodiments, the content density can be estimated by calculating the proportion of edge pixels in the image. For example, by counting the proportion of non-zero pixels (i.e., edge pixels) in the image, the content density of the image can be roughly estimated. According to the calculated content density, a threshold is set to determine whether the image is a high-density or low-density area. Then, based on this judgment, different segmentation granularities (e.g., a first size and a second size) are selected. Then, the original image is segmented according to the selected segmentation granularity.

[0068] Before dividing the drawing into blocks, calculate the content density of each area or each unit area (such as the density of lines), and adjust the size of the segmented blocks according to the different content densities. For example: high-density areas (such as fine lines in the circuit) are divided into small blocks (such as 100x100 pixels). Low-density areas (such as blank areas or large components) are divided into large blocks (such as 200x200 pixels). That is, adaptive segmentation is adopted: small blocks are used for details in high-density areas, and large blocks are used for low-density areas, so as to improve efficiency while ensuring detail retention.

[0069] In the embodiment of the present disclosure, the content density in the drawing processing can be defined as the complexity of the pixel information in a certain local area (eg, unit area). The specific calculation formula depends on the requirements.

[0070] In some embodiments, content density can be calculated based on pixel changes. Content density measures the change in pixel intensity within an area, and can be calculated using the following formula:

[0071]

[0072] In the above formula, D represents content density (degree of density); N represents the total number of pixels in the area (i.e., the total number of pixels in the unit area), and N is a positive integer greater than or equal to 1; R represents the target area, which can be a unit area or a small block after segmentation. It represents the modulus of pixel gradient, that is, the absolute value of the gradient value between different pixels, reflecting the rate of change of pixel intensity. i and j are the i-th pixel and the j-th pixel in R, and i and j are positive integers greater than or equal to 1 and less than N. This formula can be used to calculate a higher density value, that is, a higher content density, for areas with larger gradients (such as areas with dense lines and text).

[0073] In other embodiments, the content density may be calculated based on the structural complexity. If the number of specific features (such as lines, nodes, components, etc.) or the number of feature points or the number of edge feature points or features is used as an indicator, the content density may be calculated using the following formula:

[0074]

[0075] The number of feature points can be determined by the number of edge points after Canny edge detection (included in the first edge information) or the number of straight lines obtained by Hough transform. The area of ​​the region is the area (number of pixels) of the image segmentation region or small block, that is, the total number of pixels per unit area. This method is more structured and suitable for scenes with dense distribution of high-level features (such as frame lines and primitives).

[0076] In the disclosed embodiment, the high-density area is characterized by: large gradient changes, many edge feature points; concentrated distribution of features, and complex structure in the area. Exemplary criteria for determining a high-density area can be: based on gradient changes: content density D>0.5 (determined based on the normalized gradient value range of 0-1), then it is determined to be a high-density area; or, based on the number of features: the number of feature points in the area accounts for D>50 points / 100 pixels, then it is determined to be a high-density area.

[0077] In the disclosed embodiment, the low-density area is characterized by small gradient changes and sparse edge feature points. The structure in the area is simple (such as a blank area or a simple line). The criteria for determining a low-density area can be exemplified as follows: based on gradient changes: if the content density D is less than 0.2, it is determined to be a high-density area; or, based on the number of features: if the number of feature points in the area accounts for D less than 10 points / 100 pixels, it is determined to be a high-density area. In the disclosed embodiment, the portion between the above-mentioned high and low density areas can be divided into medium-density areas and adaptively adjusted according to needs.

[0078] In some embodiments, in image segmentation, fixed-size blocks (such as the first block and the second block) may cause the shape of the component in the drawing to be inconsistent with the shape of the segmented block. Especially for large components or components with irregular shapes, the segmented blocks may cut into the boundaries of the components, making subsequent feature extraction difficult. Therefore, an edge-based adaptive segmentation scheme can more effectively solve this problem.

[0079] Specifically, the outlines and boundaries of each component or node in the drawing are first detected through Canny edge detection or U-Net segmentation network. Based on this boundary information or edge information, the size and shape of the segmentation area or segmentation block are dynamically adjusted to avoid segmenting large components or components that cross the boundary into multiple small blocks. For complex drawings (such as circuit diagrams), fine-tuning can be performed near the edges to avoid incorrect segmentation of components.

[0080] In the disclosed embodiment, the high-density area is subdivided into smaller blocks to ensure that local features are accurately processed. The low-density area is coarsely divided into larger blocks to avoid wasting computing resources on blank areas. In addition, the block size and / or shape is dynamically adjusted according to the density range and / or edge information to optimize processing efficiency and accuracy.

[0081] In S204, feature extraction is performed.

[0082] In the embodiment of the present disclosure, the first feature extraction model used for high-complexity areas or high-complexity areas (i.e., first density areas) may be, for example, a ResNet model. The second feature extraction model used for low-complexity areas or low-complexity areas (i.e., second density areas) may be, for example, a MobileNet-v2 model. However, the present disclosure is not limited to this. The combined network combines the residual connections of ResNet and the depthwise separable convolutions of MobileNet-v2. Such a design can enhance the model's ability to extract features from complex images while maintaining computational efficiency.

[0083] In the embodiment of the present disclosure, in the drawing content recognition stage, nodes (such as circuit elements, flow boxes), boundary features and logical relationships in the drawing are identified. ResNet (residual network) extracts local features in the drawing, detects node categories (such as resistors, capacitors), and determines the node's boundary box and position. MobileNet-v2 (lightweight convolutional network) is used for fast feature extraction in resource-constrained environments. It provides efficient feature extraction capabilities with low computing resource consumption as a supplement to ResNet.

[0084] In S205, features are fused.

[0085] In some embodiments, taking a patent drawing containing three process step boxes as an example, the system extracts the edge features and shape information of each step box through the combined network of ResNet and MobileNet-v2. Then, the outputs of ResNet and MobileNet-v2 are feature fused to obtain fused feature information, and then the fused feature information is processed through the Softmax layer (classification layer) to obtain the probability distribution of each node classification. For example, for nodes "S101", "S102" and "S103", the system selects the best classification result based on the highest probability value, and combines the position and order to build a complete node topology. "S101" and so on are combined with the entire drawing to identify whether it is an initial step, an intermediate step or an end step. If it is a circuit diagram, the classification results may include the size type, component type, material, processing method, position (upper left, etc.), status (installed, to be installed, etc.) of the component.

[0086] In some embodiments, after the feature extraction module (including the first feature extraction model and the second feature extraction model), a global average pooling layer is used to reduce the dimension of the feature map, that is, a global average pooling operation is performed on the first feature information and the second feature information to obtain fused feature information while retaining key information.

[0087] In some embodiments, one or more fully connected layers may be connected after the global average pooling layer to further integrate and process the features extracted from the feature extraction module.

[0088] In some embodiments, a Softmax layer is connected after the fully connected layer to classify the input image. The Softmax layer outputs a probability distribution indicating the likelihood that the image belongs to each category.

[0089] In the disclosed embodiment, the classification result is determined based on the probability distribution of the model. First, the node image features are identified through ResNet and MobileNet-v2. Then a classification probability distribution is output through the classification layer, for example, P (class resistance) = 0.75, P (class capacitance) = 0.20, P (class inductance) = 0.05. According to the highest probability value (such as resistance 75%), the classification result is selected as resistance.

[0090] In some embodiments, if the highest value of the probability distribution is below a certain threshold (eg, 0.6), it may be marked as "undetermined" for manual verification.

[0091] In the embodiment of the present disclosure, for the circuit diagram node, the classification result is the type of electronic components, which may include the following categories: 1) Basic components: resistor (Resistor); capacitor (Capacitor); inductor (Inductor); diode (Diode); transistor (Transistor); integrated circuit (IC, Integrated Circuit); switch (Switch); relay (Relay); fuse (Fuse), etc.; 2) Connection points and ports: power node (Power Node); ground node (Ground Node); input port (Input Port); output port (Output Port); 3) Special symbols: test point (Test Point); connection symbols (such as jumpers or crosspoints), etc.

[0092] In the disclosed embodiment, for the nodes in the flowchart, the classification results may include: step box (ProcessStep); decision node (Decision Point); start node (Start Node); end node (End Node), etc. In the specific implementation, the classification results can be further expanded or refined to support user-defined categories or field-specific classification standards. For complex drawings, custom nodes are supported.

[0093] In the disclosed embodiment, after identifying nodes such as "S101", "S102", and "S103", the system establishes a topological relationship between the nodes by calculating the spatial relative position of each node, thereby ensuring the logical coherence of the drawing content and providing accurate context information for subsequent image and text matching.

[0094] In the disclosed embodiment, the position and order information of the nodes can be directly extracted through the ResNet and MobileNet-v2 networks. When ResNet and MobileNet-v2 recognize drawings, they extract the bounding box information (Bounding Box) and classification results of the nodes, and the coordinates of these bounding boxes can be directly used to represent the positions of the nodes. The order of the nodes can be determined by sorting the coordinates of the node bounding boxes, for example, from left to right by the x-coordinate (horizontal coordinate), or from top to bottom by the y-coordinate (vertical coordinate).

[0095] In the disclosed embodiment, the complete node topology or topological relationship may include: the positional relationship between nodes and the connection relationship between nodes. Among them, the positional relationship between nodes describes the spatial distribution of each node in the drawing, such as relative coordinates, distance or arrangement order. This relationship can be directly derived from the geometric position of the node. The connection relationship between nodes describes the logical connection between nodes, such as components in a circuit diagram are connected by wires, and components in a mechanical diagram are connected by assembly interfaces. The identification of the connection relationship can be combined with the line features in the drawing and the functional information of the node to construct a more complex relationship chain.

[0096] In the embodiment of the present disclosure, the calculation of the spatial relative position of the node depends on the features of the identified node and the geometric information of the drawing. The specific steps of implementing the spatial relative position calculation are described below by way of example.

[0097] Nodes (such as "S101", "S102", "S103") extracted from the drawing content recognition, including their bounding boxes or pixel-level segmented areas, are obtained to obtain the center point coordinates or other geometric centers of each node. For example: S101: (x1, y1); S102: (x2, y2); S103: (x3, y3). The drawing can be taken as the origin in the upper left corner, and the coordinate system is a two-dimensional plane (x, y). The positions of all nodes will be expressed in this coordinate system.

[0098] In the embodiments of the present disclosure, there may be multiple methods for calculating the spatial relative position between nodes. In some embodiments, the Euclidean distance between any two nodes may be calculated to represent their relative spatial spacing, which may be used to analyze the proximity, density, etc. of the nodes. In other embodiments, the relative direction between two nodes may be determined based on angle and direction. Angles are used to represent the relative orientation between nodes. In some applications, directions may be quantified into discrete east, south, west, north, or other partition directions. In yet other embodiments, an adjacency list may be used to construct node relationships, and nodes and their relative positions may be represented by a graph structure.

[0099] After calculating the relative positions of all nodes, the complete node topology can be further constructed. The adjacency matrix A can be constructed first. The adjacency matrix A represents the direct connection relationship between nodes: if two nodes are directly connected, the corresponding distance d is recorded. Otherwise, it is marked as infinity (inf) or 0. Then, combined with the connection rules of specific fields (such as circuit diagrams), determine which nodes should be connected: if the distance d < a certain threshold (which can be set according to the specific scenario), the nodes are considered connected. Or the node relationship can be determined by semantic information (such as node name or type).

[0100] In the disclosed embodiment, OpenCV can be used for image recognition and location extraction of nodes. NumPy is used for matrix calculations, such as adjacency matrix construction and distance calculation. NetworkX is used to construct and visualize the topological relationship diagram of nodes.

[0101] In the disclosed embodiment, the constructed complete node topology structure includes the following contents: Node position: coordinates (x, y) of the center point of each node. Node relative distance and direction: distance d and azimuth angle θ of each pair of nodes. Connection relationship: represented by adjacency matrix or adjacency table. This information can be used as input data for subsequent tasks (such as automated analysis of circuit diagrams and design verification).

[0102] In the drawing recognition task, the topological structure can be represented in the form of a graph, which includes the following elements:

[0103] 1. Components of topological structure: Node: Node is the key point in the diagram, corresponding to the important elements in the drawing, such as: electrical components (resistors, diodes, etc.) in the circuit diagram; step boxes or decision boxes in the flowchart; places or intersections in the map. Edge: Edge represents the relationship or connection between nodes, corresponding to: wires in circuit diagrams; flow lines (arrows or connecting lines) in flowcharts; roads or paths in maps.

[0104] 2. Representation, including: Mathematical form: The topological structure is represented as an undirected or directed graph: G = (V, E). V represents a set of nodes, each node has specific attributes (position, type, etc.). E represents a set of edges, each edge has connection attributes (direction, weight, etc.). Data form: Adjacency matrix: suitable for small-scale graphs, recording the connection relationship between each pair of nodes. Adjacency list: suitable for sparse graphs, listing the direct neighbors of each node.

[0105] 3. The steps of constructing the topological structure include: Node extraction: Use deep learning models (such as ResNet, U-Net, MobileNet-v2) or image-text association methods to identify key nodes in the drawings and record their: spatial coordinates (such as pixel coordinates); types (such as resistors, diodes, or step boxes in flowcharts, etc.). Edge detection and connection extraction: Use edge detection algorithms (such as Canny) to identify the locations of connecting lines in the drawings. And / or, combine morphological analysis or deep learning models to extract connection information. Connection relationship inference: Determine the connection between nodes through geometric or rule parsing. The connection relationship can be judged based on the detected lines. Alternatively, in a graphical representation, nodes in close proximity may be interconnected. Topological map generation: Organize the nodes and connection relationships into a topological map and store it as a data structure.

[0106] 4. Characteristics of topological structures, including: Spatial position relationship: In the topological graph, the position of the nodes retains spatial geometric information, which can be used to restore the drawing layout or generate visualization. Connectivity: Describes the overall structure of the system and helps analyze the relationship between components, such as: the path of the circuit; the logical flow of the process. Scalability: The topological structure can be applied to the graph algorithm (such as the shortest path, connected component analysis) for subsequent analysis.

[0107] In S206, global modeling is performed.

[0108] In an exemplary embodiment, the output results of the combined network of ResNet and MobileNet-v2 can be sequentially input into the Transformer model. For example, the category information of the node image is first extracted by ResNet and MobileNet-v2 (such as "S101" is a "step box", "S102" is a "decision node", etc.). Then, the spatial position and topological order of the node (such as the layout order from left to right, from top to bottom, etc.) are calculated. After that, the image features of the nodes in the drawing (such as "S101", "S102", "S103") and their sequence information are sorted as the input of the Transformer. The category of each node (such as a step box, a decision node) and its order in the topological structure are encoded as a sequence. For example [(S101, category 1, position 1), (S102, category 2, position 2), (S103, category 3, position 3)]. These data are converted into an embedding vector (embedding), and position encoding is added to retain the sequence information.

[0109] In the embodiment of the present disclosure, the main function of the Transformer model is to analyze the logical relationship and contextual meaning of the input node sequence and output the following content:

[0110] 1. Predict logical relationships: determine the functional or logical connections between nodes. For example: S101 → S102 means "the process connection from step 1 to step 2. S103 is an independent step and may require special processing.

[0111] 2. Generate topological structure: Output a complete node relationship chain or topological structure diagram. For example: Output: (S101→S102→S103). The functional relationship of each node is predicted and annotated by the model, which facilitates the subsequent automatic generation of description documents or description texts.

[0112] 3. Classification and recommendations: If there is a conflict in the order of nodes (such as position and logic mismatch), the model will mark the problem area. Possible classification results are recommended for uncertain nodes for manual review or further processing.

[0113] 4. Information extraction at the semantic level: In addition to topological information, the Transformer model can also extract the semantic meaning of node sequences (such as a complete process).

[0114] That is, the data input to Transformer includes information such as node category, order, and position extracted from ResNet and MobileNet-v2. The output is the logical relationship chain between nodes, the complete topological structure, and the recommended results of the predicted classification. This provides support for subsequent image-text matching and document generation.

[0115] In the disclosed embodiment, the node position and order are extracted through the ResNet network and MobileNet-v2 in the aforementioned steps, and the topological structure is preliminarily established. Afterwards, the topological relationship between the nodes is re-established through the Transformer model for more comprehensive logical analysis and verification. ResNet and MobileNet-v2 identify the nodes in the drawings (such as S101, S102, etc.). The direct connection or simple proximity relationship between the nodes is determined by the edge features or geometric relationships (such as Euclidean distance, sequential arrangement) output by ResNet and MobileNet-v2. The topological structure constructed at this stage is a connection relationship at the geometric / visual level, and pays more attention to proximity in physical space.

[0116] After constructing the preliminary geometric topological relationship, the system needs to further refine and verify the logical coherence of the drawing content. This is to take into account the needs of high-level logical analysis. The drawing content is not only a physical connection, but may also contain semantic logical relationships (such as the direction of current in a circuit diagram, the order of steps in a flowchart). The preliminary topological structure focuses on spatial adjacency, but cannot fully express the semantic rules between nodes. For example, in a flowchart, nodes A and B may be adjacent, but the logical order needs to establish a directed relationship from A to B based on the direction of the arrow or text description. This requires re-analyzing the topological structure based on context or other rules.

[0117] In addition, the geometric topology extracted by ResNet and MobileNet-v2 may be misconnected, especially when the distance between nodes is close but there is no actual connection. When the topological relationship is re-established through the Transformer model, errors can be corrected through semantic information (such as drawing annotations, connection symbols) or global rules (such as drawing specifications). For example, in a circuit diagram, two circuit elements may be initially identified as connected because of their close distance, but there is no wire symbol to actually connect them. In this case, a logical verification process is required to eliminate the misconnection.

[0118] The Transformer model in the disclosed embodiment can be used to refine node relationships. The preliminary topology may only contain simple "connected or not" relationships, but a complete drawing analysis may require the construction of a more complex topology. For example: directed graphs (clearly defining the order direction between nodes); weighted graphs (adding weights to represent properties between nodes, such as connection strength, distance, or transfer value); hierarchical structures (building hierarchical relationships based on node types).

[0119] Specifically, re-establishing the topological relationship may include:

[0120] (1) Drawing semantic analysis: Use NLP (Natural Language Processing) models (such as Transformer) to analyze semantic information such as text descriptions and symbolic tags of nodes to supplement logical coherence.

[0121] (2) Node verification and connection rule judgment: According to the predefined connection rules (such as drawing standards), it is determined whether there is a logical connection between nodes. The rules may include: circuit symbol type matching, connection direction constraints, connection mark verification, etc.

[0122] (3) Global topology optimization: Based on the preliminary topology, a more accurate logical topology is constructed: removing incorrect connections; adding missing connections; and adjusting the topological directionality or hierarchical relationship.

[0123] (4) Finally, a complete topology is generated. The output topology structure contains both geometric and logical relationships, ensuring that it can express both the physical location and reflect the logical coherence.

[0124] In the disclosed embodiment, the topological relationship initially constructed by ResNet and MobileNet-v2 focuses on physical position and geometric adjacency, and mainly relies on image features. The topological relationship is re-established through the Transformer model, focusing on the logical level, combining the semantics, rules and global consistency verification of the drawings to ensure the integrity and correctness of the topology. These two stages cooperate with each other to form a complete process from low-level image analysis to high-level logical reasoning. The Transformer model captures global information in the drawings. Identify the order and contextual relationships of nodes (such as the execution order of process boxes).

[0125] In the disclosed embodiment, the Transformer model can be used to process long sequence data (such as the order of multiple process steps). If the combined network identifies a long sequence, such as a dozen steps, the combined network is not sufficient to capture the characteristics of the long sequence. At this time, it is necessary to further use the Transformer model to process the long sequence, that is, the output of the combined network is used as the input of the Transformer model to identify the classification results, positions and order of each node. The Transformer model processes sequence data through the Self-Attention Mechanism and Positional Encoding. It can capture the relationship between any positions in the sequence, so it performs well in processing long sequence data. If the combined network identifies a short sequence, such as two steps, then the Transformer model is not needed.

[0126] In S207, the image and text are matched.

[0127] In the disclosed embodiment, when associating graphic and text content, semantic analysis and feature extraction are performed first. In an exemplary embodiment, BERT-base (e.g., 12 layers, 768 hidden units, 12 attention heads) is used to perform word segmentation, part-of-speech tagging, and semantic analysis on the text description information. The feature vector captures the main semantic information of the drawing description, laying the foundation for subsequent graphic and text matching. The text description information here describes the basic purpose of the drawing, such as describing that this is a flow chart. The BERT-base model is used to determine whether the generated text description information and its semantics match the image.

[0128] In image-text matching, the semantic feature vector of the text description information is compared with the feature vector of the drawing content (extracted by the image encoder (i.e., visual encoder) of ResNet or CLIP) to determine whether the semantics are consistent. For example, if a text description mentions "resistors connected to capacitors form an RC circuit", the system needs to verify whether there is a corresponding drawing representation.

[0129] BERT-base can measure semantic relevance, that is, assist in the calculation of matching degree. It is used to calculate the semantic similarity between text and images to ensure the logical consistency of drawings and descriptions.

[0130] The disclosed embodiments also relate to cross-modal learning. The CLIP model is a multimodal pre-trained neural network model that is pre-trained using a large amount of paired image and text data to learn the alignment relationship between images and text.

[0131] In some embodiments, the CLIP model is used for semantic matching of drawings and texts. The image and text (text description information) are respectively input into the image encoder and text encoder (i.e., language encoder) of the CLIP model to generate, for example, a 512-dimensional feature vector. The matching degree is determined by calculating the cosine similarity of the image and text feature vectors.

[0132] In some embodiments, through experimental settings, when the cosine similarity of the image and text feature vectors exceeds a first threshold (e.g., 0.85), the match is determined to be successful; for results below a second threshold (e.g., 0.5), the system marks them as manually reviewed to ensure matching accuracy. Compared with traditional matching methods, after the CLIP model sets a cosine similarity threshold of 0.7 (i.e., the first threshold mentioned above), the system's misjudgment rate is reduced by 20%, and the overall accuracy of image-text matching reaches more than 95%. No manual review is required between 0.85 and 0.5, indicating that the parameters of the CLIP model need to be adjusted, and / or the input data needs to be adjusted.

[0133] For example, three process boxes (S101, S102, S103) are identified on the drawing, and the corresponding step descriptions are generated as their text description information. The system confirms the correspondence between the image and the text through the cosine similarity of the CLIP model. The system matches the first identified box with "S101", the second with "S102", and so on, to ensure the accuracy and consistency of the matching of each step. This verifies the recognition accuracy of the model.

[0134] In the disclosed embodiment, CLIP's cross-modal contrastive learning capability can also assist the training of ResNet, MobileNet-v2, and Transformer, making the output drawing features more consistent with the text description features. In the disclosed embodiment, CLIP can also be used for error correction. If there is ambiguity between the drawing and the text description, the comparison of feature vectors can identify potential contradictions and prompt revisions.

[0135] In the disclosed embodiments, the text description information may also be referred to as the drawing description, which refers to the text description of the drawing. The feature vector extracted by BERT-base is mainly used to analyze the semantics of the text and provide support for image-text matching and semantic consistency measurement. The vector generated by CLIP's text encoder is specifically used to compare with the output of the image encoder to ensure image and text alignment. The CLIP model extracts the feature vectors of the image and text respectively through its image encoder and text encoder, and then calculates the cosine similarity to determine the degree of match (i.e., matching degree) between the two. If the similarity exceeds the set first threshold (such as 0.85), the image-text pairing is considered to be successful.

[0136] The text description information extracted by the embodiment of the present disclosure may also be referred to as drawing information, and may include the following contents:

[0137] (1) Node information of the graph, including: Nodes represent components. Through ResNet, MobileNet-v2 and Transformer models, the system has identified and classified various components in the circuit diagram (such as resistors, capacitors, switches, etc.). Each node information includes: Component category: type of component (such as resistors, capacitors); Component position: spatial coordinates or position vector of the component in the drawing; Component attributes: functional parameters related to the component (such as resistance value, capacitance value).

[0138] (2) Edge information of the graph: Edges represent the logical or functional relationship between components. The system obtains the position information of the connection lines in the circuit diagram through Canny edge detection or U-Net segmentation. Initial establishment of edges: Based on the connected lines, it is inferred which nodes have relationships and initial values ​​are assigned to these relationships.

[0139] (3) Additional features, including: semantic associations between components: the functional descriptions of the associated components are extracted using CLIP or text features; global features of the drawings: global circuit structure information extracted using the Transformer model.

[0140] Then, formatting is performed to obtain a graph structure G = (V, E), where V represents a node set, representing circuit components, and each node has a feature vector. E represents an edge set, representing possible connection relationships between components, and each edge has a preliminary weight.

[0141] In S208, a description document is generated.

[0142] In the disclosed embodiment, the NLG (Natural Language Generation) module generates a description document or description text based on the template (i.e., description template) and the above identification information (i.e., text description information), ensuring that the generated content is consistent with the drawing content to avoid ambiguity. The model identification information is not easy for users to understand, and through the template, it is converted into a description document that users can understand.

[0143] In the disclosed embodiment, templates are designed for different types of diagrams to realize automatic generation of description documents, so that after parsing the content of the drawings, the system can output clear and logically rigorous text descriptions according to the preset structure. The following uses circuit diagrams and mechanical assembly drawings as examples to illustrate the specific design of the template.

[0144] In some embodiments, a circuit diagram template is designed. The circuit diagram description template includes the following parts:

[0145] 1. Title: Drawing name and number.

[0146] 2. Description of main components: type and quantity of main components in the figure.

[0147] 3. Connection relationship description: the logical relationship between key nodes and paths.

[0148] 4. Functional description: Overview of the circuit function.

[0149] 5. Notes: Matters that require special attention during use or assembly.

[0150] The following is an example template:

[0151] Title: Circuit Diagram Drawing Number: #E12345

[0152] Main components description:

[0153] This drawing contains the following components:

[0154] Resistors: 10 (labeled as R1-R10);

[0155] Capacitors: 5 (labeled as C1-C5);

[0156] Diodes: 2 (labeled as D1 and D2).

[0157] Connection relationship description:

[0158] The circuit diagram is divided into three parts according to the signal transmission path:

[0159] The input signal enters the amplifier circuit after being limited by R1 and R2;

[0160] In the amplifier circuit, C1 is used for filtering, and R3-R5 adjusts the gain;

[0161] The output end is rectified and protected by D1-D2.

[0162] Functional description: This circuit is designed for signal amplification and rectification and is suitable for low-frequency signal processing.

[0163] Note: When soldering components, please ensure that the polarity components (such as capacitor C1, diode D1) are in the correct direction.

[0164] In some embodiments, a mechanical assembly drawing template is designed. The description template of the mechanical assembly drawing includes the following parts:

[0165] 1. Title: The name and number of the assembly drawing.

[0166] 2. Parts list: Label all the parts required for assembly.

[0167] 3. Assembly steps: Explain how to assemble each part in order.

[0168] 4. Final inspection item: quality inspection point after assembly is completed.

[0169] 5. Warning information: matters that may affect the safety of assembly or use.

[0170] The following is an example template:

[0171] Title: Mechanical Assembly Instructions Drawing Number: #M56789

[0172] Parts List:

[0173] This assembly drawing contains the following parts:

[0174] Base plate (No.: P001), 1 piece;

[0175] Fixed rod (No.: P002), 2 pcs;

[0176] Connecting bolts (No.: S001), 4 pcs;

[0177] Motor assembly (No.: M001), 1 pc.

[0178] Assembly steps:

[0179] 1. Place the base plate (P001) on a flat workbench;

[0180] 2. The fixing rod (P002) is installed on both sides of the base plate by bolts (S001);

[0181] 3. Place the motor assembly (M001) on top of the fixing rod, making sure the motor holes are aligned;

[0182] 4. Use a torque wrench to tighten the bolts to the specified torque.

[0183] Final inspection items:

[0184] Make sure all bolts are tightened;

[0185] Confirm that there is no looseness between the motor assembly and the base plate;

[0186] Check whether the overall structure is stable.

[0187] Warning: During installation, be careful not to over-tighten the bolts to avoid damaging the threads.

[0188] In the disclosed embodiment, dynamic parameters are reserved when designing the template, such as `{component type}` and `{connection relationship description}`, which are filled in by the system according to the drawing parsing results. The template is divided into independent modules (such as component description, connection relationship, precautions, etc.), which are dynamically called according to different drawing types. Templates can be used to automatically generate documents, reduce the time of manual writing of instructions, and improve efficiency; achieve standardized output, ensure that the format of the generated drawing instructions is unified, and easy to understand and review; enhance readability, and make the instructions clearer and more intuitive through the logical structure of the template.

[0189] In the disclosed embodiments, the text description is local text information generated during the drawing parsing process, which is used to describe the content of a specific part of the drawing (such as the type or location of a component). The text description can be a subset of the description document. The description document is the complete document that is finally output, which is generated based on the template organization, covers all the key information of the drawing, has more complete logic, and is more standardized in language.

[0190] In the disclosed embodiment, a diagram description can be automatically generated. Templates are designed for different types of diagrams (such as circuit diagrams, mechanical assembly drawings), and the templates contain the names, functions, positions, and relationship descriptions of the components. The template is a general template. For example, for a circuit diagram, it will describe which components are there, what the sizes of the components may be, what materials are there, etc.; descriptions of possible positions and relationships. In a circuit diagram, a "node" refers to a component or functional key point in a circuit. In the disclosed embodiment, the type of node refers to the category of circuit components, such as resistors, capacitors, diodes, transistors, etc. The position of the node is used to mark the specific location (such as geometric coordinates) of the component in the circuit diagram. The function of the node may be further refined, such as the specific resistance value of the resistor, the capacity of the capacitor, etc., depending on the annotation details of the training data. The template can be in JSON (JavaScript Object Notation, a lightweight data exchange format) format.

[0191] In S209, the final document is output.

[0192] In the disclosed embodiments, the final output document may be the above-mentioned explanatory document, or it may be a drawing description after further processing of the explanatory document. Drawing description refers to the text description corresponding to the drawing, such as the description section in the patent document, which is used to explain the content in the drawing. These text descriptions have a one-to-one correspondence or complementary relationship with the drawing content, specifically including: specific functional descriptions of drawing elements (such as components, nodes, connection relationships); explanation of the overall layout of the drawing.

[0193] In the disclosed embodiment, a drawing image is input into the model, and the output data, i.e., the final document, may include: descriptive text of the drawing (such as the technical description of the corresponding patent); a description of the function or modularization of the nodes in the drawing; and content analysis of a specific structure (such as text within a box).

[0194] The processing steps of ResNet or Transformer enhance image clarity and feature expression, making up for the shortcomings of CLIP under low resolution or complex structures. Second, the capture of complex semantic relationships is insufficient. In specific scenarios (such as the relationship between nodes and sequences in drawings), Transformer or other models (such as BERT) are required to perform semantic analysis first to provide CLIP with more logical feature vectors. The disclosed embodiment optimizes the matching effect of CLIP on domain-specific data (such as technical drawings and explanatory text) in response to high-precision matching requirements.

[0195] Figure 3 The flow of drawing input to the combinatorial network for processing is described.

[0196] In S501, an image is inputted. Before the image is inputted into the combination network, it may be standardized first, for example, the size of the standardized image data is adjusted to 224×224×3, but the present disclosure is not limited thereto.

[0197] In S502, a preprocessed image is obtained after preprocessing.

[0198] After the drawings are preprocessed (such as denoising, segmentation, etc.), the preprocessed drawings (circuit diagrams, mechanical drawings, etc.) are input into the combined network for processing. In the disclosed embodiment, region detection can be performed. The drawings are globally scanned to calculate the edge pixel density (i.e., content density) within the unit area. By comparing the density threshold (which can be set according to the actual scene) with the content density, the image is automatically divided into high-density areas and low-density areas.

[0199] In S503, a high-density area (eg, a circuit area containing complex circuits) is input into a ResNet model to obtain ResNet features (ie, first feature information).

[0200] In the disclosed embodiment, the low-level features of the drawing (such as lines, shapes, contours, etc.) are extracted through the first few layers of ResNet. A residual structure is used here, that is, the output of each layer is added to the input of the previous layer, thereby avoiding the gradient vanishing problem. ResNet is used to extract the global features of the image, especially the main information such as edges and shapes, and introduces residual blocks to reduce the gradient vanishing problem in deep network training.

[0201] Exemplarily, the initial convolution layer and the first two groups of residual blocks of ResNet-50 are used. Among them, the initial convolution layer is a 7×7 convolution with a step size of 2. The maximum pooling layer has a step size of 2. Residual block 1 (i.e., the first residual block) contains 3 convolution layers (1×1, 3×3, 1×1). Residual block 2 (i.e., the second residual block) contains 4 convolution layers (1×1, 3×3, 1×1). ResNet-50 outputs the extracted global feature map as the first feature information.

[0202] In S504, low-density areas (such as blank areas or simple annotations) are input into the MobileNet-v2 model to obtain MobileNet-v2 features.

[0203] In the disclosed embodiment, MobileNet-v2 performs fine feature extraction. MobileNet-v2 includes an inverted residual block (Inverted Residual Block), which includes a 3×3 depth-separable convolution and a ReLU6 activation function. For example, two groups of inverted residual blocks (a first inverted residual block and a second inverted residual block) are used, each group containing three inverted residual modules. Each inverted residual module gradually reduces the number of feature map channels (by reducing the dimension layer by layer). The output of MobileNet-v2 is a refined feature map as the second feature information.

[0204] The disclosed embodiment uses a combined network of ResNet and MobileNet-v2 to enhance the ability to identify the structure and details of complex parts. The network structure includes multiple convolutional layers (convolution kernel size 3×3, step size 1), and is classified by a Softmax layer at the end. ResNet is used to extract rough features, and MobileNet-v2 is used to extract more detailed features. In some embodiments, the ResNet part of the combined network uses the first few layers of ResNet (convolutional layers and downsampling layers) to extract deep features of the drawings. These features contain detailed information in the drawings and help identify the basic elements of the drawings (such as lines, edges, nodes, etc.). The low-density area is passed to MobileNet-v2 for processing. The depth-separable convolution of MobileNet-v2 is used to reduce computational complexity while maintaining a high feature expression capability. The low latency characteristics of MobileNet-v2 make it particularly suitable for providing fast response in drawing content recognition and adapting to drawings of different layouts. Through the regional blocking strategy, the input drawings are divided into high-density and low-density areas to dynamically adjust the network selection. At the same time, parallel processing of images can also be achieved.

[0205] In some embodiments, the image resolution and content density can be adapted to dynamically adjust the resolution. For example, for high-density areas, small blocks with higher resolution (e.g., 512x512) are input to avoid losing details due to low resolution. For low-density areas, small blocks with lower resolution (e.g., 256x256) are input to increase processing speed.

[0206] In S505, feature fusion is performed through a feature fusion layer.

[0207] In some embodiments, the output feature maps of ResNet and MobileNet-v2 are fused into a vector through global average pooling to form an image feature vector as fused feature information.

[0208] In other embodiments, concatenation fusion is achieved through a feature fusion layer, and the output features of the two networks, ResNet and MobileNet-v2, can be directly concatenated to form an image feature vector as fused feature information, thereby forming a richer feature vector.

[0209] In some other embodiments, the feature fusion layer can realize attention fusion, for example, dynamically assigning weights (including first weights and second weights) through the Attention mechanism, for example, enhancing the weights of ResNet features for complex areas, that is, increasing the first weights. Exemplarily, the features of both can be weighted and synthesized through a learning layer, retaining the powerful deep feature extraction capability of ResNet while also retaining the computational efficiency of MobileNet-v2.

[0210] In the disclosed embodiment, a feature fusion module (such as a concatenation or attention mechanism) or a feature fusion layer is used to merge the features of ResNet and MobileNet-v2. The contribution of the two to the final feature representation can be adjusted through a weight mechanism. For example, the following formula can be used to calculate the fused feature information F final :

[0211] F final =α·F ResNet +(1-α)·F MobileNet-v2 (3)

[0212] Among them, α is dynamically adjusted according to the regional complexity, that is, the first complexity. ResNet Indicates the first feature information; F MobileNet-v2 represents the second feature information. (1-α) represents the second complexity.

[0213] In the disclosed embodiment, the dynamic adjustment of α is to assign feature weights according to the complexity of the region. The regional complexity (the first complexity is used as an example, and the calculation of the second complexity can be referred to) can be calculated by the following indicators: edge density, texture information density and regional text information density (i.e. text information density).

[0214] In some embodiments, the Canny edge detector is used to extract the edge of the region or small block. The number of edge pixels in the region or small block (such as a 100x100 or 200x200 small block) is counted and normalized to a complexity score C. edge , which is the edge density. For example, if the total number of pixels in the region or small block is 10,000 and the number of edge pixels is 1,000, then C edge =0.1.

[0215] In some embodiments, a gradient magnitude histogram is used to count the texture changes in a region. The higher the score, the more complex the texture. texture , that is, the texture information density.

[0216] In some embodiments, OCR technology is used to detect the number or distribution of characters in the region. text (such as the density of circuit markings), that is, the density of regional text information.

[0217] In the embodiment of the present disclosure, the regional complexity C can be calculated by the following formula: region :

[0218] C region =ω1×C edge +ω2×C texture +ω3×C text (4)

[0219] Among them, the weights ω1, ω2, and ω3 can be adjusted through experiments or task requirements.

[0220] In the embodiment of the present disclosure, α can be adjusted by complexity. region The value of α is determined, and the range of α is [0,1]. Among them, α≈1 represents a high complexity area, which is more dependent on ResNet, that is, the first weight is larger. α≈0 represents a low complexity area, which is more dependent on MobileNet-v2.

[0221] For example, it can be determined by the following formula:

[0222]

[0223] C max , C min Represents the minimum and maximum values ​​of the complexity score, obtained through training data or image analysis.

[0224] In the disclosed embodiment, the value of α is linearly mapped to [0, 1] and is used to dynamically adjust the weighted contribution of the two features. In actual implementation, the adjustment of α can be operated by the following method:

[0225] (1) Online calculation of α: During the forward propagation process, the α value is calculated based on the complexity of each region. The α values ​​of different regions are dynamically input into the feature fusion module.

[0226] (2) Offline partition setting α: The complexity score C of all regions is calculated in the preprocessing stage. regionThe complexity is mapped to a discrete α value interval (e.g., the first weight of the high-complexity area is fixed to 0.8, and the second weight of the low-complexity area is fixed to 0.2) for subsequent weighted fusion.

[0227] In S506, classification is performed through a classification layer.

[0228] In the disclosed embodiment, the features of ResNet and MobileNet-v2 are combined to form a feature vector through a feature fusion strategy (such as concatenation or weighted merging). A fully connected layer is used to further compress the features. After the fully connected layer, an output layer such as Softmax classification is used to generate prediction results. For example, the classification results of each node, component or drawing are output.

[0229] In S507, the result is output.

[0230] In the embodiment of the present disclosure, the output content after Softmax classification may include: node identification results, such as the elements in the circuit diagram (such as resistors, capacitors, wires, etc.) are identified and classified; component relationship output, for example, after identifying each component, combined with the topological relationship between the components, the connection information between the components is output.

[0231] Figure 4 The following schematic diagram shows a user interface according to an embodiment of the present disclosure. Figure 4 As shown, the user interface 600 provided in the embodiment of the present disclosure includes a drawing preview area 601, an explanation document display area 602, an editing tool area 603, and an operation button area 604. This is a simplified Web interface structure for illustration. In the embodiment of the present disclosure, a Web interface is provided, which has been used for real-time feedback mechanism and structured output. Allow users to adjust the generated explanation document. User feedback will be stored in the database for optimization of the model and generated template to improve the long-term performance of the system.

[0232] Figure 4The user interface 600 is an interactive interface for users to manually intervene and adjust based on the logical relationship chain generated by the system (contained in the description document). The description document display area 602 is used to display the currently generated description document. For example, the description document is displayed by chapter or segment (such as drawing information, main components, connection relationships, functional descriptions, etc.). The user can click on a specific paragraph or field to enter the editing mode. The drawing preview area 601 is used to display the drawings corresponding to the description document. It supports zooming in, zooming out, and dragging to view the details of the drawings. The components on the drawings can be highlighted, and the associated document content can be viewed or modified by clicking on the highlighted area. When the drawing preview area 601 is used to display a circuit diagram, it can also be called a circuit diagram display area, which is used to display the circuit diagram parsed by the system, with annotated component nodes and connection relationships. The circuit diagram display area supports zooming and panning. Exemplarily, the user interface 600 can also display a relationship chain topology view, which displays nodes (components) and their logical relationships (edges) in the form of a graph structure. The relationship chain topology view supports dragging and adjusting nodes and edges. Exemplarily, the user interface 600 may also provide an operation panel to provide functions for adding, deleting and modifying nodes and edges. The operation panel displays the detailed properties of the selected element and provides modification options. Exemplarily, the user interface 600 may also include tool options, such as adding, deleting and modifying properties of connecting lines. The editing tool area 603 is used to provide editing, proofreading and adjustment tools for the content of the description document. Supports text editing, automatic grammar checking, and inserting new content. Provides style editing for specific terms and formats (such as tables and lists). The operation button area 604 is a quick entry for users to perform operations. The buttons can be used to perform: saving changes, undoing operations, submitting for manual review, and downloading generated description documents (such as PDF or Word format). Exemplarily, the user interface 600 may also include a log and feedback area to record the user's modification operations; real-time feedback on the impact of adjustments on the topological structure.

[0233] In the disclosed embodiment, the model structure can also be optimized by adjusting the network layer. In an exemplary embodiment, the convolution kernel in the model can be adjusted to a non-square convolution kernel (e.g., 1x5) to capture the linear characteristics of the circuit diagram. For another example, the pooling operation is reduced in high-density areas to retain more detailed information.

[0234] In some embodiments, the model structure used in the drawing content recognition stage is not limited to the Transformer model. For example, an LSTM (Long Short-Term Memory) model can be used to process long sequence data (such as continuous process steps) in the drawings.

[0235] The image-text association device of the embodiment of the present disclosure may be arranged on a terminal device, or may be arranged on a server, or may be arranged partly on a terminal device and partly on a server. Figure 5 As shown, the image-text association device 700 provided in the embodiment of the present disclosure includes an extraction unit 710, a segmentation unit 720, a processing unit 730 and a fusion unit 740. The extraction unit 710 is used to extract a first density area and a second density area in an image. The segmentation unit 720 is used to segment the first density area into a first block and the second density area into a second block. The processing unit 730 is used to process the first block through a first feature extraction model to obtain the first feature information of the image; and process the second block through a second feature extraction model to obtain the second feature information of the image. The fusion unit 740 is used to fuse the first feature information and the second feature information to obtain fused feature information. The processing unit 730 is also used to obtain text description information of the image according to the fused feature information. The processing unit 730 is also used to input the image and the text description information into the visual encoder and the language encoder of the multimodal model for processing respectively to obtain the matching degree of the image and the text description information. The processing unit 730 is also used to obtain the caption text of the image according to the text description information whose matching degree meets the condition, and associate the caption text with the image. The specific implementation of each module in the image-text association device provided by the embodiment of the present disclosure can refer to the content of the above-mentioned image-text association method, which will not be repeated here.

Claims

1. A method for associating images and texts, characterized in that: include: Extracting a first density region and a second density region in the image; dividing the first density region into first blocks, and dividing the second density region into second blocks; Processing the first block through a first feature extraction model to obtain first feature information of the image; processing the second block through a second feature extraction model to obtain second feature information of the image; fusing the first feature information and the second feature information to obtain fused feature information; Obtaining text description information of the image according to the fused feature information; Inputting the image and the text description information into a visual encoder and a language encoder of a multimodal model for processing respectively, to obtain a matching degree between the image and the text description information; The caption text of the image is obtained according to the text description information whose matching degree satisfies the condition, and the caption text is associated with the image.

2. The method according to claim 1, characterized in that Extracting a first density region and a second density region in an image, including: Dividing the image into unit areas; Calculating the content density of the unit area; determining the first density region and the second density region based on the content density; The step of calculating the content density of the unit area includes: Obtaining a gradient value between different pixels in the unit area, and obtaining a content density of the unit area based on the gradient value between different pixels in the unit area and the total pixels in the unit area; or, Obtain first edge information within the unit area, and obtain content density of the unit area based on the first edge information within the unit area and total pixels of the unit area; The step of dividing the first density region into a first block and dividing the second density region into a second block comprises: Divide the first density area into the first blocks according to a first size; divide the second density area into the second blocks according to a second size; Wherein, the first size is smaller than the second size.

3. The method according to claim 1, characterized in that The step of fusing the first feature information and the second feature information to obtain fused feature information includes: Determine a first weight of the first feature information and a second weight of the second feature information; obtain the fused feature information based on the first feature information and its first weight, and the second feature information and its second weight; or, Cascading the first feature information and the second feature information to obtain the fused feature information; or, The first feature information and the second feature information are fused into the fused feature information through a global average pooling operation.

4. The method according to claim 3, characterized in that Determining a first weight of the first feature information and a second weight of the second feature information includes: Obtaining a first complexity of the first block and a second complexity of the second block; determining the first weight and the second weight based on the first complexity and the second complexity; The step of obtaining the first complexity of the first block includes: Obtaining third edge information of the first block; obtaining the number of edge pixels in the first block based on the third edge information; obtaining the edge density of the first block according to the number of edge pixels in the first block and the total number of pixels in the first block; Using a gradient magnitude histogram to obtain a texture change of the first block, and obtaining a texture information density of the first block; detecting the text information density of the first block; The first complexity is obtained according to the edge density, texture information density and text information density of the first block.

5. The method according to claim 1, characterized in that Obtaining text description information of the image according to the fused feature information includes: Inputting the fused feature information into a classification layer to obtain classification results, sequence information and position information of nodes in the image; Organizing the classification results and position information of the nodes into a node sequence according to the sequence information; The node sequence is input into a converter model to obtain the text description information.

6. The method according to claim 1, characterized in that Inputting the image and the text description information into a visual encoder and a language encoder of a multimodal model for processing respectively, and obtaining a matching degree between the image and the text description information, including: Inputting the text description information into a pre-trained language model to obtain semantic information of the text description information; Preprocessing the image to obtain a preprocessed image; The preprocessed image and the semantic information are respectively input into a visual encoder and a language encoder of the multimodal model to obtain the matching degree.

7. A device for associating images and texts, characterized in that: include: An extraction unit, used for extracting a first density region and a second density region in the image; a segmentation unit, configured to segment the first density region into a first block, and segment the second density region into a second block; a processing unit, configured to process the first block through a first feature extraction model to obtain first feature information of the image; and process the second block through a second feature extraction model to obtain second feature information of the image; a fusion unit, configured to fuse the first feature information and the second feature information to obtain fused feature information; The processing unit is further used to obtain text description information of the image according to the fused feature information; The processing unit is further used to input the image and the text description information into the visual encoder and the language encoder of the multimodal model for processing, respectively, to obtain the matching degree between the image and the text description information; The processing unit is further configured to obtain the caption text of the image according to the text description information whose matching degree satisfies a condition, and associate the caption text with the image.

8. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. An electronic device, characterized in that: include: one or more processors; A storage device configured to store one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Purchase return risk early warning and management method and system based on multi-mode intelligent auditing, electronic equipment and computer readable storage medium

    CN120220158A