A schematic diagram object detection method and system based on multi-granularity text reasoning
Through the multi-grained fusion of global text inference and local text inference, the problem of sparse underlying features and high-level semantic confusion in schematic object detection is solved, and the detection accuracy is improved.
Patent Information
- Application Number
- CN202211144391.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-20
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-09-20
AI Technical Summary
The existing object detection algorithm is difficult to effectively apply to the schematic diagram expressing abstract diagrams, especially in the problems of semantic sparseness of underlying features and confusion of high-level semantics, which leads to the low accuracy of the schematic object detection.
Using a multi-grained text reasoning method, the target detection effect is improved through the fusion of global text and local OCR labeled text. Specific steps include global text inference, local text inference and multi-grained fusion. Using Bert pre-trained language model, bidirectional GRU network, OCR algorithm and convolutional neural network and other technologies, the text visual map is constructed and feature enhancement and screening are performed.
It effectively enhances the global and local context characteristics of the diagram, eliminates the influence of noisy text nodes, and improves the target detection accuracy.
Smart Images

Figure CN115393694B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of schematic diagram detection, and particularly relates to a schematic diagram object detection method and system based on multi-granularity text reasoning. Background Art
[0002] In recent years, schematic diagrams Figure 1 have been used to depict knowledge concepts, logical information, work processes, etc. of courses. From the illustrations in textbooks to the images on knowledge websites, schematic diagrams are becoming more diverse and growing continuously, and have occupied a large proportion of visual data, widely distributed in MOOC websites, open knowledge bases, technical forums, and encyclopedia websites. Such images are usually composed of simple geometric shapes such as lines, rectangles, circles, and text annotations, clearly expressing a specific theme or concept within the course, transmitting inferable logical rules or logical information, and presenting elements using abstract graphical symbols rather than real images. Applying object detection algorithms to analyze and understand such special images is the basis for tasks such as knowledge base construction and intelligent question answering, and is also an important part of cross-media intelligence.
[0003] Object detection is a key research direction in the field of computer vision, aiming to find the objects of interest in the image and determine the location and category of the objects at the same time. In recent years, the research on object detection algorithms has mainly focused on natural images, and there are a large number of studies on public image datasets. However, there are significant differences in the features between natural images and schematic diagrams in the course field, and related methods are difficult to be applied to schematic diagrams. There are also a small number of studies on schematic diagrams, such as the research on visual question answering methods on the AI2D and TQA datasets. However, the images included in such datasets are mostly images in primary school natural science courses, such as food chains, life cycles, and human physiology. The expression forms of these images are closer to natural images, and the processing methods are the same as those of natural images, without considering that the underlying visual features and high-level semantic expressions of schematic diagrams are different from those of natural images.
[0004] Therefore, it is necessary to propose an object detection algorithm for schematic diagram datasets with abstract expressions. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a schematic diagram object detection method and system based on multi-granularity text reasoning in view of the deficiencies in the above-mentioned prior art, and use global text and local OCR marked text to improve the object detection effect, so as to solve the technical problems of sparse semantics of the underlying features and confused high-level semantics of schematic diagrams.
[0006] The present invention adopts the following technical solutions:
[0007] A schematic diagram object detection method based on multi-granularity text reasoning includes the following steps:
[0008] S1. Integrate the text features and image features of the schematic diagram to obtain visually enhanced features with text features Complete global text reasoning;
[0009] S2. Based on the enhanced visual features obtained in step S1 Extract visual nodes, extract text nodes according to the characteristics of the schematic diagram itself, use the extracted text nodes and visual nodes of the schematic diagram as graph nodes, construct edges based on the relative position space and text semantic similarity between the text nodes and visual nodes, and perform fine-grained fusion of text features and visual features to obtain enhanced visual node features N v en , complete local text reasoning;
[0010] S3. Extract global text keywords; use the similarity between the text nodes obtained in step S2 and the global text keywords to filter out effective local text nodes; perform multi-grained fusion of the effective local text nodes, the global text reasoning obtained in step S1, and the local text reasoning obtained in step S2 to complete schematic diagram object detection.
[0011] Specifically, step S1 is as follows:
[0012] S101. Use the Bert pre-trained language model to perform semantic encoding on the global description text T i g Perform semantic encoding on the obtained text encoding Input it into a bidirectional GRU network to obtain semantic information in both forward and backward directions, and splice them to obtain the overall representation of the global text
[0013] S102. Input a single image d i Into the convolutional neural network ResNet, after four stages of convolutional pooling operations, obtain the overall image feature v g ;
[0014] S103. Use a transformation matrix to map the overall global text representation obtained in step S101 and the overall image feature v obtained in step S102 g Into the same feature space, then through splicing and summing, and then fuse the corresponding vectors to obtain visually enhanced features with text features Send the enhanced visual features Into the RPN and ROI networks for classification and regression operations to complete global text reasoning.
[0015] Specifically, in step S101, an update gate and a reset gate are added to the bidirectional GRU network to retain the features of the text through the update gate and the reset gate.
[0016] Specifically, step S2 is as follows:
[0017] S201, a single schematic diagram d in the CSDQA dataset i Input into the easyocr algorithm, get all the text information and its corresponding position information in the diagram, a total of M OCR marks; all the OCR marks detected in the diagram are regarded as text nodes, and the node features are The text vector and position vector in the text node are transformed by linear matrix and fused after normalization to obtain the feature vector of the text node.
[0018] S202, a single schematic diagram d in the CSDQA dataset i The first four stages of the convolutional neural network ResNet101 are input to extract the feature map; the feature map is then input into the RPN network to obtain the foreground anchor box and its position offset, and then the foreground anchor box and its position offset are combined as the candidate region. Then, the candidate regions with an area smaller than the specified threshold and exceeding the boundary are eliminated, and the remaining candidate regions are subjected to non-maximum suppression to form accurate candidate regions. Finally, the ROI layer is obtained. The ROI layer receives the original feature map and the candidate region output by the RPN network, and maps the candidate region back to the original schematic diagram d i , and then perform maximum pooling to obtain visual node N v The regional feature vector
[0019] S203, the text node N obtained in step S201 t and the visual node N obtained in step S202 v As Figure d i The similarity between text nodes and visual nodes is regarded as an edge, and the feature vector of text nodes and the feature vector of visual nodes are mapped to the same feature space using the transformation matrix. Then, the cosine function is used to calculate the semantic similarity Sim between text nodes and visual nodes. semantic ; Combined with semantic similarity Sim semantic and position relationship Sim position That is, the spatial semantic similarity Sim between the mth text node and the nth visual node is obtained mn ;
[0020] S204: Similarity Sim between the text node and the visual node obtained in step S203 mn , all visual node features N v The corresponding text node N vt Perform splicing and fusion to obtain the enhanced visual node N v en , completing local text reasoning.
[0021] Further, in step S201, the schematic diagram d i The text obtained in is input into the Bert pre-trained language model to obtain the text vector x m , and the position vector y is extracted from the position information of the text box m .
[0022] Further, in step S202, the RPN network is divided into two branches: classification and regression. The classification branch distinguishes foreground from background, and the regression branch refines the anchor box positions to determine the position offsets; the anchor box positions are the anchors generated for each point of the feature map, with aspect ratios of {1:1, 1:2, 2:1}.
[0023] Specifically, step S3 is specifically as follows:
[0024] S301. Select the StanfordCoreNLP toolkit to perform part-of-speech tagging on the global description text, and select the nouns and adjectives among them as keywords;
[0025] S302. Calculate the similarity between the keywords in the global text obtained in step S301 and the text nodes obtained in step S2 to determine the screening mechanism for local text nodes;
[0026] S303. Integrate the global text reasoning obtained in step S1 and the local text reasoning obtained in step S2, and add the screening mechanism for local text nodes obtained in step S302.
[0027] Further, in step S301, the global description text T i g is input into the toolkit, and after part-of-speech tagging, word screening, and word encoding, the keyword vector N is obtained key .
[0028] Further, in step S302, the screening mechanism for local text nodes is as follows: According to the similarity Sim m between the keywords in the global text obtained in step S301 and the local text nodes, the noise in the local text nodes is removed, and the effective local text nodes are screened out.
[0029] In a second aspect, an embodiment of the present invention provides a schematic diagram object detection system based on multi-granularity text reasoning, including:
[0030] A global text reasoning module that fuses the text features and image features of the schematic diagram to obtain visually enhanced text features to complete global text reasoning;
[0031] The local text reasoning module, according to the enhanced visual features obtained by the global text reasoning module Extract visual nodes, extract text nodes according to the characteristics of the schematic diagram itself, use the extracted text nodes and visual nodes of the schematic diagram as graph nodes, construct edges according to the relative position space between the text nodes and visual nodes and the text semantic similarity, and fuse the text features and visual features in a fine-grained manner to obtain the enhanced visual node feature N v en , and complete local text reasoning;
[0032] The fusion detection module extracts global text keywords; filters out effective local text nodes using the similarity between the text nodes obtained by the local text reasoning module and the global text keywords; performs multi-grained fusion on the effective local text nodes, the global text reasoning obtained by the global text reasoning module, and the local text reasoning obtained by the local text reasoning module to complete the schematic diagram object detection.
[0033] Compared with the prior art, the present invention has at least the following beneficial effects:
[0034] A schematic diagram object detection method based on multi-grained text reasoning, including global text reasoning, local text reasoning, and multi-grained fusion reasoning; through global text reasoning, it provides a global knowledge background for the schematic diagram, effectively enhances the semantic features of the global context, and thus avoids the problem of sparse underlying visual features; through local text reasoning, it provides the knowledge background of the local area of the schematic diagram, enhances the local context features, and avoids the problem of high-level semantic confusion; through the multi-grained fusion reasoning module, it simultaneously enhances the semantic features of the global and local contexts, eliminates the influence of noisy text nodes on the detection effect, and improves the schematic diagram object detection accuracy.
[0035] Furthermore, the global text reasoning introduces the global description text of the schematic diagram, encodes it into a text feature vector, and then performs unified fusion with the visual feature vector of the schematic diagram, enhancing the context semantic information of the visual feature vector, thereby improving the object detection effect..
[0036] Furthermore, update gates and reset gates are added to the bidirectional GRU network to retain the features of the text through the update gates and reset gates; determine what information to discard and what new information to add, and determine the degree of discarding previous information.
[0037] Furthermore, local text reasoning includes three parts: text node generation, visual node generation, and text-visual graph construction. It mainly uses OCR marking information to provide context semantic information for the local area of the schematic diagram, thereby improving the object detection effect.
[0038] Further, in step S201, the text vectors and position vectors marked by OCR are fused, taking into account both the semantic information and spatial information of the OCR marks, thus obtaining a feature vector with higher richness.
[0039] Further, the RPN network is divided into two branches: classification and regression. The classification branch is used to distinguish the schematic diagrams and give the predicted scores for each schematic diagram, and the regression branch is used to fine-tune the selected regions.
[0040] Further, since the CSDQA dataset is automatically crawled from multiple source websites, there are inevitably some OCR marks such as websites and advertisements in the schematic diagrams. In addition, not all OCR marks can effectively enhance the features of the visual regions, that is, there is noise in the local text nodes corresponding to the OCR marks, which has a negative impact on the local text reasoning of the schematic diagrams. However, step S3 can effectively reduce the input noise and improve the target detection accuracy through global text keyword extraction.
[0041] Further, most of the words directly related to the local visual regions in the global text are adjectives or nouns. Therefore, part-of-speech tagging technology is first used to select the keywords in the global text, and then encoding is performed, which can play a guiding role in screening the local text nodes.
[0042] Further, according to the similarity Sim between the keywords in the global text and the local text nodes m Remove the noise in the local text nodes and screen out the effective local text nodes, which can effectively reduce the input noise and improve the target detection accuracy.
[0043] It can be understood that the beneficial effects of the second aspect can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here.
[0044] In summary, the method of the present invention provides a global knowledge background for schematic diagrams through global text reasoning, enhances the global context features; applies the OCR mark information in the schematic diagrams through local text reasoning, and enhances the context features of the local visual regions according to the text visual graph; inherits the global reasoning and local reasoning modules through multi-granularity fusion reasoning, and adds a screening mechanism of global text for local text; improves the target detection accuracy of schematic diagrams.
[0045] The technical solutions of the present invention will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is the schematic diagram target detection framework based on multi-granularity text reasoning of the present invention;
[0047] Figure 2An example of the text visual graph of the present invention;
[0048] Figure 3 An example of the schematic diagram target detection data set. Specific embodiments
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0050] In the description of the present invention, it should be understood that the terms "including" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0051] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0052] It should be further understood that the term " / and" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally indicates that the contextually related objects have an "or" relationship.
[0053] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present invention to describe preset ranges, etc., these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0054] Depending on the context, as used herein, the term "if" can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".
[0055] Various structural schematic diagrams according to the disclosed embodiments of the present invention are shown in the accompanying drawings. These figures are not drawn to scale, where for the purpose of clear expression, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures, as well as their relative sizes and positional relationships, are merely exemplary and may deviate in practice due to manufacturing tolerances or technical limitations. Those skilled in the art can additionally design regions / layers with different shapes, sizes, and relative positions according to actual needs.
[0056] A schematic diagram is generally composed of simple geometric shapes such as lines, rectangles, circles, etc., and text annotations, clearly expressing a specific theme or concept within a course, transmitting inferable rules or logical information, and presenting elements using abstract graphical symbols rather than real images. A schematic diagram object detection dataset is mainly a dataset of schematic diagrams for object detection tasks.
[0057] Please refer to Figure 3 , which are typical examples of 10 categories of images in the CSDQA dataset, namely arrays, binary trees, deadlocks, flowcharts, network topologies, linked lists, non-binary trees, queues, stacks, and graphs. It can be seen that these images are all related images under the computer discipline. Although these schematic diagrams lack complex texture and color information, they contain rich rules and logical information. For example, the connection between the head node and the tail node in a binary tree, and the connection between the top element and the popped element in a stack. These schematic diagrams often appear together with knowledge concepts and can effectively help learners understand the knowledge concepts in the course.
[0058] The present invention provides a schematic diagram object detection method based on multi-granularity text reasoning, introducing global and local text to enhance the semantic information of the schematic diagram. In the global text reasoning stage, the global knowledge background of the schematic diagram is provided, enhancing the global context features; in the local text reasoning stage, the OCR marked information in the schematic diagram is applied to construct a text visual graph, and accordingly, the context features of the local visual region are enhanced; in the multi-granularity fusion reasoning stage, a screening mechanism of global text for local text is added.
[0059] Please refer to Figure 1 , a schematic diagram object detection method based on multi-granularity text reasoning according to the present invention, includes the following steps:
[0060] S1. Global text reasoning
[0061] Using global text to provide a global knowledge background for the schematic diagram enhances the global context features; it mainly includes three sub - modules: text feature extraction, image feature extraction, and global text reasoning. In the text feature extraction module, first, the global text is input into the Bert pre - trained language model to obtain word embeddings, and then the text features are obtained through a bidirectional GRU network. In the image feature extraction module, after the image is input into the convolutional neural network Resnet101, image features are obtained through operations such as convolutional pooling in four stages. The global text reasoning module maps the text features and image features to the same dimension, fuses the two features, and thus completes global text reasoning.
[0062] The global text reasoning stage needs to complete text feature extraction, visual feature extraction, and global text reasoning, as follows:
[0063] S101. Use the Bert pre - trained language model to perform semantic encoding on the global description text, and then input the obtained text encoding into the bidirectional GRU network to obtain semantic information in both forward and backward directions. After splicing, the overall representation of the global text is obtained.
[0064] The bidirectional GRU network introduces the concepts of update gates and reset gates. The cooperation of the update gate and the reset gate can control how much information from the previous step is remembered, updated, and reset. Through these two gate functions, the important features of the text are retained, avoiding the problem of a large amount of information loss during long - distance propagation. The bidirectional GRU network considers semantic information in both forward and backward directions, and the semantic representation is more accurate.
[0065] S102. Input a single image d i into the convolutional neural network ResNet. After four stages of convolutional pooling operations, the overall image feature v g ;
[0066] S103. Use a transformation matrix to map the overall representation of the global text obtained in step S101 and the overall image feature obtained in step S102 into the same feature space. Then, through splicing and summing, and fusing the corresponding vectors, the visual feature enhanced by text features is obtained. Send the enhanced visual feature into the RPN and ROI networks for classification and regression operations, that is, global text reasoning.
[0067] S2. Local text reasoning
[0068] The OCR text in the schematic diagram is used to enhance the contextual features of the local area of the schematic diagram. It mainly includes four sub-modules: text node generation, visual node generation, text visual graph construction, and local text reasoning. In the text node generation module, the text and the corresponding position box in the schematic diagram are identified by the OCR algorithm, and then the text and position are encoded to form a text node. In the visual node generation module, the image is input into the detection network, and after the ROI operation, all regional features are intercepted to form a visual node. In the text visual graph construction module, text nodes and visual nodes are regarded as graph nodes, and edges are constructed according to the relative position space of text nodes and visual nodes and the semantic similarity of text; finally, under the guidance of the text visual graph, text features and visual features are fine-grainedly fused to complete local text reasoning.
[0069] The local text reasoning stage includes text node generation, visual node generation, text-visual graph construction, and local text reasoning, as follows:
[0070] S201, a single schematic diagram d in the CSDQA dataset i Input into the easyocr algorithm, and get all the text information and its corresponding position information in the diagram, a total of M OCR marks; all the OCR marks detected in the diagram are used as text nodes, node features It is represented by two vectors, namely the text vector and the position vector. The text vector and the position vector are transformed by a linear matrix, normalized and fused to obtain the feature vector of the text node.
[0071] Among them, the text vector x m Is the text The position vector y obtained by inputting the Bert pre-trained language model m =[x l ,y t ,x r ,y b ]The position information of the text box Draw obtained.
[0072] S202, a single schematic diagram d in the CSDQA dataset i The first four stages of the convolutional neural network ResNet101 are input to extract the feature map; the feature map is then input into the RPN network to obtain the foreground anchor box and its position offset, and then the foreground anchor box and its position offset are combined as the candidate region. Then, the candidate regions with an area smaller than the specified threshold and exceeding the boundary are eliminated, and the remaining candidate regions are subjected to non-maximum suppression to form an accurate candidate region without interference values. Finally, the ROI layer is obtained. The ROI layer receives the original feature map and the candidate region output by the RPN network, and maps the candidate region back to the original schematic diagram d i, then perform max pooling to obtain the visual node N v regional feature vector of
[0073] Among them, the RPN network is divided into two branches: classification and regression. The classification branch distinguishes foreground and background, and the regression branch refines the position of the anchor box to determine the position offset; among them, the anchor box is the anchors generated for each point of the feature map, and the aspect ratios are three types: {1:1, 1:2, 2:1}.
[0074] S203. Consider the text node and the visual node as nodes in the graph, and consider the similarity between the text node and the visual node as an edge, which is obtained according to the relative position relationship and semantic similarity characteristics between the nodes;
[0075] In terms of calculating semantic similarity, since the feature vector spaces of the text node and the visual node are different, first use a transformation matrix to map the feature vectors of the text node and the visual node into the same feature space, and then use the cosine function to calculate the semantic similarity Sim semantic between the text node and the visual node; combining the semantic similarity Sim semantic and the position relationship Sim position to obtain the spatial semantic similarity Sim mn between the m-th text node and the n-th visual node.
[0076] Please refer to Figure 2 . The left box represents the visual node, the right box represents the text node, and the connection line between the text node and the visual node represents the similarity between the two. The thickness of the connection line represents the similarity score between the text and the visual node. It can be seen from the figure that the text node "front" has the highest similarity with the rectangular box with the internal text "1" and the lowest similarity with the rectangular box with the internal text "4".
[0077] S204. Map the text and visual nodes to the same spatial dimension, and then perform targeted visual region feature enhancement according to the similarity, and then complete the classification and regression operations.
[0078] According to the text-visual graph, obtain the text node feature set N vt corresponding to each visual node. Based on this, the feature representation of the visual node can be enriched. The specific method is to splice and fuse all visual node features with their corresponding text node features N vt to obtain the enhanced visual node features which can be sent to the subsequent fully connected network for classification and regression.
[0079] S3. Multi-granularity Fusion Inference
[0080] Add a screening mechanism of global text for local text to complete the schematic diagram object detection. It includes three parts: global text keyword extraction, local text node screening, and multi-granularity fusion inference.
[0081] For global text keyword extraction, use the StanfordCoreNLP tool to perform part-of-speech tagging and select nouns and adjectives as keywords; for the local text node screening part, use the similarity between local text and keywords to screen out valid text nodes; the multi-granularity fusion inference part fuses global text inference and local text inference; in the multi-granularity fusion inference stage, add a screening mechanism of global text for local text to complete the schematic diagram object detection; it includes global text keyword extraction, local text node screening, and multi-granularity fusion inference, specifically as follows:
[0082] S301. Select the StanfordCoreNLP toolkit to perform part-of-speech tagging on the global description text, and select nouns and adjectives as keywords;
[0083] The StanfordCoreNLP toolkit is an open-source tool developed by the Stanford NLP Group, which provides a Python interface and has been widely used in natural language processing tasks. Specifically, input the global description text T i g into the toolkit, and after part-of-speech tagging, word screening, and word encoding, the keyword vector can be obtained.
[0084] S302. Calculate the similarity between the keywords in the global text obtained in step S301 and the local text nodes, and remove the noise in the local text nodes according to the similarity comparison;
[0085] After obtaining the keywords in the global description text, calculate the similarity Sim m between the keywords in the global text and the local text nodes. According to the similarity comparison, the noise in the local text nodes can be removed, so as to screen out valid local text nodes for local text inference.
[0086] S303. Integrate global and local text inferences, and add a screening mechanism for local text nodes.
[0087] In another embodiment of the present invention, a schematic diagram object detection system based on multi-granularity text inference is provided. This system can be used to implement the above-mentioned schematic diagram object detection method based on multi-granularity text inference. Specifically, the schematic diagram object detection system based on multi-granularity text inference includes a global text inference module, a local text inference module, and a fusion detection module.
[0088] Among them, the global text reasoning module fuses the text features and image features of the schematic diagram to obtain visually enhanced features with enhanced text features to complete global text reasoning;
[0089] The local text reasoning module, based on the enhanced visual features obtained by the global text reasoning module extracts visual nodes, extracts text nodes according to the features of the schematic diagram itself, uses the extracted text nodes and visual nodes of the schematic diagram as graph nodes, constructs edges according to the relative position space between the text nodes and visual nodes and the text semantic similarity, and finely fuses the text features and visual features to obtain enhanced visual node features N v en to complete local text reasoning;
[0090] The fusion detection module extracts global text keywords; filters out valid local text nodes using the similarity between the text nodes obtained by the local text reasoning module and the global text keywords; performs multi-granularity fusion on the valid local text nodes, the global text reasoning obtained by the global text reasoning module, and the local text reasoning obtained by the local text reasoning module to complete the target detection of the schematic diagram.
[0091] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0092] The present invention conducts experiments on the CSDQA dataset, compares the multi-granularity reasoning model with the baseline algorithm with added global text introduced in Subsection 2.3.2, and uses the average detection precision AP, AP 50 and AP 75 under different intersection over union (IoU) thresholds as evaluation metrics. The experimental results are shown in Table 1.
[0093] Table 1 Comparative experimental results
[0094]
[0095] In summary, the method and system for schematic diagram object detection based on multi-granularity text reasoning of the present invention improve the detection accuracy of other benchmark methods. The average precision AP, AP 50 and AP 75 are respectively 2.9%, 3.4% and 3.1% higher than the optimal values.
[0096] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0097] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0098] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0100] The above content is only for explaining the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the claims of the present invention.
Claims
1. A schematic diagram object detection method based on multi-granularity text reasoning, characterized in that Including the following steps: S1. Fuse the text features and image features of the schematic diagram to obtain visual features enhanced by text features , and complete global text reasoning, specifically: S101. Use the Bert pre-trained language model to semantically encode the global description text and input the obtained text encoding into a bidirectional GRU network to obtain semantic information in both the forward and backward directions, and splice them to obtain the overall representation of the global text ; S102. Input a single image into the convolutional neural network ResNet. After four stages of convolutional pooling operations, obtain the overall features of the image ; S103. Use a transformation matrix to map the overall text representation obtained in step S101 and the overall image features obtained in step S102 into the same feature space, then perform splicing and summation, and then fuse the corresponding vectors to obtain visual features enhanced by text features , and send the enhanced visual features to the RPN and ROI networks for classification and regression operations to complete global text reasoning; S2. Enhanced visual features obtained according to step S1 Extract visual nodes, extract text nodes according to the characteristics of the schematic diagram itself, use the extracted text nodes and visual nodes of the schematic diagram as graph nodes, construct edges based on the relative position space between the text nodes and visual nodes and the text semantic similarity, and finely fuse the text features and visual features to obtain enhanced visual node features , and complete local text reasoning; S3. Extract global text keywords; filter out valid local text nodes using the similarity between the text nodes obtained in step S2 and the global text keywords; perform multi-granularity fusion on the valid local text nodes, the global text inference obtained in step S1, and the local text inference obtained in step S2 to complete schematic diagram object detection.
2. The method for schematic diagram object detection based on multi-granularity text reasoning according to claim 1, wherein In step S101, an update gate and a reset gate are added to the bidirectional GRU network to retain the features of the text through the update gate and the reset gate.
3. The schematic diagram object detection method based on multi-granularity text reasoning according to claim 1, characterized in that Step S2 is specifically as follows: S201. Input a single schematic diagram in the CSDQA dataset into the easyocr algorithm to obtain all the text information and its corresponding position information in the schematic diagram, a total of M OCR tags; use all the detected OCR tags in the schematic diagram as text nodes, and linearly transform the text vector and position vector in the node feature through a linear matrix, and fuse them after standardization to obtain the feature vector of the text node ; S202. Input a single schematic diagram in the CSDQA dataset into the first four stages of the convolutional neural network ResNet101 to extract feature maps; then input the feature maps into the RPN network to obtain foreground anchor boxes and their position offsets, and comprehensively use the foreground anchor boxes and their position offsets as candidate regions. Then, eliminate candidate regions with areas smaller than the specified threshold and those exceeding the boundaries, perform non-maximum suppression on the remaining candidate regions to form precise candidate regions, and finally obtain the ROI layer. The ROI layer receives the original feature maps and the candidate regions output by the RPN network, maps the candidate regions back to the original schematic diagram , and then perform max pooling to obtain visual nodes of the regional feature vectors ; S203. Take the text node obtained in step S201 and the visual node obtained in step S202 as nodes in the graph . Consider the similarity between the text node and the visual node as an edge, use a transformation matrix to map the feature vectors of the text node and the visual node into the same feature space, and then use the cosine function to calculate the semantic similarity between the text node and the visual node ; Combine the semantic similarity and the positional relationship to obtain the spatial semantic similarity between the th text node and the th visual node ; S204. According to the similarity between the text nodes and the visual nodes obtained in step S203 , splice and fuse all visual node features with their corresponding text nodes to obtain enhanced visual nodes , and complete local text reasoning.
4. The schematic diagram object detection method based on multi-granularity text reasoning according to claim 3, characterized in that, In step S201, the schematic diagram The text obtained in is input into the Bert pre-trained language model to obtain a text vector , and the position vector is extracted from the position information of the text box .
5. The method for schematic diagram object detection based on multi-granularity text reasoning according to claim 3, wherein In step S202, the RPN network is divided into two branches: classification and regression. The classification branch differentiates between foreground and background, and the regression branch refines the positions of the anchor boxes to determine the position offsets; the anchor box positions are the anchors generated for each point of the feature map, with an aspect ratio of .
6. The schematic diagram object detection method based on multi-granularity text reasoning according to claim 1, wherein, Step S3 is specifically as follows: S301. Select the StanfordCoreNLP toolkit to perform part-of-speech tagging on the global description text, and select nouns and adjectives among them as keywords; S302. Calculate the similarity between the keywords in the global text obtained in step S301 and the text nodes obtained in step S2, and determine the screening mechanism for local text nodes; S303. Integrate the global text inference obtained in step S1 and the local text inference obtained in step S2, and add the screening mechanism for local text nodes obtained in step S302.
7. The method for schematic diagram object detection based on multi-granularity text reasoning according to claim 6, characterized in that In step S301, the global description text is input into the toolkit, and after part-of-speech tagging, word screening, and word encoding, a keyword vector is obtained .
8. The schematic diagram object detection method based on multi-granularity text reasoning according to claim 6, characterized in that, In step S302, the screening mechanism for local text nodes is as follows: based on the similarity between the keywords in the global text obtained in step S301 and the local text nodes Remove the noise in the local text nodes and screen out the effective local text nodes.
9. A schematic diagram object detection system based on multi-granularity text reasoning, characterized in that, Including: Global text inference module, which fuses the text features and image features of the schematic diagram to obtain visually enhanced features with text features , and completes global text inference, specifically: S101. Use the Bert pre-trained language model to semantically encode the global description text and input the obtained text encoding into a bidirectional GRU network to obtain semantic information in both forward and backward directions, and splice them to obtain the overall representation of the global text ; S102. Input a single image into the convolutional neural network ResNet. After four stages of convolutional pooling operations, obtain the overall features of the image ; S103. Use a transformation matrix to map the global text overall representation obtained in step S101 and the image overall features obtained in step S102 into the same feature space, then perform splicing and summation, and then fuse the corresponding vectors to obtain visually enhanced features after text feature enhancement , and input the enhanced visual features into the RPN and ROI networks to perform classification and regression operations to complete global text reasoning; The local text reasoning module, according to the enhanced visual features obtained by the global text reasoning module Extract visual nodes, extract text nodes according to the characteristics of the schematic diagram itself, use the extracted text nodes and visual nodes of the schematic diagram as graph nodes, construct edges according to the relative position space between the text nodes and visual nodes and the text semantic similarity, and finely fuse the text features and visual features to obtain enhanced visual node features , and complete local text reasoning; A fusion detection module that extracts global text keywords; filters out valid local text nodes using the similarity between the text nodes obtained by the local text inference module and the global text keywords; performs multi-granularity fusion on the valid local text nodes, the global text inference obtained by the global text inference module, and the local text inference obtained by the local text inference module to complete schematic diagram object detection.