Sub-graph detection method and device and electronic equipment
The method uses a feature extraction network and multi-scale deep neural networks to accurately detect and classify subgraphs in complex images, addressing inefficiencies in existing manual and rule-based methods by enhancing recognition and reducing redundancy.
Patent Information
- Application Number
- CN202510441817.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In the prior art, the problem of low detection efficiency and poor adaptability caused by relying on manual or fixed rules for sub-graph division.
The feature extraction network and multi-scale deep neural network are used to extract multi-level feature information of images, and the sub-graph candidate areas are generated by combining regional suggestions and classification networks, and the sub-graphs are automatically detected and extracted through redundant candidate areas.
It improves the accuracy and processing efficiency of sub-graph detection, reduces false detection and multi-checking, and realizes structured extraction and orderly output of sub-graphs.
Smart Images

Figure CN120318531A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology. Specifically, this application relates to a subgraph detection method, apparatus, and electronic device. Background Art
[0002] Most of the legends added to scientific research papers are large images containing multiple subgraphs, with various types and complex arrangements of subgraphs. In order to achieve refined detection of image tampering behavior, it is usually necessary to accurately detect and segment each subgraph in the image first. Related technologies mostly perform subgraph division through manual operations or rely on fixed rules, resulting in low efficiency. Moreover, when the boundaries of subgraphs are unclear or the arrangements are irregular, problems such as inaccurate recognition and omission are likely to occur, making it difficult to meet actual needs. Summary of the Invention
[0003] Embodiments of this application provide a subgraph detection method to solve the problems in the prior art that rely on manual or fixed rules for subgraph division, resulting in low detection efficiency and poor adaptability.
[0004] Correspondingly, embodiments of this application also provide a subgraph detection apparatus, an electronic device, and a storage medium to ensure the implementation and application of the above method.
[0005] To solve the above problems, embodiments of this application disclose a subgraph detection method, and the method includes:
[0006] Obtain a target image including at least one subgraph;
[0007] Input the target image into a feature extraction network to obtain an initial feature map of the target image;
[0008] Input the initial feature map into a multi-scale deep neural network to extract feature information of the initial feature map at different spatial feature levels;
[0009] Generate corresponding multiple feature maps according to the feature information;
[0010] Based on a region proposal network, generate at least one subgraph candidate region on each of the multiple feature maps;
[0011] Based on a classification network, perform subgraph category recognition on the subgraph candidate regions in each feature map to determine the subgraph category of each subgraph candidate region; wherein, the classification network is pre-trained based on a preset subgraph category label;
[0012] Perform redundant candidate region filtering processing on multiple subgraph candidate regions with the same subgraph category to obtain at least one target subgraph candidate region;
[0013] Extract a corresponding plurality of sub - graphs from the target image based on the target candidate sub - graph region.
[0014] An embodiment of the present application also discloses a sub - graph detection device, which includes:
[0015] An acquisition module, configured to acquire a target image including at least one sub - graph;
[0016] A first processing module, configured to input the target image into a feature extraction network to obtain an initial feature map of the target image;
[0017] A second processing module, configured to input the initial feature map into a multi - scale deep neural network to extract feature information of the initial feature map at different spatial feature levels;
[0018] A third processing module, configured to generate corresponding multiple feature maps according to the feature information;
[0019] A fourth processing module, configured to generate at least one sub - graph candidate region on each of the multiple feature maps based on a region proposal network;
[0020] A fifth processing module, configured to perform sub - graph category recognition on the sub - graph candidate regions in each of the feature maps based on a classification network to determine the sub - graph category of each sub - graph candidate region; wherein, the classification network is pre - trained based on a preset sub - graph category label;
[0021] A sixth processing module, configured to perform redundant candidate region filtering processing on multiple sub - graph candidate regions with the same sub - graph category to obtain at least one target sub - graph candidate region;
[0022] A seventh processing module, configured to extract corresponding multiple sub - graphs from the target image based on the target candidate sub - graph region.
[0023] An embodiment of the present application also discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, one or more methods in the embodiments of the present application are implemented.
[0024] An embodiment of the present application also discloses a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, one or more methods in the embodiments of the present application are implemented.
[0025] An embodiment of the present application also discloses a computer program product, including a computer program. When the computer program is executed by a processor, one or more methods in the embodiments of the present application are implemented.
[0026] The beneficial effects brought by the technical solution provided in the embodiments of the present application are as follows:
[0027] In the embodiments of the present application, by obtaining a target image including at least one sub - figure and extracting multi - level feature information of the initial feature map based on a feature extraction network and a multi - scale deep neural network, the expression ability of the image under different spatial structures can be enhanced; by generating sub - figure candidate regions on multiple feature maps and combining a classification network to identify the sub - figure categories of each candidate region, the recognition accuracy of the sub - figure position and type can be effectively improved; further, by performing redundant filtering on the candidate regions with the same sub - figure category, duplicate or overlapping sub - figure regions can be removed, reducing false detection and multiple detection cases; finally, by extracting multiple sub - figures from the original image through the target candidate regions, the structured extraction and ordered output of the sub - figures are realized, thereby improving the accuracy and processing efficiency of sub - figure detection and providing more reliable image data for image analysis tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The above - mentioned and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0029] Figure 1 is a flowchart of the sub - figure detection method provided in the embodiments of the present application;
[0030] Figure 2 is a schematic diagram of the first example provided in the embodiments of the present application;
[0031] Figure 3 is a schematic structural diagram of the sub - figure detection device provided in the embodiments of the present application;
[0032] Figure 4 is a schematic structural diagram of the electronic device provided in the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The embodiments of the present application will be described below with reference to the drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute limitations on the technical solutions of the embodiments of the present application.
[0034] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations supported by the technical field of the present application. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "plurality" means two or more, in view of this, in the embodiments of the present application, "plurality" can also be understood as "at least two". The term "and / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / ", unless otherwise specified, generally means that the associated objects before and after are in an "or" relationship.
[0035] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0036] The embodiments of the present application provide a sub-graph detection method. Optionally, the embodiments of the present application can be applied to electronic devices with image processing functions, such as laptop computers, desktop computers, servers, tablet computers, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), scientific research workstations, image analysis instruments, medical image acquisition terminals, vehicle-mounted imaging systems, digital TVs, etc. devices, and are particularly suitable for the structured processing and analysis scenarios of scientific research images. Subsequently, an electronic device can be selected as the execution subject of the embodiments of the present application for description. However, this does not constitute a limitation to the embodiments of the present application.
[0037] In the process of writing scientific research papers, it is often necessary to display experimental results or analysis processes in the form of images. Related legends usually consist of multiple sub-graphs, and these sub-graphs are uniformly arranged in an image for publication, forming a so-called "composite graph". With the increasing demand for image forgery detection and the like, the structured processing of such composite graphs has become increasingly important. At present, the processing of scientific research images still mainly relies on manual identification and segmentation one by one, or on traditional image processing methods based on boundary lines, fixed arrangements or proportional rules. However, in practical applications, the sub-graphs in the image are irregularly arranged, have blurred boundaries, significant type differences, and may even have occlusion, overlap or background interference, resulting in obvious deficiencies in adaptability and accuracy in the above-mentioned methods.
[0038] Therefore, there is an urgent need for a processing method that can automatically identify, locate, and extract multiple sub-images in an image to improve the automation and accuracy of image structure analysis and meet the actual needs of scientific research image processing.
[0039] To solve the above technical problems, an embodiment of the present application provides a sub-image detection method. Refer to Figure 1 , the method may include the following steps:
[0040] Step 101, obtain a target image including at least one sub-image;
[0041] In the embodiment of the present application, a target image containing several sub-images is obtained as the input data for the entire sub-image detection process. The target image includes, but is not limited to, scientific research papers, technical reports, or other scenarios with image publishing requirements, where multiple sub-images with different structures are spliced and displayed in one image to improve the layout utilization rate or enhance the ability of graphic expression. Since such images often have problems such as mixed sub-image types and irregular arrangements, before performing image structure detection, it is necessary to first obtain and uniformly process the original image input to ensure the integrity and accuracy of subsequent steps.
[0042] Step 102, input the target image into a feature extraction network to obtain an initial feature map of the target image;
[0043] In the embodiment of the present application, the original image data is encoded and converted through a feature extraction network to obtain an initial feature map that can reflect the basic texture and edge structure of the image. Among them, the feature extraction network can be composed of a set of standard convolutional layers, activation layers, and residual connection layers. The initial feature map is usually the result of the first stage of semantic feature extraction by the neural network. While maintaining the original information of the image, it compresses unnecessary redundant details and provides high-quality input for subsequent high-level semantic modeling.
[0044] Step 103, input the initial feature map into a multi-scale deep neural network to extract feature information of the initial feature map at different spatial feature levels;
[0045] In the embodiment of the present application, the initial feature map is further modeled, and a multi-scale deep neural network structure is introduced to extract the structural features of the image at different spatial scales. The so-called "different spatial feature levels" refers to the multi-level understanding of the semantics and structure of the image by the neural network at different spatial resolutions, such as the progressive feature representation from the overall layout to local details. Through multi-scale modeling, the network's perception ability for sub-images with diverse sizes, shapes, and arrangements can be improved.
[0046] Step 104, generate corresponding multiple feature maps according to the feature information;
[0047] Based on the region proposal network, at least one sub - graph candidate region is generated on each of the multiple feature maps;
[0048] In the embodiments of the present application, the extracted multi - scale feature information is respectively mapped into multiple feature maps, and each feature map represents the image expression of a spatial level. Subsequently, the region proposal network is used to process each feature map to generate rectangular candidate regions that may contain sub - graphs. Among them, the generated candidate regions have different sizes, ratios, and position information, facilitating the coverage of various possible sub - graph positions.
[0049] Step 105: Based on the classification network, perform sub - graph category recognition on the sub - graph candidate regions in each of the feature maps to determine the sub - graph category of each sub - graph candidate region; wherein, the classification network is pre - trained based on preset sub - graph category labels;
[0050] In the embodiments of the present application, the classification network is used to classify the image content of each sub - graph candidate region. The classification network is pre - trained and can perform feature classification and judgment on the input region according to the preset sub - graph category labels (such as different image types defined in the category set), so as to assign a clear category label to each candidate region. The classification network can be a multi - layer perceptron, a convolutional neural network, or other deep structures with category discrimination ability, which is not specifically limited here. Through this process, the detection system can not only determine "where is the sub - graph", but also further identify "what type of sub - graph it is".
[0051] Step 106: Perform redundant candidate region filtering processing on multiple sub - graph candidate regions with the same sub - graph category to obtain at least one target sub - graph candidate region;
[0052] In the embodiments of the present application, the degree of overlap analysis and screening are performed on the candidate boxes with the same sub - graph category. In the actual detection process, the same sub - graph may be repeatedly hit by multiple candidate boxes. Therefore, it is necessary to identify the overlapping regions through certain spatial overlap judgment rules (such as the IOU threshold) and retain the optimal one as the target sub - graph candidate region. The filtering process can include operations such as candidate box merging and elimination to achieve the uniqueness and accuracy of the sub - graph region results.
[0053] Step 107: Based on the target candidate sub - graph regions, extract corresponding multiple sub - graphs from the target image.
[0054] In the embodiments of the present application, according to the spatial coordinate information of the target candidate region in the original image, an image region extraction operation is performed on the image, and multiple independent sub-image files are output. The extraction operation can be implemented based on a coordinate cropping method, that is, according to the boundaries of each target sub-image candidate region, the corresponding region is intercepted on the original image to form a new image file. The extracted sub-images can be used for downstream tasks such as forgery detection, image comparison, and image archiving.
[0055] In some embodiments, the multi-scale deep neural network includes multiple sub-image detection neural networks connected in series;
[0056] wherein, the sub-image detection neural network includes a first pooling layer and / or a second pooling layer for performing downsampling processing;
[0057] In the case where the sub-image detection neural network includes a first pooling layer and a second pooling layer for performing downsampling processing, the downsampling frequency of the first pooling layer is not greater than the downsampling frequency of the second pooling layer.
[0058] In the embodiments of the present application, in some implementation manners, the multi-scale deep neural network includes multiple sub-image detection neural networks connected in series, and each sub-image detection neural network serves as a feature extraction module in the network structure and sequentially processes the input feature map. To improve the diversity and effectiveness of feature representation, a first pooling layer and / or a second pooling layer for performing downsampling operations are provided in the sub-image detection neural network. The pooling operation is used to compress the spatial dimension of the feature map to reduce the subsequent computational complexity. In certain embodiments, when both the first pooling layer and the second pooling layer are provided, the downsampling frequency of the first pooling layer is not greater than the downsampling frequency of the second pooling layer, that is, the stride of the sliding window or the size of the pooling window adopted by the first pooling layer is smaller, maintaining higher-resolution feature information, while the second pooling layer emphasizes the compression effect more. Through this differential pooling structure design, the network can perform more flexible and hierarchical modeling of image features at different scales while taking into account both accuracy and efficiency.
[0059] In some embodiments, the sub-image detection neural network further includes a feedforward neural network;
[0060] wherein, the feedforward neural network is used to perform channel dimension amplification processing on the downsampled feature vector sequence.
[0061] In some implementation manners in the embodiments of the present application, the sub-graph detection neural network further includes a feed-forward neural network. In the embodiments of the present application, the feed-forward neural network is used to perform an amplification operation on the feature vector sequence that has completed downsampling processing, that is, map the channel dimension of each feature vector from a lower dimension to a higher dimension to enhance its representation ability in subsequent processing. This amplification operation can be implemented through multi-layer linear transformation or other equivalent feature mapping structures. The purpose is to enhance the semantic information carried by its channel dimension while reducing the spatial resolution of the feature map, so as to make up for the information loss caused by the reduction in resolution while maintaining the downsampling efficiency.
[0062] In some embodiments, the region proposal network is constructed based on a feature pyramid structure;
[0063] Among them, the region proposal network is pre-trained based on a plurality of sub-graph candidate regions with different aspect ratios preset.
[0064] In the embodiments of the present application, the region proposal network is constructed based on a feature pyramid structure. The feature pyramid structure is a feature organization method commonly used in image object detection, which can combine feature maps at different scales to enhance the network's perception ability of objects of different sizes. Through this structure, the region proposal network can generate candidate regions on feature maps of multiple scales to adapt to sub-graphs of different sizes or shapes existing in the image. Further, in the embodiments of the present application, the region proposal network is pre-trained based on a plurality of sub-graph candidate regions with different aspect ratios preset, so that the network can learn the typical boundary features of sub-graphs of different morphologies. These preset ratios are used to guide the network to form a generation strategy for candidate boxes during the training stage, so as to more accurately cover the real sub-graph region.
[0065] In some embodiments, the redundant candidate region filtering process for multiple sub-graph candidate regions with the same sub-graph category includes:
[0066] Based on the spatial position information between the sub-graph candidate regions, identify the sub-graph candidate regions with an overlapping relationship, and perform a deletion operation and / or a fusion operation on the sub-graph candidate regions whose overlap degree exceeds a preset threshold.
[0067] In an embodiment of the present application, the redundant candidate area filtering process for multiple sub-image candidate areas with the same sub-image category includes: based on the spatial position information between the sub-image candidate areas, identifying the candidate areas with overlapping relationships. In the process of image sub-image detection, due to the existence of feature map layered processing and multi-scale candidate box generation mechanism, multiple overlapping candidate areas of the same category may point to the same sub-image. In order to avoid redundant output caused by repeated identification, this embodiment determines the overlap relationship by calculating the spatial overlap between candidate areas (such as the bounding box intersection and union ratio). When the overlap of two or more candidate areas exceeds a preset threshold, it is considered that there is redundancy. On this basis, deletion operations and / or fusion operations can be performed on this type of area according to a preset strategy, so as to obtain a unique and more accurate target sub-image candidate area.
[0068] In some embodiments, the method further comprises:
[0069] The target sub-image candidate regions are sorted according to a preset spatial sorting rule to determine a spatial position index of each of the target sub-image candidate regions.
[0070] In the embodiment of the present application, since the sub-images are usually presented in a composite image in a two-dimensional arrangement, and their arrangement may not have strict rules, it is necessary to perform a unified spatial sorting process on them after extracting each sub-image. The preset sorting rules may include a scanning logic of "from top to bottom, and from left to right in the same row", combined with the center point coordinates or boundary positions of the candidate area to make row and column determinations. In the sorting process, by comparing the vertical and horizontal position information of each target sub-image candidate area, they are classified and numbered so that the sub-images can be processed, marked or stored in an orderly manner later.
[0071] In some embodiments, the method further comprises:
[0072] The category label and / or the spatial position index is added to the name of the extracted sub-image.
[0073] In an embodiment of the present application, after the sub-image extraction is completed, in order to achieve structured management of the results and facilitate subsequent operations, the category information corresponding to the sub-image and its spatial sorting position in the original image can be embedded in the name of each sub-image image. For example, the sub-image file can be named in the form of "type1_row2_col3.jpg", where "type1" represents the category label to which the sub-image belongs, and "row2_col3" represents the spatial position of the sub-image in the overall image. The category label comes from the aforementioned classification network recognition result, and the spatial position index is determined by the sorting step. Both can be embedded in the file name as important information for image identification.
[0074] In some embodiments, the feature extraction network includes at least one of the following:
[0075] strided convolutional layer, activation layer, residual connection layer.
[0076] In the embodiments of the present application, the feature extraction network includes at least one of the following: strided convolutional layer, activation layer, residual connection layer. The feature extraction network is a basic module for extracting an initial feature map from an original image, and its structural design directly affects the quality and integrity of subsequent feature representation. The strided convolutional layer extracts local features while downsampling by setting a stride greater than 1, which helps to reduce the spatial resolution of the feature map and increase the receptive field; the activation layer (such as ReLU) is used to introduce a non-linear mapping relationship, thereby enhancing the expressive ability of the network; the residual connection layer can be used to alleviate the vanishing gradient problem in deep networks, and at the same time promote the information flow between different layers and improve the feature retention ability. These structures can be combined according to specific implementation needs to adapt to image feature extraction tasks with different complexities and accuracy requirements.
[0077] In some embodiments, the multi-scale deep neural network includes multiple computing modules with different channel dimensions; wherein, the feature information output by the computing module includes spatial size and / or channel dimension.
[0078] In the embodiments of the present application, the multi-scale deep neural network includes multiple computing modules with different channel dimensions, which are used for multi-level modeling and expression of the input feature information. Each computing module can be regarded as a feature processing unit in the neural network and can extract the structural features of the image at different semantic levels. Among them, the feature information output by each computing module has a specific spatial size and / or channel dimension. The spatial size reflects the width and height of the feature map, and the channel dimension represents the number of feature dimensions carried at each spatial position. Taking a specific example, the multi-scale deep neural network may include four computing modules with corresponding channel dimensions of 96, 128, 256, and 512 respectively to achieve progressive enhancement of features from shallow to deep and from local to global. These channel amplifications are usually accompanied by a gradual reduction in spatial size, forming a spatial-channel hierarchical structure.
[0079] In some embodiments, the method further includes:
[0080] Performing mask generation processing on each of the sub-graph candidate regions to generate a corresponding sub-graph contour mask; wherein, the sub-graph contour mask is used to determine the sub-graph boundary region.
[0081] In the embodiments of the present application, the mask generation process refers to modeling the pixel-level distribution within an image region through a network to generate contour information corresponding to a candidate region, and the contour mask is used to further determine the boundary region of the sub-graph. In a specific implementation, a mask generation sub-network can be established, which, based on the input candidate region and its feature information, outputs a set of probability mask maps representing the pixel distribution of the sub-graph region in the image. This mask can be used to assist in the further fine positioning of the bounding box and to distinguish the boundary between the sub-graph and the background in a complex scene. For example, in some implementations, the mask generation sub-network can output masks corresponding to multiple preset sub-graph categories respectively, and the boundary contours of the masks constitute the boundary regions of the sub-graphs.
[0082] In some embodiments, the preset sub-graph category labels include at least one of the following:
[0083] Microscope image, strip chart, flow chart, visible light image, bar chart, other data analysis charts, manually drawn schematic diagram.
[0084] In the embodiments of the present application, this category label is used to guide the classification network to perform content recognition and classification on the detected sub-graph candidate regions, and the setting of the category can be based on statistical analysis and abstract induction of common diagram types in scientific research paper images. For example, microscope images are generally used to display the microscopic structure of samples, strip charts and flow charts are mostly used for the visualization of biological experiment results, visible light images reflect the actual captured content, bar charts and other data charts represent quantitative analysis results, and manually drawn schematic diagrams are used to assist in explaining theories or processes. By presetting the above typical categories and using them in the training or inference stage of the classification network, the automatic recognition ability of sub-graphs can be improved.
[0085] The following is illustrated by specific embodiments:
[0086] Embodiment 1:
[0087] As Figure 2 shown, the sub-graph detection method provided by the embodiments of the present application includes:
[0088] Step 201, input image and initial feature extraction.
[0089] In the embodiments of the present application, a target image containing at least one sub-graph is input into a feature extraction network. The target image is a composite image in a scientific research paper and may include different types of sub-graphs such as microscope images, charts, and strip charts. The feature extraction network includes structures such as a strided convolutional layer, an activation layer, and a residual connection layer, and is used to extract the initial feature map of the target image.
[0090] Step 202, multi-scale feature extraction;
[0091] In the embodiments of the present application, the initial feature map is input into a multi-scale deep neural network. The multi-scale deep neural network includes multiple computing modules with different channel dimensions, and each computing module contains multiple sub-graph detection neural networks connected in series. Each sub-graph detection neural network includes a first pooling layer, a second pooling layer, and a feed-forward neural network.
[0092] Among them, the first pooling layer is used to perform low-frequency downsampling on the query vector;
[0093] Among them, the second pooling layer performs high-frequency downsampling on the key vector and the value vector to improve the calculation efficiency;
[0094] Among them, the feed-forward neural network performs channel dimension amplification on the downsampled feature vector sequence to enhance the feature expression ability.
[0095] Finally, the multi-scale deep neural network outputs multiple feature maps, and the feature maps have differences in spatial size and channel dimension, forming an image representation with different spatial feature levels.
[0096] Step 203, generating candidate regions;
[0097] In the embodiments of the present application, the multiple feature maps are input into a region proposal network. The region proposal network is constructed based on a feature pyramid structure and generates at least one sub-graph candidate region on each scale of the feature map.
[0098] To adapt to the structural differences of different types of sub-graphs, the region proposal network sets multiple candidate regions with different aspect ratios for different sub-graph categories, such as microscope images being more square, strip images being more rectangular, etc.
[0099] Step 204, sub-graph classification and recognition;
[0100] In the embodiments of the present application, each generated sub-graph candidate region is input into a classification network, which has been pre-trained based on multiple sub-graph category labels (including microscope images, strip images, flow cytometry images, visible light images, bar charts, other data analysis charts, manually drawn schematic diagrams). The classification network performs sub-graph category recognition on each candidate region and determines the corresponding sub-graph category.
[0101] Step 205, filtering redundant candidate regions;
[0102] In the embodiments of the present application, for candidate regions with the same sub-graph category, calculate their spatial overlap degree (such as based on the intersection over union IOU). If the overlap degree between two candidate regions exceeds a preset threshold, perform redundant candidate region filtering processing. This processing includes deleting one of the duplicate regions, or fusing multiple overlapping regions into a new target sub-graph candidate region under certain conditions.
[0103] Step 206, Sub - graph sorting and numbering;
[0104] In the embodiment of the present application, for multiple target sub - graph candidate regions after filtering, according to their spatial positions in the original image, they are sorted in the arrangement order of "from top to bottom, from left to right", and a spatial position index is assigned to each region.
[0105] Step 207, Mask generation processing;
[0106] In the embodiment of the present application, to further improve the accuracy of sub - graph boundary detection, mask generation processing can be performed on each target sub - graph candidate region, generating a sub - graph contour mask based on the features within the region, which is used to more accurately define the boundary range of the sub - graph, especially suitable for the cases where the image edges are blurred or the boundaries between sub - graphs are unclear.
[0107] Step 208, Sub - graph cropping and output;
[0108] In the embodiment of the present application, according to the coordinate information of the target sub - graph candidate regions, multiple corresponding sub - graph images are cropped from the original image and output as detection results. When naming the sub - graph images, their category label fields and spatial position index fields are appended, such as "microscopy_row1_col2.jpg", to support downstream tasks such as forgery detection and duplicate checking analysis.
[0109] Embodiment Two:
[0110] The embodiment of the present application provides a sub - graph detection method, which can be applied to technical scenarios such as scientific research paper image tampering detection, and is applicable to multiple fields including but not limited to biology, medicine, materials, computer, industry, aviation, transportation, etc. Scientific research papers usually contain a composite image with several nested sub - graphs of different types. To achieve refined detection of the sub - graph content, sub - graph recognition, classification, and extraction processing need to be performed on this type of composite image.
[0111] Therefore, the embodiment of the present application provides an automated and batch sub - graph detection and extraction technical solution, which uses a multi - scale deep neural network combined with a redundant candidate region filtering strategy to achieve the positioning, classification, and segmentation output of multiple sub - graphs in the image. The technical process of this embodiment includes the following contents:
[0112] Step S1: Image normalization processing.
[0113] The input image is input into the image processing module for normalization processing. For example, the long side of the image is adjusted to 800 pixels, and the short side is scaled proportionally according to the original width - height ratio, so as to unify the image input size and improve the stability and generality of the subsequent feature extraction network.
[0114] Step S2: Initialize the class labels.
[0115] Preset 7 sub - figure class labels for the classification network, which are: microscope image, strip chart, flow chart, visible - light image, bar chart, other data - analysis charts, and manually - drawn schematic diagrams, for subsequent classification training or recognition of sub - figure candidate regions.
[0116] Step S3: Feature extraction using a multi - scale deep neural network.
[0117] Input the normalized image into the feature extraction network. This network consists of a macro - feature extraction network and a sub - figure detection neural network, and outputs feature information at different spatial feature levels:
[0118] First, extract the initial feature map through the macro - feature extraction network, which includes structures such as a strided convolutional layer, an activation layer, and a residual connection layer;
[0119] Then, unfold the initial feature map along the height and width directions to form a sequence of feature vectors, and input it into the multi - layer sub - figure detection neural network.
[0120] Among them, each layer of the sub - figure detection neural network contains the following structures:
[0121] a) Calculate the query set, key set, and value set respectively through three independent parameter matrices;
[0122] b) Set two types of pooling layers, which are respectively applied to the query set, key set, and value set. The two types of pooling layers have different window sizes and strides, so as to form down - sampled features of different lengths;
[0123] c) Among them, the down - sampling frequency of the first pooling layer is lower than that of the second pooling layer. The first pooling layer is only used for the query set, while the second pooling layer is applicable to the key set and value set. In the four computing modules of the network, the first pooling layer is only applied to the sub - figure detection network layer at the first layer of each module, while the second pooling layer runs through each layer;
[0124] d) Perform matrix multiplication on the query set and the key set to calculate the similarity weights, and perform weighted summation with the value set through the value normalization mechanism to output the sequence of feature vectors after similarity screening;
[0125] e) Input the above - mentioned sequence of feature vectors into a feed - forward neural network. The feed - forward structure includes two layers of linear transformation layers, an activation layer, a feature dropout layer, and a residual connection layer. If the query set is down - sampled by the first pooling layer, the feed - forward neural network expands the channel dimension. The output channel dimensions corresponding to the four computing modules are: 96, 128, 256, 512; the number of layers of the sub - figure detection network in the four modules are 3, 3, 7, 6 respectively;
[0126] f) Each sub - graph detection network layer applies two types of residual connections: one is to accumulate the original query set to the output result of step d, and the other is to accumulate the query set to the output result of step e to enhance the feature retention and information flow transfer capabilities.
[0127] Step S4: Feature map construction and candidate region generation.
[0128] Restore the feature vector sequences output by the above four computing modules into two - dimensional feature maps, respectively forming feature maps of four spatial feature levels, with dimensions: [56×56×96], [28×28×128], [14×14×256], [7×7×512].
[0129] Based on this multi - scale feature map, use a region proposal network based on the feature pyramid structure to generate multiple rectangular sub - graph candidate regions at each scale. Different candidate boxes have preset aspect ratios to adapt to the structural features of different sub - graph categories.
[0130] Step S5: Mask generation.
[0131] In some embodiments, establish a mask generation sub - network to generate a mask image of the corresponding category for each candidate region. Masks are generated for 7 categories respectively, corresponding to the pixel boundaries of the sub - graph, and are used to refine and mark the boundary contour region of the sub - graph.
[0132] Step S6: Sub - graph candidate box optimization.
[0133] Further establish a cascade regression network architecture for performing multi - stage localization optimization on the sub - graph candidate boxes. This architecture consists of multiple cascaded sub - graph detection neural networks. Each level of the network sets a higher IOU threshold, uses the candidate boxes of the previous stage as the input of the next stage, and continuously improves the accuracy of the bounding boxes through resampling and regression strategies.
[0134] Step S7: Sub - graph classification and recognition.
[0135] Construct a classification sub - network to perform classification and recognition on each sub - graph candidate region. The classification sub - network consists of a multi - layer perceptron and outputs confidence scores under 7 category labels.
[0136] Step S8: Output of sub - graph detection results.
[0137] Finally, output two types of information results: one is the coordinate position of the sub - graph detection box (i.e., the proposed box), and the other is the sub - graph category label corresponding to each candidate box.
[0138] Step S9: Filtering redundant candidate regions.
[0139] For the candidate regions with partial overlap, perform redundant box filtering. Specifically, it includes: sorting all candidate regions in ascending order of area, encoding their spatial positions in sequence, and calculating their IOU overlap with subsequent regions. If the overlap exceeds the threshold, record them in the redundant list, and corresponding regions can perform fusion or deletion operations to obtain a set of target candidate regions without overlap.
[0140] Step S10: Spatial sorting process.
[0141] Sort the determined subgraph candidate regions according to their spatial positions. Adopt the rule of top-down and left-to-right, and combine the size of the subgraph itself and the previous subgraph to complete the judgment and sorting of row and column attribution. Finally, assign a spatial position index to each subgraph.
[0142] Step S11: Subgraph image extraction and naming.
[0143] Extract the subgraph image region from the original image according to the coordinate information of the target candidate region. The extraction method is a cropping operation based on pixel coordinates. Each output subgraph image is appended with a category field and a spatial position index field in the naming, such as "type1_row2_col3.jpg", which is convenient for performing image classification, duplicate checking, or tampering analysis in subsequent tasks.
[0144] Based on the same principle as the method provided in the embodiments of the present application, the embodiments of the present application also provide a subgraph detection device, as Figure 3 shown, the device includes:
[0145] An acquisition module 1, configured to acquire a target image including at least one subgraph;
[0146] A first processing module 2, configured to input the target image into a feature extraction network to obtain an initial feature map of the target image;
[0147] A second processing module 3, configured to input the initial feature map into a multi-scale deep neural network to extract feature information of the initial feature map at different spatial feature levels;
[0148] A third processing module 4, configured to generate corresponding multiple feature maps according to the feature information;
[0149] A fourth processing module 5, configured to generate at least one subgraph candidate region on each of the multiple feature maps based on a region proposal network;
[0150] A fifth processing module 6, configured to perform subgraph category recognition on the subgraph candidate regions in each feature map based on a classification network to determine the subgraph category of each subgraph candidate region; wherein, the classification network is pre-trained based on a preset subgraph category label;
[0151] The sixth processing module 7 is configured to perform redundant candidate region filtering processing on multiple sub-graph candidate regions with the same sub-graph category to obtain at least one target sub-graph candidate region;
[0152] The seventh processing module 8 is configured to extract corresponding multiple sub-graphs from the target image based on the target candidate sub-graph region.
[0153] In some embodiments, the multi-scale deep neural network includes multiple sub-graph detection neural networks connected in series;
[0154] Wherein, the sub-graph detection neural network includes a first pooling layer and / or a second pooling layer for performing downsampling processing;
[0155] When the sub-graph detection neural network includes a first pooling layer and a second pooling layer for performing downsampling processing, the downsampling frequency of the first pooling layer is not greater than that of the second pooling layer.
[0156] In some embodiments, the sub-graph detection neural network further includes a feed-forward neural network;
[0157] Wherein, the feed-forward neural network is configured to perform channel dimension expansion processing on the downsampled feature vector sequence.
[0158] In some embodiments, it is characterized in that
[0159] The region proposal network is constructed based on a feature pyramid structure;
[0160] Wherein, the region proposal network is pre-trained based on a plurality of preset sub-graph candidate regions with different aspect ratios.
[0161] In some embodiments, the sixth processing module is specifically configured to:
[0162] Based on the spatial position information between the sub-graph candidate regions, identify sub-graph candidate regions with an overlapping relationship, and perform a deletion operation and / or a fusion operation on the sub-graph candidate regions whose overlap degree exceeds a preset threshold.
[0163] In some embodiments, the sub-graph detection device further includes an eighth processing module for:
[0164] Sort the target sub-graph candidate regions according to a preset spatial sorting rule to determine the spatial position index of each target sub-graph candidate region.
[0165] In some embodiments, the sub-graph detection device further includes a ninth processing module for:
[0166] Add a class label and / or the spatial location index to the naming of the extracted subgraphs.
[0167] In some embodiments, the feature extraction network includes at least one of the following:
[0168] A strided convolutional layer, an activation layer, a residual connection layer.
[0169] In some embodiments, the multi-scale deep neural network includes multiple computing modules with different channel dimensions; wherein, the feature information output by the computing module includes a spatial dimension and / or a channel dimension.
[0170] In some embodiments, the subgraph detection device further includes a tenth processing module for:
[0171] Perform a mask generation process on each of the subgraph candidate regions to generate a corresponding subgraph contour mask; wherein, the subgraph contour mask is used to determine the subgraph boundary region.
[0172] In some embodiments, the preset subgraph class labels include at least one of the following:
[0173] Microscope images, strip charts, flow charts, visible light images, bar charts, other data analysis charts, manually drawn schematic diagrams.
[0174] The subgraph detection device provided by the embodiments of the present application can implement Figure 1 or Figure 2 Each process implemented in the method embodiments of, for the sake of brevity, will not be described here again.
[0175] The subgraph detection device provided by the present application, by obtaining a target image including at least one subgraph and extracting multi-level feature information of an initial feature map based on a feature extraction network and a multi-scale deep neural network, can enhance the expression ability of the image under different spatial structures; by generating subgraph candidate regions on multiple feature maps and combining a classification network to identify the subgraph category of each candidate region, it can effectively improve the recognition accuracy of the subgraph position and type; further, by performing redundant filtering on candidate regions with the same subgraph category, it can remove duplicate or overlapping subgraph regions and reduce false detection and multiple detection situations; finally, by extracting multiple subgraphs from the original image through the target candidate regions, it realizes the structured extraction and ordered output of the subgraphs, thereby improving the accuracy and processing efficiency of subgraph detection and providing more reliable image data for image analysis tasks.
[0176] The sub - graph detection device according to the embodiments of the present application can execute the sub - graph detection method provided by the embodiments of the present application, and their implementation principles are similar. The actions performed by each module and unit in the sub - graph detection device in each embodiment of the present application correspond to the steps in the sub - graph detection method in each embodiment of the present application. For the detailed function description of each module of the sub - graph detection device, reference can be specifically made to the description in the corresponding sub - graph detection method shown above, and details will not be repeated here.
[0177] Based on the same principle as the method shown in the embodiments of the present application, the embodiments of the present application also provide an electronic device, which may include but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the sub - graph detection method shown in any optional embodiment of the present application by calling the computer program. Compared with the prior art, the sub - graph detection method provided by the present application can enhance the expression ability of an image under different spatial structures by obtaining a target image including at least one sub - graph and extracting multi - level feature information of an initial feature map based on a feature extraction network and a multi - scale deep neural network; by generating sub - graph candidate regions on multiple feature maps and combining a classification network to identify the sub - graph category for each candidate region, the recognition accuracy of the sub - graph position and type can be effectively improved; further, by performing redundant filtering on candidate regions with the same sub - graph category, duplicate or overlapping sub - graph regions can be removed, reducing false detection and multiple detection cases; finally, by extracting multiple sub - graphs from the original image through the target candidate regions, the structured extraction and ordered output of sub - graphs are realized, thereby improving the accuracy and processing efficiency of sub - graph detection and providing more reliable image data for image analysis tasks.
[0178] In an optional embodiment, an electronic device is also provided, as Figure 4 shown Figure 4 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between this electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of this electronic device 4000 does not constitute a limitation to the embodiments of the present application.
[0179] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0180] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 only a thick line is used to represent it herein, but it does not mean that there is only one bus or one type of bus.
[0181] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.
[0182] The memory 4003 is used to store the computer program for implementing the embodiments of the present application, and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0183] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The illustrated electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0184] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents shown in the foregoing method embodiments can be implemented.
[0185] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents shown in the foregoing method embodiments can be implemented.
[0186] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than that shown in the drawings or described in words.
[0187] It should be understood that although the flowchart in the embodiments of the present application indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.
[0188] The above are only alternative implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, adopting other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.
Claims
1. A subgraph detection method, characterized in that, Including: Obtain a target image including at least one sub - figure; Input the target image into a feature extraction network to obtain an initial feature map of the target image; Input the initial feature map into a multi - scale deep neural network to extract feature information of the initial feature map at different spatial feature levels; Generate corresponding multiple feature maps according to the feature information; Based on a region proposal network, generate at least one sub - figure candidate region on each of the multiple feature maps; Based on a classification network, perform sub - figure category recognition on the sub - figure candidate regions in each feature map to determine the sub - figure category of each sub - figure candidate region; wherein, the classification network is pre - trained based on a preset sub - figure category label; Perform redundant candidate region filtering processing on multiple sub - figure candidate regions with the same sub - figure category to obtain at least one target sub - figure candidate region; Based on the target candidate sub - figure region, extract corresponding multiple sub - figures from the target image.
2. The sub - figure detection method according to claim 1, wherein: The multi - scale deep neural network includes multiple sub - figure detection neural networks connected in series; Wherein, the sub - figure detection neural network includes a first pooling layer and / or a second pooling layer for performing down - sampling processing; When the sub - figure detection neural network includes a first pooling layer and a second pooling layer for performing down - sampling processing, the down - sampling frequency of the first pooling layer is not greater than the down - sampling frequency of the second pooling layer.
3. The sub-graph detection method according to claim 2, wherein The sub - figure detection neural network further includes a feed - forward neural network; Wherein, the feed - forward neural network is used to perform channel - dimension amplification processing on the down - sampled feature vector sequence.
4. The sub - figure detection method according to any one of claims 1 to 3, wherein: The region proposal network is constructed based on a feature pyramid structure; Wherein, the region proposal network is pre - trained based on a preset multiple sub - figure candidate regions with different aspect ratios.
5. The sub-graph detection method according to any one of claims 1 to 4, characterized in that The performing redundant candidate region filtering processing on multiple sub - figure candidate regions with the same sub - figure category includes: Based on the spatial position information between the sub - figure candidate regions, identify sub - figure candidate regions with an overlapping relationship, and perform a deletion operation and / or a fusion operation on sub - figure candidate regions whose overlap degree exceeds a preset threshold.
6. The sub-graph detection method according to claim 5, characterized in that, The method further includes: Sort the target sub - figure candidate regions according to a preset spatial sorting rule to determine the spatial position index of each target sub - figure candidate region.
7. The sub-graph detection method according to claim 6, wherein The method further includes: Add a category label and / or the spatial position index to the naming of the extracted sub - figures.
8. The sub - graph detection method according to any one of claims 1 to 7, characterized in that, The feature extraction network includes at least one of the following: A strided convolutional layer, an activation layer, a residual connection layer.
9. The sub-graph detection method according to any one of claims 1 to 8, characterized in that The multi - scale deep neural network includes multiple computing modules with different channel dimensions; wherein, the feature information output by the computing module includes spatial size and / or channel dimension.
10. The sub-graph detection method according to any one of claims 1 to 9, characterized in that, The method further includes: Perform mask generation processing on each sub - figure candidate region to generate a corresponding sub - figure contour mask; wherein, the sub - figure contour mask is used to determine the sub - figure boundary region.
11. The sub-graph detection method according to any one of claims 1 to 10, characterized in that, The preset sub - figure category label includes at least one of the following: Microscope images, band diagrams, flow cytometry diagrams, visible light images, bar charts, other data analysis charts, and manually drawn schematic diagrams.
12. A sub-graph detection device, characterized in that, Including: An acquisition module for acquiring a target image including at least one sub-image; A first processing module for inputting the target image into a feature extraction network to obtain an initial feature map of the target image; A second processing module for inputting the initial feature map into a multi-scale deep neural network to extract feature information of the initial feature map at different spatial feature levels; A third processing module for generating corresponding multiple feature maps according to the feature information; A fourth processing module for generating at least one sub-image candidate region on each of the multiple feature maps based on a region proposal network; A fifth processing module for performing sub-image category recognition on the sub-image candidate regions in each of the feature maps based on a classification network to determine the sub-image categories of each of the sub-image candidate regions; wherein, the classification network is pre-trained based on a preset sub-image category label; A sixth processing module for performing redundant candidate region filtering processing on multiple sub-image candidate regions with the same sub-image category to obtain at least one target sub-image candidate region; A seventh processing module for extracting corresponding multiple sub-images from the target image based on the target candidate sub-image region.
13. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method according to any one of claims 1 to 11 is implemented.
14. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the method according to any one of claims 1 to 11 is implemented.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
X-ray image foreground target extraction-based article discrimination method
CN108288279A
Multi-scale feature fusion face detection and segmentation method based on generalized intersection-to-union ratio
CN114463800A
Conversation generation method fusing basic knowledge and user information
CN116010575A
Related method and device of skin lesion detection network, equipment and storage medium
CN116912154A
Image recognition method and device, equipment and storage medium
CN118485905A