Fine-grained target component labeling method and device based on regional suggestion network
Through multimodal fusion and recursive optimization methods based on regional proposal networks, the problems of label consistency and cost in fine-grained object detection are solved, efficient and accurate fine-grained target component labeling are achieved, and the perception and operation capabilities of the robot system are improved.
Patent Information
- Application Number
- CN202510310944.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-29
AI Technical Summary
The existing fine-grained object detection technology has shortcomings in label consistency and cost, and it is difficult to accurately capture the relationship between target components, resulting in inconsistent labeling results and affecting the detection performance of the model.
A method based on region suggestion network is adopted, multimodal fusion and recursive optimization is combined with visual language models, detection chains are generated and fine-grained target components are labeled, initial candidate regions are generated through RPN, VLM is further optimized, and the annotation process is deepened layer by layer.
It realizes efficient and accurate fine-grained target component labeling, reduces the cost and time of manual labeling, improves label consistency, provides high-quality labeling data support, and enhances the perception and operation capabilities of the robot system.
Smart Images

Figure CN120388203A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a fine-grained target component annotation method and device based on a region proposal network. Background Art
[0002] Fine-grained automatic annotation systems have been widely studied and applied in recent years, especially in the fields of computer vision and deep learning. Such systems usually need to process complex and diverse data and provide accurate and detailed annotations for each data instance. Traditional annotation methods mainly rely on traditional methods and deep learning-based models. Traditional annotation methods mainly rely on manual annotation and rule-driven methods, which usually have strong domain knowledge dependence, but are inefficient in processing large-scale data and are difficult to adapt to complex and dynamic environments. Deep learning-based models often face challenges of high precision and high complexity when facing fine-grained annotation tasks. In practical applications, there are problems such as lack of annotated data, poor annotation consistency, and high annotation costs. Existing automatic annotation technologies often cannot accurately capture the relationships between target components in complex scenarios, resulting in inconsistent annotation results and affecting the quality of the training model. Especially in fine-grained object detection, the problem of annotation consistency is particularly serious, and any inconsistent annotation will lead to a decline in the quality of training data, thus affecting the detection performance of the model. Summary of the Invention
[0003] An object of the present invention is to provide a fine-grained target component annotation method based on a region proposal network, which can effectively identify fine-grained components of an object, realize automatic annotation of fine-grained target components, improve annotation consistency, reduce annotation costs and time, thereby improving the automatic annotation efficiency of fine-grained target components, and further providing strong support and data basis for subsequent perception and operation tasks of a robot system. Another object of the present invention is to provide a fine-grained target component annotation device based on a region proposal network. Still another object of the present invention is to provide a computer-readable medium. Yet another object of the present invention is to provide a computer device.
[0004] To achieve the above object, on the one hand, the present invention discloses a fine-grained target component annotation method based on a region proposal network, including:
[0005] Obtain an image to be annotated;
[0006] Through a vision-language model, perform multimodal fusion on the image to be annotated and a pre-input task description to generate a detection chain, where the detection chain includes a target component sequence;
[0007] Scan the image to be annotated through a region proposal network to generate component candidate boxes;
[0008] Recursively call the vision - language model and the region proposal network, and perform multiple rounds of recursive optimization on the component candidate boxes according to the sequence of target components in the detection chain to determine the fine - grained target components and label the fine - grained target components.
[0009] Preferably, through the vision - language model, perform multimodal fusion on the image to be labeled and the pre - input task description to generate a detection chain, including:
[0010] Perform multimodal feature extraction on the image to be labeled and the task description to generate visual features and text features;
[0011] Perform cross - modal alignment according to the visual features and text features, and perform hierarchical search on the image to be labeled to generate a detection chain.
[0012] Preferably, through the region proposal network, scan the image to be labeled to generate component candidate boxes, including:
[0013] Extract the image features of the image to be labeled through a convolutional neural network to generate a feature map;
[0014] Perform sliding scanning on the feature map according to a preset sliding window to determine the component candidate regions;
[0015] According to the component candidate regions, correspondingly determine the component candidate boxes.
[0016] Preferably, the sequence of target components includes multiple ordered target words;
[0017] Recursively call the vision - language model and the region proposal network, and perform multiple rounds of recursive optimization on the component candidate boxes according to the sequence of target components in the detection chain to determine the fine - grained target components and label the fine - grained target components, including:
[0018] Determine the first target word in the sequence of target components as the current target word;
[0019] Through the vision - language model, perform correlation analysis on the visual features corresponding to each component candidate box and the current target word to determine the intermediate component candidate boxes corresponding to the current target word;
[0020] Judge whether the current target word is the last target word in the sequence of target components;
[0021] If so, determine the candidate region corresponding to the intermediate component candidate box as the fine - grained target component and label the fine - grained target component.
[0022] Preferably, through the vision - language model, perform correlation analysis on the visual features corresponding to each component candidate box and the current target word to determine the intermediate component candidate boxes corresponding to the current target word, including:
[0023] Calculate the correlation between the visual features corresponding to each component candidate box and the current target word to generate a correlation score;
[0024] Determine the intermediate component candidate box corresponding to the current target word according to the correlation score.
[0025] Preferably, the method further includes:
[0026] If the current target word is not the last target word in the target component sequence, determine the next target word of the current target word as the current target word;
[0027] Scan the candidate region corresponding to the intermediate component candidate box through the region proposal network to generate sub-candidate boxes;
[0028] Through the vision-language model, perform a correlation analysis on the visual features corresponding to each sub-candidate box and the current target word, determine the intermediate component candidate box corresponding to the current target word, and repeat the step of determining whether the current target word is the last target word in the target component sequence until the current target word is the last target word in the target component sequence.
[0029] Preferably, after recursively calling the vision-language model and the region proposal network to perform multi-round recursive optimization on the component candidate boxes according to the target component sequence in the detection chain, determining the fine-grained target components and labeling the fine-grained target components, it further includes:
[0030] Construct a fine-grained target component dataset according to multiple labeled fine-grained target components and the target images where the fine-grained target components are located.
[0031] The present invention also discloses a fine-grained target component labeling device based on a region proposal network, including:
[0032] An image acquisition unit for acquiring an image to be labeled;
[0033] A multi-modal fusion unit for performing multi-modal fusion on the image to be labeled and a pre-input task description through a vision-language model to generate a detection chain, where the detection chain includes a target component sequence;
[0034] A scanning unit for scanning the image to be labeled through a region proposal network to generate component candidate boxes;
[0035] A recursive labeling unit for recursively calling the vision-language model and the region proposal network to perform multi-round recursive optimization on the component candidate boxes according to the target component sequence in the detection chain, determining the fine-grained target components and labeling the fine-grained target components.
[0036] Preferably, the multi-modal fusion unit is specifically configured to perform multi-modal feature extraction on the image to be annotated and the task description to generate visual features and text features; perform cross-modal alignment based on the visual features and text features, perform hierarchical search on the image to be annotated, and generate a detection chain.
[0037] Preferably, the scanning unit is specifically configured to extract the image features of the image to be annotated through a convolutional neural network to generate a feature map; perform a sliding scan on the feature map according to a preset sliding window to determine a component candidate region; and correspondingly determine a component candidate box according to the component candidate region.
[0038] Preferably, the target component sequence includes a plurality of ordered target words;
[0039] The recursive annotation unit is specifically configured to determine the first target word in the target component sequence as the current target word; perform a correlation analysis on the visual features corresponding to each component candidate box and the current target word through a vision-language model to determine an intermediate component candidate box corresponding to the current target word; determine whether the current target word is the last target word in the target component sequence; if so, determine the candidate region corresponding to the intermediate component candidate box as the fine-grained target component and annotate the fine-grained target component.
[0040] Preferably, the recursive annotation unit is specifically configured to perform a correlation calculation on the visual features corresponding to each component candidate box and the current target word to generate a correlation score; and determine an intermediate component candidate box corresponding to the current target word according to the correlation score.
[0041] Preferably, the apparatus further includes:
[0042] The recursive determination unit is configured to, if the current target word is not the last target word in the target component sequence, determine the next target word of the current target word as the current target word;
[0043] The recursive scanning unit is configured to perform a scan on the candidate region corresponding to the intermediate component candidate box through a region proposal network to generate a sub-candidate box;
[0044] The recursive correlation analysis unit is configured to perform a correlation analysis on the visual features corresponding to each sub-candidate box and the current target word through a vision-language model to determine an intermediate component candidate box corresponding to the current target word, and trigger the recursive annotation unit to repeat the step of determining whether the current target word is the last target word in the target component sequence until the current target word is the last target word in the target component sequence.
[0045] Preferably, the apparatus further includes:
[0046] A dataset construction unit for constructing a fine-grained target component dataset based on multiple annotated fine-grained target components and the target images where the fine-grained target components are located.
[0047] The present invention also discloses a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the above-described method is implemented.
[0048] The present invention also discloses a computer device, including a memory and a processor, where the memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the processor executes the program, the above-described method is implemented.
[0049] The present invention also discloses a computer program product, including computer program / instructions, and when the computer program / instructions are executed by a processor, the above-described method is implemented.
[0050] The present invention obtains an image to be annotated; through a vision-language model, performs multi-modal fusion on the image to be annotated and a pre-input task description to generate a detection chain, where the detection chain includes a target component sequence; scans the image to be annotated through a region proposal network to generate component candidate boxes; recursively calls the vision-language model and the region proposal network to perform multi-round recursive optimization on the component candidate boxes according to the target component sequence in the detection chain, determines the fine-grained target components and annotates the fine-grained target components, can effectively identify the fine-grained components of an object, realizes automatic annotation of fine-grained target components, improves annotation consistency, reduces annotation costs and time, thereby improving the automatic annotation efficiency of fine-grained target components, and further provides strong support and a data basis for subsequent perception and operation tasks of a robot system. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0052] Figure 1 It is a flowchart of a method for annotating fine-grained target components based on a region proposal network provided by an embodiment of the present invention;
[0053] Figure 2 It is a flowchart of another method for annotating fine-grained target components based on a region proposal network provided by an embodiment of the present invention;
[0054] Figure 3 It is a structural schematic diagram of a device for annotating fine-grained target components based on a region proposal network provided by an embodiment of the present invention;
[0055] Figure 4 A schematic structural diagram of a computer device provided by an embodiment of the present invention. Specific embodiments
[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0057] It should be noted that a fine-grained target component annotation method and device based on a region proposal network disclosed in this application can be used in the field of artificial intelligence technology, and can also be used in any field other than the field of artificial intelligence technology. The application field of the fine-grained target component annotation method and device based on a region proposal network disclosed in this application is not limited.
[0058] To facilitate understanding of the technical solutions provided in this application, the relevant content of the technical solutions in this application will be described first. With the rapid development of deep learning technology, methods based on deep neural networks (DNNs) have gradually become the core technology of fine-grained automatic annotation systems. Deep learning models, especially convolutional neural networks (CNNs), have been widely used in fine-grained annotation tasks. In image annotation, CNNs can automatically extract high-level features from the original image through multiple convolutional and pooling operations, thereby achieving accurate classification and localization. In addition, the Region Proposal Network (RPN) is widely used in object detection tasks, which can automatically generate candidate regions and generate corresponding annotations for each region. This method is particularly suitable for scenarios where precise localization of objects or annotations is required in fine-grained tasks. In image segmentation and instance segmentation, deep learning methods have shown great advantages, being able to perform object detection and segmentation simultaneously, generating high-precision object boundaries and labels. Instance segmentation techniques, such as Mask R-CNN, can generate accurate segmentation masks for each object by combining convolutional neural networks and segmentation tasks, thereby achieving fine-grained annotation. In addition, in recent years, methods based on graph neural networks (GNNs) have gradually attracted attention. GNNs can improve the accuracy and robustness of fine-grained annotation by modeling the spatial relationships between objects, and have strong adaptability especially in complex scenarios and multi-object environments.
[0059] To further improve the accuracy and robustness of fine-grained annotation, fine-grained annotation can also be carried out through the combination of multi-modal learning and self-supervised learning. Multi-modal learning can provide more dimensional support for fine-grained annotation by integrating information from different perceptual sources (such as vision, depth data, touch, etc.), especially having strong advantages when dealing with complex objects and dynamic environments. Self-supervised learning, on the other hand, designs proxy tasks so that the system can extract useful features by learning the latent patterns in the data without labeled data, thereby improving the generalization ability and accuracy of the annotation system.
[0060] Most current object detection technologies focus on the recognition of coarse-grained objects and cannot effectively handle the detailed parts of objects. In many tasks, robots need to identify specific components of objects, such as "door handles", "cup lids", etc. However, existing object detection systems usually can only identify the coarse-grained information of objects (such as "door" or "cup") and lack the ability to accurately describe and locate the detailed components of objects. This coarse-grained detection method has significant limitations for complex tasks. For example, in practical applications, robots need to grasp "door handles" or open "bottle caps", which requires the robot to have precise recognition and positioning of target details. However, existing object detection technologies often cannot meet this requirement, resulting in insufficient operation accuracy of the robot and even possible misoperations.
[0061] To improve the accuracy and consistency of automatic annotation, this application proposes a fine-grained object component annotation method based on RPN, aiming to generate preliminary candidate regions by introducing RPN and further optimize the annotation by combining with a Visual Language Model (VLM for short). RPN generates initial object candidate regions by extracting image features, and then uses VLM to combine task prompt information (such as "door" or "cup") and previous annotation results to further refine and optimize the annotation recursively. This process can accurately identify the fine-grained components of an object, such as "door handle" or "bottle cap", and effectively solve the problems of poor annotation consistency and high cost. In addition, the recursive mechanism of VLM enables the annotation optimization process to be deepened layer by layer, thus continuously refining the fine-grained features of the object. Each recursive call is based on the result of the previous step to further explore and optimize the position and features of the target component. For example, after identifying a "door", VLM can generate a prompt for the "door handle" and guide the detector to accurately locate the position of the "door handle". This process can be recursive until all the fine-grained components of the target (such as "door handle", "door lock", etc.) are accurately identified. In this way, fine-grained object annotation data can be automatically generated, reducing the cost and time of manual annotation. This application can effectively solve the problem of fine-grained object recognition in existing object detection technologies, and by introducing VLM, it optimizes the annotation process, improving the accuracy and consistency of annotation. The system not only provides high-quality annotation data for training high-precision object detection models, but also can significantly reduce the cost and time of manual annotation, providing strong support for the perception and operation tasks of the robot system.
[0062] This application is applicable to automatic annotation in object detection and robot perception tasks. By combining RPN and VLM, it realizes high-quality fine-grained object annotation by recursively optimizing candidate regions and the detection chain. It has the capabilities of automatic annotation and refined annotation, provides more accurate data support for object detectors, not only improves the object detail recognition ability, but also effectively reduces the cost of manual annotation, and improves the consistency and efficiency of annotation.
[0063] The core innovation of this application lies in the introduction of a Chain of Detection (CoD) mechanism, which can generate a detection chain based on the input current image and the task content to be detected. For example, when the task content is to detect the handle on the table, a detection chain of "table", "drawer", "handle" is generated first. CoD generates candidate regions using the Region Proposal Network (RPN) according to the detection steps generated by the Visual Language Model (VLM), and the VLM selects the corresponding candidate regions according to the current detection step, and then recursively repeats the above steps until the target component (handle) is found for annotation. This method models the target annotation process as a dynamic and gradually optimized decision-making process, combining task-specific hint information in each recursion to improve the performance and accuracy of the detector. This application combines the VLM and Monte Carlo Tree Search (MCTS) with the RPN to generate fine-grained annotations. This process starts from rough object detection and then recursively refines the regions to generate detailed annotations without manual intervention.
[0064] Taking the fine-grained target component annotation device based on the region proposal network as the execution subject as an example, the implementation process of the fine-grained target component annotation method based on the region proposal network provided by the embodiments of the present invention is described. It can be understood that the execution subject of the fine-grained target component annotation method based on the region proposal network provided by the embodiments of the present invention includes, but is not limited to, the fine-grained target component annotation device based on the region proposal network.
[0065] Figure 1 The flowchart of a fine-grained target component annotation method based on the region proposal network provided by the embodiments of the present invention is as Figure 1 shown, and the method includes:
[0066] Step 101, obtain the image to be annotated.
[0067] In the embodiments of the present invention, the image to be annotated is an image input by the user, and fine-grained target component annotation is performed on this image. For example: the image to be annotated is an image of a work station containing a desk.
[0068] Step 102, through the VLM, perform multimodal fusion on the image to be annotated and the pre-input task description to generate a detection chain.
[0069] In the embodiments of the present invention, the detection chain includes a target component sequence, and the target component sequence includes a plurality of ordered target words.
[0070] In the embodiments of the present invention, a detection chain, that is, a sequence of target components to be detected in sequence, is generated using the VLM according to the image to be annotated and the pre-input task description. Specifically, through C = VLM(x, S) = {c1, c2,..., c nPerform multimodal fusion on the to-be-annotated image and the pre-input task description to generate a detection chain. Among them, C is the detection chain, {c1, c2,..., c n} are multiple ordered target words, x is the to-be-annotated image, and S is the task description. For example: The to-be-annotated image is an image of a work station containing a desk, and the task description is to detect the handle on the desk; Input "an image of a work station containing a desk" and "detect the handle on the desk" into the VLM for multimodal fusion, and output a detection chain, which is {desk, drawer, handle}.
[0071] Step 103: Scan the to-be-annotated image through the RPN to generate component candidate boxes.
[0072] In the embodiment of the present invention, the RPN is used to scan the to-be-annotated image to generate an initial set of component candidate regions. The RPN extracts image features based on the to-be-annotated image and generates multiple potential component candidate boxes.
[0073] Step 104: Recursively call the VLM and the RPN, and perform multiple rounds of recursive optimization on the component candidate boxes according to the target component sequence in the detection chain to determine the fine-grained target components and label the fine-grained target components.
[0074] In the embodiment of the present invention, the chained recursive detection mechanism (CoD) is responsible for sequentially performing detection steps according to the detection chain generated by the VLM. CoD calls the RPN to generate candidate regions in the order in the detection chain, selects the candidate box corresponding to the current detection step, and recursively refines and labels the target components. Specifically, CoD realizes fine-grained annotation through the following steps:
[0075] 1. Execution of detection steps: Perform each detection step in sequence according to the order in the detection chain.
[0076] 2. Generation and selection of candidate regions: Call the RPN to generate candidate regions and select the intermediate component candidate box corresponding to the current target word of the current detection step.
[0077] 3. Recursive call: For the selected candidate box, perform the next detection step until the entire detection chain is completed.
[0078] The formula is expressed as:
[0079] D = CoD(C, x) = {d1, d2,..., d n}
[0080] Among them, D is the set of finally generated fine-grained target components, representing the detailed annotation of all detection steps; d i is the i-th fine-grained target component, C is the detection chain, and x is the to-be-annotated image.
[0081] In the embodiments of the present invention, the VLM selects a component candidate box corresponding to the target word (e.g., "table") from the component candidate regions according to the target word of the current detection step; and recursively calls the RPN and VLM according to the next step of the detection chain (e.g., "drawer") to further refine and annotate the selected candidate regions. Finally, through multiple rounds of recursive optimization, the system can accurately annotate the target component (e.g., "handle") to achieve fine-grained automatic annotation.
[0082] Through this recursive optimization mechanism, the present application can efficiently generate a fine-grained annotation dataset, significantly reduce the manual annotation cost, and improve the annotation consistency. These high-quality fine-grained datasets are further used to train the detector model, significantly improving its performance in fine-grained perception tasks, especially in complex robot grasping and operation tasks.
[0083] [[ID=**7**]]In the technical solution provided by the embodiments of the present invention, an image to be annotated is obtained; through a vision-language model, multimodal fusion is performed on the image to be annotated and a pre-input task description to generate a detection chain, and the detection chain includes a target component sequence; through a region proposal network, the image to be annotated is scanned to generate component candidate boxes; the vision-language model and the region proposal network are recursively called, and multiple rounds of recursive optimization are performed on the component candidate boxes according to the target component sequence in the detection chain to determine the fine-grained target component and annotate the fine-grained target component, which can effectively identify the fine-grained components of an object, achieve automatic annotation of the fine-grained target component, improve the annotation consistency, reduce the annotation cost and time, thereby improving the automatic annotation efficiency of the fine-grained target component, and further providing strong support and a data basis for subsequent perception and operation tasks of the robot system.
[0084] Figure 2 It is a flowchart of another fine-grained target component annotation method based on a region proposal network provided by the embodiments of the present invention. As Figure 2 shown, the method includes:
[0085] Step 201, obtain an image to be annotated.
[0086] In the embodiments of the present invention, each step is executed by a fine-grained target component annotation device based on a region proposal network.
[0087] In the embodiments of the present invention, the image to be annotated is an image input by the user, and fine-grained target component annotation is performed on the image. For example: the image to be annotated is an image of a work station including a desk.
[0088] Step 202, perform multimodal feature extraction on the image to be annotated and the task description to generate visual features and text features.
[0089] In the embodiments of the present invention, the task description is the natural language input by the user. For example, the task description is to detect the handle on the table.
[0090] Specifically, the VLM is used to extract features from the image to be annotated to generate visual features; the VLM is used to extract features from the task description to generate text features.
[0091] Step 203: Perform cross-modal alignment based on the visual features and text features, and perform hierarchical search on the image to be annotated to generate a detection chain.
[0092] In the embodiments of the present invention, the visual features and text features are mapped into the same feature space, and hierarchical search is performed on the image to be annotated to generate a detection chain. The detection chain includes a target component sequence, and the target component sequence includes a plurality of ordered target words.
[0093] For example: The image to be annotated is an image of a work station containing a desk, and the task description is to detect the handle on the table; visual features such as "desk", "water cup", "table lamp", "drawer" are extracted from the image to be annotated; text features such as "desk", "handle" are extracted from the task description; calculate the similarity between each visual feature and the text feature for cross-modal alignment, and determine that "office" matches successfully, "water cup" and "table lamp" do not match, and "drawer" partially matches (a "drawer" is usually associated with a "handle"); filter out the visual features related to the text description: "desk" and "drawer"; based on the text features and the filtered visual features, sort "handle", "desk" and "drawer" according to the hierarchical logical relationship from the whole to the part to generate a detection chain, and the detection chain is {table, drawer, handle}.
[0094] Step 204: Extract the image features of the image to be annotated through a convolutional neural network to generate a feature map.
[0095] In the embodiments of the present invention, the RPN module is responsible for generating object candidate regions from the input image. The RPN extracts the image features of the image to be annotated through a convolutional neural network to generate a feature map.
[0096] It should be noted that the image to be annotated can be a preprocessed image, and the image preprocessing includes but is not limited to size adjustment and normalization processing to improve the efficiency and accuracy of subsequent image data processing.
[0097] Step 205: Slide and scan the feature map according to a preset sliding window to determine the component candidate regions.
[0098] In the embodiments of the present invention, the size of the sliding window can be set according to actual needs, and the embodiments of the present invention do not limit this.
[0099] Specifically, a sliding window is slid on the feature map to perform a sliding scan on the feature map, predict whether there is an object and its position within each window, and output a set of potential component candidate regions, representing the possible positions of objects in the image. It is expressed by the formula:
[0100] R = RPN(x) = {r1, r2,..., r m}
[0101] where R is the set of candidate regions, r i is the i-th candidate region, and x is the image to be annotated.
[0102] Step 206: Determine the component candidate boxes corresponding to the component candidate regions.
[0103] In the embodiment of the present invention, the boundary of each component candidate region is determined as the corresponding component candidate box.
[0104] Step 207: Determine the first target word in the target component sequence as the current target word.
[0105] In the embodiment of the present invention, the detection chain is also the target component sequence. For example, if the target component sequence is {table, drawer, handle}, and the first target word is "table", then "table" is determined as the current target word.
[0106] Step 208: Through the VLM, perform a correlation analysis on the visual features corresponding to each component candidate box and the current target word, and determine the intermediate component candidate box corresponding to the current target word.
[0107] In the embodiment of the present invention, the VLM selects the intermediate component candidate box corresponding to the target word from the component candidate boxes generated by the RPN according to the current target word. Using the target word information provided by the VLM and combining the visual features of the candidate regions, the component candidate boxes that match the current target word are screened out. It is expressed by the formula:
[0108] S = Select(R, c t ) = {s1, s2,..., s k}
[0109] where S is the set of selected component candidate boxes, c t is the current target word, and s i is the i-th selected component candidate box.
[0110] In the embodiment of the present invention, Step 208 specifically includes:
[0111] Step 2081: Perform a correlation calculation on the visual features corresponding to each component candidate box and the current target word to generate a correlation score.
[0112] In the embodiments of the present invention, through various correlation analysis methods, the correlation score can be calculated for the visual features corresponding to each component candidate box and the current target word. The embodiments of the present invention do not limit the selection of specific correlation analysis. As an alternative solution, the correlation analysis method includes but is not limited to cosine similarity and multi-head attention mechanism.
[0113] Step 2082: Determine the intermediate component candidate box corresponding to the current target word according to the correlation score.
[0114] As an alternative solution, compare multiple correlation scores, screen out the highest correlation score, and determine the component candidate box corresponding to the highest correlation score as the intermediate component candidate box. Among them, each correlation score represents the correlation between the current target word and the visual features corresponding to a component candidate box.
[0115] As another alternative solution, compare multiple correlation scores with a preset correlation score threshold, screen out the correlation scores greater than the correlation score threshold, and determine the component candidate boxes corresponding to the screened correlation scores as the intermediate component candidate boxes. It should be noted that the correlation score threshold can be set according to actual needs, and the embodiments of the present invention do not limit this.
[0116] Step 209: Determine whether the current target word is the last target word in the target component sequence. If so, execute step 210; if not, execute step 211.
[0117] In the embodiments of the present invention, if the current target word is the last target word in the target component sequence, it means that the detection chain has been completed and the fine-grained target component has been detected. Continue to execute step 210; if the current target word is not the last target word in the target component sequence, it means that the detection chain has not been completed and the fine-grained target component has not been detected. Continue to execute step 211.
[0118] Furthermore, if the current target word is not the last target word in the target component sequence, label the candidate region corresponding to the intermediate component candidate box. For example, label the recognized "table" in the image to be labeled.
[0119] Step 210: Determine the candidate region corresponding to the intermediate component candidate box as the fine-grained target component, label the fine-grained target component, and continue to execute step 214.
[0120] In the embodiments of the present invention, if the fine-grained target component has been detected, it means that the candidate region corresponding to this intermediate component candidate box is the fine-grained target component, that is: determine the candidate region corresponding to the intermediate component candidate box as the fine-grained target component, and label the fine-grained target component in the image to be labeled, and continue to execute step 214.
[0121] It should be noted that the annotation method for the fine-grained target components can be set according to actual needs, and the embodiments of the present invention do not limit this. As an optional solution, the fine-grained target components can be marked by frame lines. For example: the circumscribed rectangle of the candidate area corresponding to the intermediate component candidate box.
[0122] Step 211: Determine the next target word of the current target word as the current target word.
[0123] In the embodiments of the present invention, if the detection chain is not completed, the next target word of the current target word is determined as the current target word. For example, the next target word "drawer" of "table" is determined as the current target word.
[0124] Step 212: Scan the candidate area corresponding to the intermediate component candidate box through RPN to generate sub-candidate boxes.
[0125] Specifically, through RPN, feature extraction is performed on the candidate area corresponding to the intermediate component candidate box to generate a feature map; a sliding window slides on the feature map to perform a sliding scan on the feature map, predicting whether there is an object and its position in each window, outputting a set of potential sub-candidate areas, and determining the corresponding sub-candidate boxes, indicating the possible object positions in the image.
[0126] Step 213: Through VLM, perform a correlation analysis on the visual feature corresponding to each sub-candidate box and the current target word to determine the intermediate component candidate box corresponding to the current target word, and repeat step 209 until the current target word is the last target word in the target component sequence.
[0127] Specifically, perform a correlation calculation on the visual feature corresponding to each sub-candidate box and the current target word to generate a correlation score; according to the correlation score, determine the intermediate component candidate box corresponding to the current target word, and repeat step 209 until the current target word is the last target word in the target component sequence, completing the entire detection chain.
[0128] Step 214: Construct a fine-grained target component data set according to multiple annotated fine-grained target components and the target images where the fine-grained target components are located.
[0129] In the embodiments of the present invention, the target image where the fine-grained target component is located corresponds to the input image to be annotated.
[0130] In the embodiments of the present invention, a high-quality fine-grained target component dataset is constructed by using multiple generated annotated fine-grained target components. The fine-grained target component dataset includes, but is not limited to, detailed object part annotations (such as "table", "drawer", "handle") for subsequent training of the detector model to enhance its fine-grained detection ability during the fine-grained perception process. It is expressed by the formula:
[0131]
[0132] where Dataset fine is the fine-grained target component dataset, x i is the target image where the fine-grained target component is located, D i is the annotated fine-grained target component, and N is the number of samples in the fine-grained target component dataset.
[0133] In the present invention, through VLM, the visual features of the image are combined with the natural language description to automatically generate a fine-grained annotation chain. This annotation chain can accurately identify and sequentially annotate the target components. The present invention can flexibly adjust the annotation steps according to the task description to meet complex annotation requirements. The generation of the annotation chain under the guidance of VLM ensures the efficient and accurate annotation of each target part.
[0134] The automatic annotation algorithm based on chain detection in the present invention can efficiently generate fine-grained annotation data, significantly reduce the manual annotation cost, and improve the annotation consistency. These high-quality fine-grained datasets are further used to train the detector model, significantly improving its performance in the fine-grained perception task, especially in complex robot grasping and operation tasks. This mechanism not only optimizes the detection process but also enhances the robot's perception ability of fine-grained target parts, enabling it to perform more precisely and efficiently in multi-level object recognition and operation. Finally, through the interaction process of RPN and VLM, the generated fine-grained annotations will be used to train the detector, improve the detection accuracy, and create a fine-grained annotation dataset, ensuring that each target component can obtain high-precision annotations, effectively improving the performance in complex scenarios, especially the accuracy and reliability in multi-target and fine operation tasks.
[0135] It should be noted that in the technical solutions of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of laws and regulations. The user information in the embodiments of this application is obtained through legal and compliant channels, and the acquisition, storage, use, processing, etc. of the user information have obtained the authorization and consent of the customers.
[0136] It should be noted that the information collected in this application is information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations, and standards of the relevant countries and regions. Necessary confidentiality measures are taken, which do not violate public order and good customs, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0137] It should be noted that the technical solution provided in this application provides a corresponding operation entrance for users to choose to agree or reject the automated decision-making result; if the user chooses to reject, the expert decision-making process will be entered.
[0138] In the technical solution of the fine-grained target component annotation method based on the region proposal network provided by the embodiments of the present invention, an image to be annotated is obtained; through a vision-language model, multi-modal fusion is performed on the image to be annotated and a pre-input task description to generate a detection chain, and the detection chain includes a target component sequence; through the region proposal network, the image to be annotated is scanned to generate component candidate boxes; the vision-language model and the region proposal network are recursively called, and the component candidate boxes are recursively optimized in multiple rounds according to the target component sequence in the detection chain to determine the fine-grained target components and annotate the fine-grained target components, which can effectively identify the fine-grained components of an object, realize the automatic annotation of fine-grained target components, improve the annotation consistency, reduce the annotation cost and time, thereby improving the automatic annotation efficiency of fine-grained target components, and further providing strong support and data basis for the subsequent perception and operation tasks of the robot system.
[0139] Figure 3 It is a schematic structural diagram of a fine-grained target component annotation device based on the region proposal network provided by the embodiments of the present invention. This device is used to execute the above-mentioned fine-grained target component annotation method based on the region proposal network, as Figure 3 shown, this device includes: an image acquisition unit 11, a multi-modal fusion unit 12, a scanning unit 13, and a recursive annotation unit 14.
[0140] The image acquisition unit 11 is used to acquire an image to be annotated.
[0141] The multi-modal fusion unit 12 is used to perform multi-modal fusion on the image to be annotated and a pre-input task description through a vision-language model to generate a detection chain, and the detection chain includes a target component sequence.
[0142] The scanning unit 13 is used to scan the image to be annotated through the region proposal network to generate component candidate boxes.
[0143] The recursive annotation unit 14 is used to recursively call the vision-language model and the region proposal network, perform multiple rounds of recursive optimization on the component candidate boxes according to the target component sequence in the detection chain, determine the fine-grained target components, and annotate the fine-grained target components.
[0144] In the embodiment of the present invention, the multimodal fusion unit 12 is specifically configured to perform multimodal feature extraction on the image to be annotated and the task description to generate visual features and text features; perform cross-modal alignment according to the visual features and text features, perform hierarchical search on the image to be annotated, and generate a detection chain.
[0145] In the embodiment of the present invention, the scanning unit 13 is specifically configured to extract the image features of the image to be annotated through a convolutional neural network to generate a feature map; perform a sliding scan on the feature map according to a preset sliding window to determine the component candidate regions; and correspondingly determine the component candidate boxes according to the component candidate regions.
[0146] In the embodiment of the present invention, the target component sequence includes a plurality of ordered target words; the recursive annotation unit 14 is specifically configured to determine the first target word in the target component sequence as the current target word; perform a correlation analysis on the visual features corresponding to each component candidate box and the current target word through the vision-language model to determine the intermediate component candidate box corresponding to the current target word; determine whether the current target word is the last target word in the target component sequence; if so, determine the candidate region corresponding to the intermediate component candidate box as the fine-grained target component, and annotate the fine-grained target component.
[0147] In the embodiment of the present invention, the recursive annotation unit 14 is specifically configured to perform a correlation calculation on the visual features corresponding to each component candidate box and the current target word to generate a correlation score; and determine the intermediate component candidate box corresponding to the current target word according to the correlation score.
[0148] In the embodiment of the present invention, the device further includes: a recursive determination unit 15, a recursive scanning unit 16, and a recursive correlation analysis unit 17.
[0149] The recursive determination unit 15 is configured to, if the current target word is not the last target word in the target component sequence, determine the next target word of the current target word as the current target word.
[0150] The recursive scanning unit 16 is configured to scan the candidate region corresponding to the intermediate component candidate box through the region proposal network to generate sub-candidate boxes.
[0151] The recursive correlation analysis unit 17 is used to perform a correlation analysis on the visual features corresponding to each sub-candidate box and the current target word through a vision-language model, determine the intermediate component candidate box corresponding to the current target word, and trigger the recursive annotation unit 14 to repeatedly execute the step of determining whether the current target word is the last target word in the target component sequence until the current target word is the last target word in the target component sequence.
[0152] In an embodiment of the present invention, the device further includes: a dataset construction unit 18.
[0153] The dataset construction unit 18 is used to construct a fine-grained target component dataset according to multiple annotated fine-grained target components and the target images where the fine-grained target components are located.
[0154] In the solution of the embodiment of the present invention, an image to be annotated is obtained; through a vision-language model, multimodal fusion is performed on the image to be annotated and a pre-input task description to generate a detection chain, and the detection chain includes a target component sequence; through a region proposal network, the image to be annotated is scanned to generate component candidate boxes; the vision-language model and the region proposal network are recursively called, and the component candidate boxes are recursively optimized in multiple rounds according to the target component sequence in the detection chain to determine and annotate the fine-grained target components, which can effectively identify the fine-grained components of an object, realize automatic annotation of fine-grained target components, improve annotation consistency, reduce annotation costs and time, thereby improving the automatic annotation efficiency of fine-grained target components, and further providing strong support and a data basis for subsequent perception and operation tasks of a robot system.
[0155] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by a computer chip or an entity, or by a product with a certain function. A typical implementation device is a computer device. Specifically, the computer device can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0156] An embodiment of the present invention provides a computer device, including a memory and a processor. The memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, the steps of the above embodiments of the fine-grained target component annotation method based on a region proposal network are implemented. For specific descriptions, reference can be made to the embodiments of the fine-grained target component annotation method based on a region proposal network.
[0157] Reference is made below Figure 4 , which shows a schematic structural diagram of a computer device 600 suitable for implementing the embodiments of the present application.
[0158] As Figure 4 shown, the computer device 600 includes a central processing unit (CPU) 601, which can perform various appropriate operations and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computer device 600 are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0159] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage section 608 as needed.
[0160] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 609, and / or installed from the removable medium 611.
[0161] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0162] For convenience of description, the above-described apparatus is described by functionally dividing it into various units. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.
[0163] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0164] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0165] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 or steps for implementing the functions specified in one block or multiple blocks.
[0166] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the said element.
[0167] In the technical solution of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.
[0168] It should be noted that in the embodiments of this application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be considered exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.
[0169] Those skilled in the art should understand that the embodiments of this application can be provided as a method, a system or a computer program product. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0170] This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that execute specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are executed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0171] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the relevant part of the method embodiment for the relevant content.
[0172] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A fine-grained target component annotation method based on a region proposal network, characterized in that The method includes: Obtaining an image to be annotated; Performing multimodal fusion on the image to be annotated and a pre-input task description through a vision-language model to generate a detection chain, where the detection chain includes a target component sequence; Scanning the image to be annotated through a region proposal network to generate component candidate boxes; Recursively calling the vision-language model and the region proposal network, and performing multi-round recursive optimization on the component candidate boxes according to the target component sequence in the detection chain to determine fine-grained target components and annotate the fine-grained target components.
2. The fine-grained target component annotation method based on the region proposal network according to claim 1, wherein The performing multimodal fusion on the image to be annotated and the pre-input task description through the vision-language model to generate a detection chain includes: Performing multimodal feature extraction on the image to be annotated and the task description to generate visual features and text features; Performing cross-modal alignment according to the visual features and text features, and performing hierarchical search on the image to be annotated to generate the detection chain.
3. The method for fine-grained target component annotation based on a region proposal network according to claim 1, wherein The scanning the image to be annotated through the region proposal network to generate component candidate boxes includes: Extracting image features of the image to be annotated through a convolutional neural network to generate a feature map; Performing sliding scanning on the feature map according to a preset sliding window to determine component candidate regions; Correspondingly determining the component candidate boxes according to the component candidate regions.
4. The method for fine-grained target component annotation based on region proposal network according to claim 1, wherein The target component sequence includes a plurality of ordered target words; The recursively calling the vision-language model and the region proposal network, and performing multi-round recursive optimization on the component candidate boxes according to the target component sequence in the detection chain to determine fine-grained target components and annotate the fine-grained target components includes: Determining the first target word in the target component sequence as the current target word; Performing correlation analysis on the visual features corresponding to each component candidate box and the current target word through the vision-language model to determine intermediate component candidate boxes corresponding to the current target word; Judging whether the current target word is the last target word in the target component sequence; If so, determining the candidate region corresponding to the intermediate component candidate box as the fine-grained target component and annotating the fine-grained target component.
5. The method for fine-grained target component annotation based on region proposal network according to claim 4, wherein The performing correlation analysis on the visual features corresponding to each component candidate box and the current target word through the vision-language model to determine intermediate component candidate boxes corresponding to the current target word includes: Performing correlation calculation on the visual features corresponding to each component candidate box and the current target word to generate a correlation score; Determining intermediate component candidate boxes corresponding to the current target word according to the correlation score.
6. The method for fine-grained target component annotation based on region proposal network according to claim 4, wherein The method further includes: If the current target word is not the last target word in the target component sequence, determining the next target word of the current target word as the current target word; Scanning the candidate region corresponding to the intermediate component candidate box through the region proposal network to generate sub-candidate boxes. Through the visual language model, perform a correlation analysis on the visual features corresponding to each sub-candidate box and the current target word, determine the intermediate component candidate box corresponding to the current target word, and repeatedly execute the step of determining whether the current target word is the last target word in the target component sequence until the current target word is the last target word in the target component sequence.
7. The method for fine-grained target component annotation based on region proposal network according to claim 1, wherein After recursively calling the visual language model and the region proposal network, and performing multiple rounds of recursive optimization on the component candidate boxes according to the target component sequence in the detection chain to determine the fine-grained target components and label the fine-grained target components, it further includes: Construct a fine-grained target component dataset according to multiple labeled fine-grained target components and the target images where the fine-grained target components are located.
8. An apparatus for fine-grained target component annotation based on a region proposal network, characterized in that The device includes: An image acquisition unit for acquiring an image to be labeled. A multi-modal fusion unit for performing multi-modal fusion on the image to be labeled and a pre-input task description through a visual language model to generate a detection chain, where the detection chain includes a target component sequence. A scanning unit for scanning the image to be labeled through a region proposal network to generate component candidate boxes. A recursive labeling unit for recursively calling the visual language model and the region proposal network, and performing multiple rounds of recursive optimization on the component candidate boxes according to the target component sequence in the detection chain to determine the fine-grained target components and label the fine-grained target components.
9. The fine-grained target component annotation device based on the region proposal network according to claim 8, wherein The multi-modal fusion unit is specifically configured to perform multi-modal feature extraction on the image to be labeled and the task description to generate visual features and text features; perform cross-modal alignment according to the visual features and text features, and perform hierarchical search on the image to be labeled to generate the detection chain.
10. The fine-grained target component annotation device based on the region proposal network according to claim 8, wherein, The scanning unit is specifically configured to extract image features of the image to be labeled through a convolutional neural network to generate a feature map. Perform a sliding scan on the feature map according to a preset sliding window to determine a component candidate region; correspondingly determine the component candidate box according to the component candidate region.
11. The fine-grained target component annotation device based on the region proposal network according to claim 8, characterized in that, The target component sequence includes a plurality of ordered target words. The recursive labeling unit is specifically configured to determine the first target word in the target component sequence as the current target word; through the visual language model, perform a correlation analysis on the visual features corresponding to each component candidate box and the current target word to determine the intermediate component candidate box corresponding to the current target word. Determine whether the current target word is the last target word in the target component sequence; if so, determine the candidate region corresponding to the intermediate component candidate box as the fine-grained target component and label the fine-grained target component.
12. The fine-grained target component annotation device based on the region proposal network according to claim 11, wherein, The recursive labeling unit is specifically configured to perform a correlation calculation on the visual features corresponding to each component candidate box and the current target word to generate a correlation score; determine the intermediate component candidate box corresponding to the current target word according to the correlation score.
13. The fine-grained target component annotation device based on the region proposal network according to claim 11, characterized in that, The device further includes: A recursive determination unit, configured to determine the next target word of the current target word as the current target word if the current target word is not the last target word in the target component sequence; A recursive scanning unit, configured to scan the candidate region corresponding to the intermediate component candidate box through the region proposal network to generate sub-candidate boxes; A recursive correlation analysis unit, configured to perform a correlation analysis on the visual feature corresponding to each sub-candidate box and the current target word through the vision-language model, determine the intermediate component candidate box corresponding to the current target word, and trigger the recursive annotation unit to repeatedly execute the step of determining whether the current target word is the last target word in the target component sequence until the current target word is the last target word in the target component sequence.
14. The fine-grained target component annotation device based on the region proposal network according to claim 8, characterized in that, The apparatus further includes: A dataset construction unit, configured to construct a fine-grained target component dataset according to a plurality of annotated fine-grained target components and the target images where the fine-grained target components are located.
15. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the fine-grained target component annotation method based on a region proposal network according to any one of claims 1 to 7.
16. A computer device, comprising a memory and a processor, the memory being used for storing information including program instructions, and the processor being used for controlling the execution of the program instructions, characterized in that, When the program instructions are loaded and executed by a processor, it implements the fine-grained target component annotation method based on a region proposal network according to any one of claims 1 to 7.
17. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, it implements the fine-grained target component annotation method based on a region proposal network according to any one of claims 1 to 7.
Citation Information
Cited By
Cross-granularity target detection method and device and electronic equipment
CN120411465A
Cross-granularity target detection method and device, and electronic device
CN120411465B