Cross-granularity target detection method and apparatus, and electronic device
Patent Information
- Application Number
- PCT/CN2025/146809
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2025-12-29
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025146809_01102026_PF_FP_ABST
Abstract
Description
Cross-granularity target detection method and device, and electronic device
[0001] Related Applications
[0002] This application claims priority to Chinese Patent Application No. 202510369539.2, filed on March 26, 2025, and incorporates by reference the entire disclosure of the aforementioned patent application as part of the present application. TECHNICAL FIELD
[0003] The present disclosure relates to the field of artificial intelligence, and in particular to a cross-granularity target detection method, device and electronic equipment. BACKGROUND
[0004] Robot perception technology has made significant progress in recent years, greatly promoting the development of many fields such as industrial manufacturing, medical health, and human-machine collaboration. However, current robot detection technology mainly focuses on coarse-grained target recognition, i.e., identifying and positioning the overall object, and shows obvious deficiencies when dealing with tasks that require fine-grained understanding. In scenarios such as fine assembly, complex human-machine interaction, and flexible object manipulation, robots not only need to understand the overall structure of the target object, but also need to accurately identify and locate the local details and structural relationships of the object to complete high-precision tasks. This demand poses higher challenges to the accuracy, generalization ability, and context understanding of existing perception technology. SUMMARY
[0005] To address the problems in the prior art, the embodiments of the present disclosure provide a cross-granularity target detection method, device and electronic equipment, which can at least partially solve the problems in the prior art.
[0006] According to a first aspect of the present disclosure, a cross-granularity target detection method is provided. The method can include: obtaining a target image and target task information; generating a detection target sequence according to the target image and the target task information, wherein each detection target is contained in the target image, and in the detection target sequence, the detection targets are sequentially contained in order; and identifying the position of each detection target in the target image step by step in order according to the order of the detection targets in the detection target sequence by a recursive method.
[0007] In some embodiments, generating the detection target sequence according to the target image and the target task information includes: generating the detection target sequence according to the target image and the target task information by a target visual language model, wherein the target visual language model is obtained by training a preset visual language model based on first annotation data, the first annotation data includes an image and task information used as model input, and further includes a detection target sequence used as a label.
[0008] In some embodiments, the target visual language model generates the detection target sequence according to the target image and the target task information, including: performing first visual encoding on the target image to generate a first visual feature vector; performing first language encoding on the target task information to generate a first language feature vector; fusing the first visual feature vector and the first language feature vector by using a first cross-modal attention mechanism to obtain a first joint feature representation; and performing language decoding on the first joint feature representation to obtain the detection target sequence.
[0009] In some embodiments, in the detection target sequence, each detection target is presented in a textual description.
[0010] In some embodiments, in the detection target sequence, each detection target is presented in a textual description.
[0011] In some embodiments, the target detector processes the textual description and the image of each input detection target as follows: performing second visual encoding on the image to generate a second visual feature vector; performing second language encoding on the textual description of the detection target to generate a second language feature vector; fusing the second visual feature vector and the second language feature vector by using a second cross-modal attention mechanism to obtain a second joint feature representation; and determining the position of the detection target in the image according to the second joint feature representation.
[0012] In some embodiments, before inputting the target image and the textual description of the first detection target in the detection target sequence into the target detector, the method further includes: obtaining second annotation data, wherein the second annotation data includes the image and the textual description of the detection target used as the model input, and also includes the position information used as the label, and the position information meets the target accuracy requirement; and training the preset detector according to the second annotation data to obtain the target detector.
[0013] In some embodiments, the method further comprises: according to the position information of each identified detection target, performing motion control on the mechanical arm to enable the mechanical arm to complete the target task.
[0014] According to a second aspect of the present disclosure, a cross-granularity target detection apparatus is provided. The cross-granularity target detection apparatus comprises: an acquisition module configured to acquire a target image and target task information; a generation module configured to generate a detection target sequence according to the target image and the target task information, wherein the target image contains each detection target, and in the detection target sequence, the detection targets are sequentially contained in order; and an identification module configured to identify the position of each detection target in the target image in a recursive manner according to the order of the detection targets in the detection target sequence.
[0015] According to a third aspect of the present disclosure, an electronic device is provided. The electronic device can comprise a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements the steps of the cross-granularity target detection method of any of the above embodiments when executing the program.
[0016] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps of the cross-granularity target detection method of any of the above embodiments.
[0017] According to a fifth aspect of the present disclosure, a computer program product is provided. The computer program product comprises a computer program stored on a non-transitory computer-readable storage medium, and the computer program comprises program instructions, and when the program instructions are executed by a computer, the computer can execute the steps of the cross-granularity target detection method of any of the above embodiments.
[0018] The cross-granularity target detection method, the electronic device, the computer-readable storage medium, and the computer program product provided by the embodiments of the present disclosure realize layer-by-layer accurate identification from coarse-grained targets to fine-grained targets, solve the limitations of current target detection techniques in robot perception systems, and particularly solve the challenges in processing fine-grained target identification and accurate positioning in complex environments. BRIEF DESCRIPTION OF DRAWINGS
[0019] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure.
[0020] FIG. 1 shows a flowchart of a cross-granularity target detection method of an example of the present disclosure.
[0021] FIG. 2 shows a flowchart of processing of a target image and target task information by a target visual language model of an example of the present disclosure.
[0022] FIG. 3 shows a flowchart of a process of training a visual language model to obtain a target visual language model according to an example of the present disclosure.
[0023] FIG. 4 shows a partial flowchart of a cross-granularity object detection method according to an example of the present disclosure.
[0024] FIG. 5 is a schematic diagram of a process of generating a textual description and an image of a detected object for each input according to an example of the present disclosure.
[0025] FIG. 6 is a schematic diagram of a training process of an object detector according to an example of the present disclosure.
[0026] FIG. 7 is a schematic diagram of a structure of a cross-granularity object detection system according to an example of the present disclosure.
[0027] FIG. 8 and FIG. 9 are schematic diagrams of a process of detecting an object using a cross-granularity object detection method according to an example of the present disclosure.
[0028] FIG. 10 is a structural block diagram of a cross-granularity object detection apparatus according to an example of the present disclosure.
[0029] FIG. 11 is a schematic diagram of a physical structure of an electronic device according to an example of the present disclosure. DETAILED DESCRIPTION
[0030] With specific reference now to the drawings in detail, illustrative examples of the embodiments are depicted. The subject matter of the following description and drawings is illustrative of the embodiments and is not in any way limiting. Numerous specific details are described to provide a thorough understanding of the embodiments. However, in certain instances, well known or
[0031] The terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used in the description of the application and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It also will be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0032] It should be understood that although the terms "first", "second", "third", etc. can be used herein to describe various information, these information should not be limited by these terms. These terms are only used to distinguish the categories of information. For example, without departing from the scope of the present application, the first information can be referred to as the second information, and similarly, the second information can be referred to as the first information. The term "if" used herein can be understood as "when", "at this time" or "in response to judgment" according to the context.
[0033] The traditional target detection method usually only focuses on single-granularity target recognition, and it is difficult to cope with the precise positioning task involving the detailed components of the object, resulting in insufficient perception accuracy and operation ability of the robot when performing complex operations. Therefore, the present disclosure aims to provide a cross-granularity target detection method to realize layer-by-layer accurate recognition from coarse-granularity targets to fine-granularity targets, solve the limitations of current target detection technology in the robot perception system, and especially solve the challenges faced when dealing with fine-granularity target recognition and precise positioning in complex environments.
[0034] The execution subject of the cross-granularity target detection method provided by the present disclosure includes but is not limited to a computer.
[0035] FIG. 1 shows a flowchart of the cross-granularity target detection method of the present disclosure example. As shown in FIG. 1, the cross-granularity target detection method provided by the embodiment of the present disclosure includes:
[0036] S1, obtaining a target image and target task information;
[0037] In step S1, the target image is generally associated with the target task. The target task can be a detection task for one or more detection targets in the target image, or an operation task for one or more detection targets in the target image. The target task information includes but is not limited to a text form. When the target task is to detect or operate an object with an entity structure, the detection target can specifically refer to the object, the component or assembly of the object. For example, the target task is "detecting the handle of the refrigerator", and the target image contains the refrigerator (which can contain the local or the whole of the refrigerator); or the target task is to open the refrigerator, and the target image contains the refrigerator (which can contain the local or the whole of the refrigerator). The target image can be obtained by using an image acquisition device such as a camera, a camera, etc., and the target task information can be received by using a task input interface. The method of the embodiment of the present disclosure can be executed by a processor, which is in communication connection with the image acquisition device and the task input interface.
[0038] S2, generating a detection target sequence according to the target image and the target task information, wherein the target image contains each detection target, and in the detection target sequence, the detection targets are sequentially contained in order;
[0039] In step S2, the target task can be understood by using artificial intelligence technology, and the visual features of the target image can be identified, and then a detection target sequence suitable for the target image and the target task information can be generated. For example, the semantic features of the target task information can be extracted by using a language model (for example, a visual model based on a Transformer architecture, a visual model based on a convolutional neural network (CNN), a language model based on a hybrid architecture, etc.), the visual features of the target image can be extracted by using a visual model (for example, a visual model based on a Transformer architecture, a visual model based on a convolutional neural network (CNN), etc.), and then the semantic features and the visual features can be fused, and a detection target sequence can be generated based on a language decoder according to the joint feature representation.
[0040] In the detection target sequence, each detection target can be presented in the form of a textual description, and of course, can also be presented in the form of a picture, code, symbol, etc., which is not limited in the embodiment.
[0041] In generating the detection target sequence, the target task and the target image are comprehensively considered to ensure that each detection target in the generated detection target sequence exists in the target image, and in order to complete the detection from coarse granularity to fine granularity, it is also necessary to ensure that each detection target in the detection target sequence is in a sequentially contained relationship. It can be understood that the cross-granularity target detection method of the present disclosure example is to detect each detection target in the detection target sequence as a target.
[0042] S3, according to the order of each detection target in the detection target sequence, each detection target in the target image is identified step by step by a recursive method.
[0043] In step S3, specifically, for the first detection target in the detection target sequence, the position of the first detection target in the target image is identified; according to the position of the first detection target, the target image is cropped to obtain a first sub-image containing the first detection target; the position of the second detection target in the detection target sequence in the first sub-image is identified; according to the position of the second detection target, the first sub-image is cropped to obtain a second sub-image containing the second detection target; and so on, until the position of the last detection target in the detection target sequence is obtained, and the accurate positioning of the target is realized.
[0044] Specifically, the workflow of detection starts from coarse-grained target detection, which first identifies the approximate type of the object, such as "door" or "cup", and generates a preliminary detection hint. The semantic information of the object is combined with the visual features to generate more detailed target detection hints. These hints are not just a preliminary description of the object, but they can guide the system to further explore the fine-grained features of the object. For example, after identifying the coarse-grained target "door", the semantic relationship of the object is analyzed to generate a detection hint for the fine-grained target "door handle", thereby guiding the subsequent detection to focus on the precise location of the door handle. Each step in this process is based on the results of the previous step, gradually refining the location and features of the target components, ensuring high-precision positioning of the target.
[0045] As can be seen, the cross-grained target detection method provided in the examples of the present disclosure realizes higher-precision target positioning from coarse-grained target detection to fine-grained target identification.
[0046] When applied to the field of robotics, this method can significantly improve the perception ability of the robot system, especially in tasks that require precise control and grasping, such as grasping door handles, operating bottle caps, etc. This enables the robot to achieve more efficient and accurate target recognition and operation in complex environments, not only improving the adaptability of the robot in tasks, but also increasing its execution precision.
[0047] In some embodiments, the above step S2 can include generating a detection chain according to the target image and the target task information through a target visual language model. Specifically, the target visual language model (VLM, Visual Language Model) can not only identify the visual features of the detected target in the image, but also generate more detailed detection hints by analyzing the semantic information of the detected target, further guiding the positioning of fine-grained detection targets. The target visual language model provides multi-level semantic and visual support for target detection, which can help understand and identify the location and features of fine-grained detection targets. For example, after identifying "door", the target visual language model generates a detection hint for the "door handle" component, significantly improving the accuracy and robustness of target detection.
[0048] FIG. 2 shows a flowchart of the process of the target visual language model of the examples of the present disclosure processing the target image and the target task information. As shown in FIG. 2, the target visual language model 01 can include a first visual encoder 11, a first language encoder 12, a first cross-modal feature fusion module 13, and a language decoder 14. The specific process of the target visual language model 01 generating a detection target sequence according to the target image and the target task information is as follows:
[0049] Step one, the first visual encoder 11 visually encodes the target image x to generate a first visual feature vector φ(x), which can be a high-dimensional visual feature vector.
[0050] Step two, the first language encoder 12 generates a first language feature vector ψ(S) according to the target task information S, which can be an embedding vector.
[0051] Step three, the first cross-modal feature fusion module 13 fuses the first visual feature vector φ(x) and the first language feature vector ψ(S) through a cross-modal attention mechanism to generate a first joint feature representation ξ(x, S).
[0052] Step four, the language decoder 14 decodes the first joint feature representation ξ(x, S) to generate a detection target sequence C = {c1, c2, …, c n}。
[0053] The order of the above steps one and two is not important, and they can be executed simultaneously. The formula representation of the above process is: φ(x) = VisionEncoder(x), ψ(S) = LanguageEncoder(S) ξ(x, S) = CrossModalFusion(φ(x), ψ(S)) C = LanguageDecoder(ξ(x, S)) = {c1, c2, …, c n}
[0054] In the formula, C is the generated detection target sequence, c i represents the i-th detection target in the detection target sequence, S is the natural language description (text description) of the target task, and x is the target image.
[0055] A target visual language model can be obtained by special training of a visual language model, further improving the adaptability to the detection target sequence generation task. FIG. 3 shows a flowchart of obtaining a target visual language model by special training of a visual language model according to an example of the present disclosure. As shown in FIG. 3, the training process is as follows:
[0056] S01, obtain a first labeled data set, wherein the first labeled data set includes at least two first labeled data, and each first labeled data includes an image, a task information and a detection target sequence;
[0057] In step S01, the first annotated data set is used for training, i.e., the first annotated data, and each first annotated data contains multiple granularity detection targets, such as "table", "drawer", "handle", etc. The target visual language model trained by the annotated data can more accurately generate a detection target sequence containing different granularity detection targets in the multi-modal reasoning process.
[0058] S02, training the preset visual language model according to each first annotated data in the first annotated data set to obtain the target visual language model.
[0059] In step S02, the preset visual language model can be any visual language model, and each first annotated data in the first annotated data set is used for recursive training of the preset visual language model. In the training process, the loss function is used to constrain the matching degree between the detection target sequence C generated by the visual language model and the real annotated detection target sequence C GT (First annotated data detection target sequence). GT
[0060] Through training, the visual language model can more accurately generate a multi-granularity target detection target sequence that meets the requirements. It can be seen that, by combining the powerful multi-modal understanding ability of the visual language model and the special optimization of the annotated data, a detection target sequence containing multi-granularity targets can be efficiently generated, which provides accurate task guidance for subsequent target accurate positioning.
[0061] In the detection target sequence generated by the above step S20, each detection target in the detection target sequence can be presented in a textual description, and the textual description of each detection target can only include the name of the detection target, or can also include other information of the detection target. For example, the specific form of the detection target sequence can be a sequence of detection target names. For example, a detection target sequence can be "table", "drawer", "handle".
[0062] As shown in FIG. 4, in some embodiments, the above step S3 can include:
[0063] S31, inputting the target image and the textual description of the first detection target in the detection target sequence into the target detector to obtain the position information of the first detection target in the target image output by the target detector;
[0064] S32, cropping the target image according to the position information of the first detection target in the target image to obtain a first sub-image containing the first detection target;
[0065] S33, input the first sub-image and the textual description of the second detection target in the detection target sequence into the target detector to obtain position information of the second detection target in the first sub-image output by the target detector;
[0066] S34, according to the position information of the second detection target in the first sub-image, crop the first sub-image to obtain a second sub-image containing the second detection target;
[0067] S35, and so on until the position information of the last detection target in the detection target sequence in the Nth sub-image is obtained, where N is a positive integer greater than 2.
[0068] Specifically, the target detector (Detector) can be flexibly adapted to different detection tasks through language prompts (Prompt). The target detector can be an open-vocabulary detection model (Open-Vocabulary Object Detection, OVOD). According to the order of the detection targets in the detection target sequence, the target detector is called to recognize the target image and its sub-image, recursively refine and detect the detection targets, and realize cross-granularity target detection.
[0069] Fig. 5 is a schematic diagram of the processing flow of the target detector in the example of the present disclosure for each input detection target and image. As shown in Fig. 5, the target detector 02 includes a second visual encoder 21, a second language encoder 22, a second cross-modal feature fusion module 23, and a detection head 24. The target detector 02 can simultaneously encode the input image and the textual description of the detection target. Specifically, the second visual encoder 21 performs second visual encoding on the input image to generate a second visual feature vector, and the second language encoder 22 performs second language encoding on the textual description of the detection target to generate a second language feature vector. The second cross-modal feature fusion module 23 fuses the second visual feature vector and the second language feature vector using a second cross-modal attention mechanism to obtain a second joint feature representation. The detection head 24 determines the position and category of the detection target in the image according to the second joint feature representation, where the category can be the name of the detection target recognized by the detection head. In this example, the target detector uses a unified detection head to generate the position and category of the detection target, supporting open-vocabulary detection tasks. The above process can be represented by the following formula:
[0070] wherein, is the detection result generated by the target detector on the input image x i and the textual description c i of the detection target.
[0071] For specific tasks, the detector can be fine-tuned using a labeled dataset with custom precision to obtain the aforementioned target detector, thereby enhancing its adaptability to specific tasks. Figure 6 is a schematic diagram of the training process of an example target detector provided in this disclosure. As shown in Figure 6, the target detector can be trained by performing the following steps:
[0072] S03. Obtain the second annotation dataset, wherein the second annotation dataset includes at least two second annotation datasets, each second annotation dataset including an image, a text description of a detected target and the location information of the detected target in the image, and the location information meets the target accuracy requirements;
[0073] In step S03, the second annotation dataset refers to the annotation dataset with custom precision. The second annotation data in the second annotation dataset is data annotated according to the custom precision. Specifically, the position information of the detected target in the image in each piece of second annotation data is annotated according to the custom position precision.
[0074] S04. Train the preset detector based on each second-annotated data in the second-annotated dataset to obtain the target detector.
[0075] In step S04, the preset detector can be any open-vocabulary detection model. The preset detector is recursively fine-tuned (or trained) using each second-labeled data point in the second-labeled dataset. The training process is optimized using the following loss function:
[0076] In the formula, The classification loss is the target category. The regression loss is the target bounding box (a type of location information).
[0077] In some embodiments, when the cross-granularity target detection method is applied to the field of robotics, the method may further include: performing motion control on the robotic arm based on the position information of each identified target, so that the robotic arm completes the target task. Specifically, robot motion path planning can be performed based on the position information of each identified target; for example, for the position information of each detected target, the center point coordinates of the finest-grained target are calculated and sent to the motion planning module of the robotic arm. The robotic arm performs kinematic planning based on the center point coordinates to complete the target operation (such as grasping, moving, etc.).
[0078] To better understand this disclosure, the following detailed description of the cross-granularity target detection system and method provided in this disclosure is provided through a specific embodiment.
[0079] I. System Architecture
[0080] FIG. 7 is a structural schematic diagram of the cross-granularity target detection system provided by the present embodiment. As shown in FIG. 7, the cross-granularity target detection system 03 of the present disclosure comprises the following main modules: a visual language model module 31, a chain of detection (CoD) module 32, and a target detector module 33.
[0081] 1. The visual language model module is responsible for generating a detection target sequence according to an input image and a detection task description. The visual language model module has strong multi-modal reasoning capability, including a visual encoder, a language encoder, a cross-modal interaction module, and a language decoder. The visual encoder is used to extract high-dimensional features of the image, the language encoder is used to extract language features of the detection task description, the cross-modal interaction module performs deep fusion of the visual features and the language features to understand the detection task and generate the detection target sequence, and the language decoder is used to generate a semantic detection target sequence. In the present system, the visual language model is further trained to improve its adaptability to the detection target sequence generation task. The special training uses a labeled detection target sequence dataset, which contains detection task samples of multi-granularity target components, such as "table", "drawer", "handle", etc. Through these labeled data, the model can more accurately generate detection target sequences containing different granularity components in the multi-modal reasoning process. The specific reasoning and training process is as follows:
[0082] (1) Image feature extraction: the input image x is processed by the visual encoder to generate a high-dimensional visual feature vector φ(x).
[0083] (2) Task description embedding: the detection task description S is input into the language encoder to generate an embedding vector ψ(S).
[0084] (3) Cross-modal feature fusion: the visual feature φ(x) and the task description feature ψ(S) are fused through a cross-modal attention mechanism to obtain a joint feature representation ξ(x, S).
[0085] (4) Detection target sequence generation: based on the joint feature representation ξ(x, S), the language decoder generates a target detection target sequence C.
[0086] The above process can be expressed by the following formulas: φ(x) = VisionEncoder(x), ψ(S) = LanguageEncoder(S) ξ(x, S) = CrossModalFusion(φ(x), ψ(S)) C = LanguageDecoder(ξ(x, S)) = {c1, c2, …, c n}
[0087] In the formula, C is the generated detection target sequence, c iLet S represent the i-th target component in the target sequence to be detected, and let S be a natural language description of the detection requirements. The input image is x. During specialized training, an annotated target sequence dataset is used, and the target sequence C generated by the model is constrained by a loss function to compare it with the ground truth annotated target sequence Ci. GT The degree of matching between them. The loss function can be defined as: L = CrossEntropy(C, C GT )
[0088] Through training, the model can more accurately generate multi-granularity target detection sequences that meet task requirements. This module combines the powerful multimodal understanding capabilities of visual language models with specialized optimization of labeled data, enabling efficient generation of target detection sequences containing multi-level target components, providing accurate task guidance for subsequent detector modules.
[0089] 2. The recursive detection module, a chain-like recursive detection module, is responsible for executing detection steps sequentially based on the target sequence generated by the visual language model module. The CoD module calls the target detector module to identify the image and its sub-images according to the order in the target sequence, recursively refining and detecting target components. Specifically, the CoD module achieves cross-granularity through the following steps:
[0090] (1) Execution of detection steps: Each detection step is performed in sequence according to the order in the detection target sequence.
[0091] (2) Target recognition: Call the target detector module to recognize the target components of the current detection step.
[0092] (3) Recursive call: For the identified target region, the next detection step is carried out until the entire target sequence is detected.
[0093] 3. Object Detector Module: The object detector module is responsible for identifying and locating each object component in the target sequence within the image or sub-image. This module employs an open-vocabulary detection model, which can flexibly adapt to different detection tasks through language prompts. The working principle of the open-vocabulary detection model is as follows:
[0094] (1) Multimodal feature extraction: The open vocabulary detection model encodes both the input image and the target component description. The image generates feature representations through a visual encoder, and the target component description generates embedding vectors through a language encoder.
[0095] (2) Cross-modal feature fusion: The open vocabulary detection model utilizes a cross-modal attention mechanism to deeply fuse image features with language features, thereby locating and classifying the target in the joint feature space.
[0096] (3) Target detection and localization: Based on the fused features, the open vocabulary detection model uses a unified detection head to generate the location and category of the target, supporting the detection task of open vocabulary.
[0097] The formula for the above process is expressed as follows:
[0098] in, For open vocabulary detection models in input subgraph x i and target component description c i The generated detection results indicate the location and category of the target area.
[0099] In this system, the open vocabulary detection model was fine-tuned using a labeled dataset with custom precision to enhance its adaptability to specific tasks. The fine-tuning process was optimized using the following loss function:
[0100] In the formula, The classification loss is the target category. The regression loss is the target bounding box.
[0101] II. Detailed Work Process
[0102] The entire testing process may include the following steps:
[0103] 1. Target Sequence Generation: The system utilizes a Visual Language Model (VLM) to generate a target sequence C = {c1, c2, ..., cn} based on the input image x and the detection task description S. The Visual Language Model extracts visual features from the input image and combines them with language cues to generate the target component sequence through multimodal feature fusion and contextual understanding.
[0104] The specific process of generating the target sequence using a visual language model is as follows:
[0105] (1) Image x is processed by a visual encoder to extract high-dimensional features φ(x).
[0106] (2) The detection task description S is used to generate an embedding vector ψ(S) by a language encoder.
[0107] (3) By combining image features φ(x) and language features ψ(S) through the cross-modal feature fusion module, a joint feature ξ(x,S) is generated.
[0108] (4) Based on joint features, generate the target detection sequence C = {c1, c2, ..., cn}. The formula is expressed as: C = VLM(x, S) = {c1, c2, ..., cn}, ξ(x, S) = CrossModalFusion(φ(x), ψ(S))
[0109] 2. Cross-granularity target recognition: For each detection cue (ct) in the target sequence, the system calls the target detector module to perform recognition in the image or sub-image. Based on the detection cue (ct) and the current image (xt), the target detector outputs the location information and category label of the target region. Through fine-tuning with open-vocabulary detection and labeled datasets, the target detector can support multi-granularity target detection tasks.
[0110] The detailed workflow of the target detector is as follows:
[0111] (1) Input the current image or sub-image xt and the detection prompt ct;
[0112] (2) The target detector extracts image features and linguistic features of the target description, and performs feature fusion through a cross-modal attention mechanism;
[0113] (3) Based on the fused multimodal features, the target detector outputs the bounding box of the target region. and category labels The target detector is fine-tuned using the following loss function:
[0114] in, The cross-entropy loss is the target class. The bounding box regression loss.
[0115] The formula for the above process is expressed as follows:
[0116] 3. Recursive calls and refinement: The system determines the next step c based on the detected target sequence. t+1 The target detector is recursively called to process the currently detected content d. t Further detailed testing will be conducted until all components are identified.
[0117] 4. Failure Handling and Remediation: If a detection step fails (e.g., the detector fails to identify the target), the system will attempt the following remedial measures: (a) adjust the camera position and retake the image; (b) increase the light source to improve image quality; (c) if the target still cannot be detected after multiple attempts, stop the operation and output a failure message to avoid erroneous actions.
[0118] 5. Robot Path Planning: For each detected target area, the system calculates the coordinates of the object's center point and sends them to the robotic arm's motion planning module. The robotic arm performs kinematic planning based on the target position to complete the target operation (such as grasping or moving).
[0119] III. Examples
[0120] 1. The cross-granularity target detection method provided in this disclosure is described in detail by means of cross-granularity detection of handles on a table.
[0121] (1) Input image: The system receives an input image x containing a table.
[0122] (2) Detection target sequence generation: VLM generates the detection target sequence C = {table, drawer, handle}.
[0123] (3) Cross-granularity detection: (a) Detect the "table" and extract the table region sub-image x1; (b) Detect the "drawer" and extract the drawer region sub-image x2; (c) Detect the "handle" and identify the target region. 4. Path planning: Send the coordinates of the center point of the identified "handle" to the robotic arm motion module to complete the grasping task.
[0124] 2. The cross-granularity target detection method provided in this disclosure is described in detail with reference to the illustrations.
[0125] Figures 8 and 9 are schematic diagrams illustrating the target detection process using a cross-granularity target detection method in an example of this disclosure. Figures 8 and 9 demonstrate how the recursive detection mechanism progressively detects object components from coarse-grained to fine-grained levels. Specifically, as shown in Figure 8, starting with detecting a "refrigerator," the system identifies the "refrigerator door" and "refrigerator handle," guiding the robot to accurately perform specific operational tasks. Similarly, as shown in Figure 9, for a plastic bag, the process progressively identifies the handle and the functional components required for grasping. Therefore, the cross-granularity target detection system and method provided by this disclosure possess powerful autonomous reasoning capabilities, allowing for progressive optimization of targets at different levels, thereby significantly improving the accuracy and robustness of the robot's perception capabilities. In other words, this invention can achieve progressive refinement of targets during the detection process, dynamically adjust the detection path, provide efficient visual and language support for the robot, and complete multi-granularity target detection and precise localization in complex tasks.
[0126] As can be seen, the cross-granularity target detection method and system provided in this disclosure have the following characteristics:
[0127] Visual Language Model (VLM) Guidance: This disclosure combines a Visual Language Model (VLM) with traditional object detection methods. During detection, the VLM not only provides visual features of objects but also generates more detailed detection cues by analyzing the semantic information of objects, further guiding the detector to locate fine-grained targets. The VLM provides multi-layered semantic and visual support for object detection, helping the system understand and recognize the location and features of fine-grained components. For example, after recognizing a "door," the VLM generates detection cues for the component "door handle," thereby guiding the detector to accurately locate and recognize detailed features. This technique significantly improves the accuracy and robustness of object detection.
[0128] Chained Recursive Detection (CoD): This disclosure proposes a chained recursive detection algorithm. The detection process generates a sequence of detection targets and progressively optimizes the target recognition path based on the hierarchical relationship of the target components. The recognition result of each layer serves as the input for the next layer, thereby progressively refining the target detection. Chained recursive detection can progressively refine the localization and recognition of target components according to different levels of the detection task, ensuring layer-by-layer optimization from the overall to the detailed. This process greatly improves the perception accuracy of robots in complex tasks and can effectively handle the detection and localization of fine-grained targets.
[0129] Dynamically Adjusting the Detection Path: The chain-based detection method in this disclosure dynamically adjusts the detection path based on the relationship between the overall structure and detailed components of an object, thereby optimizing the target recognition process. Dynamically adjusting the detection path allows for adjustments to the focus based on actual needs, shifting from coarse-grained overall targets to fine-grained component targets, ensuring the robot can flexibly handle different tasks in complex environments. For example, upon detecting a "door," the system will prioritize focusing on the detailed component "door handle" based on semantic information, thereby improving the efficiency and accuracy of task execution.
[0130] Recursive Refinement of Target Recognition: This disclosure employs a recursive refinement approach. The system recursively calls detectors according to the order of the detected target sequence, further refining each identified target region until all target components of the entire detected target sequence are identified. Recursive refinement ensures that each layer of target detection is optimized based on the previous step, progressively improving the accuracy of target recognition. This technique effectively avoids the accuracy deficiencies of traditional detection methods, especially when robots perform high-precision operations, ensuring the precise positioning of each target component.
[0131] Therefore, the cross-granularity target detection method and system provided in this disclosure have the following significant advantages:
[0132] Automated Detection: By combining a Visual Language Model (VLM) with a detector, this disclosure achieves an automated detection process that requires no human intervention. This not only significantly reduces the cost of manual detection but also improves the speed and accuracy of target detection. Traditional methods require extensive manual annotation and adjustments, while this invention greatly reduces the need for human intervention through intelligent reasoning and adaptive optimization. This disclosure utilizes a Visual Language Model to automatically generate detection cues and automatically identifies and locates target components through detection algorithms. Because VLM can simultaneously combine the visual and semantic information of an object, the detection process becomes more intelligent, reducing manual operations and improving the level of automation.
[0133] Refined Detection Capability: This disclosure, through the detection of target sequences and a recursive optimization mechanism, enables the system to generate fine-grained detection results, significantly improving the accuracy of target recognition and the ability to identify details. This is crucial for handling complex tasks, such as accurately grasping and locating object parts (e.g., "door handles" or "bottle caps"). The recursive detection algorithm allows the detection process to be progressively optimized, gradually moving from coarse-grained target recognition to precise localization of fine-grained parts. The recognition results at each layer serve as input for recursive refinement, ensuring accurate identification at the detail level, thereby improving the overall accuracy and detail recognition capability.
[0134] High Consistency and Efficiency: This disclosure ensures the consistency and efficiency of detection results through recursive interactive optimization, demonstrating a strong advantage, especially in large-scale detection tasks. The multi-level detection and recursive refinement process enables the system to perform large-scale target recognition and localization tasks quickly and stably, avoiding common false positives and false negatives. The recursive refinement mechanism ensures that the detection of each target is based on the optimization results of the previous step, allowing each round of detection to supplement and correct previous results, thus maintaining consistency. Furthermore, the system's layer-by-layer optimization and path adjustment improve detection efficiency, enabling it to maintain high performance in large-scale tasks.
[0135] High robustness: The fine-grained detection results generated by this disclosure cover the precise localization and spatial relationships of various object components, improving the robustness of the detector in different environments and complex scenes. Even in some extremely complex or varied scenes, the system can still maintain high detection accuracy and stability. By combining semantic information provided by the Visual Language Model (VLM), the detector can better understand the structure and details of objects, thereby achieving greater adaptability in complex scenes. In addition, the layer-by-layer optimization strategy of chain-recursive detection ensures that the target detection process is more robust, effectively avoiding misjudgments and omissions in complex environments.
[0136] Improving Detector Performance: This disclosure significantly improves the detector's performance in object recognition and cross-granularity perception tasks by training it with high-quality, fine-grained detection results. Through continuous optimization and iteration, the detector's performance is enhanced, especially in tasks requiring fine manipulation. High-quality, fine-grained labeled data is used to train the detector, allowing the model to achieve higher accuracy in target recognition and part localization. Fine-grained detection results provide richer training samples, enhancing the detector's adaptability to different target levels. By continuously optimizing the target detection path and improving detail recognition capabilities, the detector's performance in practical tasks is significantly improved.
[0137] Experimental verification shows that the cross-granularity target detection method and system provided in this disclosure perform excellently in robot operation tasks in cross-granularity scenarios. Compared with traditional detection methods, the success rate of operation on conventional objects is improved by 17.31%, while the success rate of operation on larger objects is improved by 51.39%, fully demonstrating its effectiveness and reliability in practical applications.
[0138] In summary, the cross-granularity target detection method and system provided in this disclosure achieve efficient and accurate automatic detection through the collaborative work of VLM and detector, combined with a chain recursive detection mechanism, and has broad application prospects and significant technical advantages.
[0139] It should be noted that the cross-granularity target detection system and method provided in this disclosure can be used in the field of robotics, or in any field other than robotics. This disclosure does not limit the application field of the cross-granularity target detection system and method.
[0140] Based on the same inventive concept, this disclosure also provides a cross-granularity target detection device, which can be used to implement the method described in the above embodiments, as shown in the following embodiments. Since the principle of the cross-granularity target detection device in solving the problem is similar to that of the above method, the implementation of the cross-granularity target detection device can refer to the implementation of the above method, and repeated details will not be elaborated further. As used below, the terms "unit" or "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0141] Figure 10 is a structural block diagram of a cross-grain size target detection device provided in an embodiment of the present disclosure. As shown in Figure 10, the cross-grain size target detection device 04 provided in this embodiment of the present disclosure includes:
[0142] Acquisition module 41 is used to acquire target image and target task information;
[0143] The generation module 42 is used to generate a detection target sequence based on the target image and target task information, wherein the target image contains each detection target, and the detection targets are sequentially contained in the detection target sequence.
[0144] The recognition module 43 is used to identify the location of each detection target in the target image step by step in a recursive manner, according to the order of each detection target in the detection target sequence.
[0145] Figure 11 is a schematic diagram of the physical structure of an electronic device provided in an embodiment of this disclosure. As shown in Figure 11, the electronic device 05 may include: a processor 51, a communication interface 52, a memory 53, and a communication bus 54. The processor 51, the communication interface 52, and the memory 53 communicate with each other through the communication bus 54. The processor 51 may call logical instructions in the memory 53 to execute the method of any of the above embodiments, such as: acquiring a target image and target task information; generating a detection target sequence based on the target image and target task information, wherein the target image contains each detection target, and the detection targets in the detection target sequence are sequentially contained within each other; and recursively identifying the location of each detection target in the target image according to the order of the detection targets in the detection target sequence.
[0146] Furthermore, the logical instructions in the aforementioned memory 53 can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0147] This embodiment of the application provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can perform the methods provided in the above-described method embodiments, such as: acquiring a target image and target task information; generating a detection target sequence based on the target image and target task information, wherein the target image contains each detection target, and in the detection target sequence, the detection targets are sequentially contained within each other; and recursively identifying the location of each detection target in the target image according to the order of the detection targets in the detection target sequence.
[0148] This embodiment provides a computer-readable storage medium storing a computer program that enables a computer to execute the methods provided in the above-described method embodiments. These methods include, for example,: acquiring a target image and target task information; generating a detection target sequence based on the target image and target task information, wherein each detection target is contained in the target image, and the detection target sequence contains each detection target sequentially; and recursively identifying the location of each detection target in the target image according to the order of the detection targets in the detection target sequence.
[0149] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0150] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams.
[0151] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0152] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0153] In the description of this specification, the references to terms such as "an embodiment," "a specific embodiment," "some embodiments," "for example," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0154] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of this disclosure. It should be understood that the above are only specific embodiments of this disclosure and are not intended to limit the scope of protection of this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for cross-granularity target detection, characterized in that, include: Acquire target image and target task information; A detection target sequence is generated based on the target image and target task information. Each detection target is contained in the target image, and the detection targets are sequentially contained within each other in the detection target sequence. Following the order of the detected targets in the target sequence, the location of each detected target in the target image is identified step by step using a recursive method.
2. The method according to claim 1, characterized in that, Generate a target detection sequence based on the target image and target task information, including: The target visual language model generates a detection target sequence based on the target image and target task information. The target visual language model is obtained by training a preset visual language model on first labeled data. The first labeled data includes the image and task information used as input to the model, and also includes the detection target sequence used as labels.
3. The method according to claim 2, characterized in that, The target visual language model generates a sequence of detected targets based on the target image and target task information, including: Perform a first visual encoding on the target image to generate a first visual feature vector; The target task information is encoded in the first language to generate a first language feature vector. The first visual feature vector and the first language feature vector are fused using the first cross-modal attention mechanism to obtain the first joint feature representation; Language decoding is performed on the first joint feature representation to obtain the detected target sequence.
4. The method according to claim 1, characterized in that, In the target detection sequence, each target is presented with a textual description; Following the order of the detected targets in the target sequence, the location of each detected target in the target image is identified step by step using a recursive method, including: Input the target image and the text description of the first detected target in the target detection sequence into the target detector to obtain the position information of the first detected target in the target image output by the target detector; Based on the position information of the first detected target in the target image, the target image is cropped to obtain the first sub-image containing the first detected target; Input the text description of the second detected target in the first sub-image and the detected target sequence into the target detector to obtain the position information of the second detected target in the first sub-image output by the target detector; Based on the position information of the second detected target in the first sub-image, the first sub-image is cropped to obtain a second sub-image containing the second detected target; This process continues until the position information of the last detected target in the target sequence in the Nth sub-image is obtained, where N is a positive integer greater than 2.
5. The method according to claim 4, characterized in that, The target detector processes the text description and image of the target input as follows: Perform second visual encoding on the image to generate a second visual feature vector; The textual description of the detected target is encoded in a second language to generate a second language feature vector; The second visual feature vector and the second language feature vector are fused using the second cross-modal attention mechanism to obtain the second joint feature representation; The location of the detected target in the image is determined based on the second joint feature representation.
6. The method according to claim 4, characterized in that, Before inputting the target image and the textual description of the first detected target in the target detection sequence into the target detector, the method further includes: Obtain second labeled data, which includes the image and text description of the detected target used as input to the model, and also includes position information used as labels, wherein the position information meets the target accuracy requirements; The target detector is obtained by training the preset detector based on the second labeled data.
7. A cross-particle size target detection device, characterized in that, include: The acquisition module is used to acquire target images and target task information; The generation module is used to generate a detection target sequence based on the target image and target task information. The target image contains each detection target, and the detection targets are sequentially contained within each other in the detection target sequence. The recognition module is used to identify the location of each detected target in the target image step by step using a recursive method, according to the order of each detected target in the target sequence.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the steps of any one of claims 1 to 6 when executing a computer program.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When a computer program is executed by a processor, it implements the steps of any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of any one of claims 1 to 6.