Target detection method and device, computer readable storage medium and electronic device

CN121392235BActive Publication Date: 2026-09-11UBTECH ROBOTICS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511341714.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2026-09-11
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

[0004]有鉴于此,本申请实施例提供了一种目标检测方法、装置、计算机可读存储介质及电子设备,以解决现有的目标检测方法中存在的不同任务类型之间相互干扰,导致目标检测结果的准确性较差的问题

Benefits of technology

[0051] The beneficial effects of this application embodiment compared with the prior art are as follows: This application embodiment obtains a task to be detected; wherein, the task to be detected includes an image to be detected and a corresponding task identifier, the task identifier being used to distinguish different task types; target detection is performed on the task to be detected using a preset target detection model to obtain a target detection result corresponding to the task to be detected; wherein, the target detection model includes a task adapter group, and different task adapter groups in the task adapter group correspond to different task identifiers, used to extract features related to the task type. Through this application embodiment, different task adapter groups corresponding to different task types are added to the target detection model, thereby enabling more targeted extraction of features related to the task type, effectively mitigating mutual interference between different task types, and obtaining more accurate target detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121392235B_ABST
    Figure CN121392235B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of target detection, and particularly relates to a target detection method and device, a computer readable storage medium and an electronic device. The method comprises: obtaining a to-be-detected task; wherein the to-be-detected task comprises a to-be-detected image and a corresponding task identifier, and the task identifier is used to distinguish different task types; performing target detection on the to-be-detected task through a preset target detection model to obtain a target detection result corresponding to the to-be-detected task; wherein the target detection model comprises a task adapter group, and different task adapter groups in the task adapter group correspond to different task identifiers. Through the application, different task adapter groups corresponding to different task types are added in the target detection model, so that features related to the task type can be extracted more specifically, mutual interference between different task types is effectively alleviated, and a more accurate target detection result can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of target detection technology, and in particular relates to a target detection method, apparatus, computer-readable storage medium and electronic device. Background Technology

[0002] Object detection is a key task in computer vision. Its core objective is to accurately locate and identify instances of objects of interest in a given image. Its output is usually a combination of bounding boxes and categories.

[0003] In existing technologies, there are many mature object detection methods that can achieve good detection results in general scenarios. However, in practical object detection applications, there are often situations where different task types interfere with each other. For example, for one task type, it is necessary to identify object A but not object B, while for another task type, it is necessary to identify object B but not object A. These interferences between different task types lead to poor accuracy of object detection results. Summary of the Invention

[0004] In view of this, embodiments of this application provide a target detection method, apparatus, computer-readable storage medium, and electronic device to solve the problem that interference between different task types in existing target detection methods leads to poor accuracy of target detection results.

[0005] A first aspect of this application provides a target detection method, which may include:

[0006] Obtain the task to be detected; wherein, the task to be detected includes the image to be detected and the corresponding task identifier, and the task identifier is used to distinguish different task types;

[0007] The target detection result is obtained by performing target detection on the target task using a preset target detection model.

[0008] The target detection model includes a task adapter group, in which different task adapter groups correspond to different task identifiers and are used to extract features related to the task type.

[0009] In one specific implementation of the first aspect, the target detection model may further include a shared image feature extraction network, a task mask feature extraction network, and a detection head; the step of performing target detection on the task to be detected through a preset target detection model to obtain a target detection result corresponding to the task to be detected may include:

[0010] The shared image feature extraction network is used to extract features from the image to be detected, thereby obtaining the shared image features of the detection task.

[0011] The task mask feature is obtained by extracting features from the task identifier through the task mask feature extraction network.

[0012] The shared image features and the task mask features are fused to obtain the task fusion features of the task to be detected.

[0013] Select the task adapter group corresponding to the task identifier from the task adapter group, and extract the task fusion features through the selected task adapter group to obtain the task adaptation features of the task to be detected.

[0014] The detection head performs target detection on the task adaptation features to obtain the target detection result corresponding to the task to be detected.

[0015] In one specific implementation of the first aspect, before performing target detection on the task to be detected using a preset target detection model, the method may further include:

[0016] The target detection model is trained in stages to obtain the trained target detection model.

[0017] In one specific implementation of the first aspect, the step of progressively training the target detection model in stages to obtain the trained target detection model may include:

[0018] During the first stage of training, the task mask feature extraction network is pruned, and the shared image feature extraction network, the task adapter group, and the detection head are trained to obtain the target detection model after the first stage of training.

[0019] In one specific implementation of the first aspect, the step of progressively training the target detection model in stages to obtain the trained target detection model may further include:

[0020] During the second stage of training, the parameters of the shared image feature extraction network are frozen, and the task mask feature extraction network is added. The task mask feature extraction network, the task adapter group, and the detection head are trained to obtain the target detection model after the second stage of training.

[0021] In one specific implementation of the first aspect, the step of progressively training the target detection model in stages to obtain the trained target detection model may further include:

[0022] In the third stage of training, the parameters of the shared image feature extraction network are unfrozen, and the target detection model is trained as a whole to obtain the trained target detection model.

[0023] In one specific implementation of the first aspect, the target detection method may further include:

[0024] During each stage of training, the training loss corresponding to different task types is determined.

[0025] The average training loss is determined based on the training loss corresponding to different task types;

[0026] Gradient backpropagation training is performed based on the average training loss.

[0027] A second aspect of the embodiments of this application provides a target detection device, which may include:

[0028] The task acquisition module is used to acquire the task to be detected; wherein, the task to be detected includes the image to be detected and the corresponding task identifier, and the task identifier is used to distinguish different task types;

[0029] The target detection module is used to perform target detection on the task to be detected using a preset target detection model, and obtain the target detection result corresponding to the task to be detected.

[0030] The target detection model may include a task adapter group, in which different task adapter groups correspond to different task identifiers and are used to extract features related to the task type.

[0031] In one specific implementation of the second aspect, the target detection model may further include a shared image feature extraction network, a task mask feature extraction network, and a detection head;

[0032] The target detection module may include:

[0033] A shared image feature extraction unit is used to extract features from the image to be detected through the shared image feature extraction network to obtain the shared image features of the detection task.

[0034] The task mask feature extraction unit is used to extract features from the task identifier through the task mask feature extraction network to obtain the task mask features of the task to be detected.

[0035] The feature fusion unit is used to fuse the shared image features and the task mask features to obtain the task fusion features of the task to be detected.

[0036] The task adaptation unit is used to select a task adapter group corresponding to the task identifier from the task adapter group, and to extract features from the task fusion features through the selected task adapter group to obtain the task adaptation features of the task to be detected.

[0037] The target detection unit is used to perform target detection on the task adaptation features through the detection head to obtain the target detection result corresponding to the task to be detected.

[0038] In one specific implementation of the second aspect, the target detection device may further include:

[0039] The phased training module is used to perform phased progressive training on the target detection model to obtain the trained target detection model.

[0040] In one specific implementation of the second aspect, the phased training module may include:

[0041] The first-stage training unit is used to prune the task mask feature extraction network during the first-stage training process, and train the shared image feature extraction network, the task adapter group, and the detection head to obtain the target detection model after the first-stage training.

[0042] In one specific implementation of the second aspect, the phased training module may further include:

[0043] The second-stage training unit is used to freeze the parameters of the shared image feature extraction network and add the task mask feature extraction network during the second-stage training process, and train the task mask feature extraction network, the task adapter group and the detection head to obtain the target detection model after the second-stage training.

[0044] In one specific implementation of the second aspect, the phased training module may further include:

[0045] The third-stage training unit is used to unfreeze the parameters of the shared image feature extraction network during the third-stage training process, and to train the target detection model as a whole to obtain the trained target detection model.

[0046] In one specific implementation of the second aspect, the phased training module may further include:

[0047] The gradient backpropagation training unit is used to determine the training loss corresponding to different task types in each stage of training; determine the average training loss based on the training loss corresponding to different task types; and perform gradient backpropagation training based on the average training loss.

[0048] A third aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the above-described target detection methods.

[0049] A fourth aspect of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the above-described target detection methods.

[0050] A fifth aspect of this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the steps of any of the above-described target detection methods.

[0051] The beneficial effects of this application embodiment compared with the prior art are as follows: This application embodiment obtains a task to be detected; wherein, the task to be detected includes an image to be detected and a corresponding task identifier, the task identifier being used to distinguish different task types; target detection is performed on the task to be detected using a preset target detection model to obtain a target detection result corresponding to the task to be detected; wherein, the target detection model includes a task adapter group, and different task adapter groups in the task adapter group correspond to different task identifiers, used to extract features related to the task type. Through this application embodiment, different task adapter groups corresponding to different task types are added to the target detection model, thereby enabling more targeted extraction of features related to the task type, effectively mitigating mutual interference between different task types, and obtaining more accurate target detection results. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 This is a flowchart of one embodiment of a target detection method in this application.

[0054] Figure 2 This is a flowchart illustrating the process of performing target detection on a task using a pre-defined target detection model.

[0055] Figure 3 This is a schematic diagram of the first stage of the training process;

[0056] Figure 4This is a schematic diagram of the second phase of the training process;

[0057] Figure 5 This is a structural diagram of one embodiment of a target detection device according to the present application.

[0058] Figure 6 This is a schematic block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0059] To make the inventive objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0060] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0061] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0062] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0063] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [the described condition or event] is detected," or "in response to detection of [the described condition or event]."

[0064] Furthermore, in the description of this application, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0065] Object detection is a key task in computer vision. Its core objective is to accurately locate and identify instances of objects of interest in a given image. Its output is usually a combination of bounding boxes and categories.

[0066] In existing technologies, there are many mature object detection methods that can achieve good detection results in general scenarios. However, in practical object detection applications, there are often situations where different task types interfere with each other. For example, for one task type, it is necessary to identify object A but not object B, while for another task type, it is necessary to identify object B but not object A. These interferences between different task types lead to poor accuracy of object detection results.

[0067] In view of this, embodiments of this application provide a target detection method, apparatus, computer-readable storage medium, and electronic device to solve the problem that interference between different task types in existing target detection methods leads to poor accuracy of target detection results.

[0068] In this embodiment, different task adapter groups corresponding to different task types are added to the target detection model, which can extract task-type related features more specifically, effectively alleviate the mutual interference between different task types, and obtain more accurate target detection results.

[0069] Please see Figure 1 One embodiment of a target detection method in this application may include:

[0070] Step S101: Obtain the task to be detected.

[0071] The task to be detected includes the image to be detected and the corresponding task identifier (task_id). The task identifier is used to distinguish different task types. For example, task identifier 1 (Task 1) corresponds to the first task type, task identifier 2 (Task 2) corresponds to the second task type, and so on. The task identifiers for different task types are different.

[0072] Step S102: Perform target detection on the target to be detected using a preset target detection model to obtain the target detection result corresponding to the target to be detected.

[0073] The object detection model includes task adapter groups, where different task adapter groups correspond to different task identifiers and are used to extract features related to the task type.

[0074] In one specific implementation of this application embodiment, the target detection model may also include, but is not limited to, a shared image feature extraction network, a task mask feature extraction network, and a detection head, such as... Figure 2 As shown, step S102 can specifically include the following process:

[0075] Step S1021: Extract features from the image to be detected using a shared image feature extraction network to obtain the shared image features of the detection task.

[0076] The shared image feature extraction network may include, but is not limited to, a backbone network and a neck network.

[0077] The specific network structure of the backbone network can be flexibly configured according to actual conditions, and this application embodiment does not impose specific limitations on it. As an example, the YOLOv8 network is used here for illustration. The processing of the YOLOv8 network includes three stages, referred to as the first stage (stage1), the second stage (stage2), and the third stage (stage3). In the first stage, feature extraction is performed on the image to be detected to obtain the image features of the first stage, whose feature size can be 1×80×80×128; in the second stage, further feature extraction is performed on the image features of the first stage to obtain the image features of the second stage, whose feature size can be 1×40×40×256; in the third stage, further feature extraction is performed on the image features of the second stage to obtain the image features of the third stage, whose feature size can be 1×20×20×512.

[0078] In this embodiment, the neck network can further fuse the image features (first-stage, second-stage, and third-stage image features) output by the backbone network at each stage. The specific network structure of the neck network can be flexibly configured according to actual conditions, and may include, but is not limited to, Path Aggregation Network (PANet), etc. This embodiment does not impose specific limitations on this. The output of the neck network can be three different sizes of image features (e.g., 1 / 8 resolution, 1 / 16 resolution, and 1 / 32 resolution), which are referred to here as shared image features.

[0079] Step S1022: Extract features from the task identifier using the task mask feature extraction network to obtain the task mask features of the task to be detected.

[0080] The task mask feature extraction network may include, but is not limited to, a task embedding module (task_emb), a mapping module (reflect_block), and an extension module (ex_block).

[0081] The `task_emb` parameter converts the task identifier into a low-dimensional, dense, continuous vector representation, denoted as the task embedding vector. This vector captures the feature information of different task types, allowing comparisons between different task types within a shared vector space. Let N be the total number of task types, then a trainable `task_embedding` of size N×64 can be defined for each of these N task types.

[0082] Corresponding to the three stages in the shared image feature extraction network, reflect_block processes task_embedding sequentially using one-dimensional convolution (Conv1d), group normalization (GroupNorm), and rectified linear function (ReLU) to obtain N×128, N×256, and N×512 features. Then, through unsqueeze processing in the height (H) and width (W) directions, it is further expanded to N×80×80×128, N×40×40×256, and N×20×20×512 features, which are then input into ex_block for further processing.

[0083] The `ex_block` can contain five 3×3 convolutions. The parameters of the last convolution can be initialized with all zeros (Zero-init) to ensure a smooth transition from training without `task_id` to training with `task_id`. After processing by the five 3×3 convolutions, the output features are N×80×80×128, N×40×40×256, and N×20×20×512, which are denoted here as the task mask features (`task_mask`).

[0084] Step S1023: Perform feature fusion on the shared image features and task mask features to obtain the task fusion features of the task to be detected.

[0085] The specific feature fusion method used can be flexibly set according to the actual situation, and this application does not impose specific limitations on it. As an example, shared image features and task mask features can be directly concatenated to obtain the task fusion features of the task to be detected.

[0086] Step S1024: Select the task adapter group corresponding to the task identifier in the task adapter group, and extract the task fusion features through the selected task adapter group to obtain the task adaptation features of the task to be detected.

[0087] The task adapter group consists of N task adapter groups, each corresponding to one of the N task identifiers. Corresponding to the three different sizes of input features, each task adapter group can have three task adapters, and each adapter can process one of the input feature sizes.

[0088] The structural modules in the Adapter can include, but are not limited to, convolutional modules, C2f modules (CSP Bottleneck with 2 Convolutions), etc. Considering that the core of the Adapter is to extract channel features, channel attention modules can also be added after the aforementioned modules, including but not limited to squeeze-and-excitation (SE) networks.

[0089] Step S1025: Perform target detection on the task adaptation features using the detection head to obtain the target detection result corresponding to the task to be detected.

[0090] In this embodiment, a corresponding detection head can be set for each task adapter group. The task adaptation features output by the selected task adapter group are input into its corresponding detection head for processing, thereby obtaining the target detection result corresponding to the task to be detected, including the target bounding box and category.

[0091] In one specific implementation of this application, before performing target detection on the target detection task using the target detection model, the target detection model can be pre-trained in stages to obtain the trained target detection model.

[0092] Specifically, during the first phase of training, such as Figure 3 As shown, the task mask feature extraction network can be pruned, and the shared image feature extraction network, task adapter group, and detection head can be trained to obtain the target detection model after the first stage of training. Since the task mask feature extraction network is pruned, the output of the shared image feature extraction network can be directly used as the input of the task adapter group.

[0093] During the first phase of training, the training loss can be determined for different task types. This training loss may include, but is not limited to, classification loss. cls ) and target bounding box regression loss (Loss objFor each training batch, the training loss for Task 1 can be denoted as Loss1, the training loss for Task 2 as Loss2, ..., the training loss for Task N as LossN, and so on. Then, the average training loss can be determined based on the training losses corresponding to different task types. For example, the mean of Loss1, Loss2, ..., LossN can be determined as the average training loss, and gradient backpropagation training can be performed based on the average training loss to obtain the object detection model after the first stage of training.

[0094] During the second phase of training, such as Figure 4 As shown, the parameters of the shared image feature extraction network can be frozen, and a task mask feature extraction network can be added. The task mask feature extraction network, task adapter group, and detection head are trained to obtain the target detection model after the second stage of training.

[0095] In the third stage of training, the parameters of the shared image feature extraction network can be unfrozen, and the object detection model can be trained as a whole to obtain the trained object detection model.

[0096] It is easy to understand that the calculation of training loss and gradient backpropagation training in the second and third training stages are similar to the training process in the first stage. For details, please refer to the detailed process mentioned above, which will not be repeated here.

[0097] In the first training phase, the entire network is difficult to train sufficiently, often converging after only a few epochs. Based on this insufficiently trained model, in the second training phase, the parameters of the shared image feature extraction network can be frozen, and a task_id can be introduced for each task to help the Adapter train fully, enabling it to learn the ability to extract features relevant to the corresponding task. Once the Adapter has completed sufficient training, in the third training phase, the parameters of the shared image feature extraction network can be enabled, and the initial learning rate can be appropriately lowered. This allows the entire network to undergo sufficient fine-tuning under relatively optimal initial conditions, effectively mitigating task conflicts and gradually achieving the desired model where the shared image feature extraction network is responsible for extracting all task-related features, while the task mask feature extraction network and the Adapter are responsible for extracting relevant task features.

[0098] After the target detection model is trained in stages, the trained target detection model can be used to perform target detection on the target task, thereby obtaining the target detection result corresponding to the target task.

[0099] In one specific implementation of this application, after obtaining the target detection result, a preset object grasping mechanism can be controlled to perform a grasping task on the target object based on the target detection result. The object grasping mechanism may include, but is not limited to, a robot's end effector.

[0100] In summary, this application embodiment obtains a task to be detected; wherein, the task to be detected includes an image to be detected and a corresponding task identifier, the task identifier being used to distinguish different task types; target detection is performed on the task to be detected using a preset target detection model to obtain a target detection result corresponding to the task to be detected; wherein, the target detection model includes a task adapter group, different task adapter groups in the task adapter group corresponding to different task identifiers, used to extract features related to the task type. Through this application embodiment, different task adapter groups corresponding to different task types are added to the target detection model, thereby enabling more targeted extraction of features related to the task type, effectively mitigating mutual interference between different task types, and obtaining more accurate target detection results.

[0101] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0102] Corresponding to the target detection method described in the above embodiments, Figure 5 This diagram illustrates a structural diagram of one embodiment of a target detection device provided in this application.

[0103] In this embodiment, a target detection device may include:

[0104] The task acquisition module 501 is used to acquire the task to be detected; wherein, the task to be detected includes the image to be detected and the corresponding task identifier, and the task identifier is used to distinguish different task types;

[0105] The target detection module 502 is used to perform target detection on the task to be detected using a preset target detection model, and obtain the target detection result corresponding to the task to be detected.

[0106] The target detection model may include a task adapter group, in which different task adapter groups correspond to different task identifiers and are used to extract features related to the task type.

[0107] In one specific implementation of this application embodiment, the target detection model may further include a shared image feature extraction network, a task mask feature extraction network, and a detection head;

[0108] The target detection module may include:

[0109] A shared image feature extraction unit is used to extract features from the image to be detected through the shared image feature extraction network to obtain the shared image features of the detection task.

[0110] The task mask feature extraction unit is used to extract features from the task identifier through the task mask feature extraction network to obtain the task mask features of the task to be detected.

[0111] The feature fusion unit is used to fuse the shared image features and the task mask features to obtain the task fusion features of the task to be detected.

[0112] The task adaptation unit is used to select a task adapter group corresponding to the task identifier from the task adapter group, and to extract features from the task fusion features through the selected task adapter group to obtain the task adaptation features of the task to be detected.

[0113] The target detection unit is used to perform target detection on the task adaptation features through the detection head to obtain the target detection result corresponding to the task to be detected.

[0114] In one specific implementation of this application embodiment, the target detection device may further include:

[0115] The phased training module is used to perform phased progressive training on the target detection model to obtain the trained target detection model.

[0116] In one specific implementation of this application embodiment, the staged training module may include:

[0117] The first-stage training unit is used to prune the task mask feature extraction network during the first-stage training process, and train the shared image feature extraction network, the task adapter group, and the detection head to obtain the target detection model after the first-stage training.

[0118] In one specific implementation of this application embodiment, the phased training module may further include:

[0119] The second-stage training unit is used to freeze the parameters of the shared image feature extraction network and add the task mask feature extraction network during the second-stage training process, and train the task mask feature extraction network, the task adapter group and the detection head to obtain the target detection model after the second-stage training.

[0120] In one specific implementation of this application embodiment, the phased training module may further include:

[0121] The third-stage training unit is used to unfreeze the parameters of the shared image feature extraction network during the third-stage training process, and to train the target detection model as a whole to obtain the trained target detection model.

[0122] In one specific implementation of this application embodiment, the phased training module may further include:

[0123] The gradient backpropagation training unit is used to determine the training loss corresponding to different task types in each stage of training; determine the average training loss based on the training loss corresponding to different task types; and perform gradient backpropagation training based on the average training loss.

[0124] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0125] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0126] Figure 6 A schematic block diagram of an electronic device provided in an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiment of this application are shown.

[0127] like Figure 6 As shown, the electronic device 6 in this embodiment includes: a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60. When the processor 60 executes the computer program 62, it implements the steps in the various target detection method embodiments described above, for example... Figure 1 Steps S101 to S102 are shown. Alternatively, when the processor 60 executes the computer program 62, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 5 The functions of modules 501 to 502 are shown.

[0128] For example, the computer program 62 may be divided into one or more modules / units, which are stored in the memory 61 and executed by the processor 60 to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program 62 in the electronic device 6.

[0129] The electronic device 6 may include, but is not limited to, computing devices such as mobile phones, tablets, desktop computers, laptops, handheld computers, robots, and servers. Those skilled in the art will understand that... Figure 6 This is merely an example of electronic device 6 and does not constitute a limitation on electronic device 6. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device 6 may also include input / output devices, network access devices, buses, etc.

[0130] The processor 60 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0131] The memory 61 can be an internal storage unit of the electronic device 6, such as a hard disk or memory. The memory 61 can also be an external storage device of the electronic device 6, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory 61 can include both internal and external storage units of the electronic device 6. The memory 61 is used to store the computer program and other programs and data required by the electronic device 6. The memory 61 can also be used to temporarily store data that has been output or will be output.

[0132] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0133] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0134] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0135] In the embodiments provided in this application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0136] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0137] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0138] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable storage medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.

[0139] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A target detection method, characterized in that, include: Obtain the task to be detected; wherein, the task to be detected includes the image to be detected and the corresponding task identifier, and the task identifier is used to distinguish different task types; The shared image features of the target detection task are obtained by extracting features from the image to be detected through the shared image feature extraction network in the preset target detection model. The task identifier is extracted using the task mask feature extraction network in the target detection model to obtain the task mask feature of the task to be detected. The shared image features and the task mask features are fused to obtain the task fusion features of the task to be detected. In the task adapter group of the target detection model, a task adapter group corresponding to the task identifier is selected, and the task fusion feature is extracted by the selected task adapter group to obtain the task adaptation feature of the task to be detected; different task adapter groups in the task adapter group correspond to different task identifiers and are used to extract features related to the task type. The target detection head in the target detection model is used to perform target detection on the task-adaptive features to obtain the target detection result corresponding to the task to be detected.

2. The target detection method according to claim 1, characterized in that, Also includes: The target detection model is trained in stages to obtain the trained target detection model.

3. The target detection method according to claim 2, characterized in that, The step of progressively training the target detection model in stages to obtain the trained target detection model includes: During the first stage of training, the task mask feature extraction network is pruned, and the shared image feature extraction network, the task adapter group, and the detection head are trained to obtain the target detection model after the first stage of training.

4. The target detection method according to claim 3, characterized in that, The step of progressively training the target detection model in stages to obtain the trained target detection model includes: During the second stage of training, the parameters of the shared image feature extraction network are frozen, and the task mask feature extraction network is added. The task mask feature extraction network, the task adapter group, and the detection head are trained to obtain the target detection model after the second stage of training.

5. The target detection method according to claim 4, characterized in that, The step of progressively training the target detection model in stages to obtain the trained target detection model includes: In the third stage of training, the parameters of the shared image feature extraction network are unfrozen, and the target detection model is trained as a whole to obtain the trained target detection model.

6. The target detection method according to any one of claims 2 to 5, characterized in that, Also includes: During each stage of training, the training loss corresponding to different task types is determined. The average training loss is determined based on the training loss corresponding to different task types; Gradient backpropagation training is performed based on the average training loss.

7. A target detection device, characterized in that, include: The task acquisition module is used to acquire the task to be detected; wherein, the task to be detected includes the image to be detected and the corresponding task identifier, and the task identifier is used to distinguish different task types; The target detection module is used to extract features from the image to be detected using a shared image feature extraction network in a preset target detection model to obtain shared image features of the task to be detected; to extract features from the task identifier using a task mask feature extraction network in the target detection model to obtain task mask features of the task to be detected; to fuse the shared image features and the task mask features to obtain task fusion features of the task to be detected; to select a task adapter group corresponding to the task identifier from the task adapter group in the target detection model, and to extract features from the task fusion features using the selected task adapter group to obtain task adaptation features of the task to be detected; different task adapter groups in the task adapter group correspond to different task identifiers and are used to extract features related to the task type; and to perform target detection on the task adaptation features using the detection head in the target detection model to obtain target detection results corresponding to the task to be detected.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the target detection method as described in any one of claims 1 to 6.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the target detection method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-task target detection method and device, automatic driving system and storage medium

    CN114821269A

  • SAM2 multi-task perception binary segmentation method based on hybrid expert adapter

    CN120526149A