Detection model training method and device and target detection method and device
Unifying multi-level data sets through the label mapping method solves the shortcomings of detection models in the power industry in terms of target recognition and accuracy, and realizes a more comprehensive target detection and a more efficient reasoning process.
Patent Information
- Application Number
- CN202510339721.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the detection model based on a single mode has problems such as incomplete target recognition, insufficient detection accuracy, and low model inference efficiency in automated inspections in the power industry, especially in the identification of multiple categories.
Through the label mapping method, the sample labels of multi-level data sets are unified to obtain joint data sets, thereby reducing the requirements for full data annotation of each data set, saving labor costs, and enabling the model to conduct more comprehensive target detection.
It is realized that without increasing the annotation cost, it can cover various types of information more comprehensively, thereby improving the overall accuracy and inference efficiency of the detection model.
Smart Images

Figure CN119992070A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to technical fields such as large model training, and in particular to a detection model training method and device and a target detection method and device. Background Art
[0002] With the development of artificial intelligence and drone technology, automated inspection of distribution network lines has become an important task in the power industry. However, in existing technologies, detection models based on a single modality (such as images) face problems such as incomplete target recognition, insufficient detection accuracy, and low model reasoning efficiency. In the power sector, due to the incompleteness of data sets and the limitation of annotation costs, how to identify multiple categories in one reasoning is still a difficult problem to be solved.
[0003] In traditional target detection tasks, dataset construction often adopts a divide-and-conquer approach, with each level of data independently labeled. For example, for the tower layer dataset, only cement poles, steel pipe poles, etc. are labeled. In this way, the training set is also constructed separately, and each level of the dataset is trained independently, and finally multiple different models are generated. However, due to the different features learned by the category detection algorithms at different levels, the difficulty of training a task learning algorithm alone increases, and the overall accuracy of the algorithm is often low. If all categories are fully labeled, not only will the labeling cost be huge, but there are too many labeled categories, and it is difficult to fully cover all category information. Summary of the invention
[0004] To this end, the purpose of the implementation mode of the present application is to propose a detection model training method, device, target detection method, device, electronic device, storage medium and computer program product, which can unify the sample labels of multi-level types of data sets through a label mapping method to obtain a joint data set. This reduces the requirements for full data labeling of each data set, saves labor costs, and can comprehensively cover various types of information, so that the trained model can perform more comprehensive target detection.
[0005] An embodiment of the present application provides a method for training a detection model, the method comprising: acquiring a multi-level data set; performing label mapping on sample labels of the multi-level data set to obtain a target mapping label set; obtaining a joint data set based on the target mapping label set; training a preset detection model based on the joint data set and a joint loss function corresponding to the joint data set to obtain a trained detection model.
[0006] Exemplarily, each hierarchical data set includes multiple category data sets, and label mapping is performed on sample labels of the multi-level data sets to obtain a target mapping label set, including: establishing a sample index for each sample; determining the sample category label ID based on the sample index, and merging sample label IDs of the same category to obtain a multi-category label set; label mapping is performed on the multi-category label set to obtain a first mapping label set corresponding to the hierarchical data set; and a target mapping label set is obtained based on the first mapping label set.
[0007] Exemplarily, the multi-level data set includes at least one of a device layer data set, a component layer data set and a tower layer data set, and obtaining a target mapping label set based on the first mapping label set includes: merging the first mapping label set corresponding to the device layer data set, the first mapping label set corresponding to the component layer data set, and the first mapping label set corresponding to the tower layer data set to obtain a target mapping label set.
[0008] Exemplarily, obtaining a joint dataset based on the target mapping label set includes: determining labels after sample mapping based on the target mapping label set and the sample index; and obtaining a joint dataset based on samples in the category dataset and the mapped labels corresponding to the samples.
[0009] Exemplarily, the joint data set includes multiple sub-data sets, and the preset detection model is trained based on the joint data set and the joint loss function corresponding to the joint data set to obtain a trained detection model, including: predicting samples of each sub-data set in the joint data set to obtain a prediction result; determining the loss function corresponding to the sub-data set based on the prediction result; weighting the loss function corresponding to the sub-data set to obtain a joint loss function corresponding to the joint data set; training the preset detection model based on the joint loss function corresponding to the joint data set to obtain a trained detection model.
[0010] Exemplarily, the prediction result includes a category prediction result and a detection box prediction result; determining the loss function corresponding to the sub-dataset based on the prediction result includes: obtaining a classification loss function based on the category prediction result and the category reference result, and obtaining a detection box loss function based on the detection box prediction result and the detection box reference result; determining the loss function corresponding to the sub-dataset based on the classification loss function and the detection box loss function.
[0011] Exemplarily, obtaining a classification loss function based on the category prediction result and the category reference result includes: determining a prediction probability corresponding to a category based on the category prediction result and the category reference result; and weighting the prediction probability corresponding to the category to obtain a classification loss function.
[0012] Exemplarily, the detection frame prediction result includes a detection frame center point prediction result and a detection frame size prediction result, the detection frame reference result includes a detection frame center point reference result and a detection frame size reference result, and obtaining a detection frame loss function based on the detection frame prediction result and the detection frame reference result includes: determining a detection frame distance loss function based on the detection frame center point prediction result and the detection frame center point reference result; determining a detection frame matching loss function based on the detection frame size prediction result and the detection frame size reference result; and obtaining a detection frame loss function based on the detection frame distance loss function and the detection frame matching loss function.
[0013] Exemplarily, the detection model includes a shared backbone network submodel and multiple task detection head submodels, and the task detection head model is coupled with the shared backbone network submodel to obtain a trained detection model. The method further includes: decoupling the shared backbone network submodel and the multiple task detection head submodels according to the category length corresponding to the task detection head submodel to obtain multiple detection head submodels; wherein the multiple detection head submodels correspond one-to-one to the multiple task detection head submodels, and the detection head submodels include one of the shared backbone network submodels and one of the task detection head submodels, and the detection head submodels share the shared backbone network submodel.
[0014] Another embodiment of the present application provides a target detection method, which includes: obtaining a picture to be detected; detecting the picture to be detected based on a detection model trained according to the above-mentioned detection model training method to obtain a target detection result.
[0015] Exemplarily, the detection model includes a shared backbone network sub-model and multiple task detection head sub-models, and the multiple task detection head sub-models are all connected to the shared backbone network sub-model. The detection model-based detection of the image to be detected to obtain a target detection result includes: extracting features of the image to be detected based on the shared backbone network sub-model to obtain image features; detecting the image features based on the multiple task detection head sub-models to obtain multiple first detection results corresponding to the task detection head sub-models; and merging the first detection results to obtain the target detection result.
[0016] Another embodiment of the present application provides a training device for a detection model, the device comprising: a first acquisition module, used to acquire a multi-level data set; a mapping module, used to perform label mapping on sample labels of the multi-level data set to obtain a target mapping label set; an acquisition module, used to obtain a joint data set based on the target mapping label set; a training module, used to train a preset detection model based on the joint data set and a joint loss function corresponding to the joint data set to obtain a trained detection model.
[0017] Another embodiment of the present application provides a target detection device, which includes: a second acquisition module, used to acquire a picture to be detected; a detection module, used to detect the picture to be detected based on a detection model trained by the training device according to the above-mentioned detection model, to obtain a target detection result.
[0018] Another embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method of any of the above embodiments when executing the computer program.
[0019] Another embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method of any of the above embodiments are implemented.
[0020] Another embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed by a processor of a computer device, the computer device is enabled to perform the steps of the method of any of the above embodiments.
[0021] In the above implementation, the training method of the detection model includes: obtaining a multi-level data set; performing label mapping on the sample labels of the multi-level data set to obtain a target mapping label set; obtaining a joint data set based on the target mapping label set; training a preset detection model based on the joint data set and a joint loss function corresponding to the joint data set to obtain a trained detection model. The training method of the detection model of the present invention can unify the sample labels of multi-level types of data sets through a label mapping method to obtain a joint data set, which reduces the requirements for full data annotation of each data set, saves labor costs, and can comprehensively cover various types of information, so that the trained model can perform more comprehensive target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A flowchart of a method for training a detection model provided in an embodiment of the present application;
[0023] Figure 2 A flowchart of label mapping for sample labels of a multi-level data set provided in an embodiment of the present application;
[0024] Figure 3 A schematic diagram of a label mapping solution provided for an implementation manner of the present application;
[0025] Figure 4 A flowchart of obtaining a joint data set based on a target mapping label set provided in an embodiment of the present application;
[0026] Figure 5 A flowchart for training a preset detection model provided in an embodiment of the present application;
[0027] Figure 6 A flowchart of determining a loss function corresponding to a sub-data set provided in an embodiment of the present application;
[0028] Figure 7 A schematic diagram of obtaining a classification loss function provided in an embodiment of the present application;
[0029] Figure 8 A flowchart of obtaining a detection box loss function provided in an embodiment of the present application;
[0030] Fig. 9 A schematic diagram of the forward propagation of the model provided in the embodiment of the present application during the training phase;
[0031] Fig.10 A schematic diagram of the detection model conversion provided in the embodiment of the present application;
[0032] Fig.11 A flowchart of the overall technical solution provided for the implementation method of this application;
[0033] Fig.12 A flow chart of a target detection method provided in an embodiment of the present application;
[0034] Fig.13 A flowchart for detecting an image to be detected provided in an embodiment of the present application;
[0035] Fig.14 A schematic diagram of a detection model provided in an embodiment of the present application detecting a picture set;
[0036] Fig.15 A schematic diagram of a training device for a detection model provided in an embodiment of the present application;
[0037] Fig.16 A schematic diagram of a target detection device provided in an embodiment of the present application;
[0038] Fig.17 A block diagram of an electronic device provided for an embodiment of the present application. DETAILED DESCRIPTION
[0039] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0040] In traditional target detection tasks, dataset construction often adopts a divide-and-conquer approach, with each level of data independently labeled. For example, for the tower layer dataset, only cement poles, steel pipe poles, etc. are labeled. In this way, the training set is also constructed separately, and each level of the dataset is trained independently, and finally multiple different models are generated. However, due to the different features learned by the category detection algorithms at different levels, the difficulty of training a task learning algorithm alone increases, and the overall accuracy of the algorithm is often low. If all categories are fully labeled, not only will the labeling cost be huge, but there are too many labeled categories, and it is difficult to fully cover all category information.
[0041] To achieve this goal, traditional detection algorithms usually need to collect huge data sets to cover all categories for training. However, in the power industry, data sets with complete category information are very scarce, and supplementing this annotation information often requires a lot of manpower. In addition, such data sets may contain a lot of noise, which reduces the detection effect of the model.
[0042] In some examples, the complete set of categories can be split into multiple sub-datasets, each of which contains only part of the category information. By increasing the number of data sets, the label recognition range of the detector is gradually expanded. For example, the corresponding data sets for each category are collected separately, and a detection model is trained separately for each data set. In the inference stage, the detection output covering all categories is obtained by fusing the prediction results of multiple models. Although this method of multi-model training by dividing the data set into categories is feasible, its disadvantages are long training cycle, high algorithm complexity, low detection accuracy, and the need to load multiple models during inference, which requires high computing power and is difficult to meet the efficiency requirements in practical application.
[0043] In some examples, the training process can also be divided into two steps, training a detector for each sub-dataset and using the detector to infer other datasets. Then, the prediction results (pseudo labels) of all datasets are integrated with their true labels to form a comprehensive dataset containing all categories. A new model is trained on this comprehensive dataset so that it can learn information from all categories. However, differences in labels and data distribution between different datasets may cause the fused dataset to be inconsistent, and the inference results may contain false detections, which will affect the training effect of the final model. In addition, if the same category is in multiple datasets, generating duplicate outputs for the same object will affect the next round of model training for the full category dataset.
[0044] Based on this, this application proposes a method for training a detection model, introduces multimodal visual large model technology into the UAV inspection task of the power industry, and combines the technical knowledge of the power industry to propose an improved data set construction method and training method to reduce the requirements for full data labeling. It achieves the unification of data at all levels without repeated labeling, so that all target categories in the data set participate in one-time training and efficiently identify objects.
[0045] Figure 1 It is a flowchart of a method for training a detection model according to an embodiment of the present application.
[0046] As an example, Figure 1 As shown, the training method of the detection model includes:
[0047] S101, obtaining a multi-level data set.
[0048] S102, performing label mapping on sample labels of the multi-level data set to obtain a target mapping label set.
[0049] S103, obtaining a joint data set based on the target mapping label set.
[0050] S104, training a preset detection model based on the joint data set and a joint loss function corresponding to the joint data set to obtain a trained detection model.
[0051] Exemplarily, a multi-level data set is first obtained. The data set can be obtained through inspection equipment, such as pictures taken by drones. To construct a training data set, the pictures need to be annotated in advance. Since it is difficult to annotate a picture in full and it is easy to miss positive samples, a large amount of image data obtained in the distribution network drone patrol mission needs to be manually annotated for subsequent target detection model training. Among them, the manual annotation can use the LabelBee annotation tool to divide the data into component layer, equipment layer and tower layer according to the target type for parallel annotation to obtain a multi-level data set.
[0052] Exemplarily, the sample label categories of each hierarchical data set are different. For example, the tower layer data set includes annotations related to the materials and functions of power towers, and the equipment layer data set includes annotations covering equipment such as transformers and distribution switchgear. The distribution of labels of different categories varies greatly. The present application proposes a simple label mapping method to merge the label spaces of multiple data sets into a unified label space, thereby effectively utilizing different data sets and simplifying the multi-dataset training process. That is, label mapping is performed on the sample labels of the multi-level data sets to obtain a target mapping label set, and a joint data set is obtained based on the target mapping label set. It can be understood that the labels in the target mapping label set are assigned back to the original data sets of the samples to obtain multiple data sets with a unified label space, namely, a joint data set. The preset detection model is trained based on the joint data set and the joint loss function corresponding to the joint data set, and finally a trained detection model is obtained.
[0053] The training method of the detection model of the present application reduces the need for full data annotation and significantly reduces labor costs. And through a simple label mapping method, the label spaces of multiple data sets are merged into a unified label space, thereby effectively utilizing different data sets, so that the training set can comprehensively cover various types of information, simplifying the multi-dataset training process.
[0054] The label mapping scheme for multiple datasets is described in detail below.
[0055] As an example, Figure 2 As shown, each hierarchical dataset includes multiple category datasets. Label mapping is performed on the sample labels of the multi-level datasets to obtain a target mapping label set, including:
[0056] S201, establishing a sample index for each sample.
[0057] S202, determining the sample category label ID based on the sample index, and merging the sample label IDs of the same category to obtain a multi-category label set.
[0058] S203, performing label mapping on the label sets of multiple categories to obtain a first mapped label set corresponding to the hierarchical data set.
[0059] S204: Obtain a target mapping tag set based on the first mapping tag set.
[0060] Exemplarily, each hierarchical dataset includes multiple category datasets. For example, the component layer dataset includes category datasets of small components such as conductors and insulators; the equipment layer dataset includes category datasets of equipment such as transformers and distribution switchgear; the tower layer dataset includes category datasets related to the materials and functions of power towers. Each level of the dataset contains the target categories of the corresponding level. During the annotation process of the new dataset, it can be divided into different levels according to the type of target to be annotated. After the annotation is completed, the labels are first mapped to unify the datasets from multiple sources to ensure the consistency of the labels.
[0061] Exemplarily, multiple data sets are first received, and the size and cumulative size of each data set are calculated to facilitate subsequent indexing. It should be noted that not one data set corresponds to one category, and some samples may be both component layer and device layer. For example, one data set has 100 pictures, and the other data set has 300 pictures, with a cumulative size of 400 pictures. Of course, there are more than two data sets. An index is established for each sample (a sample refers to each picture in the data set), and the sample is obtained according to the given index, and the source data set of the sample and its index in the data set are determined. Determine the sample category label ID based on the sample index, traverse all data sets, merge the sample label IDs of the same category, and obtain a multi-category label set. This step ensures that each category has a unique label ID.
[0062] Exemplarily, label mapping is performed on a multi-category label set to obtain a first mapping label set corresponding to the hierarchical data set. For example, the recordable category label set is x1, x2, x3, etc. For example, the equipment layer data set has a total of 15 categories of labels. Then label mapping is performed on the multi-category label set to obtain a first mapping label set S1 = [x1, x2, ...., x15] corresponding to the equipment hierarchical data set. Similar processing is performed on other hierarchical data sets to obtain a first mapping label set corresponding to the hierarchical data set, and the number of the first mapping label set is the same as that of the hierarchical data set. For example, the hierarchical data set includes an equipment layer, a component layer, and a tower layer. Then the data of the first mapping label set is three. Based on the first mapping label set, a target mapping label set is obtained, and multiple first mapping label sets are combined to obtain a final target mapping label set.
[0063] As an example, a multi-level data set includes at least one of a device layer data set, a component layer data set, and a tower layer data set, and obtaining a target mapping label set based on a first mapping label set includes: merging the first mapping label set corresponding to the device layer data set, the first mapping label set corresponding to the component layer data set, and the first mapping label set corresponding to the tower layer data set to obtain a target mapping label set.
[0064] Exemplarily, the multi-level data set includes at least one of a device layer data set, a component layer data set, and a tower layer data set. Each category label set in the device layer data set can be recorded as x1, x2, x3, etc., each category label set in the component layer data set can be recorded as y1, y2, y3, etc., and each category label set in the tower layer data set can be recorded as z1, z2, z3, etc. For example, there are 15 categories of labels in the device layer, 26 categories of labels in the component layer, and 18 categories of labels in the tower layer. Then the first mapping label sets from the device layer, component layer, and tower layer data sets are obtained respectively, as shown in the following formula:
[0065] S1=[x1,x2,...,x 15 ],
[0066] S2=[y1,y2,...,y 26 ],
[0067] S3=[z1,z2,…,z 18 ],
[0068] S all =S1∪S2∪S3
[0069] Among them, S1, S2 and S3 are the first mapping label sets from the equipment layer, component layer and tower layer data sets. The first mapping label set corresponding to the equipment layer data set, the first mapping label set corresponding to the component layer data set and the first mapping label set corresponding to the tower layer data set are merged to obtain the target mapping label set, S all Map the label set to the target, which represents the set of all 59 labels.
[0070] Of course, the above number of labels is only an example in this application, and the number of labels corresponding to each level of data set can also be adjusted according to actual conditions.
[0071] Figure 3 It is a schematic diagram of a label mapping solution according to an embodiment of the present application.
[0072] like Figure 3 As shown in the figure, the equipment layer dataset has 15 types of labels, the component layer dataset has 26 types of labels, and the tower layer dataset has 18 types of labels. The label IDs corresponding to each layer of the dataset may be the same or different. To facilitate the subsequent model training, the label space is unified to ensure that each category has a unique label ID. For example, ID0 corresponds to the transformer category, ID1 corresponds to the arrester category, ID2 corresponds to the paralleling clamp category, ID3 corresponds to the porcelain column insulator category, ID4 corresponds to the cement pole category, ID5 corresponds to the steel pipe pole category, and so on. Finally, the unified label space has 59 categories.
[0073] As an example, Figure 4As shown, a joint dataset is obtained based on the target mapping label set, including:
[0074] S401, determining a label after sample mapping based on a target mapping label set and a sample index.
[0075] S402: Obtain a joint data set based on samples in the category data set and mapped labels corresponding to the samples.
[0076] Exemplarily, after obtaining the target mapping label set, the label of the sample after mapping is determined according to the target mapping label set and the sample index. The sample index is established for each sample, and the sample index indicates the source data set of each sample. That is, according to the above target mapping label set (also called category dictionary, category dictionary, such as 0-transformer category, 1-lightning arrester category, 2-parallel line clamp category), the label of each sample after mapping is determined. The label of the sample after mapping is assigned back to each data set based on the sample index. Finally, a joint data set is obtained according to the samples in the category data set and the mapped labels corresponding to the samples.
[0077] This application obtains a joint dataset through the above label mapping scheme, which can integrate the information of multiple datasets. By merging samples from different datasets, the subsequent model training can be exposed to richer features and categories, improving its generalization ability. At the same time, ensuring that each category has a unique ID helps reduce label conflicts and simplify the subsequent data processing process.
[0078] The specific training process of the preset detection model is described in detail below.
[0079] As an example, Figure 5 As shown, the joint data set includes multiple sub-data sets, and the preset detection model is trained based on the joint data set and the joint loss function corresponding to the joint data set to obtain a trained detection model, including:
[0080] S501, predicting the samples of each sub-dataset in the joint data set to obtain a prediction result.
[0081] S502: Determine a loss function corresponding to the sub-dataset based on the prediction result.
[0082] S503, performing weighted processing on the loss functions corresponding to the sub-datasets to obtain a joint loss function corresponding to the joint data set.
[0083] S504: Train the preset detection model based on the joint loss function corresponding to the joint data set to obtain a trained detection model.
[0084] Exemplarily, the joint data set includes multiple sub-data sets, and the multiple sub-data sets are input into a preset detection model. The preset detection model predicts the samples of each sub-data set in the joint data set to obtain a prediction result.
[0085] Exemplarily, during the training process, for each pair of image and text, the text encoder and image encoder are first used to extract image features and text features respectively. Before entering the text encoder, the category names in the training set are filled into the text template, such as "a photo of {category name}" or "a painting of {category name}", and the annotation information is converted into a text prompt. The text list of the target to be detected is [t1, t2, t3, ..., tn], and the corresponding encoding result [e1, e2, e3, ..., en] is obtained after passing through the text encoder. The framework selected by the text encoder is the Chinese pre-trained BERT as the encoder. The image encoder uses Swin Transformer for feature extraction. After that, these two common features are input into the feature enhancement module to realize the fusion of multimodal cross-space features. After obtaining cross-modal text and image features, the query selection module guided by text language is used to select cross-modal queries from image features. Finally, these cross-modal queries will be input into the cross-modal decoder to detect the required features and update themselves. After multiple rounds of updates, the decoder outputs the prediction results.
[0086] Exemplarily, the loss function corresponding to the sub-dataset is determined based on the prediction result, and the loss function corresponding to the sub-dataset is weighted to obtain the joint loss function corresponding to the joint data set. It can be understood that each sub-dataset corresponds to a loss function, and different weights can be set for the loss function of the sub-dataset according to the difference in importance of the sub-dataset. The loss function corresponding to the sub-dataset is weighted to obtain the joint loss function corresponding to the joint data set. The preset detection model is trained based on the joint loss function to finally obtain a trained detection model.
[0087] As an example, in the process of modal feature extraction, Swin Transformer is first used to extract features from the input image to generate a feature map of dimension D×H×W, which contains rich local detail information and global semantic information. At the same time, the text input is processed by the pre-trained BERT model to extract a text feature sequence of dimension L×D, where L is the length of the text sequence and D is the feature dimension. These two Transformer-based models are selected as feature extractors mainly because they have superior feature expression capabilities and hierarchical information processing capabilities in their respective fields and have greater advantages in multimodal information fusion.
[0088] Exemplarily, the extracted bimodal features are then input into a text-guided query selection module, which implements deep interactive modeling of image and text features through a multi-head cross-modal attention mechanism. Specifically, the module contains three key structures: a text-guided query selection structure, a visual-guided text query selection structure, and a multimodal information alignment structure. Taking the text-guided query selection structure as an example, the structure uses text features as queries, image features as keys and values, and generates weighted image features through self-attention operations. At the same time, the L2 distance is introduced in the feature projection space to ensure the semantic alignment of the two modal features and enhance the cross-modal expression ability. Subsequently, the structure combines the text features to perform a weighted score on each local feature of the image, calculates its semantic relevance to the text features through cosine similarity, and selects the cross-modal query most relevant to the text description as a candidate feature. Similarly, the visual-guided text query selection structure uses the same attention mechanism, but uses image features as a guide to select relevant text features. Finally, the features of these two directions are fused and optimized through a multimodal information alignment structure, which uses a bidirectional cross-attention mechanism to achieve efficient integration of modal information.
[0089] Exemplarily, after feature interaction, the fused features are input into the cross-modal decoder to generate candidate targets. The decoder first calculates the cosine similarity matrix between text and image features, and selects the top K positions with the highest similarity after Softmax normalization (the K value is dynamically adjusted according to the input image size, and K can be set to 100) to construct candidate targets. Each candidate target contains two parts: position information and content features. In each round of iteration, the decoder progressively optimizes the target features through the attention mechanism and the feedforward network, and finally outputs the prediction results, including the target category set cls = [c1, c2, c3, ..., cm] and the bounding box set bboxes = [b1, b2, b3, ..., bm], where each bounding box represents the center point coordinates and size of the target with bm = (cx, cy, w, h). In addition, the model also outputs a prediction score, which is obtained by calculating the cosine similarity between the L2-normalized content features and the text feature targets. A threshold can also be set for each category to remove the target box with a prediction score less than the threshold, so as to obtain the final prediction result. This method based on text semantics guidance and multi-round feature optimization achieves accurate classification and positioning of cross-modal targets.
[0090] As an example, Figure 6 As shown, the prediction results include category prediction results and detection box prediction results; the loss function corresponding to the sub-dataset is determined based on the prediction results, including:
[0091] S601, obtaining a classification loss function based on the category prediction result and the category reference result, and obtaining a detection box loss function based on the detection box prediction result and the detection box reference result.
[0092] S602: Determine a loss function corresponding to the sub-dataset based on the classification loss function and the detection box loss function.
[0093] Exemplarily, after multiple data sets are trained through unified category mapping, the model can learn the common features of similar categories in different data sets. In this model that jointly trains multiple data sets and supports flexible export, the present application considers multiple aspects in the design of the loss function. The loss function corresponding to each sub-dataset includes classification loss and bounding box regression loss, that is, classification loss function and detection box loss function. The classification loss function is obtained based on the category prediction result and the category reference result, and the detection box loss function is obtained based on the detection box prediction result and the detection box reference result. The detection box reference result and the category reference result are obtained in the previous manual annotation of the data set.
[0094] Exemplarily, the loss function corresponding to each sub-dataset can be obtained according to the above method. The present application also takes into account the independent loss of each specific data set detection head when designing the loss function, which is similar to the independent loss designed for each task in multi-task learning, and is balanced through learnable weights, that is, weighted processing is performed based on the loss function corresponding to each sub-dataset to obtain a joint loss function.
[0095] Exemplarily, the formula of the joint loss function is as follows:
[0096]
[0097] Among them, Loss total represents the joint loss function, D represents the total number of data sets, Loss d represents the loss function of the dth data set, including classification loss and bounding box loss, etc., λ d Represents the weight of the d-th dataset, which is used to balance the importance of different datasets.
[0098] For example, the loss function Loss for the dth data set is d , and its calculation formula is shown as follows:
[0099] Loss d =L cls +L loc
[0100] Among them, L loc represents the detection box loss, L cls represents the classification loss.
[0101] As an example, Figure 7 As shown, the classification loss function is obtained based on the category prediction results and the category reference results, including:
[0102] S701, determining a prediction probability corresponding to a category based on the category prediction result and the category reference result.
[0103] S702, weighting the predicted probabilities corresponding to the categories to obtain a classification loss function.
[0104] Exemplarily, for the classification loss function, the prediction probability corresponding to the category is determined based on the category prediction result and the category reference result, and the prediction probability corresponding to the category represents the accuracy of the detection model for the prediction of a single category. The prediction probability corresponding to the category is weighted to obtain the prediction probability of all categories, which is the classification loss function.
[0105] Exemplarily, the formula of the classification loss function is as follows:
[0106]
[0107] Among them, L cls represents the classification loss function, C is the number of categories, α i represents the weight of category i, p t,i Represents the model's predicted probability for category i.
[0108] As an example, Figure 8 As shown, the detection box prediction result includes the detection box center point prediction result and the detection box size prediction result, and the detection box reference result includes the detection box center point reference result and the detection box size reference result. The detection box loss function is obtained based on the detection box prediction result and the detection box reference result, including:
[0109] S801, determining a detection box distance loss function based on a detection box center point prediction result and a detection box center point reference result.
[0110] S802, determining a detection box matching loss function based on the detection box size prediction result and the detection box size reference result.
[0111] S803: Obtain a detection box loss function based on the detection box distance loss function and the detection box matching loss function.
[0112] Exemplarily, the detection model finally outputs a prediction result, including a target category set cls = [c1, c2, c3, ..., cm] and a bounding box set bboxes = [b1, b2, b3, ..., bm], wherein each bounding box represents the center point coordinates and size of the target with bm = (cx, cy, w, h). That is, the detection box prediction result includes the detection box center point prediction result and the detection box size prediction result. The detection box distance loss function is determined based on the detection box center point prediction result and the detection box center point reference result, and the detection box distance loss function can use the L1 distance matching function. The detection box matching loss function is determined based on the detection box size prediction result and the detection box size reference result, and the detection box matching loss function can use the GIOU matching function. The detection box loss function is obtained according to the detection box distance loss function and the detection box matching loss function.
[0113] Exemplarily, the formula of the detection box loss function is as follows:
[0114] L loc =L1+L GIOU
[0115] Among them, L loc represents the detection box loss function, L1 represents the distance matching loss, L GIOU represents the GIOU matching loss.
[0116] The above completes the calculation process of the entire multimodal visual detection large model. Next, the parameters of the large model are updated through derivation and gradient backpropagation. Multiple data sets are successfully integrated in the training stage. During training, the activated Head is dynamically switched according to the label mapping dictionary to achieve parameter optimization of the Backbone shared by multiple data sets. At the same time, it supports independent optimization of different Heads to obtain a trained detection model.
[0117] In the inference process of the existing multi-task large model, some parameters are repeatedly loaded, resulting in redundant calculations and memory usage, wasting algorithm resources. This application also splits the target detection model obtained by the above training method, and uses inference acceleration technology to implement a parallel inference algorithm for multiple tasks sharing the backbone network, thereby improving the performance of algorithm reasoning. In addition, by decoupling tasks, specific tasks can be decoupled from the backbone, and tasks can be tuned separately, which significantly improves the reasoning efficiency and accuracy of the model.
[0118] As an example, the detection model includes a shared backbone network sub-model and multiple task detection head sub-models. The task detection head model is coupled with the shared backbone network sub-model. After the trained detection model is obtained, the method further includes:
[0119] The shared backbone network submodel and multiple task detection head submodels are decoupled according to the category length corresponding to the task detection head submodel to obtain multiple detection head submodels; wherein the multiple detection head submodels correspond one-to-one to the multiple task detection head submodels, the detection head submodel includes a shared backbone network submodel and a task detection head submodel, and the detection head submodels share the shared backbone network submodel.
[0120] Exemplarily, this application achieves the decoupling of specific modules of multi-dataset tasks through a parallel reasoning algorithm of a backbone network shared by multiple tasks, which can effectively improve the reasoning speed and performance of large models in practical applications. To achieve this goal, in terms of model structure design, the model training structure is innovatively split: the original single forward propagation structure is deconstructed into separable modules, so that during the model export process, the model can be flexibly separated into a shared backbone feature extraction network and multiple detection heads (Head) for specific tasks according to the number of categories in the input data set.
[0121] For example, Fig. 9 As shown in the figure, the solid line represents the forward propagation process of the model in the training phase: the input features are first extracted through the backbone network, and then the extracted features are input into the corresponding detection head to complete the training process. The dotted line represents the processing flow of the model conversion: although the basic path is consistent with the training process, the flexible calling mechanism for the backbone and head layer features is innovatively implemented, so that the model can show differentiated computing characteristics at different stages. Through the runtime dynamic method replacement technology, the structural adaptability problem in the traditional model export is effectively solved, and the precise decoupling and independent export of the backbone and head modules are realized.
[0122] For example, the specific implementation is as follows Fig.10 As shown in the figure, this method can flexibly convert the trained unified model into a combination of a shared backbone and multiple task-specific heads according to the length of the category list of different data sets. This design significantly improves the computational efficiency while ensuring the model performance, providing a practical technical solution for large-scale multi-dataset model deployment. Experimental results show that this decoupled model structure not only reduces the reuse of computing resources, but also enables flexible switching of the model between different tasks, providing a new technical idea for improving reasoning performance.
[0123] For example, the target detection model can be converted into ONNX (Open Neural Network Exchange) format, which is an open deep learning model exchange standard that enables interoperability between different frameworks. The formula for ONNX model conversion is as follows:
[0124]
[0125] Among them, M is the original model, Backbone ONNX To share the backbone network sub-model Backbone, the detection head is dynamically generated according to the category list length Li of each data set during the conversion process Therefore, three different groups of detection head sub-models and backbones are obtained. Each detection head sub-model is specially designed for a category of a data set while maintaining the consistency of backbone parameters.
[0126] Exemplarily, the formula of each detection head sub-model is as follows:
[0127]
[0128] Among them, g(Li) is a function of the dynamic detection head generated according to the category list length Li for each dataset, i∈{S1, S2, S3} represents three datasets, and Li represents the length of the category list of the i-th dataset.
[0129] Finally, the complete ONNX model is in the following form:
[0130]
[0131] This application retains a shared Backbone in the integrated reasoning solution, which not only reduces computational redundancy, but also improves the reasoning speed of the model on multiple data sets. At the same time, an independent detection head is configured for each data set, so that the model can flexibly cope with different categories of detection tasks. This design makes the model more efficient and accurate during reasoning. After completing these steps, this integrated model can be further converted to TensorRT format. TensorRT is an efficient reasoning optimization tool provided by NVIDIA, specifically designed to improve the reasoning speed of deep learning models on NVIDIA GPUs.
[0132] Fig.11 It is a flow chart of the overall technical solution of an embodiment of the present application.
[0133] like Fig.11As shown in the figure, the solution first inputs a multi-level dataset and extracts the text features and image visual features of the labeled categories respectively, and then obtains the basic model through multimodal feature interactive training. In order to improve the reasoning efficiency of the model, the model is first converted to the ONNX format. Then, the model is split into two parts: Backbone and detection head (Head), and then converted to TensorRT format for deep optimization. Through TensorRT's optimization technology, the model has been significantly improved in terms of reasoning speed, memory utilization, and latency, thereby maximizing its reasoning performance on NVIDIA GPUs and meeting efficient reasoning requirements.
[0134] This application also proposes a target detection method.
[0135] As an example, Fig.12 As shown, the target detection method includes:
[0136] S1201, obtaining a picture to be detected.
[0137] S1202, detecting the image to be detected based on the detection model trained according to the above-mentioned detection model training method to obtain a target detection result.
[0138] As an example, Fig.13 As shown in the figure, the detection model includes a shared backbone network sub-model and multiple task detection head sub-models. The multiple task detection head sub-models are all connected to the shared backbone network sub-model. The detection model is used to detect the image to be detected to obtain the target detection results, including:
[0139] S1301, extracting features of the image to be detected based on the shared backbone network sub-model to obtain image features.
[0140] S1302: Detect image features based on multiple task detection head sub-models to obtain multiple first detection results corresponding to the task detection head sub-models.
[0141] S1304: Combine the first detection results to obtain a target detection result.
[0142] For example, when the detection model trained by the above detection model training method detects the picture set, Fig.14 As shown, the image first passes through the shared backbone network sub-model to extract features of the image to be detected to obtain image features. The shared backbone network sub-model is connected to multiple task detection head sub-models, and the image features are detected based on the multiple task detection head sub-models to obtain multiple first detection results corresponding to the task detection head sub-models, such as Fig.14As shown, the first detection result can be understood as the detection result output by the equipment detection head sub-model, the component detection head sub-model, or the tower detection head sub-model, and the first detection result is combined to obtain the final target detection result. The target detection result includes multiple categories of detection results, and the detection result is more comprehensive.
[0143] This application also proposes a training device for a detection model.
[0144] As an example, Fig.15 As shown, the training device of the detection model includes: a first acquisition module 1501, used to acquire a multi-level data set; a mapping module 1502, used to perform label mapping on the sample labels of the multi-level data set to obtain a target mapping label set; an acquisition module 1503, used to obtain a joint data set based on the target mapping label set; a training module 1504, used to train a preset detection model based on the joint data set and a joint loss function corresponding to the joint data set to obtain a trained detection model.
[0145] The present application also proposes a target detection device.
[0146] As an example, Fig.16 As shown, the target detection device includes: a second acquisition module 1601, used to acquire the image to be detected; a detection module 1602, used to detect the image to be detected based on the detection model trained by the training device according to the above-mentioned detection model, and obtain the target detection result.
[0147] The application also proposes a computer-readable storage medium.
[0148] In this embodiment, a computer program is stored on a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned detection model training method and target detection method are implemented.
[0149] Fig.17 A block diagram of an electronic device provided for an embodiment of the present application.
[0150] An embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the above-mentioned detection model training method and target detection method are implemented.
[0151] like Fig.17 As shown, for ease of understanding, the embodiment of the present application shows a specific electronic device.
[0152] Electronic devices are intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Electronic devices may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0153] like Fig.17 As shown, the device includes a computing unit 1701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1702 or a computer program loaded from a storage unit 1708 into a random access memory (RAM) 1703. In RAM 1703, various programs and data required for the operation of the electronic device can also be stored. The computing unit 1701, ROM 1702, and RAM 1703 are connected to each other via a bus 1704. An input / output (I / O) interface 1705 is also connected to the bus 1704.
[0154] Multiple components in the electronic device are connected to the I / O interface 1705, including: an input unit 1706, such as a keyboard, a mouse, etc.; an output unit 1707, such as various types of displays, speakers, etc.; a storage unit 1708, such as a disk, an optical disk, etc.; and a communication unit 1709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1709 allows the electronic device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0155] The computing unit 1701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1701 performs the various methods described above, such as the training method of the detection model and the target detection method. For example, in some embodiments, the training method of the detection model and the target detection method may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 1708. In some embodiments, part or all of the computer program may be loaded and / or installed on an electronic device via ROM 1702 and / or a communication unit 1709. When the computer program is loaded into RAM 1703 and executed by the computing unit 1701, the training method of the detection model and the target detection method described above may be executed. Alternatively, in other embodiments, the computing unit 1701 may be configured to execute the detection model training method and the target detection method in any other appropriate manner (eg, by means of firmware).
[0156] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this application, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing in a suitable manner if necessary, and then stored in a computer memory.
[0157] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0158] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In the present application, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0159] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.
[0160] In addition, the terms "first", "second", etc. used in the embodiments of the present application are only used for descriptive purposes and should not be understood as indicating or implying relative importance, or implicitly indicating the number of technical features indicated in the present embodiment. Therefore, the features defined by the terms "first", "second", etc. in the embodiments of the present application can explicitly or implicitly indicate that at least one of the features is included in the embodiment. In the description of the present application, the word "multiple" means at least two or two or more, such as two, three, four, etc., unless otherwise clearly and specifically defined in the embodiments.
[0161] In this application, unless otherwise clearly specified or limited in the embodiments, the terms "installed", "connected", "connected" and "fixed" etc. appearing in the embodiments should be understood in a broad sense. For example, the connection can be a fixed connection, a detachable connection, or an integrated connection. It can be understood that it can also be a mechanical connection, an electrical connection, etc.; of course, it can also be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal connection of two elements, or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to the specific implementation situation.
[0162] In the present application, unless otherwise clearly specified and limited, a first feature being “above” or “below” a second feature may mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, a first feature being “above”, “above”, and “above” a second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. A first feature being “below”, “below”, and “below” a second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.
[0163] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A method for training a detection model, characterized in that: The method comprises: Obtain multi-level datasets; Performing label mapping on sample labels of the multi-level data set to obtain a target mapping label set; Obtaining a joint dataset based on the target mapping label set; The preset detection model is trained based on the joint data set and the joint loss function corresponding to the joint data set to obtain a trained detection model.
2. The method for training a detection model according to claim 1, characterized in that: Each hierarchical data set includes multiple category data sets, and label mapping is performed on sample labels of the multi-level data sets to obtain a target mapping label set, including: Create a sample index for each sample; Determine the sample category label ID based on the sample index, and merge the sample label IDs of the same category to obtain a multi-category label set; Performing label mapping on the multi-category label sets to obtain a first mapped label set corresponding to the hierarchical data set; A target mapping tag set is obtained based on the first mapping tag set.
3. The method for training a detection model according to claim 2, characterized in that: The multi-level data set includes at least one of a device layer data set, a component layer data set, and a tower layer data set. The step of obtaining a target mapping tag set based on the first mapping tag set includes: The first mapping label set corresponding to the equipment layer data set, the first mapping label set corresponding to the component layer data set, and the first mapping label set corresponding to the tower layer data set are merged to obtain a target mapping label set.
4. The method for training a detection model according to claim 2, characterized in that: The obtaining of a joint data set based on the target mapping label set includes: Determine a label after sample mapping based on the target mapping label set and the sample index; A joint dataset is obtained based on the samples in the category dataset and the mapped labels corresponding to the samples.
5. The method for training a detection model according to claim 1, characterized in that: The joint data set includes a plurality of sub-data sets, and the training of a preset detection model based on the joint data set and a joint loss function corresponding to the joint data set to obtain a trained detection model includes: Predicting the samples of each sub-dataset in the joint data set to obtain a prediction result; Determine a loss function corresponding to the sub-dataset based on the prediction result; Performing weighted processing on the loss functions corresponding to the sub-datasets to obtain a joint loss function corresponding to the joint data set; The preset detection model is trained based on the joint loss function corresponding to the joint data set to obtain a trained detection model.
6. The method for training a detection model according to claim 5, characterized in that: The prediction result includes a category prediction result and a detection box prediction result; and determining a loss function corresponding to the sub-dataset based on the prediction result includes: Obtaining a classification loss function based on the category prediction result and the category reference result, and obtaining a detection box loss function based on the detection box prediction result and the detection box reference result; A loss function corresponding to the sub-dataset is determined based on the classification loss function and the detection box loss function.
7. The method for training a detection model according to claim 6, characterized in that: The obtaining of a classification loss function based on the category prediction result and the category reference result includes: Determining a prediction probability corresponding to a category based on the category prediction result and the category reference result; The predicted probabilities corresponding to the categories are weighted to obtain a classification loss function.
8. The method for training a detection model according to claim 6, characterized in that: The detection frame prediction result includes a detection frame center point prediction result and a detection frame size prediction result, the detection frame reference result includes a detection frame center point reference result and a detection frame size reference result, and obtaining a detection frame loss function based on the detection frame prediction result and the detection frame reference result includes: Determine a detection frame distance loss function based on the detection frame center point prediction result and the detection frame center point reference result; Determine a detection frame matching loss function based on the detection frame size prediction result and the detection frame size reference result; A detection box loss function is obtained based on the detection box distance loss function and the detection box matching loss function.
9. The method for training a detection model according to claim 1, characterized in that: The detection model includes a shared backbone network sub-model and a plurality of task detection head sub-models. The task detection head model is coupled with the shared backbone network sub-model. After obtaining a trained detection model, the method further includes: Decoupling the shared backbone network submodel and the multiple task detection head submodels according to the category length corresponding to the task detection head submodel to obtain multiple detection head submodels; Among them, the multiple detection head sub-models correspond one-to-one to the multiple task detection head sub-models, the detection head sub-models include one of the shared backbone network sub-models and one of the task detection head sub-models, and the detection head sub-models share the shared backbone network sub-model.
10. A target detection method, characterized in that: The method comprises: Get the image to be detected; The image to be detected is detected based on the detection model trained according to the detection model training method according to any one of claims 1 to 9 to obtain a target detection result.
11. The target detection method according to claim 10, characterized in that: The detection model includes a shared backbone network sub-model and multiple task detection head sub-models, and the multiple task detection head sub-models are all connected to the shared backbone network sub-model. The detection model is used to detect the image to be detected to obtain a target detection result, including: Extracting features of the image to be detected based on the shared backbone network sub-model to obtain image features; Detect the image features based on the multiple task detection head sub-models to obtain multiple first detection results corresponding to the task detection head sub-models; The first detection results are combined to obtain the target detection result.
12. A detection model training device, characterized in that: The device comprises: A first acquisition module is used to acquire a multi-level data set; A mapping module, used to perform label mapping on sample labels of the multi-level data set to obtain a target mapping label set; An obtaining module, used for obtaining a joint data set based on the target mapping label set; A training module is used to train a preset detection model based on the joint data set and a joint loss function corresponding to the joint data set to obtain a trained detection model.
13. A target detection device, characterized in that: The device comprises: The second acquisition module is used to acquire the image to be detected; A detection module is used to detect the image to be detected based on the detection model trained by the detection model training device according to claim 12 to obtain a target detection result.
14. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the steps of the method described in any one of claims 1 to 11 when executing the computer program.
15. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 11 are implemented.