Road element detection method and device, electronic device and storage medium
By using a multi-task detection network in multi-task learning and combining a multi-scale aggregation-separated attention module, the complex problems of model design and optimization in multi-task learning are solved, the generalization ability and performance of the model are improved, and the occurrence of "negative transfer" phenomenon is reduced.
Patent Information
- Application Number
- CN202411356924.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-09-27
AI Technical Summary
In multi-task learning, model design and optimization are complex, especially conflicts or interferences may occur between different tasks, making it difficult for the model to optimize the performance of all tasks at the same time, and the so-called "negative migration" phenomenon occurs.
By obtaining a multi-task training sample set for road feature detection, the multi-task detection network is iteratively trained using the multi-task training sample set, where the multi-task detection network includes an output layer independently configured for each detection task, each output layer includes a multi-scale aggregation-separated attention module.
This method can comprehensively utilize feature information at different scales, unify attention distribution, improve the expression ability and richness of features, thereby enhancing the generalization ability and performance of the model and reducing the occurrence of "negative migration" phenomenon.
Smart Images

Figure CN118865308B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of deep learning, and in particular relates to a road element detection method and device, an electronic device, and a computer-readable storage medium. Background Art
[0002] In recent years, deep learning technology has been widely used and achieved remarkable results in various fields, especially in image recognition, natural language processing and speech recognition. Multi-task learning (MTL) is a learning paradigm in deep learning that effectively uses task-specific and shared information to solve multiple related tasks simultaneously.
[0003] Although multi-task learning can complete training tasks, it also has some problems. For example, there may be conflicts or interferences between different tasks, which makes it difficult for the model to optimize the performance of all tasks at the same time, resulting in the so-called "negative transfer" phenomenon. In multi-task learning, the importance and difficulty of different tasks may be different, and it is necessary to reasonably select tasks and adjust weights to ensure that each task can be effectively trained. This increases the complexity of model design and optimization. Summary of the invention
[0004] The embodiments of the present application provide a road feature detection method and device, an electronic device, and a computer-readable storage medium, which can solve the complex problems of model design and optimization in related technologies.
[0005] In a first aspect, an embodiment of the present application provides a road feature detection method, the method comprising: obtaining a multi-task training sample set for road feature detection, the multi-task training sample set comprising multiple multi-task training samples; using the multi-task training sample set to iteratively train a multi-task detection network, wherein the multi-task detection network comprises an output layer independently configured for each detection task, and each output layer comprises a multi-scale aggregation-separation attention module.
[0006] In a second aspect, an embodiment of the present application provides a road element detection method, the method comprising: acquiring a road image; inputting the road image into a multi-task detection network to obtain a multi-task detection result of the road element, wherein the multi-task detection network is trained using the method described in the first aspect above.
[0007] In a third aspect, an embodiment of the present application provides a road feature detection device, which includes: a first acquisition module, used to acquire a multi-task training sample set for road feature detection, the multi-task training sample set including multiple multi-task training samples; a training module, used to iteratively train a multi-task detection network using the multi-task training sample set, wherein the multi-task detection network includes an output layer independently configured for each detection task, and each output layer includes a multi-scale aggregation-separation attention module.
[0008] In a fourth aspect, an embodiment of the present application provides a road element detection device, which includes: a second acquisition module for acquiring a road image; a detection module for inputting the road image into a multi-task detection network to obtain a multi-task detection result of the road element, wherein the multi-task detection network is trained using the device described in the third aspect above.
[0009] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the road element detection method described in the first aspect or the second aspect when executing the computer program.
[0010] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the road element detection method described in the first or second aspect above.
[0011] In a seventh aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device executes the road element detection method described in the first aspect or the second aspect above.
[0012] Compared with the prior art, the embodiments of the present application have the following beneficial effects: by obtaining a multi-task training sample set for road feature detection, the multi-task training sample set includes multiple multi-task training samples; using the multi-task training sample set to iteratively train a multi-task detection network, wherein the multi-task detection network includes an output layer independently configured for each detection task, each output layer includes a multi-scale aggregation-separation attention module, and the multi-scale aggregation-separation attention module can comprehensively utilize feature information of different scales, unify attention distribution, and improve the expressiveness and richness of features, thereby enhancing the generalization ability and performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0014] Figure 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application;
[0015] Figure 2 It is a flowchart of a road element detection method provided by an embodiment of the present application;
[0016] Figure 3 yes Figure 2A specific flow diagram of S11;
[0017] Figure 4 is a schematic diagram of obtaining a multi-task training sample set in a specific example of the present application;
[0018] Figure 5 is a schematic diagram of the structure of a multi-task detection network in a specific example of the present application;
[0019] Figure 6 This is a schematic diagram of the structure of a multi-scale aggregation-separation attention module in a specific example of the present application;
[0020] Figure 7 is a flowchart of a road element detection method provided by another embodiment of the present application;
[0021] Figure 8 is a structural schematic diagram of a road element detection device provided in one embodiment of the present application;
[0022] Fig. 9 It is a structural schematic diagram of a road feature detection device provided in another embodiment of the present application. DETAILED DESCRIPTION
[0023] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0024] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0025] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0026] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0027] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0028] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0029] The road element detection method provided in the embodiment of the present application can be applied to electronic devices, including but not limited to servers, server clusters, mobile phones, tablet computers, laptop computers, desktop computers, personal digital assistants, wearable devices and other electronic devices with computing functions. The embodiment of the present application does not impose any restrictions on the specific type of electronic devices.
[0030] Figure 1 FIG. 1 is a block diagram of a partial structure of an electronic device provided in an embodiment of the present application. Figure 1 , the electronic device includes: a processor 10, a memory 20, a bus 30, an input device 40, an output device 50 and a communication device 60. The processor 10 and the memory 20 are connected to each other through the bus 30, and the input device 40, the output device 50 and the communication device 60 are also connected to the bus 30. Those skilled in the art can understand that Figure 1 The structure of the electronic device shown in the figure does not constitute a limitation of the electronic device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0031] Combine the following Figure 1 A detailed introduction to the various components of electronic equipment:
[0032] The processor 10 is the control center of the electronic device, and can run the program stored in the memory 20 to perform various functions and process data. The processor 10 can be a central processing unit (CPU), and the processor 10 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In some embodiments, the processor 10 may include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0033] The memory 20 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as program codes of computer programs. The memory 20 can also be used to temporarily store data required for executing the program and generated. The memory 20 may include a high-speed random access memory, and may also include a non-volatile memory, such as a flash memory, a hard disk, a multimedia card, a card-type memory, etc. The memory 20 may include a storage unit disposed inside the electronic device, such as a hard disk of the electronic device, and / or a removable external storage unit, such as a mobile hard disk, a USB flash drive, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, etc.
[0034] The input device 40 may include at least one of a keyboard, a mouse, a touch panel, a joystick, etc., and is used to collect user input operations to generate corresponding operation instructions.
[0035] The output device 50 is used to output information to be provided to the user. The output device 50 generally includes a display, and optionally, a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. can be used. In addition, the output device can further include a speaker.
[0036] The communication device 60 may include a modem, a network card, etc., for establishing a network connection with other electronic devices and communicating with each other.
[0037] The road element detection method provided by the embodiment of the present application can be implemented as a computer software program. For example, the embodiment of the present application provides a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 60, and / or installed from a detachable external storage unit. When the computer program is executed by the processor 10, the various functions defined in the road element detection method provided by the embodiment of the present application are implemented.
[0038] Figure 2 A schematic flowchart of a road feature detection method provided in an embodiment of the present application is shown. As an example but not a limitation, the method can be applied to the above-mentioned electronic device.
[0039] S11: Obtain a multi-task training sample set for road feature detection.
[0040] The multi-task training sample set includes multiple multi-task training samples, each of which includes a road sample image and its multi-task annotation. The multi-task annotation includes annotation information of each detection task of road feature detection. The annotation information of each detection task includes information of a detection box whose target is the detection task. The information of the detection box may specifically include the coordinates of the center point, width, height and confidence (also referred to as a score).
[0041] Specifically, the detection task can include at least two of people, vehicles, signs and dynamic elements. People and vehicles include pedestrians and vehicles, covering a wide range of examples in urban environments, which can help the model learn the characteristics of pedestrians and vehicles under different conditions; signs include traffic signs, ground markings, traffic lights and electronic monitoring equipment, which are crucial for autonomous vehicles to understand traffic rules; dynamic elements focus on temporary obstacles at road construction and accident sites, such as water barriers, cones, barriers, crash barrels, construction fences and accident warning tripods, which help vehicles make safe decisions under complex road conditions.
[0042] For ease of understanding, the following takes the detection tasks including people and vehicles, signboards, and dynamic elements as examples for explanation.
[0043] The labeling information in the multi-task training sample set may be obtained by manual labeling, for example, manually performing multi-task labeling on a plurality of collected road images.
[0044] Currently, the most common data sets are single-task sample sets, which cannot be directly used as multi-task training sample sets. It is necessary to manually add labels to the newly added categories to obtain multi-task training sample sets. However, manual labeling is not only time-consuming and labor-intensive, but also prone to introducing labeling errors, affecting the model training effect. To solve this problem, the embodiment of the present application proposes a method that can effectively integrate multiple data sets with inconsistent labels without manual labeling to obtain a multi-task training sample set.
[0045] like Figure 3 As shown, in a specific embodiment of the present application, S11 may include the following parts.
[0046] S111: Acquire multiple single-task training sample sets for road feature detection.
[0047] Each single-task training sample set includes multiple single-task training samples, each single-task training sample includes a road sample image and its original annotation. Different single-task training sample sets correspond to different detection tasks, and the original annotation targets of the single-task training samples are different.
[0048] S112: Iteratively train a single-task detection network for the corresponding detection task using the single-task training sample sets respectively.
[0049] The specific architecture of the single-task detection network is not restricted here.
[0050] S113: Using the single-task detection network to perform detection on single-task training sample sets with different detection tasks, respectively, to obtain newly added annotations for the road sample images.
[0051] S114: merging the original annotations of each road sample image with the newly added annotations, and merging each single-task training sample set to obtain a multi-task training sample set.
[0052] The detection tasks of the multi-task training sample set are composed of the detection tasks of each single-task training sample set.
[0053] For example, Figure 4As shown, first, datasets 1, 2, and 3 are obtained, including a, b, and c samples respectively. Dataset 1 is a human-vehicle dataset, dataset 2 is a signboard element dataset, and dataset 3 is a dynamic element dataset. Optionally, the original annotation information of the samples in each dataset is stored in an easy-to-process txt format, which is convenient for researchers to perform data preprocessing, model training, and performance evaluation. Subsequently, the popular target detection framework YOLO is used to train specialized detection models A, B, and C based on datasets 1, 2, and 3, respectively. Model A is specifically used to detect pedestrians and vehicles; model B can only recognize signboards; and model C focuses on detecting dynamic elements. Since each model is optimized for the categories in its specific dataset, they all show excellent detection performance. Then, the three detection models A, B, and C are used to add annotations to the other two datasets, that is, datasets 2 and 3 are annotated using model A, datasets 1 and 3 are annotated using model B, and datasets 1 and 2 are annotated using model C. Finally, merge datasets 1, 2, and 3, and merge the newly added annotations of each sample with the original annotations, that is, record the scores of the newly added annotations (i.e., the confidence output by the single-task detection network) and the category information and location information together with the original annotations in a txt file. Optionally, in order to distinguish between the original annotations and the pre-annotations, we set the score mark (i.e., confidence) of the original annotations to 1. In this way, a multi-task training sample set with annotation information for all tasks including pedestrians, vehicles, signs, and dynamic elements is obtained, and the number of samples is a+b+c.
[0054] S12: Iteratively train the multi-task detection network using the multi-task training sample set.
[0055] Specifically, in each round of iterative training, the multi-task training samples can be input into the multi-task detection network, the loss function is calculated according to the output of the multi-task detection network and the multi-task training samples, and the parameters of the multi-task detection network are adjusted according to the loss function.
[0056] The multi-task detection network includes an output layer independently configured for each detection task, and each output layer includes a multi-scale aggregation-separation attention (MSASA) module. This module is an advanced multi-scale aggregation-separation attention mechanism that can dynamically adjust the model's attention distribution to features of different scales, allowing the model to more keenly capture information that is critical to the success of the task. In this way, not only is the synergy between tasks enhanced, but the independence and learning efficiency of each task are also guaranteed, effectively alleviating the problem of task conflict in multi-task learning.
[0057] The multi-scale aggregation-separation attention module includes the first part and the second part. The first part is used to pool the multiple input features of the multi-scale aggregation-separation attention module to obtain the pooling results, and then splice the pooling results through the fully connected layer, and then reversely segment them with reference to the splicing method to obtain the global attention map of each input feature. Different input features have different spatial scales, which enables the model to use information from different spatial scales at the same time. The low-level feature map contains detailed information, and the high-level feature map contains abstract information. By splicing these feature maps, the model can better capture the relationship between fine-grained and high-level features, realize feature complementarity, and ensure that the features processed by the fully connected layer are consistent and coherent, helping the model to better understand the relationship between features of different scales.
[0058] The second part is used to convolve the input features separately to obtain the convolution results. Separate convolution operations are performed on input features of different scales, and fine-grained processing can be performed on different feature maps. Input features of different scales contain information at different levels. By operating separately, the details in each input feature can be better optimized and the expressiveness of the features can be improved. Avoid conflicts and interference that may arise when processing input features of different scales on the same channel, ensure that each input feature can be optimized independently, and improve overall performance. The convolution results of each input feature are then combined with the corresponding global attention map to obtain the output features, which can comprehensively utilize feature information of different scales, unify attention distribution, improve the expressiveness and richness of features, and enhance the generalization ability and performance of the model.
[0059] For example, Figure 5 As shown in the figure, the multi-task detection network can be called MT-YOLO. The network includes a backbone, a neck, and a head (also called an output layer).
[0060] The backbone is responsible for extracting basic features, including 4 stage layers. The input image passes through a convolution module (ConvModule) and then passes through stage layers 1, 2, 3, and 4 in sequence. Each stage layer includes a convolution module and a cross-stage partial layer (CSPLayer). Stage layer 4 also includes a fast spatial pyramid pooling layer (SPPF).
[0061] The neck is responsible for further integration and processing of the basic features extracted by the backbone, including upsampling layer (Upsample), concatenation module (concat), cross-stage partial layer and convolution module.
[0062] The specific structures of the cross-stage partial layers, residual blocks (ResBlock) and convolution modules are as follows Figure 5As shown in the upper right part, the cross-stage partial layers include convolution modules at the input and output positions, a splicing module connecting the convolution modules at the output position, multiple residual blocks and a two-dimensional convolution layer (Conv2d); the residual block includes a convolution module and a splicing module; the convolution module includes a two-dimensional convolution layer, a normalization layer (BN) and an activation layer (ReLu).
[0063] The backbone and neck are common to all tasks. To meet the specific needs of each task, we configure independent output layers for each task, which ensures that personalized features for each task are deeply learned and fine-tuned.
[0064] The neck outputs three feature maps of different scales. In the head, each feature map passes through three convolution modules (ConvModule) to obtain three C5 layer features, three C4 layer features and three C3 layer features. These features are input into three MSASA modules respectively. The input of each MSASA module includes one C5 layer feature, one C4 layer feature and one C3 layer feature, and the inputs of different MSASA modules do not overlap.
[0065] The MSASA module is embedded in each output layer of MT-YOLO, and its structure is as follows: Figure 6 As shown. Specifically, the MSASA module includes the first part and the second part. The first part includes three average pooling layers (AvgPool), which are used to pool the input C5 layer features, C4 layer features and C3 layer features respectively. The output features of the three average pooling layers are spliced by the splicing module and then input into the multi-layer perceptron (MLP), which can also be called a fully connected layer. The output features of the MLP are processed by the segmentation module (Split), and the segmentation is reversed according to the splicing method to obtain the global attention map of each input feature.
[0066] In the second part, the input C5 layer features, C4 layer features and C3 layer features are respectively multiplied with the corresponding global attention map after passing through two two-dimensional convolutional layers. The multiplication results are then added to the input features. The addition results pass through the convolution module to obtain the output P5 layer features, P4 layer features and P3 layer features.
[0067] The design of MT-YOLO cleverly balances the relationship between feature sharing and task independence. By optimizing the attention mechanism, it improves the overall performance and adaptability of the model when facing multiple parallel tasks.
[0068] The loss function used to train the multi-task detection network is a multi-task unified loss function. This loss function enables the model to make full use of the rich information from different tasks during the training phase. By intelligently adjusting the loss weights between different tasks, it avoids redundant calculations and improves training efficiency and model convergence speed.
[0069] For the multi-task detection network, due to the decoupled output head design, the tasks correspond to the output heads one by one, and the loss function of each output head contains the classification loss related to the confidence and the regression loss related to the penalty of the center point distance and aspect ratio. For the classification loss, we use an enhanced version of Focal Loss, which is specially designed to deal with the problem of class imbalance. By reducing the weight of easy-to-classify samples, the model focuses more on those difficult-to-distinguish samples, thereby improving the overall classification accuracy. For the regression loss, we use an improved version of the CIoU loss function, which not only considers the overlapping area between the predicted box and the true box, but also adds the penalty of the center point distance and aspect ratio, which makes the model more accurate in the positioning of the bounding box. The improved version of CIoU loss further enhances this feature, ensuring that the model performs more robustly in regression tasks, especially in application scenarios such as object detection that require high-precision positioning.
[0070] Specifically, the loss function of each task can be the weighted sum of the regression loss function and the classification loss function, as shown below:
[0071]
[0072] in, , The value of can be determined according to actual experiments, for example , .
[0073] The regression loss function UCIoUL is as follows:
[0074]
[0075]
[0076] in, is a hard label associated with the annotation score, and the positive sample generated by this sample In keeping with this sample, is the square of the Euclidean distance between the center point of the real box and the predicted box, is the square of the diagonal length of the rectangle surrounding the real box and the predicted box, , is the width and height of the real box, w and h are the width and height of the predicted box.
[0077] The classification loss function UFL is as follows:
[0078]
[0079] in, is a soft label related to the intersection over union (IoU) with a value between [0, 1]. For all positive samples, the larger the IoU with the true annotation box, the greater the impact on the loss function. , and is a hyperparameter, The default value is 1.0. The higher the annotation score, the greater the impact on the loss. For all negative samples, use the hyperparameter and Reduce the contribution of negative samples to the loss and prevent over-suppression.
[0080] Optionally, the PyTorch deep learning framework is used to build and optimize the proposed MT-YOLO network model. In the preparation stage, we collected and cleaned the dataset, converted its annotated labels into the format required for model training, and divided the training set and validation set in a ratio of 8:2 to ensure that the model can show good generalization ability on unseen data.
[0081] To enhance the robustness and adaptability of the model, we implemented a series of data augmentation techniques, including but not limited to randomly resizing images, adding random noise, and simulating different weather conditions to simulate various scenarios that may be encountered in the real world. These augmented images were then fed into the network in the hope that the model could stably learn under a wider range of conditions.
[0082] During the training process, we use the gradient descent method to iteratively update the model parameters and find the optimal solution by minimizing the loss function. This process is essentially to find the point with the best performance in the model parameter space. As the training progresses, the model gradually learns how to make accurate predictions in various tasks. After the training is completed, we save the final model file, which not only contains the network structure information, but also the optimized weight parameters to ensure that it can be directly loaded and used in the future when the application is deployed. The entire process embodies a series of rigorous steps from data preprocessing to model training to model preservation, aiming to build a high-performance deep learning model that can cope with complex environments and task requirements.
[0083] During the model deployment phase, we convert the trained model from PyTorch's .pt or .pth format to ONNX format, which is a key step in realizing model service. ONNX (Open Neural Network Exchange) is an open format that allows models to be seamlessly migrated between different platforms and frameworks, thereby simplifying the model deployment process. Through this conversion, the model can be integrated into various server-side applications to provide real-time reasoning services.
[0084] For mobile deployment, the situation is more complicated. First, the model needs to be converted into a specific format suitable for processing by the NPU (Neural Processing Unit) chip of the mobile device, which usually involves optimizing and lightweighting the model to adapt to the limited computing power and storage space of the mobile device. In addition, the model needs to undergo a precision quantization process, that is, converting floating-point operations to fixed-point operations to reduce computing resource consumption while maintaining the model's prediction accuracy as much as possible.
[0085] After completing the format conversion and optimization, the performance of the model on the mobile terminal must be fully verified, including whether the prediction accuracy of the model meets expectations and whether the running speed of the model meets the real-time requirements.
[0086] Through the implementation of this embodiment, a multi-task training sample set for road feature detection is obtained, and the multi-task training sample set includes multiple multi-task training samples; the multi-task detection network is iteratively trained using the multi-task training sample set, wherein the multi-task detection network includes an output layer independently configured for each detection task, and each output layer includes a multi-scale aggregation-separation attention module, and the multi-scale aggregation-separation attention module can comprehensively utilize feature information of different scales, unify attention distribution, and improve the expressiveness and richness of features, thereby enhancing the generalization ability and performance of the model.
[0087] Figure 7 A schematic flowchart of a road feature detection method provided in another embodiment of the present application is shown. As an example but not a limitation, the method can be applied to the above-mentioned electronic device.
[0088] S21: Acquire a road image.
[0089] S22: Input the road image into the multi-task detection network to obtain the multi-task detection results of the road elements.
[0090] The multi-task detection network adopts Figure 3 The corresponding embodiment provides a method for training and then deploys it in the execution subject of this embodiment. Figure 3 The execution entities of the corresponding embodiments may be the same electronic device or different electronic devices.
[0091] Figure 8 A schematic structural diagram of a road feature detection device provided in an embodiment of the present application is shown. The road feature detection device includes a first acquisition module 11 and a training module 12.
[0092] The first acquisition module 11 is used to acquire a multi-task training sample set for road feature detection, where the multi-task training sample set includes a plurality of multi-task training samples.
[0093] The training module 12 is used to iteratively train the multi-task detection network using the multi-task training sample set, wherein the multi-task detection network includes an output layer independently configured for each detection task, and each output layer includes a multi-scale aggregation-separation attention module.
[0094] Fig. 9 A schematic structural diagram of a road feature detection device provided in another embodiment of the present application is shown. The road feature detection device includes a second acquisition module 21 and a detection module 22.
[0095] The second acquisition module 21 is used to acquire a road image.
[0096] The detection module 22 is used to input the road image into a multi-task detection network to obtain a multi-task detection result of road elements, wherein the multi-task detection network adopts Fig. 9 The corresponding embodiment provides the device for training.
[0097] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / modules / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0098] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0099] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0100] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0101] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the camera / electronic device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0102] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0103] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0104] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0105] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0106] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A road element detection method, characterized in that: The method comprises: Acquire a multi-task training sample set for road feature detection, wherein the multi-task training sample set includes a plurality of multi-task training samples; Iteratively train a multi-task detection network using the multi-task training sample set, wherein the multi-task detection network includes an output layer independently configured for each detection task, and each output layer includes a multi-scale aggregation-separation attention module; The loss function used in the training includes a classification loss function related to confidence and a regression loss function related to the penalty term of the center point distance and the aspect ratio. The regression loss function is as follows: in, is the hard label associated with the annotation score, is the square of the Euclidean distance between the center point of the real box and the predicted box, is the square of the diagonal length of the rectangle surrounding the real box and the predicted box, , is the width and height of the real box, w and h are the width and height of the predicted box.
2. The method according to claim 1, characterized in that The multi-scale aggregation-separation attention module includes a first part and a second part. The first part is used to pool the multiple input features of the multi-scale aggregation-separation attention module respectively to obtain pooling results, splice the pooling results and pass them through a fully connected layer, and then reversely segment them with reference to the splicing method to obtain a global attention map of each input feature; the second part is used to convolve the input features respectively to obtain convolution results, and combine the convolution results of each input feature with the corresponding global attention map to obtain output features.
3. The method according to claim 1, characterized in that The loss function is the weighted sum of the regression loss function and the classification loss function. The classification loss function is as follows: Among them, p is the network prediction value, is a soft label related to the intersection-over-union ratio, , and is a hyperparameter.
4. The method according to claim 1, characterized in that The method of obtaining a multi-task training sample set for road feature detection includes: Acquire multiple single-task training sample sets for road feature detection, each of the single-task training sample sets includes multiple single-task training samples, each of the single-task training samples includes a road sample image and its original annotation, different single-task training sample sets correspond to different detection tasks, and the original annotation targets of the single-task training samples are different; Iteratively training a single-task detection network for the corresponding detection task using the single-task training sample sets respectively; Using the single-task detection network to perform detection on the single-task training sample sets with different detection tasks, respectively, to obtain newly added annotations of the road sample images, wherein the confidence of the original annotations is 1, and the confidence of the newly added annotations is the confidence output by the single-task detection network; The original annotations of each of the road sample images are merged with the newly added annotations, and each of the single-task training sample sets is merged to obtain the multi-task training sample set.
5. The method according to any one of claims 1 to 4, characterized in that: The detection tasks include at least two of people and vehicles, signboards and dynamic elements.
6. A road element detection method, characterized in that: The method comprises: Acquire road images; The road image is input into a multi-task detection network to obtain a multi-task detection result of road elements, wherein the multi-task detection network is trained using the method according to any one of claims 1 to 5.
7. A road element detection device, characterized in that: The device comprises: A first acquisition module is used to acquire a multi-task training sample set for road feature detection, wherein the multi-task training sample set includes a plurality of multi-task training samples; A training module, configured to iteratively train a multi-task detection network using the multi-task training sample set, wherein the multi-task detection network includes an output layer independently configured for each detection task, and each of the output layers includes a multi-scale aggregation-separation attention module; The loss function used in the training includes a classification loss function related to confidence and a regression loss function related to the penalty term of the center point distance and the aspect ratio. The regression loss function is as follows: in, is the hard label associated with the annotation score, is the square of the Euclidean distance between the center point of the real box and the predicted box, is the square of the diagonal length of the rectangle surrounding the real box and the predicted box, , is the width and height of the real box, w and h are the width and height of the predicted box.
8. A road element detection device, characterized in that: The device comprises: A second acquisition module is used to acquire a road image; A detection module is used to input the road image into a multi-task detection network to obtain a multi-task detection result of road elements, wherein the multi-task detection network is trained using the device as described in claim 7.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Multi-task learning method, system and equipment suitable for roadside automatic driving scene
CN118429921A
Small target detection method based on attention guidance fusion
CN118447230A