Scene classification and object detection method and system based on cross-task cooperation, and medium
Through the Dynamic Information Socialized Collaboration Mechanism (DISC) to achieve efficient coordination between scene classification and object detection tasks in the autonomous driving system, the problem of insufficient coordination and flexibility among tasks is solved, the system's adaptability and recognition accuracy are improved, and the waste of computing resources and operational costs are reduced.
Patent Information
- Application Number
- CN202510441764.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-11
AI Technical Summary
In complex and dynamically changing environments, existing autonomous driving systems lack coordination and flexibility between scene classification and object detection tasks, resulting in waste of computing resources and degradation of performance.
The dynamic information social collaboration mechanism (DISC) is adopted, and through dynamic hierarchical cooperation (DHC) and dynamic selection cooperation (DSC) mechanisms, the coordination and division of labor between tasks are dynamically adjusted, the underlying features are shared, and the task-specific layer is optimized to ensure that the scene classification and object detection tasks are efficiently collaboratively evolved in the same model.
It improves the adaptability and robustness of the autonomous driving system in complex environments, reduces waste of computing resources, improves recognition accuracy and generalization capabilities, enhances the system's real-time decision-making ability and long-term adaptability, and reduces operating costs.
Smart Images

Figure CN120298797A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning for autonomous driving, and more specifically, to a method, system, and medium for scene classification and object detection based on cross-task collaboration. Background Art
[0002] With the rapid development of autonomous driving technology, autonomous driving systems are gradually applied to various types of transportation, such as private cars, shared mobility, logistics transportation, and other fields. Their safety, reliability, and traffic management efficiency are closely related. One of the core tasks of an autonomous driving system is to accurately perceive and classify the surrounding environment, especially scene classification and object detection. Scene classification refers to the autonomous driving system's identification of the current road environment type, such as urban roads, highways, rural roads, etc. Driving strategies vary greatly in different scenarios, so accurate scene classification is crucial for autonomous driving decisions. Object detection is the autonomous driving system's identification and localization of key objects in traffic, such as pedestrians, other vehicles, traffic signs, obstacles, etc. Object detection helps the system avoid obstacles, determine the driving route, and make safe decisions. To ensure the stable performance of the autonomous driving system in complex and dynamically changing environments, traditional methods usually need to separately process scene classification and object detection tasks. Although these tasks can be carried out independently, there are often coordination difficulties in dynamic environments, and their complementary advantages cannot be fully utilized. Therefore, how to effectively combine these two tasks so that the autonomous driving system can simultaneously complete efficient classification and detection in different traffic scenarios has become one of the hotspots in current technology research.
[0003] In traditional autonomous driving systems, scene classification and object detection are usually processed by using different models separately. The scene classification model focuses on identifying different road environment types, such as urban roads, highways, rural roads, etc., while the object detection model focuses on identifying and localizing key objects in traffic, such as pedestrians, other vehicles, traffic signs, etc. This method can achieve good performance by ensuring that each model can focus on its own task. However, although this method can achieve good results, it also has significant limitations. First, since different tasks usually use different models for processing, this leads to a waste of computing resources. Especially in real-time processing autonomous driving systems, how to efficiently utilize computing resources has become an important issue. Second, this separate processing method ignores the close connection between scene information and object information. For example, the behavior or position of an object may be affected by the current scene, and this connection is difficult to reflect in traditional methods. Therefore, effective collaboration between tasks cannot be achieved, resulting in poor flexibility and adaptability of the system when dealing with complex environments.
[0004] Multi-task learning (MTL) enables different tasks to share underlying features through a shared network structure, thereby reducing computational overhead and improving overall performance. Multi-task learning improves the synergy between tasks by sharing knowledge. Ideally, it can benefit both the scene classification and object detection tasks simultaneously, thus enhancing the performance of the system. However, although multi-task learning can achieve information sharing between tasks to a certain extent and reduce computational resource consumption, it also faces significant deficiencies. First, conflicts may occur when multi-task learning shares features between tasks. Especially when the goals or requirements between tasks are quite different, the shared features extracted may not effectively support the independence of each task, resulting in a decline in task performance. Second, when multi-task learning faces complex and dynamically changing environments, it may not be able to dynamically adjust the shared and independent parameters according to the different requirements of tasks, leading to poor adaptability of the system to new environments and insufficient flexibility in practical applications. Therefore, existing multi-task learning methods still have certain limitations when facing complex and changing scenarios in autonomous driving.
[0005] To address this issue, existing methods have designed some methods to solve the conflicts of shared features and improve flexibility, such as dynamic task weight adjustment, adaptive feature selection, and adding task-specific sharing mechanisms. However, they still fail to fully solve the conflicts of feature sharing between tasks. Especially when facing complex and changing environments, the model often cannot flexibly adjust the independence and synergy of tasks. Therefore, how to achieve efficient cooperation between tasks in a dynamically changing complex scenario and avoid performance degradation between tasks is the key to improving the perception accuracy and decision-making ability of autonomous driving systems in a changing traffic environment. Summary of the Invention
[0006] To overcome the deficiencies in the prior art and considering the problems of insufficient task collaboration and flexibility in autonomous driving systems in complex environments, the present invention proposes a method, system, and medium for scene classification and object detection based on cross-task collaboration. Through the Dynamic Information Socialization Collaboration mechanism (DISC), introducing Dynamic Hierarchical Cooperation (DHC) and Dynamic Selection Cooperation (DSC), the present invention dynamically adjusts the collaboration and division of labor between tasks, promoting efficient cooperation between tasks while ensuring the independence of each task. This method enables the scene classification and object detection tasks to co-evolve through sharing underlying features and an optimized structure of task-specific layers, effectively avoiding conflicts in feature sharing between tasks when dealing with new and old tasks, thereby enhancing the adaptability and real-time decision-making ability of the system. By introducing a flexible task adjustment mechanism, the present invention can incrementally introduce new scene and object detection tasks without losing the performance of previous tasks, enabling the autonomous driving system to better cope with dynamically changing complex traffic environments and enhancing the robustness and generalization ability of the autonomous driving system.
[0007] The object of the present invention can be achieved by the following technical solutions.
[0008] A method for scene classification and object detection based on cross-task collaboration, comprising the following steps:
[0009] S1: The vehicle travels along a pre-determined route and collects raw images in real time;
[0010] S2: Preprocess the raw images collected in step S1 to ensure that the image data meets the input requirements of the deep learning network in the autonomous driving system;
[0011] S3: The image data preprocessed in step S2 is used as the input image of the deep learning network in the autonomous driving system. During the training process, a dynamic hierarchical cooperation mechanism is used to interactively optimize the image features;
[0012] The preprocessed image data is input into the deep learning network in the autonomous driving system. The backbone part of the deep learning network in the autonomous driving system includes a scene classification backbone network and an object detection backbone network. The scene classification backbone network and the object detection backbone network respectively extract image feature information from the input image, and interactively optimize the image feature information of the scene classification task and the image feature information of the object detection task extracted through the dynamic hierarchical cooperation mechanism;
[0013] S4: The image feature information of the scene classification task and the image feature information of the object detection task optimized in step S3 are respectively sent to their respective task heads, and are respectively converted into the original scores of the scene classification task and the original scores of the object detection task. A dynamic selection cooperation mechanism is used to interactively optimize the original scores of the scene classification task and the original scores of the object detection task;
[0014] S5: The original scores of the scene classification task optimized in step S4 are processed by an activation function to generate the final classification output. The original scores of the object detection task optimized in step S4 are respectively processed by an activation function and a linear transformation to generate the class probability of each candidate box and the bounding box coordinates of the object; Then, the errors are respectively calculated through the loss functions of the scene classification task and the object detection task, and the model parameters of the scene classification backbone network and the task head, and the object detection backbone network and the task head are respectively optimized according to the gradients of their respective loss functions;
[0015] S6: Deploy the network model trained in step S5 into the autonomous driving system of the vehicle. During the driving process of the vehicle, images are collected in real time, image features are extracted, and classification and detection are performed in real time.
[0016] Further, preferably, in step S1, the vehicle travels along a pre-determined route and collects raw images in real time. The specific process is as follows: Start the vehicle and enter the autonomous driving mode, and travel along the pre-determined route. During the travel, the vehicle collects raw image data in real time through the sensors carried by itself.
[0017] Further, preferably, in step S2, the collected raw images are pre-processed to ensure that the image data meets the input requirements of the deep learning network in the autonomous driving system. The specific process is as follows: Perform necessary pre-processing steps on the raw images collected in step S1, including image normalization, cropping, and size adjustment, to ensure that the image data meets the input requirements of the deep learning network in the autonomous driving system.
[0018] Further, preferably, in step S3, the dynamic hierarchical cooperation mechanism: During the training process, the dynamic hierarchical cooperation makes the scene classification task and the object detection task share the underlying features by dynamically adjusting the proportion of the shared layer and the dedicated layer between tasks;
[0019] The dynamic hierarchical cooperation mechanism dynamically adjusts the information flow between layers, gradually introduces the feature information from the auxiliary model, and enhances the learning ability of the main model; Here, if the scene classification backbone network is used as the main model, then the object detection backbone network is used as the auxiliary model, and vice versa, the object detection backbone network is used as the main model, and the scene classification backbone network is used as the auxiliary model;
[0020] Specifically, the input of the current layer of the main model is constructed by adding the weighted sum of the output of the previous layer of the main model and the output of the auxiliary model from the first layer to the previous layer l-1. The formula is defined as:
[0021]
[0022] Among them, represents the input of the l-th layer of the main model, represents the output of the (l-1)-th layer of the main model, and DHC(l-1) represents the hierarchical output of the auxiliary model;
[0023]
[0024] Among them, E n represents the current training cycle, E t represents the threshold cycle, E loops represents the total number of training cycles, represents the output of the i-th layer of the auxiliary model.
[0025] Further, preferably, in step S4, a cooperation mechanism is dynamically selected: during the model processing, the original scores output by the scene classification task and the object detection task cooperate with each other, and through weighted adjustment, it is ensured that the output of each task not only reflects its own goal but also can obtain valuable information from other tasks;
[0026] When calculating the original scores, the dynamically selected cooperation mechanism dynamically adjusts the weights of the main model and the auxiliary model according to the task requirements; here, if the scene classification backbone network is the main model, then the object detection backbone network is the auxiliary model, and vice versa, the object detection backbone network is the main model and the scene classification backbone network is the auxiliary model;
[0027] Specifically, the final original score output of the main model is generated by adding the original score of the main model and the original score weighted by the auxiliary model according to the task situation, and the formula is defined as:
[0028]
[0029] where DSC(m) represents the final original score obtained after the m-th model is optimized by dynamic selection cooperation, p m represents the predicted value of the m-th model, represents the original score output of the m-th model, represents the total number of models.
[0030] Further, preferably, the original score of the optimized scene classification task in step S5 is processed by an activation function to generate the final classification output. The original scores of the optimized object detection tasks are processed by an activation function and a linear transformation respectively to generate the class probability of each candidate box and the bounding box coordinates of the object; then the errors are calculated respectively through the loss functions of the scene classification task and the object detection task, and the model parameters of the scene classification backbone network and task head, and the object detection backbone network and task head are optimized according to the gradients of their respective loss functions. The specific process is as follows:
[0031] First, the original score of the optimized scene classification task is transformed into the probability distribution of each category through the softmax activation function; at the same time, the original score of the optimized object detection task is used to calculate the class probability of each candidate box through the sigmoid activation function or the softmax activation function, and the bounding box coordinates of the object are output through regression processing;
[0032] Then, the classification loss L cls and the regression loss L reg are calculated respectively. The cross-entropy loss is used as the classification loss L cls , and the L1 norm loss is used as the regression loss L reg , and the classification loss L clsThe error between the predicted category of the scene classification backbone network model and the actual label is calculated, and the regression loss L reg The bounding box regression error is calculated; for the scene classification task, only the classification loss L is applied cls , at this time the regression loss L reg is 0, and for the object detection task, both the classification loss L cls and the regression loss L reg are applied. The total losses of the scene classification task and the object detection task are calculated respectively according to the following loss function formula, and finally, through the backpropagation algorithm, the parameters of the scene classification backbone network and the object detection backbone network models are optimized according to the gradients of their respective total loss functions;
[0033] L overall = L cls (y, f DSC (f DHC (x))) + L reg (y, f DSC (f DHC (x)))
[0034] where L overall is the total loss, x is the input, y is the label, L cls represents the classification loss, L reg represents the regression loss, f DHC (·) represents the collaboration with the dynamic hierarchical cooperation mechanism, and f DSC (·) represents the collaboration with the dynamic selection cooperation mechanism.
[0035] The object of the present invention can also be achieved by the following technical solutions.
[0036] A scene classification and object detection system based on cross-task collaboration includes a preprocessing module, a feature extraction module, a dynamic hierarchical cooperation module, a dynamic selection cooperation module, a classifier, and a detector;
[0037] The preprocessing module is used to preprocess the image data collected during the vehicle driving process, including image normalization, cropping, and size adjustment, to ensure that the image data meets the input requirements of the deep learning network in the autonomous driving system;
[0038] Two feature extraction modules are provided, and both use the ResNet50 network as the feature extractor, including a scene classification task feature extraction module and an object detection task feature extraction module. The inputs of the two feature extraction modules are both the preprocessed image data. The output of the scene classification task feature extraction module is the image feature information of the scene classification task after being interactively optimized by the dynamic hierarchical cooperation module, and the output of the object detection task feature extraction module is the image feature information of the object detection task after being interactively optimized by the dynamic hierarchical cooperation module;
[0039] The dynamic hierarchical cooperation module acts between two feature extraction modules and is used to interactively optimize the image feature information of the scene classification task and the image feature information of the object detection task extracted;
[0040] The dynamic selection cooperation module acts between the classifier and the detector, and interactively optimizes the original score of the scene classification task extracted by the classifier and the original score of the object detection task extracted by the detector;
[0041] The classifier is composed of a fully connected network layer and an activation function layer. After the image feature information output by the scene classification task feature extraction module is flattened, it is converted into the original score of the scene classification task after interactive optimization by the dynamic selection cooperation module through this fully connected network layer. The optimized original score of the scene classification task is converted into the probability distribution of each category through the activation function layer, and the category with the highest probability is selected as the final prediction result of the scene classification task;
[0042] The detector is composed of a classifier and a regression network. The classifier is composed of a fully connected network layer and an activation function layer. The regression network is composed of a fully connected network layer and a regression layer; After the image feature information output by the object detection task feature extraction module is flattened, it is converted into the original score of the object detection task after interactive optimization by the dynamic selection cooperation module through the fully connected network layer of this classifier. The optimized original score of the object detection task is converted into the probability distribution of each category through the activation function layer of this classifier, and the category with the highest probability is selected as the final prediction result of the object detection task; At the same time, after the image feature information output by the object detection task feature extraction module is flattened, it is input into the fully connected network layer of the regression network, and the output of this fully connected layer is used as the input of the regression layer, and the final object bounding box position is obtained through the regression layer.
[0043] The object of the present invention can also be achieved by the following technical solutions.
[0044] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned scene classification and object detection method based on cross-task collaboration is implemented.
[0045] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the above-mentioned scene classification and object detection method based on cross-task collaboration is implemented.
[0046] Compared with the prior art, the beneficial effects brought by the technical solution of the present invention are:
[0047] (1) The present invention fully considers the collaboration and independence between scene classification and object detection tasks, and further improves the task learning ability of the system in complex autonomous driving environments. By dynamically adjusting the collaboration and division of labor between tasks, it reduces the conflicts in feature sharing between tasks, improves the overall performance of the system, and enhances the adaptability and robustness of the autonomous driving system to complex and dynamic traffic scenes.
[0048] (2) The present invention proposes a new cross-task co-evolution learning method for task collaboration in autonomous driving scenarios. Through a dynamic hierarchical cooperation mechanism and a dynamic selection cooperation mechanism, this method effectively avoids the performance degradation between tasks, ensuring that the scene classification and object detection tasks can co-evolve efficiently in the same model. The system can not only handle known tasks but also flexibly perform incremental learning of new tasks, improving the recognition accuracy and generalization ability of the autonomous driving system.
[0049] (3) The present invention designs a framework based on underlying feature sharing and task-specific layer optimization, which can dynamically adjust the structure of the shared layer and the specific layer between tasks, enabling different tasks to selectively share features according to their specific needs. Through this method, the system can achieve efficient cooperation between tasks while maintaining the independence of tasks, avoiding the problem of shared feature conflicts in cross-task learning.
[0050] (4) By optimizing the feature sharing and information interaction mechanism, the present invention enables the autonomous driving system to flexibly adjust and share valuable features between different tasks, thereby improving the system's perception ability of various traffic scenes, especially performing more excellently in complex road conditions and multi-object environments.
[0051] (5) The present invention allows the autonomous driving system to quickly adapt to new tasks through cross-task co-evolution when facing a dynamically changing road environment, reducing the need to retrain the entire model, thereby greatly improving the system's response speed and real-time decision-making ability, especially suitable for rapid deployment and emergency scenarios.
[0052] (6) The present invention effectively reduces the waste of computing resources caused by mutual interference between tasks, and reduces the dependence on a large amount of labeled data through incremental learning, reducing the cost of model update and training. The long-term adaptability and low computational overhead of the system help to reduce the operating cost of the autonomous driving system and improve its long-term operating efficiency and stability. Brief Description of the Drawings
[0053] Figure 1 is a flowchart of the method for scene classification and object detection based on cross-task collaboration of the present invention;
[0054] Figure 2 is a schematic diagram of the effect after applying the cross-task co-evolution learning algorithm;
[0055] Figure 3 This is the comparison schematic diagram in Embodiment 3 of the present invention;
[0056] Figure 4 This is the schematic diagram of the principle of the electronic device in Embodiment 4 of the present invention. Detailed implementation manners
[0057] To make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.
[0058] Embodiment 1
[0059] Due to the differences in task objectives between the scene classification and object detection tasks in the autonomous driving mission, and as the environmental complexity increases, the system needs to continuously introduce new scene categories and object types. Traditional methods often cannot maintain the recognition ability of the learned tasks when incrementally learning new tasks. To address this issue, the embodiments of the present invention propose a method for scene classification and object detection based on cross-task collaboration, which can maintain or even improve the model's ability in the learned tasks while incrementally learning new tasks. Refer to Figure 1 、 2 , and this method includes the following steps:
[0060] S1: The vehicle travels along a pre-determined route and collects raw images in real time.
[0061] Preferably, start the vehicle and enter the autonomous driving mode. The vehicle travels along a pre-determined route. During the driving process, the vehicle uses its own sensors (such as cameras, lidar, etc.) to collect raw image data in real time. During the driving process of the vehicle, the collected raw images will be transmitted to the data processing platform in real time, such as in-vehicle or remote servers, etc.
[0062] S2: Preprocess the raw images to ensure that the image data meets the input requirements of the deep learning network in the autonomous driving system.
[0063] Preferably, perform necessary preprocessing steps on the raw images collected by the vehicle, including image normalization, cropping, size adjustment, etc., to ensure that the image data meets the input requirements of the deep learning network in the autonomous driving system.
[0064] S3: The preprocessed image data is used as the input image of the deep learning network in the autonomous driving system, and a dynamic hierarchical cooperation mechanism is used to interactively optimize the image features during the training process.
[0065] The preprocessed image data is input into the deep learning network in the autonomous driving system, and the backbone part of the network extracts features from the input image. According to different tasks, the scene classification task and the object detection task use independent backbone networks for feature learning respectively. During the feature extraction process, the backbone network extracts feature information suitable for their respective tasks.
[0066] Here, the backbone part of the deep learning network in the autonomous driving system includes a scene classification backbone network and an object detection backbone network. The scene classification backbone network and the object detection backbone network respectively extract image feature information from the input image, and obtain the image feature information of the scene classification task and the object detection task respectively. Then, through the dynamic hierarchical cooperation mechanism, the image feature information of the scene classification task and the object detection task extracted is interactively optimized.
[0067] Dynamic hierarchical cooperation mechanism: During the training process, dynamic hierarchical cooperation dynamically adjusts the proportion of the shared layer and the dedicated layer between tasks, enabling the scene classification task and the object detection task to share underlying features. This mechanism ensures that the learning objectives of each task are independently optimized, thus avoiding interference between tasks. The dynamic hierarchical cooperation mechanism automatically adjusts the weighted ratio of the task head according to the needs of the task, improving the efficiency of multi-task learning and the collaborative effect between tasks.
[0068] When performing feature extraction, the dynamic hierarchical cooperation mechanism dynamically adjusts the information flow between layers, gradually introducing feature information from the auxiliary model to enhance the learning ability of the main model. Here, if the scene classification backbone network is used as the main model, then the object detection backbone network is used as the auxiliary model, and vice versa, the object detection backbone network is used as the main model, and the scene classification backbone network is used as the auxiliary model.
[0069] Specifically, the input of the current layer of the main model is constructed by adding the weighted sum of the output of the previous layer of the main model and the output of the auxiliary model from the first layer to the previous layer l-1. The calculation of this weighted sum dynamically adjusts the information dependence degree of the auxiliary model, enabling the main model to gradually absorb and integrate the knowledge in the auxiliary model during the learning process. As the training progresses, the influence of the auxiliary model gradually increases, helping the main model to learn the auxiliary task without forgetting the tasks it is good at. In this way, the dynamic hierarchical cooperation mechanism achieves the goal of multi-task collaboration, optimizes the learning process of the main model, and effectively improves the collaborative effect between tasks.
[0070] Mathematically, the dynamic hierarchical cooperation mechanism is defined as:
[0071]
[0072] where, represents the input of the l-th layer of the main model, represents the output of the (l-1)-th layer of the main model, and DHC(l-1) represents the hierarchical output of the auxiliary model.
[0073]
[0074] Among them, E n represents the current training epoch, E t represents the threshold epoch, E loops represents the total number of training epochs, represents the output of the i-th layer of the auxiliary model.
[0075] S4: The image feature information of the optimized scene classification task and the image feature information of the object detection task are respectively fed into their respective task heads (such as classifiers and detectors), and are respectively converted into the raw scores of the scene classification task and the raw scores of the object detection task. The dynamic selection cooperation mechanism is used to interactively optimize the raw scores of the scene classification task and the raw scores of the object detection task.
[0076] Dynamic selection cooperation mechanism: The optimized features will be sent to their respective task heads for further processing, and the features will be converted into the raw scores of the scene classification task and the raw scores of the object detection task. The dynamic selection cooperation mechanism enhances the cooperation between tasks by dynamically adjusting the degree of information sharing between different tasks during this process. Specifically, during the model processing, the output raw scores of the scene classification task and the output raw scores of the object detection task cooperate with each other, and through weighted adjustment, it is ensured that the output raw scores of each task can not only reflect its own goals, but also obtain valuable information from other tasks.
[0077] Among them, the dynamic selection cooperation mechanism dynamically determines the cooperation weights of the task heads according to the correlation between tasks, so as to interactively optimize the results. Through this co-evolution, the system can avoid conflicts between tasks when simultaneously processing scene classification and object detection, and improve the accuracy and robustness of the autonomous driving system in complex environments.
[0078] When calculating the raw scores, the dynamic selection cooperation mechanism will dynamically adjust the weights of the main model and the auxiliary model according to task requirements. Here, if the scene classification backbone network is used as the main model, then the object detection backbone network is used as the auxiliary model, and vice versa, the object detection backbone network is used as the main model, and the scene classification backbone network is used as the auxiliary model.
[0079] Specifically, the final raw score output of the main model is generated by adding the raw score of the main model and the raw score of the auxiliary model weighted according to the task situation. The dynamic adjustment of the weighting factor enables the main model to flexibly select the contribution degree of the auxiliary model according to the characteristics of the current task and samples during the inference process, so as to better combine the advantages of different models.
[0080] Mathematically, the dynamic selection cooperation mechanism is defined as:
[0081]
[0082] where DSC(m) represents the final raw score obtained after the m-th model is optimized by dynamic selection cooperation, p m represents the predicted value of the m-th model, represents the raw score output of the m-th model, represents the total number of models.
[0083] S5: The raw scores of the scene classification task optimized in step S4 are processed by the activation function to generate the final classification output. The raw scores of the object detection task optimized in step S4 are processed by the activation function and linear transformation respectively to generate the class probability of each candidate box and the bounding box coordinates of the object; then the errors are calculated through the loss functions of the scene classification task and the object detection task respectively, and the model parameters of the scene classification backbone network and task head, and the object detection backbone network and task head are optimized according to the gradients of their respective loss functions.
[0084] In the scene classification and object detection tasks, the generated raw scores are processed by the activation function and transformed into the final classification and detection outputs. Specifically, first, the raw scores of the optimized scene classification task are transformed into the probability distribution of each category through the softmax activation function; at the same time, the raw scores of the optimized object detection task are used to calculate the class probability of each candidate box through the sigmoid activation function or softmax activation function, and the bounding box coordinates of the object are output through regression processing.
[0085] Then, the loss function is calculated to evaluate the difference between the prediction result and the true label. The present invention designs two loss functions: classification loss L cls and regression loss L reg , and the cross-entropy loss is used as the classification loss L cls , and the L1 norm loss is used as the regression loss L reg . Only the classification loss L cls is applied to the scene classification task. At this time, the regression loss L reg is 0. At this time, the classification loss L clsThe error between the predicted category of the scene classification backbone network model and the actual label was calculated. For the object detection task, the classification loss L cls and the regression loss L reg were applied. At this time, the classification loss L cls calculated the error between the predicted category of the object detection backbone network model and the actual label, and the regression loss L reg calculated the bounding box regression error (i.e., the difference between the predicted bounding box and the ground truth bounding box).
[0086] Next, the total losses of the scene classification task and the object detection task were calculated according to the following loss function formula (4). Finally, through the backpropagation algorithm, the parameters of the scene classification backbone network and the object detection backbone network models were optimized according to the gradients of their respective total loss functions.
[0087] The Dynamic Hierarchical Cooperation mechanism (DHC) and the Dynamic Selection Cooperation mechanism (DSC) have a hierarchical organizational structure, a progressive interaction pattern, and a strongly oriented communication mechanism. The Dynamic Hierarchical Cooperation mechanism (DHC) and the Dynamic Selection Cooperation mechanism (DSC) together constitute the Dynamic Information Socialization Collaboration mechanism (DISC) of the present invention. The above two mechanisms (DHC and DSC) were integrated into the overall loss function, and the loss function is as follows:
[0088] L overall = L cls (y, f DSC (f DHC (x))) + L reg (y, f DSC (f DHC (x)))(4)
[0089] where, L overall is the total loss, x is the input, y is the label, L cls represents the classification loss, L reg represents the regression loss, f DHC (·) represents the collaboration with the Dynamic Hierarchical Cooperation mechanism, and f DSC (·) represents the collaboration with the Dynamic Selection Cooperation mechanism.
[0090] By minimizing the total loss (the result of updating the network parameters), the model can optimize the performance of both the scene classification and the object detection tasks simultaneously. The total loss function is the sum of the classification loss and the regression loss, and these losses jointly guide the collaborative learning of the model among different tasks, thereby improving the accuracy, robustness, and adaptability of the model in complex environments.
[0091] S6: Deploy the network model trained in step S5 to the vehicle's autonomous driving system. During the vehicle's driving process, sensors such as cameras and lidar can be used to collect images in real time, extract image features, and perform classification and detection in real time.
[0092] After completing model training and optimization, deploy the trained model to the vehicle's autonomous driving system. During the vehicle's driving process, sensors such as cameras and lidar are used to collect images and environmental data in real time. The collected data is preprocessed and then input into the model. The model extracts features and performs real-time classification and object detection. During the driving process, the model can identify objects such as roads, obstacles, and traffic signs, and provide support for vehicle decision-making.
[0093] In summary, through the above steps S1 - S6, the embodiments of the present invention achieve the collaborative optimization of scene classification and object detection tasks in the autonomous driving system. Through the dynamic hierarchical cooperation mechanism and the dynamic selection cooperation mechanism, the model can maintain the recognition ability of the learned tasks while incrementally learning new tasks, and effectively improve the synergistic effect between tasks. Finally, the trained model can collect and process sensor data in real time during the vehicle's driving process, perform accurate classification and object detection, and ensure the accuracy and robustness of the autonomous driving system in complex environments.
[0094] Embodiment 2
[0095] Based on the principle of the method for scene classification and object detection based on cross-task collaboration in Embodiment 1, this embodiment proposes a system for scene classification and object detection based on cross-task collaboration, which mainly includes a preprocessing module, a feature extraction module, a dynamic hierarchical cooperation module, a dynamic selection cooperation module, a classifier, and a detector.
[0096] (1) Preprocessing module
[0097] The preprocessing module is used to preprocess the image data collected during the vehicle's driving process, including image normalization, cropping, and size adjustment, to ensure that the image data meets the input requirements of the deep learning network in the autonomous driving system.
[0098] (2) Feature extraction module
[0099] In this embodiment, two feature extraction modules are set up, and both use a residual network (ResNet) as the feature extractor, such as ResNet50. It can be divided into a feature extraction module for scene classification tasks and a feature extraction module for object detection tasks. The input of both feature extraction modules is the preprocessed image data. The output of the feature extraction module for scene classification tasks is the image feature information of the scene classification task after being interactively optimized by the dynamic hierarchical cooperation module. The output of the feature extraction module for object detection tasks is the image feature information of the object detection task after being interactively optimized by the dynamic hierarchical cooperation module.
[0100] (3) Dynamic hierarchical cooperation module
[0101] The dynamic hierarchical cooperation module acts between the two feature extraction modules and is used to interactively optimize the extracted image feature information of the scene classification task and the object detection task. The dynamic hierarchical cooperation module adopts a carefully designed dynamic hierarchical cooperation mechanism during the feature extraction process. Through hierarchical and progressive interaction methods, it effectively integrates the information embedded in the auxiliary model, aiming to balance the data dependence and auxiliary information dependence in the model learning process, thereby helping the main model fully absorb and effectively utilize the knowledge in the auxiliary model.
[0102] (4) Dynamic selection cooperation module
[0103] The dynamic selection cooperation module acts between the classifier and the detector and interactively optimizes the original scores of the scene classification task extracted by the classifier and the original scores of the object detection task extracted by the detector. The dynamic selection cooperation module adopts a carefully designed dynamic selection cooperation mechanism during the calculation process of the model output layer. By selectively weighting the original scores and gradually introducing the contribution of the auxiliary model, it ensures the efficient cooperation of the main model and the auxiliary model in the output stage.
[0104] (5) Classifier
[0105] In this embodiment, the classifier consists of a fully connected network layer and an activation function layer, and is used to map the feature vector to the final category space. After the image feature information output by the feature extraction module for scene classification tasks is flattened, it is converted into the original score of the scene classification task after being interactively optimized by the dynamic selection cooperation module through this fully connected network layer. The optimized original score of the scene classification task is then processed by the activation function layer (softmax activation function) and converted into the probability distribution of each category, thereby realizing the final prediction of the multi-classification task. By normalizing the prediction values of all categories, the softmax function can ensure that the probability value of each category is between 0 and 1, and the sum is 1. Finally, the classifier will select the category with the highest probability as the final prediction result of the scene classification task according to these probability values;
[0106] (6) Detector
[0107] In this embodiment, on the basis of the classifier, a regression network for predicting the target position is added. Specifically, the detector consists of a classifier and a regression network. The classifier consists of a fully connected network layer and an activation function layer, and the regression network consists of a fully connected network layer and a regression layer.
[0108] After the image feature information output by the object detection task feature extraction module is flattened, it is converted into the original score of the object detection task optimized by the dynamic selection cooperation module through the fully connected network layer of the classifier. The optimized original score of the object detection task is converted into the probability distribution of each category through the activation function layer of the classifier, and the category with the highest probability is selected as the final prediction result of the object detection task to complete the classification task.
[0109] In addition, the detector also predicts the bounding box of the target object through the regression layer. After the image feature information output by the object detection task feature extraction module is flattened, it is input into the fully connected network layer of the regression network. The output of this fully connected layer is used as the input of the regression layer, and the final position of the object bounding box is obtained through the output of the regression layer. The output of the regression layer is the coordinates of the bounding box of the target object, including the center coordinates, width, and height of the bounding box. By predicting these bounding box parameters, the detector can locate the target in the image. Finally, through post-processing steps such as non-maximum suppression (NMS), the detector will screen out the best prediction boxes and finally achieve the recognition and positioning of the target object.
[0110] Embodiment 3
[0111] The following combines Figure 3 with specific experimental data to verify the feasibility of the solutions in Embodiments 1 and 2. See the following description for details.
[0112] This embodiment uses the CIFAR100 dataset and the VOC07+12 dataset to verify the method. Figure 3 It shows the improvement in the accuracy of all categories when applying the dynamic information social collaboration mechanism (DISC) compared with the basic model on the CIFAR100 and VOC07+12 datasets. It can be seen that the accuracy of the vast majority of categories has been significantly improved after adopting the dynamic information social collaboration mechanism (DISC).
[0113] The CIFAR100 dataset is an image classification dataset containing 100 classes, with 600 color images of 32×32 pixels for each class, totaling 60,000 images. Among them, 50,000 are for the training set and 10,000 are for the test set. The classes in the dataset include animals, vehicles, etc., and each class is further divided into 5 subclasses. In addition, the dataset is organized by superclasses, with a total of 20 superclasses. The CIFAR100 dataset is often used for benchmark testing of image classification algorithms, especially in the field of deep learning, and is a challenging standard dataset.
[0114] The VOC07+12 dataset is part of the PASCAL VOC challenge, combining the VOC 2007 and VOC 2012 datasets, and is widely used in computer vision tasks such as object detection, semantic segmentation, and image classification. This dataset contains more than 21,000 images, covering 20 object classes, and annotates the bounding boxes of objects and pixel-level classification labels. The VOC07+12 dataset provides an important benchmark for evaluating vision algorithms, especially in the fields of object detection and image segmentation, and is widely used in deep learning and computer vision research.
[0115] For the above CIFAR100 dataset and VOC07+12 dataset, 10% of the training data and all of the test data are used. The training set of CIFAR100(10%) contains 5000 images, and the test set contains 10000 images; the training set of VOC07+12(10%) contains 4696 images, and the test set contains 14976 images.
[0116] For models with classification as the main task, the parameters are optimized using stochastic gradient descent, the batch size is set to 64, and the total number of training epochs E loops is set to 1500, the training threshold epoch E t is set to 300, the learning rate is 0.01, the momentum of stochastic gradient descent is set to 0.9, and the weight decay is 0.001; for models with detection as the main task, the parameters are optimized using stochastic gradient descent, the batch size is set to 32, and the total number of training epochs E loops is set to 600, the training threshold epoch E t is set to 300, the learning rate is 0.03, the momentum of stochastic gradient descent is set to 0.9, and the weight decay is 0.001.
[0117] In the experiments of the embodiments of the present invention, the accuracy Acc is used as the evaluation index for the classification task, and the mean average precision mAP is used as the evaluation index for the detection task.
[0118] Accuracy is the most commonly used index in the classification task. It represents the proportion of the number of samples correctly predicted by the model to the total number of samples, and the calculation formula is as follows:
[0119]
[0120] Wherein, TP is the number of samples that the model correctly predicts as the positive class, TN is the number of samples that the model correctly predicts as the negative class, FP is the number of samples that the model incorrectly predicts as the positive class, and FN is the number of samples that the model incorrectly predicts as the negative class.
[0121] The mean average precision is a commonly used evaluation metric in object detection tasks. It calculates the average precision (AP) for each class and then takes the average of the APs for all classes. The calculation formula is as follows:
[0122]
[0123] Where N is the total number of classes, and AP i is the average precision of the i-th class.
[0124] Table 1 shows a comparison of the fine-tuning (FT), knowledge distillation (KD), and multi-task learning (MTL) methods based on classification and detection models and the dynamic information social collaboration (DISC) method proposed in the embodiments of the present invention. It is a comparison of the accuracy and mean average precision before and after evolution for the above several different methods on the CIFAR100 dataset and the VOC07+12 dataset. DISC can maintain or even further improve the performance of the model on the original proficient tasks while learning new tasks.
[0125] Table 1 Comparison results of different methods before and after evolution on two datasets
[0126]
[0127] In addition, the experimental results are compared with 7 existing methods, including: SwAV, DeepClusterV2, MoCo v2, CLIP, LSKD, CrossKD, PPAL. For all methods, 5 experiments are run and the average results are adopted, as shown in Table 2. Table 2 shows the comparison of the accuracy and mean average precision for different methods on the CIFAR100 dataset and the VOC07+12 dataset.
[0128] Table 2 Performance comparison of different methods on two datasets
[0129]
[0130] In addition, the embodiments of the present invention verify the effectiveness of each module in its own design, as shown in Table 3. Table 3 is an ablation experiment conducted on the CIFAR100 dataset and the VOC07+12 dataset, indicating the effectiveness of each component of the present method.
[0131] Ablation experiment results on two datasets in Table 3
[0132]
[0133] Among them, DHC represents the dynamic hierarchical cooperation module, and DSC represents the dynamic selection cooperation module.
[0134] Example 4
[0135] An electronic device, see Figure 4 , including a processor 1 and a memory 2. A computer program that can run on the processor 1 is stored in the memory 2. When the processor 1 executes the computer program, the method for scene classification and object detection based on cross-task collaboration in the above Example 1 can be implemented.
[0136] It should be noted here that the device descriptions in the above embodiments correspond to the method descriptions in the embodiments. The embodiments of the present invention will not be elaborated here. The execution subjects of the above processor 1 and memory 2 can be devices with computing functions such as a computer, a single-chip microcomputer, and a microcontroller. When specifically implemented, the embodiments of the present invention do not limit the execution subject and can be selected according to the needs in actual applications. Data signals are transmitted between the memory 2 and the processor 1 through a bus 3. The embodiments of the present invention will not be elaborated here.
[0137] Example 5
[0138] Based on the same inventive concept, an embodiment of the present invention also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for scene classification and object detection based on cross-task collaboration in the above Example 1 is implemented.
[0139] The computer-readable storage medium includes but is not limited to flash memory, hard disk, solid-state drive, etc.
[0140] It should be noted here that the description of the readable storage medium in the above embodiments corresponds to the method description in the embodiments. The embodiments of the present invention will not be elaborated here.
[0141] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part.
[0142] The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that incorporates one or more available media. The available medium can be a magnetic medium or a semiconductor medium, etc.
[0143] In the embodiments of the present invention, unless otherwise specified for the models of each device, the models of other devices are not limited, and any device that can perform the above functions can be used.
[0144] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0145] Although the functions and working processes of the present invention are described above in conjunction with the drawings, the present invention is not limited to the above specific functions and working processes. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.
Claims
1. A method for scene classification and object detection based on cross-task collaboration, characterized in that, It includes the following steps: S1: The vehicle travels along a pre-determined route and collects raw images in real time; S2: Preprocess the raw images collected in step S1 to ensure that the image data meets the input requirements of the deep learning network in the autonomous driving system; S3: The image data preprocessed in step S2 is used as the input image of the deep learning network in the autonomous driving system. During the training process, a dynamic hierarchical cooperation mechanism is used to interactively optimize the image features; The preprocessed image data is input into the deep learning network in the autonomous driving system. The backbone part of the deep learning network in the autonomous driving system includes a scene classification backbone network and an object detection backbone network. The scene classification backbone network and the object detection backbone network respectively extract image feature information from the input image, and interactively optimize the image feature information of the scene classification task and the image feature information of the object detection task through a dynamic hierarchical cooperation mechanism; S4: The image feature information of the scene classification task and the image feature information of the object detection task optimized in step S3 are respectively sent to their respective task heads, and are respectively converted into the raw scores of the scene classification task and the raw scores of the object detection task. A dynamic selection cooperation mechanism is used to interactively optimize the raw scores of the scene classification task and the raw scores of the object detection task; S5: The raw scores of the scene classification task optimized in step S4 are processed by an activation function to generate the final classification output. The raw scores of the object detection task optimized in step S4 are respectively processed by an activation function and a linear transformation to generate the class probability of each candidate box and the bounding box coordinates of the object; Then, the errors are respectively calculated through the loss functions of the scene classification task and the object detection task, and the model parameters of the scene classification backbone network and the task head, and the object detection backbone network and the task head are respectively optimized according to the gradients of their respective loss functions; S6: Deploy the network model trained in step S5 into the autonomous driving system of the vehicle. During the driving process of the vehicle, images are collected in real time, image features are extracted, and classification and detection are performed in real time.
2. The method for scene classification and object detection based on cross-task collaboration according to claim 1, wherein In step S1, the vehicle travels along a pre-determined route and collects raw images in real time. The specific process is as follows: Start the vehicle and enter the autonomous driving mode, and travel along the pre-determined route. During the driving process, the vehicle collects raw image data through the sensors carried by itself.
3. The method for scene classification and object detection based on cross-task collaboration according to claim 1, wherein, In step S2, the collected raw images are preprocessed to ensure that the image data meets the input requirements of the deep learning network in the autonomous driving system. The specific process is as follows: Perform necessary preprocessing steps on the raw images collected in step S1, including image normalization, cropping, and size adjustment, to ensure that the image data meets the input requirements of the deep learning network in the autonomous driving system.
4. The method for scenario classification and object detection based on cross-task collaboration according to claim 1, characterized in that The dynamic hierarchical cooperation mechanism in step S3: During the training process, dynamic hierarchical cooperation dynamically adjusts the proportion of the shared layer and the dedicated layer between tasks, enabling the scene classification task and the object detection task to share underlying features; The dynamic hierarchical cooperation mechanism dynamically adjusts the information flow between layers, gradually introducing the feature information from the auxiliary model to enhance the learning ability of the main model. Here, if the scene classification backbone network is used as the main model, then the object detection backbone network is used as the auxiliary model, and vice versa, the object detection backbone network is used as the main model, and the scene classification backbone network is used as the auxiliary model. Specifically, the input of the current layer of the main model is constructed by adding the weighted sum of the output of the previous layer of the main model and the outputs of the auxiliary model from the first layer to the previous layer l-1. The formula is defined as: Among them, represents the input of the l-th layer of the main model, represents the output of the (l - 1)-th layer of the main model, and DHC(l - 1) represents the hierarchical output of the auxiliary model; Among them, E n represents the current training cycle, E t represents the threshold cycle, E loops represents the total number of training cycles, represents the output of the i-th layer of the auxiliary model.
5. The method for scene classification and object detection based on cross-task collaboration according to claim 1, wherein The dynamic selection cooperation mechanism in step S4: During the model processing, the original scores of the outputs of the scene classification task and the object detection task cooperate with each other, and through weighted adjustment, it is ensured that the output of each task not only reflects its own target but also can obtain valuable information from other tasks. When calculating the original scores, the dynamic selection cooperation mechanism dynamically adjusts the weights of the main model and the auxiliary model according to the task requirements. Here, if the scene classification backbone network is used as the main model, then the object detection backbone network is used as the auxiliary model, and vice versa, the object detection backbone network is used as the main model, and the scene classification backbone network is used as the auxiliary model. Specifically, the final original score output of the main model is generated by adding the original score of the main model and the original score of the auxiliary model weighted according to the task situation. The formula is defined as: Among them, DSC(m) represents the final original score obtained after the m-th model is optimized by dynamic selection and cooperation, and p m represents the predicted value of the m-th model, represents the original score output of the m-th model, represents the total number of models.
6. The method for scene classification and object detection based on cross-task collaboration according to claim 1, wherein, In step 5, the original scores of the optimized scene classification task are processed by the activation function to generate the final classification output. The original scores of the optimized object detection task are processed by the activation function and linear transformation respectively to generate the class probability of each candidate box and the bounding box coordinates of the object. Then, the errors are calculated through the loss functions of the scene classification task and the object detection task respectively, and the model parameters of the scene classification backbone network and task head, and the object detection backbone network and task head are optimized according to the gradients of their respective loss functions. The specific process is as follows: First, the original scores of the optimized scene classification task are transformed into the probability distribution of each category through the softmax activation function. At the same time, the original scores of the optimized object detection task are used to calculate the class probability of each candidate box through the sigmoid activation function or softmax activation function, and the bounding box coordinates of the object are output through regression processing. Then, calculate the classification loss L cls and the regression loss L reg . Use the cross-entropy loss as the classification loss L cls , and the L1-norm loss as the regression loss L reg . The classification loss L cls calculates the error between the predicted category of the scene classification backbone network model and the actual label. The regression loss L reg calculates the bounding box regression error; The scene classification task only applies the classification loss L cls , at this time the regression loss L reg is 0. The object detection task applies both the classification loss L cls and the regression loss L reg . Calculate the total losses of the scene classification task and the object detection task respectively according to the following loss function formula. Finally, through the backpropagation algorithm, optimize the parameters of the scene classification backbone network and the object detection backbone network model according to the gradients of their respective total loss functions; L overall = L cls (y, f DSC (f DHC (x))) + L reg (y, f DSC (f DHC (x))) Among them, L overall is the total loss, x is the input, y is the label, and L cls represents the classification loss, and L reg represents the regression loss. f DHC (·) represents the collaboration with the dynamic hierarchical cooperation mechanism, and f DSC (·) represents the collaboration with the dynamic selection cooperation mechanism.
7. An autonomous driving system based on the cross-task collaborative scenario classification and object detection method according to any one of the above claims 1 to 6, characterized in that, Including a preprocessing module, a feature extraction module, a dynamic hierarchical cooperation module, a dynamic selection cooperation module, a classifier, and a detector. The preprocessing module is used to preprocess the image data collected during the vehicle driving process, including image normalization, cropping, and size adjustment, to ensure that the image data meets the input requirements of the deep learning network in the autonomous driving system. Two feature extraction modules are provided, both of which use the ResNet50 network as the feature extractor, including a scene classification task feature extraction module and an object detection task feature extraction module. The input of both feature extraction modules is the preprocessed image data. The output of the scene classification task feature extraction module is the image feature information of the scene classification task optimized by the dynamic hierarchical cooperation module. The output of the object detection task feature extraction module is the image feature information of the object detection task optimized by the dynamic hierarchical cooperation module. The dynamic hierarchical cooperation module acts between the two feature extraction modules and is used to interactively optimize the extracted image feature information of the scene classification task and the object detection task. The dynamic selection cooperation module acts between the classifier and the detector to interactively optimize the original scores of the scene classification task extracted by the classifier and the original scores of the object detection task extracted by the detector. The classifier consists of a fully connected network layer and an activation function layer. After the image feature information output by the scene classification task feature extraction module is flattened, it is converted into the original score of the scene classification task optimized by the dynamic selection cooperation module through this fully connected network layer. The optimized original score of the scene classification task is converted into the probability distribution of each category through the activation function layer, and the category with the highest probability is selected as the final prediction result of the scene classification task. The detector consists of a classifier and a regression network. The classifier consists of a fully connected network layer and an activation function layer. The regression network consists of a fully connected network layer and a regression layer. After the image feature information output by the object detection task feature extraction module is flattened, it is converted into the original score of the object detection task optimized by the dynamic selection cooperation module through the fully connected network layer of this classifier. The optimized original score of the object detection task is converted into the probability distribution of each category through the activation function layer of this classifier, and the category with the highest probability is selected as the final prediction result of the object detection task. At the same time, after the image feature information output by the object detection task feature extraction module is flattened, it is input into the fully connected network layer of the regression network, and the output of this fully connected layer is used as the input of the regression layer, and the final position of the object bounding box is obtained through the regression layer.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the cross-task collaborative based scene classification and object detection method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the cross-task collaborative based scene classification and object detection method according to any one of claims 1 to 6.
Citation Information
Cited By
Adaptive model optimization method and system based on closed-loop dynamic feedback
CN121787492A