3D target detection method, system and device based on roadside monocular camera and medium

Through a two-stage training strategy and pseudo-label generation mechanism based on the roadside monocular camera, the calibration dependence, insufficient generalization capability and lack of labeling data of the roadside monocular 3D object detection system are solved, and efficient long-distance object detection and scene adaptability are achieved.

CN120375352APending Publication Date: 2025-07-25BEIJING SINOITS TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510241920.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The roadside monocular 3D object detection system has shortcomings in camera calibration dependence, insufficient generalization capability, lack of labeling data and difficulty in long-distance object detection, resulting in degradation of detection performance and poor adaptability.

Method used

The 3D object detection method based on a roadside monocular camera is adopted, through a two-stage training strategy combined with labelless data, the pseudo-label generation mechanism is used to improve the generalization ability of depth features, and the dependence on camera calibration parameters is reduced. Combined with relative depth and absolute depth modeling, the robustness and accuracy of object detection are achieved.

Benefits of technology

It significantly improves the performance and robustness of 3D object detection on the roadside monocular, enhances the adaptability and accuracy of the model in different scenarios, reduces the dependence on calibration parameters, and expands the training data scale.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375352A_ABST
    Figure CN120375352A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D target detection method, system and device based on a roadside monocular camera and a medium, and relates to the technical field of vehicle detection, and the method comprises the steps: obtaining at least one target image of a to-be-detected region based on the roadside monocular camera, inputting the target image into a pre-training model, and obtaining a 3D target detection result, the pre-training model is trained through a two-stage training strategy, and the two-stage training strategy comprises a relative deep learning stage and a combined fine tuning stage. According to the method, a two-stage training scheme is combined, and the unlabeled data is fully utilized for supervision, so that the generalization ability of the model is remarkably improved under the condition of no large-scale labeled data. Especially in an unlabeled scene with a large distribution difference between training data and test data, the method shows high adaptability and robustness, the dependence on camera calibration parameters is reduced through modular design, and the performance and robustness of roadside monocular 3D target detection are comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the rapid development of intelligent transportation and autonomous driving technologies, roadside perception technology has become the core of intelligent transportation and vehicle networking systems due to its crucial role in enhancing traffic safety, efficiency, and intelligence. Compared with in-vehicle perception systems, roadside perception systems, through high-position cameras or sensors fixedly installed on both sides of the road, can not only provide a broader field of view and global perception capabilities but also effectively make up for the perception defects of in-vehicle perception systems in blind or occluded scenarios.

[0003] However, as one of the core tasks of roadside perception systems, roadside monocular 3D object detection faces the following technical problems due to the lack of direct depth information:

[0004] 1. Dependence on camera calibration: The installation positions of roadside cameras are diverse, and they are significantly affected by calibration parameters. Existing methods highly rely on accurate camera calibration (such as focal length, optical axis, installation angle, etc.). When the calibration deviates (such as equipment aging or external environmental interference), the detection performance drops severely.

[0005] 2. Insufficient generalization ability: The installation configurations (height, viewing angle, etc.) and scene conditions of roadside cameras vary greatly in different scenarios (such as cities, suburbs, highways). Existing models are difficult to adapt to unseen new scenarios. Especially when the distribution differences between training data and test data are large, the model performance drops significantly.

[0006] 3. Lack of roadside annotation data: The training of existing 3D detection models usually requires a large amount of high-quality annotation data, while the annotation cost of roadside scenes is extremely high. Existing datasets (such as Rope3D and DAIR-V2X) are insufficient in both scale and diversity and are difficult to meet the training requirements of generalization models.

[0007] 4. Difficulty in detecting distant objects: Roadside cameras are installed at high positions, covering a much larger range than in-vehicle sensors, but the target resolution is low. Existing depth estimation and object detection methods are difficult to handle distant objects, resulting in a decrease in accuracy. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to address the deficiencies of the prior art, specifically aiming at problems such as slow development cycle and long debugging time. Specifically, a 3D object detection method, system, device, and medium based on a roadside monocular camera are provided as follows:

[0009] 1) In the first aspect, the present invention provides a 3D object detection method based on a roadside monocular camera, and the specific technical solution is as follows:

[0010] Based on a roadside monocular camera, at least one target image of the area to be detected is obtained, and the target image is input into a pre-trained model to obtain a 3D target detection result;

[0011] The pre-trained model includes a backbone network, a depth prediction module, a decoder, and a 3D detection head. The backbone network is used to extract visual features from the target image, generate a first extraction result, and input the first extraction result into the depth prediction module. The depth prediction module extracts geometric information features from the first extraction result, generates a second extraction result, inputs the second extraction result into the decoder to generate a target embedding. The target embedding includes depth information, visual features, and target-to-target relationships in the geometric information features. The target embedding is input into the 3D detection head to obtain a 3D target detection result;

[0012] The pre-trained model is trained through a two-stage training strategy, and the two-stage training strategy includes a relative deep learning stage and a joint fine-tuning stage.

[0013] The beneficial effects of a 3D target detection method based on a roadside monocular camera provided by the present invention are as follows:

[0014] Combined with a two-stage training scheme, unlabeled data is fully utilized to improve the generalization ability of depth features, and the dependence on camera calibration parameters is reduced through modular design, comprehensively improving the performance and robustness of roadside monocular 3D target detection.

[0015] On the basis of the above solution, the present invention can also be improved as follows.

[0016] Further, the depth prediction module includes a relative depth estimation and an absolute depth conversion module;

[0017] The relative depth estimation module discretely models the depth of each pixel point through a depth classification head, generates a relative depth distribution, and uses a pseudo-label generation mechanism to supervise the relative depth. The pseudo-label is generated by the DepthAnything model;

[0018] The absolute depth conversion module uses a multi-layer perceptron to convert the relative depth feature into an absolute depth value, generating a global depth map, thereby providing the actual depth information of the target;

[0019] The depth prediction module can generate pseudo-labels through unlabeled data, effectively expanding the scale of training data and enhancing the generalization ability of the model;

[0020] Relative depth represents the depth difference between targets or pixels, and is used to describe the relative distance of targets within a local range, without involving absolute physical distances. It is processed through discrete classification modeling and is used to infer the relative positional relationship between targets;

[0021] Absolute depth provides the actual depth value of each pixel point, representing the physical distance between the target and the camera, and providing accurate spatial information for target positioning and 3D modeling;

[0022] Absolute depth is obtained by converting relative depth, thereby achieving the consistency of the global depth scale.

[0023] Further, the decoder includes: a self-attention module, a cross-attention module, and a feed-forward network.

[0024] Further, the relative deep learning stage includes:

[0025] Freeze the backbone network, decoder, and 3D detection head, and only optimize the depth prediction module; at this stage, through the pseudo-label supervision mechanism, combined with unlabeled data, optimize the learning of relative depth features, and generate the geometric relationship between targets through relative depth.

[0026] Further, the joint fine-tuning stage includes:

[0027] Unfreeze the backbone network, depth prediction module, decoder, and 3D detection head, and through the joint optimization of depth features and target detection tasks, fuse relative depth and absolute depth features to further improve the accuracy and generalization ability of the model in target detection.

[0028] 2) In the second aspect, the present invention also provides a 3D target detection system based on a roadside monocular camera, and the specific technical solution is as follows:

[0029] The detection unit is used to: based on the roadside monocular camera, obtain at least one target image of the area to be detected, and input the target image into the pre-trained model to obtain the 3D target detection result;

[0030] The pre-trained model includes a backbone network, a depth prediction module, a decoder, and a 3D detection head. The backbone network is used to extract visual features from the target image, generate a first extraction result, and input the first extraction result into the depth prediction module. The depth prediction module extracts geometric information features from the first extraction result, generates a second extraction result, inputs the second extraction result into the decoder to generate a target embedding, and the target embedding includes depth information, visual features, and the relationship between targets in the geometric information features. Input the target embedding into the 3D detection head to obtain the 3D target detection result;

[0031] The pre-trained model is trained through a two-stage training strategy, and the two-stage training strategy includes a relative deep learning stage and a joint fine-tuning stage.

[0032] Based on the above solution, the present invention can also be improved as follows.

[0033] Furthermore, the depth prediction module includes a relative depth estimation and an absolute depth conversion module;

[0034] The relative depth estimation module discretely models the depth of each pixel point through a depth classification head, generates a relative depth distribution, and supervises the relative depth using a pseudo-label generation mechanism, where the pseudo-labels are generated by the DepthAnything model;

[0035] The absolute depth conversion module uses a multi-layer perceptron to convert the relative depth features into absolute depth values, generating a global depth map, thereby providing the actual depth information of the target;

[0036] The depth prediction module can generate pseudo-labels through unlabeled data, effectively expanding the scale of training data and enhancing the generalization ability of the model;

[0037] Relative depth represents the depth difference between objects or pixels, used to describe the relative distance of objects within a local range without involving absolute physical distances, and is processed through discrete classification modeling for inferring the relative positional relationship between objects;

[0038] Absolute depth provides the actual depth value of each pixel point, representing the physical distance between the object and the camera, and providing accurate spatial information for object localization and 3D modeling;

[0039] Absolute depth is obtained by converting the relative depth, thereby achieving consistency in the global depth scale.

[0040] Furthermore, the decoder includes: a self-attention module, a cross-attention module, and a feed-forward network.

[0041] Furthermore, the relative deep learning stage includes:

[0042] Freeze the backbone network, decoder, and 3D detection head, and only optimize the depth prediction module; at this stage, through the pseudo-label supervision mechanism, combined with unlabeled data, optimize the learning of relative depth features, and generate the geometric relationship between objects through relative depth.

[0043] Furthermore, the joint fine-tuning stage includes:

[0044] Unfreeze the backbone network, depth prediction module, decoder, and 3D detection head, and through joint optimization of depth features and object detection tasks, fuse relative depth and absolute depth features to further improve the accuracy and generalization ability of the model in object detection.

[0045] 3) In a third aspect, the present invention further provides an electronic device, which includes a processor coupled to a memory. At least one computer program is stored in the memory and is loaded and executed by the processor to enable the electronic device to implement any one of the above methods.

[0046] 4) In a fourth aspect, the present invention further provides a computer-readable storage medium in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to enable a computer to implement any one of the above methods.

[0047] It should be noted that for the beneficial effects obtained by the technical solutions of the second to fourth aspects of the present invention and the corresponding possible implementation manners, reference may be made to the technical effects of the first aspect and its corresponding possible implementation manners described above, and details will not be elaborated here. Description of the Drawings

[0048] By reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:

[0049] Figure 1 It is a flowchart of a 3D object detection method based on a roadside monocular camera according to an embodiment of the present invention;

[0050] Figure 2 It is a schematic diagram of a model architecture of a 3D object detection method based on a roadside monocular camera according to an embodiment of the present invention;

[0051] Figure 3 It is a structural framework diagram of an electronic device. Detailed Embodiments

[0052] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0053] As Figure 1 and Figure 2 shown, a 3D object detection method based on a roadside monocular camera according to an embodiment of the present invention has the following specific technical solutions:

[0054] S1, based on the roadside monocular camera, obtain at least one target image of the area to be detected, and input the target image into a pre-trained model to obtain a 3D object detection result;

[0055] The pre-trained model includes a backbone network, a depth prediction module, a decoder, and a 3D detection head. The backbone network is used to extract visual features from the target image, generate a first extraction result, and input the first extraction result into the depth prediction module. The depth prediction module extracts geometric information features from the first extraction result, generates a second extraction result, inputs the second extraction result into the decoder to generate a target embedding. The target embedding includes depth information, visual features, and target relationship in the geometric information features. The target embedding is input into the 3D detection head to obtain a 3D target detection result;

[0056] The pre-trained model is trained by a two-stage training strategy, and the two-stage training strategy includes a relative deep learning stage and a joint fine-tuning stage.

[0057] The beneficial effects of a 3D target detection method based on a roadside monocular camera provided by the present invention are as follows:

[0058] Combined with a two-stage training scheme, it fully utilizes unlabeled data to improve the generalization ability of depth features, and reduces the dependence on camera calibration parameters through modular design, comprehensively improving the performance and robustness of roadside monocular 3D target detection.

[0059] The area to be detected refers to a traffic intersection or a section of a road where vehicles are traveling.

[0060] The 3D target refers to at least one of a motor vehicle, a non-motor vehicle, and a pedestrian.

[0061] (1) Backbone network:

[0062] The backbone network is responsible for extracting visual features from the input monocular image and generating multi-scale visual feature representations for subsequent modules to use. The present invention uses a pre-trained ResNet-50 as the backbone network and further fuses features of different scales through a feature pyramid structure. The visual features generated by the backbone network have the following characteristics:

[0063] Multi-scale fusion: Align features to the same resolution through upsampling and downsampling, enhancing the model's adaptability to target size changes.

[0064] High-dimensional visual representation: Provides rich semantic information for subsequent depth modeling and target detection.

[0065] The output visual features of the backbone network provide input for the subsequent depth prediction module.

[0066] (2) Depth prediction module

[0067] The depth prediction module is one of the core modules of the present invention, responsible for generating the geometric information of the scene, including relative depth and absolute depth. Among them, the relative depth describes the depth relationship between pixels, indicating the depth proximity of the target relative to other pixel points. The relative depth reflects the local geometric relationship of the target, but cannot provide the absolute distance value. The absolute depth provides the global depth value for each pixel point, indicating the actual distance from the target to the camera. The absolute depth is obtained through the conversion of the relative depth.

[0068] The module receives the multi-scale visual features generated by the backbone network, predicts the relative depth of pixel points through the depth classification head, and uses the pseudo-labels generated by the DepthAnything model for relative depth estimation to supervise the relative depth, ensuring that the module can accurately capture the geometric relationship between targets. The pseudo-labels are derived from unlabeled data, significantly expanding the scale of the training data.

[0069] The relative depth is converted into the absolute depth through a multi-layer perceptron (MLP). During the conversion process, the MLP uses the geometric relationship of the relative depth combined with the global features to generate the depth value of each pixel point, ensuring the geometric consistency of the global depth scale and the relative depth.

[0070] The depth features (including relative depth and absolute depth) generated by the depth prediction module will be used for depth-aware embedding in the subsequent Transformer decoder.

[0071] (3) Transformer decoder

[0072] The Transformer decoder combines the depth features and visual features through a multi-layer decoding structure to optimize the target embedding. The decoder captures the context relationship between targets and fuses the relative depth and absolute depth features to enhance the geometric perception ability of the target embedding.

[0073] The input of the module includes visual features and depth features (including relative depth and absolute depth), and the input is mapped into a unified target embedding through the embedding layer.

[0074] The decoder consists of a multi-layer structure, including:

[0075] Self-attention module: Captures the context relationship between targets and enhances the target embedding.

[0076] Cross-attention module: Combines the input embeddings (visual and depth features) to optimize the feature representation of the target.

[0077] Feed-forward network: Performs a non-linear transformation on the target embedding to further enhance the feature representation ability.

[0078] The target embeddings output by the decoder contain depth information, visual features, and the contextual relationships between targets, providing support for the 3D attribute prediction of the detection head.

[0079] (4) 3D Detection Head

[0080] The 3D detection head is responsible for parsing the target embeddings output by the decoder to generate the 3D attributes of the targets, including position, size, orientation, and depth.

[0081] Summary of the technical process:

[0082] Input: Receive RGB images from a roadside monocular camera;

[0083] Feature extraction: Generate multi-scale features through the backbone network;

[0084] Depth modeling: Generate target geometric features by combining relative depth and absolute depth;

[0085] Feature decoding: Optimize the target embeddings through Transformer decoding;

[0086] Object detection: Generate 3D attribute predictions of the targets.

[0087] To give full play to the role of the relative depth module and achieve the collaborative optimization of relative deep learning and 3D object detection tasks, the present invention designs a two-stage training strategy. This strategy effectively solves the conflict between depth estimation and object detection and significantly improves the generalization ability of the model by optimizing depth feature learning and 3D detection tasks in stages.

[0088] First stage: Relative deep learning stage

[0089] Training objectives:

[0090] Generate high-quality relative depth features to capture the depth differences between targets; use pseudo-labels generated from unlabeled data to expand the scale of training data and reduce the dependence on high-cost labeled data.

[0091] Implementation method:

[0092] Freeze unnecessary modules: In the first stage, only the depth prediction module, including the depth classification head and the absolute depth conversion module, is trained. The Transformer decoder and the 3D detection head are both frozen.

[0093] Pseudo-label supervision: Use the DepthAnything model to generate pseudo-labels to supervise the relative depth. The pseudo-labels cover complex scenarios (such as occlusion and dynamic targets), significantly improving the generalization ability of the relative depth estimation module.

[0094] By independently training the depth prediction module in the first stage, the model can capture high-quality relative depth features, provide geometric perception support for subsequent 3D detection tasks, and effectively reduce the mutual interference of multi-task training.

[0095] The second stage: Joint fine-tuning stage

[0096] In the second stage, all modules are unfrozen, and the depth features and 3D detection tasks are jointly optimized, enabling the depth features to be fully integrated into the visual features and object detection tasks, thereby improving the overall model performance.

[0097] Training objectives:

[0098] Fuse relative depth features and absolute depth features with visual features to enhance the geometric perception ability of objects; optimize the overall performance of the model and improve object detection accuracy and scene generalization ability.

[0099] Implementation method:

[0100] Unfreeze all modules: including the backbone network, depth prediction module, Transformer decoder, and 3D detection head. At the same time, optimize depth feature learning (including relative depth and absolute depth) and object detection tasks.

[0101] In the second stage, by jointly optimizing the depth prediction and object detection modules, the synergy between depth features and visual features is significantly enhanced. Especially in cross-source scene tests, the model shows stronger generalization ability.

[0102] Training in two stages can achieve the following effects:

[0103] 1. Task decoupling and reducing training conflicts: By independently training the depth prediction module in the first stage, the optimization conflicts common in multi-task joint training are avoided, and the quality of depth features is significantly improved.

[0104] 2. Making full use of unlabeled data: Using the pseudo-labels generated by DepthAnything expands the training data scale and makes up for the shortage of labeled data.

[0105] 3. Enhancing the model's generalization ability: The high-quality relative depth features captured in the first stage provide geometric perception support for object detection, while the joint optimization strategy in the second stage significantly improves the robustness and adaptability of the model in cross-source scenarios.

[0106] In another embodiment of this solution, after the target image is collected by the roadside monocular camera, the target image needs to be preprocessed to ensure that the input data into the pre-trained model is unified and the features are obvious.

[0107] The process of preprocessing the target image is specifically as follows:

[0108] Perform edge detection processing on any target image to obtain a detection result. Based on the detection result, determine the areas where edges cannot be recognized and the occluded areas, and integrate the above two areas to obtain a first image set. If the first image set is an empty set, directly input the target image into the pre-trained model for subsequent processing. If the first image set is not an empty set, process any unit in the first image set through image enhancement processing, and perform edge detection processing on the obtained second image set again. Determine whether the ratio of the area corresponding to the edge detection result in any target image to the total area of all targets exceeds a preset value. If it does not exceed, generate a label and perform manual recognition and adjustment processing on the target image with the generated label.

[0109] Among them, the total area of all targets is: the sum of the areas corresponding to all edge detection results and the rough areas of the targets where edges are not detected. The rough area is determined by manual recognition or by roughly determining the outline through color recognition. The method of determining the outline through color recognition is: determine the range of the area (quasi-annular) where there is an obvious difference between the color corresponding to the target and the color corresponding to the background, determine the midline of this area range, and use the area enclosed by the midline as the rough area. The midline is determined by connecting the midpoints of the boundary points of the inner circle and the outer circle of the quasi-annular.

[0110] The process of processing any unit in the first image set through image enhancement processing is specifically as follows:

[0111] Determine the target shooting time of the image to be enhanced and the target weather condition corresponding to the shooting location. Retrieve at least one historical image corresponding to the target weather condition from the historical database, and select the first historical image with the same shooting time as the target shooting time from all historical images. If not, select the historical image closest to the target shooting time. If there are at least two first historical images with the same shooting time as the target shooting time, further determine whether there is a first historical image with a similarity to the target in the image to be enhanced that exceeds the preset similarity. If so, use the enhancement processing scheme corresponding to this first historical image as the target enhancement processing scheme for the image to be enhanced. If the similarity is lower than the preset similarity, perform an average processing on the enhancement processing schemes of all first historical images, that is, perform a summation and averaging processing on the adjusted parameters, and use the average processing result as the target enhancement processing scheme.

[0112] It should be noted that the similarity between two images is determined through a preset model.

[0113] Based on the above solution, the present invention can also be improved as follows.

[0114] Furthermore, the depth prediction module includes a relative depth estimation and an absolute depth conversion module;

[0115] The relative depth estimation module discretely models the depth of each pixel point through a depth classification head, generates a relative depth distribution, and uses a pseudo-label generation mechanism to supervise the relative depth, where the pseudo-label is generated by the DepthAnything model;

[0116] The absolute depth conversion module uses a multi-layer perceptron to convert the relative depth features into absolute depth values, generating a global depth map, thereby providing the actual depth information of the target;

[0117] The depth prediction module can generate pseudo-labels through unannotated data, effectively expanding the scale of training data and enhancing the generalization ability of the model;

[0118] Relative depth represents the depth difference between objects or pixels, used to describe the relative distance of an object within a local range without involving absolute physical distance, and is processed through discrete classification modeling for inferring the relative positional relationship between objects;

[0119] Absolute depth provides the actual depth value of each pixel point, representing the physical distance between the object and the camera, and providing accurate spatial information for object localization and 3D modeling;

[0120] Absolute depth is obtained by converting the relative depth, thereby achieving the consistency of the global depth scale.

[0121] Furthermore, the decoder includes: a self-attention module, a cross-attention module, and a feed-forward network.

[0122] Furthermore, the relative deep learning stage includes:

[0123] Freeze the backbone network, decoder, and 3D detection head, and only optimize the depth prediction module; at this stage, through the pseudo-label supervision mechanism, combined with unannotated data, optimize the learning of relative depth features, and generate the geometric relationship between objects through relative depth.

[0124] Furthermore, the joint fine-tuning stage includes:

[0125] Unfreeze the backbone network, depth prediction module, decoder, and 3D detection head, and through joint optimization of depth features and object detection tasks, fuse relative depth and absolute depth features to further improve the accuracy and generalization ability of the model in object detection.

[0126] The effects that this solution can achieve compared with the prior art are:

[0127] 1. A joint modeling strategy for relative depth and absolute depth is proposed, which captures the local geometric relationships between objects and generates depth information with consistent global scale. Relative depth is modeled through discrete classification, enhancing the local depth perception ability of distant objects; absolute depth is converted into continuous values through MLP to make up for the lack of global depth scale information.

[0128] 2. Combining with the DepthAnything model to generate high-quality pseudo-labels to supervise relative depth effectively expands the scale of training data. The pseudo-labels cover complex scenarios (such as occlusion, dynamic objects, etc.), significantly enhancing the generalization ability of the model.

[0129] 3. Through a two-stage training strategy, in the first stage, focus on relative deep learning and only optimize the depth prediction module to avoid conflicts in multi-task training; in the second stage, jointly optimize the depth features with the 3D detection task to achieve the coordinated improvement of depth features and visual features, significantly improving the object detection accuracy.

[0130] 4. In the first stage, through pseudo-label supervision, learn depth features with generalization ability; in the second stage, through the joint optimization strategy, significantly improve the performance of the model in heterogeneous scenarios.

[0131] 5. The depth prediction module significantly reduces the dependence on calibration parameters through the joint modeling of relative depth and absolute depth; even in scenarios where the calibration parameters are inaccurate or change dynamically, the model can still maintain stable detection performance.

[0132] 6. Using the pseudo-label generation mechanism to cover complex scenarios significantly enhances the adaptability of the model to dynamic objects and multi-object occlusion; combined with the Transformer decoder to capture the context relationships between objects and improve the detection performance in complex scenarios.

[0133] In the above embodiments, although the steps are numbered S1, S2, etc., these are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation, which is also within the protection scope of the present invention. It can be understood that in some embodiments, it may include some or all of the above embodiments.

[0134] This solution also provides a 3D object detection system based on a roadside monocular camera. The specific technical solution is as follows:

[0135] The detection unit is used to: based on the roadside monocular camera, obtain at least one target image of the area to be detected, and input the target image into the pre-trained model to obtain the 3D object detection result;

[0136] The pre-trained model includes a backbone network, a depth prediction module, a decoder, and a 3D detection head. The backbone network is used to extract visual features from the target image, generate a first extraction result, and input the first extraction result into the depth prediction module. The depth prediction module extracts geometric information features from the first extraction result, generates a second extraction result, inputs the second extraction result into the decoder to generate a target embedding. The target embedding includes depth information, visual features, and target relationship in the geometric information features. The target embedding is input into the 3D detection head to obtain a 3D object detection result;

[0137] The pre-trained model is trained by a two-stage training strategy, and the two-stage training strategy includes a relative deep learning stage and a joint fine-tuning stage.

[0138] Based on the above solution, the present invention can also be improved as follows.

[0139] Further, the depth prediction module includes a relative depth estimation and an absolute depth conversion module;

[0140] The relative depth estimation module discretely models the depth of each pixel point through a depth classification head, generates a relative depth distribution, and uses a pseudo-label generation mechanism to supervise the relative depth. The pseudo-label is generated by the DepthAnything model;

[0141] The absolute depth conversion module uses a multi-layer perceptron to convert the relative depth feature into an absolute depth value, generates a global depth map, and thus provides the actual depth information of the target;

[0142] The depth prediction module can generate pseudo-labels through unlabeled data, effectively expanding the scale of training data and enhancing the generalization ability of the model;

[0143] Relative depth represents the depth difference between objects or pixels, and is used to describe the relative distance of objects within a local range without involving absolute physical distances. It is processed through discrete classification modeling and is used to infer the relative positional relationship between objects;

[0144] Absolute depth provides the actual depth value of each pixel point, representing the physical distance between the object and the camera, and providing accurate spatial information for object positioning and 3D modeling;

[0145] Absolute depth is obtained by converting the relative depth, thus achieving the consistency of the global depth scale.

[0146] Further, the decoder includes: a self-attention module, a cross-attention module, and a feed-forward network.

[0147] Further, the relative deep learning stage includes:

[0148] Freeze the backbone network, decoder, and 3D detection head, and only optimize the depth prediction module; at this stage, through the pseudo-label supervision mechanism, combined with unlabeled data, optimize the learning of relative depth features, and generate the geometric relationship between objects through relative depth.

[0149] Furthermore, the joint fine-tuning stage includes:

[0150] Unfreeze the backbone network, depth prediction module, decoder, and 3D detection head, and further improve the accuracy and generalization ability of the model in object detection by jointly optimizing depth features and object detection tasks, and fusing relative depth and absolute depth features.

[0151] It should be noted that the beneficial effects of the 3D object detection system based on the roadside monocular camera provided in the above embodiments are the same as those of the 3D object detection method based on the roadside monocular camera, and will not be elaborated here. In addition, when the system provided in the above embodiments realizes its functions, only the above-mentioned division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be elaborated here.

[0152] As Figure 3 shown, an electronic device 300 according to an embodiment of the present invention, the electronic device 300 includes a processor 320, the processor 320 is coupled to a memory 310, and at least one computer program 330 is stored in the memory 310. The at least one computer program 330 is loaded and executed by the processor 320 to enable the electronic device 300 to implement any one of the above methods. Specifically:

[0153] The electronic device 300 may vary greatly due to configuration or performance differences, and may include one or more processors 320 (Central Processing Units, CPUs) and one or more memories 310. Among them, at least one computer program 330 is stored in the one or more memories 310, and the at least one computer program 330 is loaded and executed by the one or more processors 320 to enable the electronic device 300 to implement a 3D object detection method based on the roadside monocular camera provided in the above embodiments. Of course, the electronic device 300 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The electronic device 300 may also include other components for implementing the functions of the device, which will not be elaborated here.

[0154] A computer-readable storage medium according to an embodiment of the present invention stores at least one computer program, and the at least one computer program is loaded and executed by a processor to enable a computer to implement any of the above methods.

[0155] Optionally, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0156] In an exemplary embodiment, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the electronic device to execute any of the above methods.

[0157] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and do not represent a limitation on a specific order or sequence. In appropriate cases, the order of use of similar objects may be interchanged so that the embodiments of the present application described herein can be implemented in an order other than the illustrated or described order.

[0158] Those skilled in the art of the present technology know that the present invention can be implemented as a system, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: it can be completely hardware, can be completely software (including firmware, resident software, microcode, etc.), or can be a combination of hardware and software, which is generally referred to as "circuit", "module", or "system" herein. In addition, in some embodiments, the present invention can also be implemented in the form of a computer program product in one or more computer-readable media, and the computer-readable media contain computer-readable program codes.

[0159] Any combination of one or more computer-readable media may be employed. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present document, a computer-readable storage medium may be any tangible medium that contains or stores a program which can be used by or in connection with an instruction execution system, apparatus, or device.

[0160] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art may make variations, modifications, substitutions, and alterations to the above embodiments within the scope of the present invention.

Claims

1. A 3D object detection method based on a roadside monocular camera, characterized in that, Including: Based on a roadside monocular camera, at least one target image of the area to be detected is obtained, and the target image is input into a pre-trained model to obtain a 3D target detection result; The pre-trained model includes a backbone network, a depth prediction module, a decoder, and a 3D detection head. The backbone network is used to extract visual features from the target image, generate a first extraction result, and input the first extraction result into the depth prediction module. The depth prediction module extracts geometric information features from the first extraction result, generates a second extraction result, inputs the second extraction result into the decoder to generate a target embedding. The target embedding includes depth information, visual features, and the relationship between targets in the geometric information features. The target embedding is input into the 3D detection head to obtain a 3D target detection result; The pre-trained model is trained through a two-stage training strategy, and the two-stage training strategy includes a relative deep learning stage and a joint fine-tuning stage.

2. The 3D object detection method based on a roadside monocular camera according to claim 1, wherein The depth prediction module includes a relative depth estimation module and an absolute depth conversion module; The relative depth estimation module discretely models the depth of each pixel point through a depth classification head, generates a relative depth distribution, and uses a pseudo-label generation mechanism to supervise the relative depth. The pseudo-label is generated by the DepthAnything model; The absolute depth conversion module uses a multi-layer perceptron to convert the relative depth feature into an absolute depth value, generating a global depth map, thereby providing the actual depth information of the target; The depth prediction module can generate pseudo-labels through unlabeled data, effectively expanding the scale of training data and enhancing the generalization ability of the model; Relative depth represents the depth difference between targets or pixels, and is used to describe the relative distance of targets within a local range without involving absolute physical distances. It is processed through discrete classification modeling and is used to infer the relative position relationship between targets; Absolute depth provides the actual depth value of each pixel point, representing the physical distance between the target and the camera, and providing accurate spatial information for target positioning and three-dimensional modeling; Absolute depth is obtained by converting relative depth, thereby achieving the consistency of the global depth scale.

3. A 3D object detection method based on a roadside monocular camera according to claim 1, characterized in that, The decoder includes: a self-attention module, a cross-attention module, and a feed-forward network.

4. The 3D object detection method based on a roadside monocular camera according to claim 1, wherein The relative deep learning stage includes: Freeze the backbone network, decoder, and 3D detection head, and only optimize the depth prediction module; at this stage, through the pseudo-label supervision mechanism, combined with unlabeled data, optimize the learning of relative depth features, and generate the geometric relationship between targets through relative depth.

5. A 3D object detection method based on a roadside monocular camera according to claim 1, characterized in that, The joint fine-tuning stage includes: Unfreeze the backbone network, depth prediction module, decoder, and 3D detection head, and further improve the accuracy and generalization ability of the model in target detection by jointly optimizing the depth features and the target detection task, and fusing relative depth and absolute depth features.

6. A 3D object detection system based on a roadside monocular camera, characterized in that, Including: The detection unit is used to: based on a roadside monocular camera, obtain at least one target image of the area to be detected, and input the target image into a pre-trained model to obtain a 3D target detection result; The pre-trained model includes a backbone network, a depth prediction module, a decoder, and a 3D detection head. The backbone network is used to extract visual features from the target image, generate a first extraction result, and input the first extraction result into the depth prediction module. The depth prediction module extracts geometric information features from the first extraction result, generates a second extraction result, inputs the second extraction result into the decoder to generate a target embedding. The target embedding includes depth information, visual features, and target-interaction relationships in the geometric information features. Input the target embedding into the 3D detection head to obtain a 3D object detection result; The pre-trained model is trained by a two-stage training strategy, and the two-stage training strategy includes a relative deep learning stage and a joint fine-tuning stage.

7. The 3D object detection system based on a roadside monocular camera according to claim 6, wherein The depth prediction module includes a relative depth estimation and an absolute depth conversion module; The relative depth estimation module discretely models the depth of each pixel point through a depth classification head, generates a relative depth distribution, and supervises the relative depth using a pseudo-label generation mechanism. The pseudo-label is generated by the DepthAnything model; The absolute depth conversion module uses a multi-layer perceptron to convert the relative depth features into absolute depth values, generating a global depth map, thereby providing the actual depth information of the target; The depth prediction module can generate pseudo-labels through unlabeled data, effectively expanding the scale of training data and enhancing the generalization ability of the model; Relative depth represents the depth difference between objects or pixels, and is used to describe the relative distance of an object within a local range, without involving absolute physical distances. It is processed through discrete classification modeling and is used to infer the relative position relationship between objects; Absolute depth provides the actual depth value of each pixel point, representing the physical distance between the object and the camera, and providing accurate spatial information for object localization and three-dimensional modeling; Absolute depth is obtained by converting the relative depth, thereby achieving the consistency of the global depth scale.

8. The 3D object detection system based on a roadside monocular camera according to claim 7, wherein The decoder includes: a self-attention module, a cross-attention module, and a feed-forward network.

9. An electronic device, characterized in that, The electronic device includes a processor, the processor is coupled to a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by the processor so that the electronic device implements the method according to any one of claims 1 to 5.

10. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor so that a computer implements the method according to any one of claims 1 to 5.