Road vehicle real-time twinning method and system based on monocular camera
By combining target detection, target tracking and depth estimation technologies, using a monocular camera to perceive and map vehicle information in real time, the shortcomings of the existing road twin system in dynamic vehicle twin support are solved, and accurate twinning and rapid adaptation to multi-target vehicles are achieved.
Patent Information
- Application Number
- CN202510367536.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The existing road twin system is difficult to achieve effective twinning of real-time moving vehicles when only monocular video streaming data is available, and cannot meet real-time requirements.
Combining target detection, target tracking and depth estimation technologies, a monocular camera is used to perceive vehicle changes in real time, and extract features through visual question-and-answer big models to map vehicle locations and features into virtual twin scenes in real time.
It realizes the precise twinning of multi-target vehicles in dynamic scenarios under monocular video streaming data, reduces dependence on specific task data and scenarios, supports generative reasoning and multi-round interactions, and quickly adapts to new scenarios.
Smart Images

Figure CN120260000A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to digital twins, and particularly relates to a method and system for real-time road vehicle twinning based on a monocular camera. Background Art
[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] As an important bridge connecting the physical world and the virtual world, Digital Twin has been widely applied in fields such as industrial manufacturing, smart cities, and intelligent transportation in recent years. By synchronizing the state of physical entities to the virtual environment in real time, Digital Twin can achieve dynamic monitoring, prediction, and optimization of physical objects and systems. Among them, the twinning application of local road scenes is very extensive, such as traffic flow monitoring, traffic congestion analysis, road planning, etc. These applications have greatly improved the road management efficiency and traffic safety level.
[0004] However, in real application scenarios, most roads are only equipped with fixed-position monitoring cameras and do not have expensive sensing devices such as lidar. They can only provide single monocular video stream data to feedback the real-time change state of road traffic. This leads to the fact that most current road twinning systems mainly focus on the modeling and management of static scenes, and there are certain deficiencies in supporting the twinning of real-time moving vehicles, making it difficult to fully meet the real-time requirements. Therefore, how to complete the real-time twinning of vehicles in road scenes on the premise of only having monocular video stream data has become an urgent problem to be solved in road twinning systems. Summary of the Invention
[0005] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a method and system for real-time road vehicle twinning based on a monocular camera. By combining object detection, object tracking, and depth estimation technologies, it can real-time sense the changes of multi-target vehicles in dynamic scenes and accurately map their positions and features to the virtual twin scene, making up for the deficiencies of existing road twin scenes in supporting dynamic vehicle twinning.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] In the first aspect, the present invention provides a method for real-time road vehicle twinning based on a monocular camera, including:
[0008] Using an object detection network to detect vehicles in the video stream and obtain the rectangular bounding boxes of the detected vehicles; wherein, the video stream is captured by a monocular camera;
[0009] Identify the color and category of the vehicles in the video stream based on the large vision question answering model, update the target tracker based on the rectangular bounding boxes of the detected vehicles obtained by object detection, and use the updated tracking ID as the detected vehicle ID;
[0010] Use the depth estimation model to calculate the depth map of the vehicles in the video stream, and combine it with the rectangular bounding boxes of the detected vehicles to obtain the depth information of the detected vehicles;
[0011] Based on the depth information of the detected vehicles and the geographical coordinates of the monocular camera, obtain the real geographical coordinates of the detected vehicles;
[0012] Update the real-time obtained detected vehicle ID, color, category, and the real geographical coordinates of the detected vehicles to the road twin scenario.
[0013] In a second aspect, the present invention provides a real-time road vehicle twin system based on a monocular camera, including:
[0014] An object detection module, which is configured to: use an object detection network to detect vehicles in the video stream to obtain the rectangular bounding boxes and contours of the detected vehicles; wherein, the video stream is captured by a monocular camera;
[0015] A target tracking module, which is configured to: identify the color and category of the vehicles in the video stream based on the large vision question answering model, update the target tracker based on the rectangular bounding boxes of the detected vehicles obtained by object detection, and use the updated tracking ID as the detected vehicle ID;
[0016] A depth estimation module, which is configured to: use the depth estimation model to calculate the depth map of the vehicles in the video stream, and combine it with the rectangular bounding boxes of the detected vehicles to obtain the depth information of the detected vehicles;
[0017] A coordinate conversion module, which is configured to: based on the depth information of the detected vehicles and the geographical coordinates of the monocular camera, obtain the real geographical coordinates of the detected vehicles;
[0018] A real-time twin module, which is configured to: update the real-time obtained detected vehicle ID, color, category, and the real geographical coordinates of the detected vehicles to the road twin scenario.
[0019] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.
[0020] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.
[0021] In a fifth aspect, the present invention provides a computer program product, including a computer program which, when executed by a processor, implements the method described in the first aspect.
[0022] The above one or more technical solutions have the following beneficial effects:
[0023] In the present invention, by combining object detection, object tracking, and depth estimation technologies, it is possible to real-time perceive the changes of multi-target vehicles in a dynamic scene and accurately map their positions and features to the virtual twin scene, making up for the deficiencies of the existing road twin scene in supporting dynamic vehicle twins. Using a visual question answering large model to extract vehicle features reduces the dependence on specific task data and scenarios, supports generative reasoning and multi-round interaction, can dynamically adjust questions, and quickly adapt to new scenarios.
[0024] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0026] Figure 1 It is a flowchart of a real-time road vehicle twin method based on a monocular camera in the first embodiment of the present invention;
[0027] Figure 2 It is a network structure diagram of FC3k2, FC3k, FasterNet Block, and PConv modules in the first embodiment of the present invention;
[0028] Figure 3 It is a network structure diagram of the OAA Head module in the first embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0030] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments of the present invention.
[0031] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0032] Embodiment 1
[0033] This embodiment discloses a real-time road vehicle digital twin method based on a monocular camera, including:
[0034] Using an object detection network to detect vehicles in the video stream and obtain the rectangular bounding boxes of the detected vehicles; wherein, the video stream is captured by a monocular camera;
[0035] Based on a visual question answering large model, identify the colors and categories of the vehicles in the video stream, update the object tracker based on the rectangular bounding boxes of the detected vehicles obtained by object detection, and use the updated tracking ID as the detected vehicle ID;
[0036] Use a depth estimation model to calculate the depth map of the vehicles in the video stream, and combine it with the rectangular bounding boxes of the detected vehicles to obtain the depth information of the detected vehicles;
[0037] According to the depth information of the detected vehicles and the geographical coordinates of the monocular camera, obtain the real geographical coordinates of the detected vehicles;
[0038] Update the detected vehicle ID, color, category, and the real geographical coordinates of the detected vehicles obtained in real time to the road digital twin scenario.
[0039] The solution of this embodiment combines object detection, object tracking, and depth estimation technologies, can real-time perceive the changes of multi-target vehicles in a dynamic scene, and accurately map their positions and features to the virtual digital twin scenario, making up for the deficiencies of the existing road digital twin scenario in supporting dynamic vehicle digital twins. Using a visual question answering large model to extract vehicle features, reduces the dependence on specific task data and scenarios, and supports generative reasoning and multi-round interaction, can dynamically adjust questions and quickly adapt to new scenarios.
[0040] Next, in combination with Figure 1 A real-time road vehicle digital twin method based on a monocular camera proposed in this embodiment will be described in detail.
[0041] Step 1: Use an object detection network to detect vehicles in the video stream and obtain the rectangular bounding boxes of the detected vehicles.
[0042] In this embodiment, the improved YOLOv11 network built is used as the object detection network. Based on the FC3k2 module, the inference speed is improved through partial convolution and point convolution; by introducing the OAA detection head, the detection accuracy of occluded targets is improved using a hybrid attention mechanism; by using a loss function, the target box regression is optimized, and the detection accuracy of medium-quality targets is improved.
[0043] Next, the network structure of the improved YOLOv11 network will be described in detail.
[0044] Step 11: Construct the FC3k2 module, i.e., Faster C3k2, to replace the C3k2 module in the YOLOv11 network, thereby improving the overall network operation speed.
[0045] The FC3k2 in this embodiment is improved based on the C3k2 module. Specifically, the FasterNet Block module with PConv (partial convolution) is introduced to replace the Bottleneck module based on traditional convolution in the C3k2 module.
[0046] The structure of the FasterNet Block module consists of multiple core components, aiming to improve the inference speed and maintain high precision by optimizing computational efficiency and memory access. First, the input feature map is initially processed through a conventional convolution operation. Then, the PConv operation is introduced, which only performs convolution calculations on some channels of the input feature map, thereby reducing redundant calculations and memory bandwidth consumption and significantly improving computational efficiency. Next, PWConv (point convolution) is used to fuse the information of the channels of the feature map, further strengthening the feature interaction between channels. To avoid information loss, the FasterNet Block uses a residual connection to add the input feature map to the output after convolution processing, ensuring that the information flow is not blocked. In addition, after each layer of convolution operation, batch normalization (BN) and ReLU activation function are used to help accelerate training and improve the stability of the network.
[0047] Benefiting from the introduction of PConv to replace the traditional convolution operation, the FC3K2 module reduces the computational amount and memory bandwidth requirements while increasing the floating-point operations per second (FLOPS), accelerating the inference process while maintaining high precision. In the traditional convolution operation, each channel of the input feature map needs to be operated with the convolution kernel, and the convolution kernel slides through the entire input feature map. This means that for each convolution kernel, all pixel points of each channel of the input feature map need to be accessed, and the results are stored in the corresponding positions of the output feature map, which will lead to redundant computational amounts, especially when some channels contain a large amount of repetitive information. Because all the information of each channel must be read and processed, a large amount of redundant calculations not only increase the floating-point operation amount (FLOPs), but also increase the memory access amount. For a large input feature map and multiple convolution kernels, this will generate a large number of memory read and write operations. Especially in the case of multiple channels, the memory access requirements increase sharply. This high-frequency memory access not only increases the memory bandwidth requirements, but also may lead to a memory access bottleneck, thereby affecting the overall computational efficiency. While PConv can better solve the operation performance problems brought by traditional convolution by only performing convolution calculations on some channels of the input feature map.
[0048] First, the behavior of PConv that only performs convolution operations on some channels reduces redundant calculations and decreases the computational amount. In the PConv operation, some channels of the input feature map participate in the convolution calculation while other channels remain unchanged. Therefore, the number of memory accesses is greatly reduced. Assuming the size of the feature map is w×h, the total number of channels is c, the convolution kernel size is k, and the number of channels that need to perform convolution calculation is c p , then the FLOPs of PConv are only Under normal circumstances, taking a partial ratio then the FLOPs of PConv are only of the traditional convolution. In addition, the PConv operation significantly reduces the amount of memory that needs to be accessed by only performing convolution calculations on a part of the channels of the input feature map. PConv only selects a part of the channels for calculation, thus avoiding accessing all channels. By reducing the amount of data required for each convolution operation, PConv reduces the consumption of memory bandwidth, alleviates the burden of memory access, makes PConv more efficient during the calculation process, and can reduce the latency caused by frequent memory accesses.
[0049] Therefore, by introducing the FC3k2 module to replace the C3k2 module in the original YOLOv11 network, the purpose of improving the inference speed of object detection can be achieved.
[0050] The SPPF module enhances the multi-scale information of the feature map through the feature pyramid technology. The module processes the input feature map through pooling layers of multiple different sizes such as 1*1, 2*2, 3*3, etc., and finally concatenates these pooling results together.
[0051] The PSABloc module combines the attention mechanism and improves the model's attention to important features by weighting the feature map. The module first extracts features through 1 convolutional layer, then applies the attention mechanism to weight the feature map, and finally outputs the enhanced feature map through 1 convolutional layer.
[0052] The C2PSA module (C2P Spatial Attention) extracts the features of the input image through the convolutional layer and enhances the spatial attention of the features through the PSABloc module. The module first extracts features through 1 convolutional layer, then inputs these features into multiple PSABloc modules for spatial attention processing, and finally outputs the enhanced feature map through 1 convolutional layer.
[0053] The Backbone network of the YOLOv11 network has 11 layers, which are, in order: 1 layer of Conv module, 2x2 layers of Conv-FC3k2 (c3k = false) combined module, 2x2 layers of Conv-FC3k2 (c3k = true) combined module, 1 layer of SPPF module, 1 layer of C2PSA module, and the output channel numbers are 64, 128, 256, 256, 512, 512, 512, 1024, 1024, 1024, 1024 respectively.
[0054] The Neck network of the YOLOv11 network is used for feature fusion and upsampling. The Neck network has 3 branches and a total of 9 layers, which are, in order: 1 layer of Upsample module, 1 layer of C3k2 (c3k = False) module, 1 layer of Upsample module, 1 layer of C3k2 (c3k = False) module, 1 layer of Conv module, 1 layer of C3k2 (c3k = False) module, 1 layer of Conv module, 1 layer of C3k2 (c3k = True) module. The 3 branches finally output 3 feature maps with sizes of 80*80*64, 40*40*128, and 20*20*256 respectively.
[0055] In this embodiment, an OAAHead (Occlusion-Aware Attention Head) module is constructed to replace the Head module in YOLOv11 to improve the anti-occlusion ability of the overall network.
[0056] Specifically, the OAA Head module includes two processing branches for classification and regression loss. The first branch includes a Conv module, an OAA module, and a Con2d module connected in sequence; the second branch includes a DWConv module, a Conv module, an OAA module, and a Con2d module connected in sequence; the outputs of the first branch and the second branch are concatenated to obtain the output result of the OAA Head module.
[0057] The first part of the OAA module is the depthwise separable convolution with residual connections. The depthwise separable convolution is performed layer by layer, that is, through channel-separable convolution. Although the depthwise separable convolution can learn the importance of different channels and reduce the number of parameters, it ignores the information relationship between channels. To make up for this loss, the outputs of different depth convolutions are then combined through PWConv (pointwise convolution); then, referring to the classic channel attention mechanism SENet, average pooling operation is used to aggregate the feature map along the channel direction to obtain a global description of the channels; then, the information of each channel is fused through a designed two-layer fully connected network, and the multi-channel association is learned through non-linear transformation between the two-layer fully connected networks; finally, the output of the OAA module is used as the attention weight and multiplied by the original features input to the OAA module. The key to the OAA module's ability to effectively solve the object occlusion problem lies in its introduction of a hybrid attention mechanism, which enhances the model's feature learning ability for occluded regions. In the case of object occlusion, the occluding object will cause some features of the target to not be extracted normally, thus affecting the overall object detection accuracy.
[0058] The traditional convolutional network used in the original detection head of the YOLOv11 network is difficult to process the lost features in the occluded region, while the OAA module can extract the features of the occluded region from multiple dimensions by introducing depthwise separable convolution (DWConv) and fully connected transformation. Specifically, the depthwise separable convolution in OAA convolves the spatial positions of each channel independently, which can construct a spatial attention distribution for spatial recalibration. And the fully connected transformation can learn the complex relationships between channels, which is a specific implementation of the channel attention mechanism and helps to capture the important features between channels.
[0059] In this embodiment, a DAS (Dynamic-Attentive-Scale) loss function is designed to replace the CIoU loss function in the original YOLOv11 network to improve the overall detection accuracy of the network.
[0060] The formula of the DAS loss function is:
[0061]
[0062]
[0063] where λ is a hyperparameter that controls the behavior of the non-monotonic attention mechanism; I represents the intersection between the predicted box and the ground truth box; U represents the union of the predicted box and the ground truth box; dw1 represents the horizontal distance between the left edge of the predicted box and the left edge of the target box; dw2 represents the horizontal distance between the right edge of the predicted box and the right edge of the target box; dh1 represents the vertical distance between the upper edge of the predicted box and the upper edge of the target box; dh2 represents the vertical distance between the lower edge of the anchor box and the lower edge of the target box. w gtand h gt respectively represent the width and height of the target bounding box; w pred and h pred respectively represent the width and height of the anchor box.
[0064] The DAS loss introduces a penalty factor that adapts to the size of the target bounding box, avoiding the unnecessary expansion of the anchor box during the regression process, thereby accelerating the convergence process. At the same time, the DAS loss dynamically adjusts the gradient according to the quality of the anchor box through a gradient adjustment mechanism, making the regression more stable, especially effectively improving the optimization speed of medium-quality anchor boxes when dealing with anchor boxes with unbalanced quality. In addition, the DAS loss also improves the attention to medium-quality anchor boxes by introducing a non-monotonic attention mechanism, enhancing the adaptability of the network; finally, a scale matching penalty term is added to the loss to guide the regression process to consider the consistency of the aspect ratio and area ratio simultaneously, enabling anchor boxes of different scales to obtain reasonable gradients and focusing intensities, and avoiding gradient imbalance between small and large targets. Therefore, introducing the DAS loss to measure the similarity between the predicted bounding box and the true target bounding box can achieve the purpose of improving the object detection accuracy of the overall network.
[0065] Load the pre-trained weights of the Coco dataset, input the RGB three-channel frame images intercepted from the monocular camera video stream, and after calculation by the trained object detection network, output the bounding box information and vehicle contour information of the vehicles in the picture.
[0066] Step 2: Based on the large vision question answering model, identify the color and category appearance features of the vehicles in the video stream, update the target tracker based on the rectangular bounding box of the detected vehicles, and use the updated tracking ID as the detected vehicle ID.
[0067] Create a deep learning tracker and a large vision question answering model to obtain the tracking ID and feature information such as color and category of the detected vehicles.
[0068] In Step 2, the specific steps are as follows:
[0069] Step 21: Initialize the deep learning object tracker, update the target tracker using the vehicle detection bounding box information, and use the updated tracking ID as the vehicle ID.
[0070] Step 22: For vehicle IDs that have not appeared before, use the corresponding vehicle contour information to crop the transparent background image matrix of the target vehicle from the original frame image.
[0071] Step 23: Create and initialize the large vision image question answering model.
[0072] Pre-define the appearance feature candidate set F:
[0073] F = {f1, f2, …, fm}
[0074] Among them, f k represents a candidate for color or category appearance features.
[0075] Design problem template:
[0076] Is this vehicle [X]?
[0077] Among them, the placeholder [X] represents the color or category feature description part.
[0078] Fill each feature f in the candidate feature set F k one by one into the placeholder [X] of the template question to form a candidate question set Q set :
[0079] Q set = {Q k = Is this vehicle f k ? | k = 1, 2,..., m}
[0080] Input each candidate question Q k together with the target vehicle image i into the visual image question answering large model to calculate the confidence score S of each target vehicle for each question k :
[0081] S k = M VQA (Q k , i)
[0082] Among them, M VQA represents the visual question answering model.
[0083] Select the feature corresponding to the highest score from the confidence scores of the candidate features as the final recognition result.
[0084] Step 24: According to the object detection result of the image, intercept the local image of each detected vehicle to obtain the target set:
[0085] T = {t1, t2,..., t n}
[0086] Among them, t i represents the local image of the i-th target vehicle.
[0087] Input the local image of the target vehicle into the visual question answering large model, and cache the extracted color and appearance features associated with the corresponding vehicle's tracking ID.
[0088] Step 25, for the vehicle ID that has appeared before, directly use the cached vehicle feature extraction result.
[0089] Step 3: Calculate the depth map of the vehicle in the video stream using a depth estimation model, and combine it with the rectangular bounding box of the detected vehicle to obtain the depth information of the detected vehicle.
[0090] Calculate the depth map of the video stream frame in real time, and calculate the vehicle depth in combination with the vehicle pixel coordinates.
[0091] In step 3, the specific steps are as follows:
[0092] Step 31: Create and initialize a monocular absolute depth estimation deep learning model.
[0093] Step 32: Input the RGB three-channel frame image intercepted from the monocular camera video stream, and calculate the depth map through the absolute depth estimation model.
[0094] Step 33: Combine the depth map and the vehicle pixel coordinates obtained from object detection to obtain the depth information of the detected vehicle.
[0095] Step 34: Use a mean smoother to smooth the depth value obtained from the latest estimate in combination with historical data.
[0096] Step 4, vehicle spatial coordinate conversion calculation: Calculate the true geographical coordinates of the vehicle according to the vehicle pixel coordinates, vehicle depth, and camera information.
[0097] In step 4, the specific steps are as follows:
[0098] Step 41: Use the target camera to take multiple checkerboard pictures, confirm that the pose of the checkerboard and the camera angle are different each time, and grayscale each image.
[0099] Step 42: Use the cornerSubPix function to perform sub-pixel level refinement on the detected corner coordinates to improve the corner positioning accuracy.
[0100] Step 43: Set the checkerboard plane as the XY plane of the world coordinate system, Z = 0, and construct the true three-dimensional coordinates of each corner of the checkerboard in the world coordinate system. For a checkerboard of m×n (m rows and n columns), the coordinates of each corner can be expressed as:
[0101] (x,y,z) = (j·d,j·d,0)
[0102] where d is the side length of each grid of the checkerboard, and i, j are the row and column numbers of the corner points.
[0103] Step 44: Call the calibrateCamera function of OpenCV, and solve the internal parameter matrix K and distortion coefficient dst of the camera based on the corner data (the corresponding relationship between pixel coordinates and true world coordinates) of multiple calibration images.
[0104]
[0105] dst = [k1, k2, p1, p2, k3]
[0106] where f x , f y is the representation of the focal length in pixel units in the image width and height directions; c x , c y is the position of the principal point (the intersection of the optical axis and the image plane) in the image coordinates; the fixed value 1 is used for the transformation of the homogeneous coordinate system, representing the normalization of the linear projection matrix; k1, k2, k3 are the radial distortion coefficients; p1, p1 are the tangential distortion coefficients.
[0107] Step 45: Optimize the original intrinsic matrix K according to the distortion coefficient dst, and use OpenCV to calculate the new intrinsic matrix K'.
[0108]
[0109] Step 46: Combine the vehicle tracking ID, vehicle pixel coordinates, vehicle depth, and the intrinsic parameters after camera calibration to obtain the spatial coordinates of the vehicle in the camera coordinate system.
[0110] Step 47: Combine the spatial coordinates of the vehicle in the camera coordinate system, the real geographical location of the camera, and the orientation angle of the camera to obtain the real geographical coordinates of the vehicle.
[0111] Step 5: Vehicle twin and scene update. According to the vehicle ID, vehicle type, vehicle color, and real geographical coordinates of the vehicle calculated by detection, twin the corresponding vehicle in real time and update the road twin scene.
[0112] In Step 5, the specific steps are as follows:
[0113] Step 51: Construct twin models representing different types of vehicles according to the preset vehicle types in the object detection dataset, and put them into the model pool for unified management.
[0114] Step 52: Use WEBGL to construct the road twin scene, which includes a road model and a camera model.
[0115] Step 53: Use the vehicle category information to retrieve the corresponding vehicle model from the model pool.
[0116] Step 54: Use the vehicle color information to adjust the texture of the vehicle model to make its color consistent with that in the real scene.
[0117] Step 55: Combine the historical positions of the vehicle to calculate the vehicle orientation and apply it to the vehicle model.
[0118] Step 56: Use the real geographical coordinate information of the vehicle to add the vehicle to the appropriate position in the twin scene. If a vehicle with the same tracking ID already exists in the twin scene, use interpolation animation to smoothly move the vehicle from the old position to the new position.
[0119] Embodiment 2
[0120] The purpose of this embodiment is to provide a real-time twin system for road vehicles based on a monocular camera, including:
[0121] A target detection module, which is configured to: use a target detection network to detect vehicles in the video stream, and obtain the rectangular bounding box and contour of the detected vehicle; wherein, the video stream is captured by a monocular camera;
[0122] A target tracking module, which is configured to: identify the color and category of the vehicle in the video stream based on a visual question answering large model, update the target tracker based on the rectangular bounding box of the detected vehicle obtained by target detection, and use the updated tracking ID as the detected vehicle ID;
[0123] A depth estimation module, which is configured to: use a depth estimation model to calculate the depth map of the vehicle in the video stream, and combine the rectangular bounding box of the detected vehicle to obtain the depth information of the detected vehicle;
[0124] A coordinate conversion module, which is configured to: obtain the real geographical coordinates of the detected vehicle according to the depth information of the detected vehicle and the geographical coordinates of the monocular camera;
[0125] A real-time twin module, which is configured to: update the detected vehicle ID, color, category, and the real geographical coordinates of the detected vehicle obtained in real time to the road twin scene.
[0126] In more embodiments, there is also provided:
[0127] An electronic device, including a memory and a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.
[0128] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0129] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0130] A computer-readable storage medium for storing computer instructions, which when executed by a processor, implement the method described in the first embodiment.
[0131] The method in the first embodiment can be directly embodied as being executed by a hardware processor, or implemented by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0132] A computer program product including a computer program, which when executed by a processor, implements the method described in the first embodiment.
[0133] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to execute the process / method as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided as needed. The machine-executable instructions for program modules can be executed locally or within a distributed device. In a distributed device, program modules can be located in local and remote storage media.
[0134] The computer program code for implementing the method of the present invention can be written in one or more programming languages. These computer program codes can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program code is executed by the computer or other programmable data processing devices, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as an independent software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.
[0135] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that a device, apparatus or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals can include electrical, optical, radio, sound, or other forms of propagated signals, such as carrier waves, infrared signals, etc.
[0136] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0137] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A real-time road vehicle digital twin method based on a monocular camera, characterized in that, Including: Using a target detection network to detect vehicles in the video stream and obtain rectangular bounding boxes of the detected vehicles; wherein, the video stream is captured by a monocular camera; Based on a visual question answering large model, identifying the colors and categories of vehicles in the video stream, updating the target tracker based on the rectangular bounding boxes of the detected vehicles obtained by target detection, and using the updated tracking ID as the detected vehicle ID; Using a depth estimation model to calculate the depth map of vehicles in the video stream, and combining with the rectangular bounding boxes of the detected vehicles to obtain the depth information of the detected vehicles; According to the depth information of the detected vehicles and the geographical coordinates of the monocular camera, obtaining the real geographical coordinates of the detected vehicles; Updating the real-time obtained detected vehicle IDs, colors, categories, and the real geographical coordinates of the detected vehicles to the road twin scenario.
2. The real-time digital twin method for road vehicles based on a monocular camera according to claim 1, wherein Using an improved YOLOv11 network as the target detection network. The improvement of the YOLOv11 network is specifically as follows: Using the FC3k2 module to replace the C3k2 module in the YOLOv11 network; wherein, the FC3k2 module uses the FasterNet Block module with partial convolution introduced to perform convolution calculations on some channels of the input feature map; Using the OAAHead module to replace the Head module in the YOLOv11 network; wherein, the OAAHead module introduces depthwise separable convolution and fully connected transformation to extract features of occluded regions from multiple dimensions.
3. A real-time road vehicle digital twin method based on a monocular camera according to claim 1, characterized in that, Using a depth estimation model to calculate the depth map of vehicles in the video stream, and combining with the rectangular bounding boxes of the detected vehicles to obtain the depth information of the detected vehicles, specifically: Using a depth estimation model to calculate the depth map of vehicles in the video stream; Calculating the pixel coordinates of the center point of the bottom edge of the rectangular bounding box of the detected vehicle as the pixel coordinates of the detected vehicle; Combining the pixel coordinates of the detected vehicle with the depth map of the detected vehicle to obtain the depth information of the detected vehicle.
4. A real-time digital twin method for road vehicles based on a monocular camera according to claim 1, characterized in that, Based on a visual question answering large model, identifying the colors and categories of vehicles in the video stream, specifically: Designing question templates and candidate feature sets; wherein, the question templates include placeholders for vehicle color or vehicle category feature description parts; the candidate feature sets are pre-defined candidates for vehicle color or vehicle category appearance features; Filling each feature in the candidate feature set into the placeholder of the designed question template one by one to form a candidate question set; Inputting each candidate question and the image of the vehicle in the video stream into the visual question answering large model to calculate the confidence score of the target vehicle for each question; Selecting the feature corresponding to the highest score from the confidence scores of the candidate features as the final recognition result.
5. The real-time digital twin method for road vehicles based on a monocular camera according to claim 2, wherein In the training of the target detection network, using the DAS function as the loss function. The DAS loss function is specifically: Among them, λ is a hyperparameter that controls the behavior of the non-monotonic attention mechanism; I represents the intersection between the predicted box and the ground truth box; U represents the union between the predicted box and the ground truth box; dw1 represents the horizontal distance between the left edge of the predicted box and the left edge of the target box; dw2 represents the horizontal distance between the right edge of the predicted box and the right edge of the target box; dh1 represents the vertical distance between the upper edge of the predicted box and the upper edge of the target box; dh2 represents the vertical distance between the lower edge of the anchor box and the lower edge of the target box; w gt and h gt represent the width and height of the target box respectively; w pred and h pred represent the width and height of the anchor box respectively.
6. The real-time digital twin method for road vehicles based on a monocular camera according to claim 1, wherein According to the depth information of the detected vehicles and the geographical coordinates of the monocular camera, obtaining the real geographical coordinates of the detected vehicles, specifically: Using a monocular camera to take multiple checkerboard pictures; wherein, the pose of the checkerboard and the camera angle are different each time; Setting the checkerboard plane as the world coordinate system and using the camera calibration function to obtain the checkerboard corner points; Based on the corner data of multiple calibrated images, solve the internal parameter matrix and distortion coefficients of the camera, and optimize the internal parameter matrix of the camera according to the distortion coefficients; Based on the optimized internal parameter matrix of the camera, combine the vehicle tracking ID, vehicle pixel coordinates, and vehicle depth to obtain the true geographical coordinates of the vehicle.
7. A real-time digital twin system for road vehicles based on a monocular camera, characterized in that, It includes: A target detection module, which is configured to: use a target detection network to detect vehicles in the video stream and obtain the rectangular bounding boxes and contours of the detected vehicles; wherein, the video stream is captured by a monocular camera; A target tracking module, which is configured to: identify the colors and categories of vehicles in the video stream based on a large vision question answering model, update the target tracker based on the rectangular bounding boxes of the detected vehicles obtained by target detection, and use the updated tracking ID as the detected vehicle ID; A depth estimation module, which is configured to: use a depth estimation model to calculate the depth map of vehicles in the video stream, and combine the rectangular bounding boxes of the detected vehicles to obtain the depth information of the detected vehicles; A coordinate conversion module, which is configured to: obtain the true geographical coordinates of the detected vehicles according to the depth information of the detected vehicles and the geographical coordinates of the monocular camera; A real-time twin module, which is configured to: update the detected vehicle ID, color, category, and the true geographical coordinates of the detected vehicles obtained in real time to the road twin scenario.
8. An electronic device, characterized in that, It includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method according to any one of claims 1-6 is completed.
9. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the method according to any one of claims 1-6 is completed.
10. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Automatic docking method of automatic driving marshalling vehicle
CN112363510A
Vehicle detection and tracking method based on radar signal and visual fusion
CN112991391A
Weakly supervised vehicle feasible region segmentation method fusing road space prior and region-level features
CN114359873A
X-ray image contraband detection method
CN117058606A
Traffic flow digital twin mapping method based on video trajectory completion
CN118430240A
Cited By
Floating object high-precision identification early warning method and system for unmanned aerial vehicle
CN120853066A
Patrol vehicle path planning and dynamic obstacle avoidance method in combination with virtual-real hybrid positioning
CN121612331A
Patrol vehicle path planning and dynamic obstacle avoidance method combining virtual-real hybrid positioning
CN121612331B