A spatiotemporal linkage engineering geological annotation method and system and storage medium
By installing multi-view cameras and radar on drones or vehicle-mounted platforms, and combining multi-target echo models and viewpoint consistency constraints, the problem of paper-and-pen records in field investigations during the survey and design of low-grade highways has been solved, realizing intelligent engineering geological mapping and target identification, and improving the quality of survey and design and the construction period.
Patent Information
- Application Number
- CN202311658407.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2043-12-05
AI Technical Summary
In the current technology for surveying, designing and maintaining low-grade roads, on-site investigations still rely mainly on paper and pen records, which makes it difficult to guarantee the quality and schedule of the survey and design, and makes it impossible to build an overall route model, thus failing to meet the needs of modern highway surveying.
By installing multi-view cameras and radar on drones or vehicle-mounted platforms, multi-view image data and radar echo signals are acquired. Through multi-target echo model decomposition, target recognition model training, and view consistency constraints, a spatiotemporal linkage engineering geological mapping method and system are generated to achieve intelligent digital mapping of multiple types of targets.
It enables intelligent digital mapping of complex scenarios, generates information-rich 3D environment models, improves the level of intelligence in engineering construction, and supports engineering design analysis and construction safety assessment.
Smart Images

Figure CN117746263B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of engineering surveying and mapping and target identification, and in particular to a spatiotemporal linkage engineering geological mapping method and system and a storage medium. BACKGROUND
[0002] The time limit, quality and cost contradictions of low-grade road survey and design, maintenance, etc. are more prominent than high-grade highways.
[0003] To solve these problems is a very large, super complex system that must be phased, graded, and point by point broken through. In this way, the present scheme mainly solves the problem of collecting and displaying video information of the existing highway reconstruction site. The surveying means of the highway industry has made considerable progress, introducing satellite remote sensing, low-altitude photography, InSAR, and airborne LIDAR and a series of advanced technologies to promote the continuous progress of modern highway surveying technology. However, these advanced technologies and equipment still cannot replace the field investigation work of survey and design personnel, and some work can only be used as a reference at present. Field investigation data is still very important design basis data.
[0004] At present, field investigation is still mainly based on paper and pen recording, and some units have gradually changed to digital field mapping, but it is still only a single-point recording, shooting, video recording, and measurement. For highway and other route engineering, single-point intermittent recording is not conducive to building a whole line model, nor is it conducive to experts and technical personnel who have not been to the site to establish a whole concept. The mapping information recorded in the time line includes different spatial positions, spatial forms, and audio and video displays at different times, which can greatly enhance the sense of real scene of the participating team, reduce the time, energy and cost wasted by repeated site reconnaissance, and thus enable designers to better focus on the points that need to be treated and designed, and better ensure the quality and design time limit of survey and design. Therefore, it is necessary to propose a spatiotemporal linkage engineering geological mapping method and system. SUMMARY
[0005] Therefore, the present application provides a spatiotemporal linkage engineering geological mapping method, which comprises:
[0006] Step 1, on a UAV or vehicle-mounted platform, a multi-view camera and a radar are installed to collect target scenes and obtain multi-view image data and radar echo signals, respectively;
[0007] Step 2, the radar echo signals are decomposed into multiple single-target signals according to a multi-target echo model M and Doppler shift;
[0008] Step 3: Calculate the projection position of each single target signal on the image plane of the multi-view camera; extract the image patch of the corresponding target under each camera in the multi-view camera; calculate the depth information of each image patch and register it with the original point cloud data in the decomposed single target signal to obtain the three-dimensional point cloud of each target;
[0009] Step 4: Define the target recognition task, with each target labeled as its true category; perform fusion recognition based on images of the target from different perspectives taken by multiple cameras to determine the target's category;
[0010] Step 5: Construct a regularized empirical risk loss function, which includes classification loss and perspective consistency constraints;
[0011] Step 6: Train the target recognition model to minimize the loss function until training is complete;
[0012] Step 7: Perform multi-view target recognition on the newly collected target scene data using the trained target recognition model, and aggregate the recognition results together with the images captured by the multi-view camera, the shooting time, and the positioning information at the time of shooting to generate the final mapping file and output it.
[0013] Specifically, the multi-target echo model M in step 2 is as follows:
[0014]
[0015] Among them, A n V represents the magnitude of the nth target. n Let fn represent the radial velocity of the nth target, c represent the speed of light, f0 represent the radar carrier frequency, t represent the radar echo time, PRI represent the pulse repetition interval, and τ represent the pulse repetition interval. n The delay of the nth target is represented by μ; μ represents the frequency modulation. The complex target function of the nth target describes the scattering characteristics of the target. This radar echo signal model represents the combined echo of multiple moving targets. Based on the Doppler frequency shift caused by the motion of the moving targets, the echo signal is decomposed into multiple single target signals.
[0016] Specifically, step 3 includes: for each decomposed single-target echo signal, calculating its projection position on each camera image plane based on its distance and direction angle information; associating each image region corresponding to the single-target echo in each camera image; and selecting image blocks corresponding to the same target on each camera image based on the radar echo parameters.
[0017] In particular, the step 4 specifically comprises: determining the target categories to be identified according to the actual requirements of the target identification task, and jointly detecting the classification accuracy and the positioning accuracy of each target category.
[0018] Obtaining image or video data containing different categories of targets; performing target detection frame labeling and category label labeling on the collected target scene data;
[0019] Extracting features for each image block or extracting multi-scale or multi-modal features; concatenating the image block features extracted under different perspectives into longer feature vectors or learning new representations of multiple perspective features to obtain fusion features; and training a target classification model using the fusion features.
[0020] In particular, the target identification task includes geological bodies, roads, signs, and ecological identification; and the joint detection parameter mAP is calculated according to the following formula for the joint detection of the classification accuracy and the positioning accuracy of each target category:
[0021]
[0022] Wherein, mAP is a joint detection parameter for representing the positioning accuracy and the classification accuracy, wherein l*, r* are the left and right coordinates of the target detection frame, l, r are the left and right coordinates of the ground truth frame, t*, b* are the upper and lower coordinates of the target detection frame, t, b are the upper and lower coordinates of the ground truth frame; represents the positioning accuracy of target detection, and is the proportion of the overlapping area of the target detection frame and the ground truth frame, reflecting the positioning accuracy of the target detection frame to the target position; In the formula, d is the target detection frame, and dgt is the ground truth frame. is the intersection-over-union between a single target detection frame and a ground truth frame, ρ represents the center point coordinates of the ground truth frame, c represents the center point coordinates of the target detection frame, and |ρ-c| 2 represents the square of the distance between the center point of the target detection frame and the center point of the ground truth frame, represents the classification accuracy between a single target detection frame and a ground truth frame, wherein IoU is the overall intersection-over-union index of the detection result and the truth; the classification accuracy of a single target frame is represented by the overall intersection-over-union index IoU of the detection result and the truth, the intersection-over-union between a single target detection frame and a ground truth frame, and the distance between the center points, β represents a classification accuracy adjustment coefficient, which is greater than 0; λ represents a weight factor for adjusting the positioning accuracy and the classification accuracy.
[0023] In particular, in the step 5, the loss function comprises: Wherein L(θ) represents the loss function of the model parameter θ; m represents the number of training samples, h θ (xi represents the predicted output of the model for the i-th sample x i ; y i represents the true label of the i-th sample x i ; alpha represents a regularization parameter for controlling the regularization strength; H(v) is a regularization term, which takes the L1 norm or L2 norm of theta.
[0024] In particular, in step 7, all operations, data acquisition actions, collected data and spatial positions during data collection are connected in series by using a time-space table.
[0025] In particular, in step 7, audio and video, pictures, LAS data and GIS plane positions are synchronously linked by a time-space database, and spatial position changes and corresponding audio and video images and three-dimensional data are obtained through the time line of an event; or the space on the route is selected by a mouse or other pointing devices.
[0026] The application further provides an engineering geology surveying and mapping system based on time-space linkage, which comprises: a multi-view image data and echo signal acquisition module, which is used for installing a multi-view camera and a radar on a UAV or a vehicle platform, collecting a target scene, and respectively acquiring multi-view image data and radar echo signals;
[0027] A multi-target decomposition module is used for decomposing the radar echo signals into a plurality of single-target signals according to a multi-target echo model M and Doppler frequency shift;
[0028] A target image point cloud registration module is used for calculating the projection position of each single-target signal on the image plane of the multi-view camera, extracting the image block of the corresponding target under each camera in the multi-view camera, calculating the depth information of each image block, and registering the original point cloud data in the single-target signal to obtain the three-dimensional point cloud of each target;
[0029] A target category recognition module is used for defining a target recognition task, labeling the real category of each target, and judging the category of the target based on the different view images of the corresponding target under a plurality of cameras;
[0030] A loss function construction module is used for constructing a regularized empirical risk loss function, which comprises a classification loss and a view consistency constraint;
[0031] A target recognition model training module is used for training the target recognition model to minimize the loss function until the training is completed.
[0032] The mapping file generation module is configured to perform multi-view target recognition on newly collected target scene data by using the trained target recognition model, aggregate the recognition result together with images captured by the multi-view camera, capture time and positioning information at the capture time to generate a final mapping file, and output the mapping file.
[0033] The application further provides a computer readable storage medium, wherein a computer program is stored on the computer readable storage medium, and the computer program is executed by a processor to implement the space-time linkage engineering geological mapping method.
[0034] The application achieves close combination of the engineering geological mapping and the target recognition algorithm, adopts a multi-target echo model to decompose a complex scene, and constructs a view consistency regularization loss function to train a robust recognition model.
[0035] Specifically, the technical effects of the application are embodied in the following aspects.
[0036] 1) The multi-target echo model is used to separate geological bodies, roads, signs and ecological features in a scene, and to perform individual recognition and state analysis.
[0037] 2) The view consistency constraint loss function is constructed by using multi-view images, and a target recognition model with strong adaptability to the same target can be learned.
[0038] 3) The fusion of the mapping and the recognition technology generates an information-rich digital three-dimensional environment model, which has strong scene analysis and condition perception capabilities.
[0039] 4) Based on the environment model, various intelligent decision supports such as engineering design analysis, construction safety evaluation and working condition monitoring can be performed, which greatly improves the intelligent level of engineering construction.
[0040] 5) Compared with the traditional mapping of a single modality, the application makes the perception of a complex scene more vivid and intuitive, and provides more rich and effective information support for decision-making.
[0041] 6) The application creates a new idea of deep combination of engineering geological mapping and computer vision recognition algorithm, and has an important demonstration and leading role in the field of engineering construction. BRIEF DESCRIPTION OF DRAWINGS
[0042] Figure 1 FIG. 1 is a flow chart of the space-time linkage engineering geological mapping method in the application;
[0043] FIG. 2 is a horizontal and vertical view of the distortion-free camera module in the application;
[0044] Figure 3 For the device module in the application, the software system function block diagram;
[0045] Figure 4 For the thought block diagram of generating the final video file in the application;
[0046] Figure 5 For the field schematic diagram of setting the final video file in the application;
[0047] Figure 6 For the schematic diagram of the spatio-temporal linkage engineering geological mapping system in the application. DETAILED DESCRIPTION
[0048] The application will be described in detail below with reference to the accompanying drawings and examples.
[0049] Some terms in the target recognition field are explained as follows:
[0050] Target detection box-Object Detection Box;
[0051] Ground-truth box-Ground-truth Box;
[0052] Detection result-DetectionResult (DR);
[0053] Ground truth-GroundTruth (GT);
[0054] IoU is the ratio of the intersection and the union between the predicted box and the ground-truth box, and the IoU value belongs to [0, 1]. The closer to 1, the closer to the true value, and the better the prediction effect.
[0055] The application provides a spatio-temporal linkage engineering geological mapping method, as shown in the figure, the method specifically comprises the following steps: Figure 1
[0056] Step 1, on the unmanned aerial vehicle or vehicle platform, install multi-view camera and radar, collect target scene, obtain multi-view image data and echo signal; wherein wide-angle, distortion-free multi-view camera records space position while recording video, the multi-view camera is a 180° distortion-free camera module mainly used for recording left, right and front three direction distortion-free images, the horizontal field of view angle is not less than ± 90° (the distortion-free image angle obtained by each camera is 60°), the vertical field of view angle is not less than ± 60°, such as Figure 2a and Figure 2b The radar can be installed on the bottom of the rack to avoid interference from the propeller. Vehicle-mounted: generally installed on the roof or front of the vehicle to ensure the range of the field of view. Fixed mode: requires a stable support to prevent vibration from affecting accuracy during work. Field of view requirements: consider the platform motion pattern to ensure that the required range of scenes can be scanned. The size and weight can be selected according to the load capacity of the platform. A small radar can provide a high-speed data port to transmit radar data in real time to the computing unit. At the same time, effective heat dissipation is ensured to ensure the working environment of the radar, and shielding measures are taken to avoid mutual interference.
[0057] Step 2, the radar echo signal is decomposed into a plurality of single-target signals according to a multi-target echo model M and Doppler shift; the multi-target echo model M in step 2 includes:
[0058]
[0059] wherein A n represents the amplitude of the nth target, V n represents the radial velocity of the nth target, c represents the speed of light, f0 represents the carrier frequency of the radar, t represents time, PRI represents the pulse repetition interval, τ n represents the delay of the nth target; μ represents the frequency modulation, represents the complex target function of the nth target, which describes the scattering characteristics of the target;
[0060] The radar echo signal model represents the combined echo of multiple moving targets. According to the Doppler shift caused by the motion, the echo signal is decomposed into a plurality of single-target signals.
[0061] Step 3, calculate the projection position of each single-target signal in the multi-camera image plane; extract the image block of the corresponding target under each camera in the multi-camera; calculate the depth information of the image block and register with the original point cloud data to obtain the three-dimensional point cloud of the target. The step 3 specifically includes: for each decomposed single-target echo signal, according to its distance and direction angle information, calculate its projection position on each camera image plane; associate each image region corresponding to the single-target echo in each camera image; according to the parameters of the radar echo, screen out the image blocks corresponding to the same target on each camera image.
[0062] Step 4, defining the target recognition task, and labeling the true category of each target; based on the different perspective images of the corresponding target under multiple cameras, fusion recognition is performed to judge the category of the target; in the target recognition task in the present application, the target can be a geological body, and geological disaster areas such as collapse, landslide and debris flow can be identified, which belongs to the semantic understanding range of the landform scene; it can also be the identification of the road, the diseases such as pits, cracks and oil on the surface of the road, and the ecological identification, the identification of vegetation and other targets can help to analyze and remove non-ground objects, so as to perform vegetation removal. Therefore, in geological interpretation, multi-target recognition can play a very important auxiliary role to provide more rich environmental information, and this technology should be fully utilized to improve the intelligent level of geological interpretation.
[0063] The step 4 specifically comprises: according to the actual demand of the target recognition task, determining the target categories to be recognized, and jointly detecting the classification accuracy and positioning accuracy of each target category of each target category; obtaining image or video data containing different categories of targets; performing target detection frame labeling and category label labeling on the collected data; determining the corresponding area of the same target on different images according to the extrinsic parameters of each camera to extract multi-view image blocks of the target; extracting features for each image block or extracting multi-scale or multi-modal features; splicing the features under different perspectives into longer feature vectors; or learning new representations of multiple perspective features; and training a target classification model using the fusion features. Set accuracy indicators to evaluate the labeling quality and ensure that it reaches a usable level. The data set is arranged, and the original data, boundary box coordinates and category labels are arranged into formats required for model training, such as XML, JSON, etc. Optionally, the data is enhanced, and the labeled data is enhanced through mirror image, rotation, disturbance and other operations to improve the robustness of the model.
[0064] The joint detection of the classification accuracy and the positioning accuracy of each target category comprises:
[0065] The target recognition task comprises: geological body, road, identification and ecological identification; the joint detection of the classification accuracy and the positioning accuracy of each target category comprises calculating the joint detection parameter mAP according to the following formula:
[0066]
[0067] Wherein, mAP is a joint detection parameter for representing positioning accuracy and classification accuracy, wherein l*, r* are the left and right coordinates of the target detection frame, l, r are the left and right coordinates of the ground truth frame, t*, b* are the upper and lower coordinates of the target detection frame, t, b are the upper and lower coordinates of the target truth frame; Positioning accuracy of target detection, calculated is the proportion of overlapping area of target detection frame and ground truth frame, reflects the positioning accuracy of target detection frame to target position; wherein d is the target detection frame, and dgt is the ground truth frame. is the intersection over union between a single target detection frame and a ground truth frame, and p represents the center point coordinate of the ground truth frame, and c represents the center point coordinate of the target detection frame, and |p-c| 2 represents the square of the distance between the center point of the target detection frame and the center point of the ground truth frame, is the classification accuracy between a single target detection frame and a ground truth frame, wherein IoU is the overall intersection over union index of the detection result and the true value; the classification accuracy of a single target frame is represented by the overall intersection over union index IoU of the detection result and the true value, the intersection over union between a single target detection frame and a ground truth frame, and the distance between the center points, and β represents a classification accuracy adjustment coefficient, which is greater than 0; λ represents a weight factor for adjusting the positioning accuracy and the classification accuracy.
[0068] Step 5, constructing a regularized empirical risk loss function, including a classification loss and a view consistency constraint; the loss function comprises:
[0069]
[0070] wherein L(θ) represents the loss function of the model parameter θ; m represents the number of training samples, h θ (x i ) represents the prediction output of the model for the i-th sample x i ; y i represents the true label of the i-th sample x i ; (h θ (x i ), y i ) represents the overall classification cross-entropy loss, which measures the gap between the predicted category and the true category; α represents a regularization parameter, which is used to control the regularization strength; H(v) is a regularization term, which takes the L1 norm or L2 norm of θ, and represents the view consistency constraint term.
[0071] Optionally, wherein j and k represent two different view indices, x ij represents the i-th target image under the j-th view; x ik represents the i-th target image under the k-th view; h θ (x ij ), h θ (x ik ) are respectively the prediction output of the model for the i-th image under the j-th view and the prediction output of the model for the i-th image under the k-th view; Ph θ (xij )-h θ (x ik )P represents the L2 norm of the prediction output difference of the jth and kth view pair for the same target image, that is, the square sum of the elements of the difference vector between the vectors is squared and then taken the square root; H(v) takes the average of the sum of the prediction differences between all view pairs.
[0072] Step 6, training the target recognition model, minimizing the loss function until the training is completed; in this embodiment, a model structure such as a convolutional neural network is designed, such as VGG, ResNet, etc., and a model prediction function h θ (x) is provided.
[0073] Prepare the training data, which includes a labeled multi-view target image data set.
[0074] Define the optimization goal: use the constructed loss function as the optimization goal, for example, the cross-entropy loss + view consistency constraint term used in this embodiment. Select the optimizer: such as SGD, Adam, etc., determine the learning rate and parameters. Determine the training parameters: the number of iterations, the batch size, and other hyperparameters. Model training: input the training data and train iteratively in batches, update the parameters θ through backpropagation to minimize the defined loss function, and evaluate the model training effect on the validation set to ensure that the accuracy is improved and the view consistency constraint is effectively learned. After training, save the model to obtain a target recognition model that fits well on the training data and meets the set constraint conditions.
[0075] Step 7, multi-view target recognition of newly collected target scene data is performed through the trained target recognition model, the recognition result is aggregated with the image captured by the multi-view camera, the shooting time, and the positioning information at the time of shooting to generate a final mapping file, and is output. In step 7, all operations, data collection actions, collected data, and spatial positions at the time of collecting data are linked in series through a time-space table. Through a time-space database, audio and video, pictures, LAS data, and GIS plane positions are established in synchronization, and through the timeline of events, the spatial position changes and the corresponding audio and video images and three-dimensional data are obtained; or the spatial positions on the route are selected through a mouse or other pointing device.
[0076] In this embodiment, the device module and the software system functional block diagram are as shown in Figure 3 The spatial position recording device is used: GNSS module positioning is used to output CGCS2000 or WCS84 coordinates, which are fused and recorded into the video file.
[0077] Time position recording: use the time module to output the time in East Eight Zone, which is fused and recorded into the video file.
[0078] The data collected above is fused to form a video+space+time video database file, which can adopt a video file+spatial and temporal data file mode, because the video file format cannot include the position information of each frame.
[0079] A PC terminal program-GimsFMV is developed for revisiting the collected video, and the multimedia player function includes:
[0080] (1) Audio and video space-time line interaction function
[0081] (2) Three-dimensional data space-time line interaction function
[0082] (3) Fusion GIS function window, display satellite map base map and route kml and play space point position, and can be interactively operated (that is, video playing, the progress point can be positioned in the time line, and the spatial position can be displayed on the GIS platform at the same time, or the spatial position on the GIS platform can be pulled to control the video to be positioned and displayed to the corresponding point)
[0083] (4) Fusion CAD function window, which can display the position of video display on the dwg drawing of CAD simultaneously, which is only operated on the CAD platform compared with the GIS platform function;
[0084] (5) The image in the left, middle and right frames should be able to be zoomed in and out, and the image details can be viewed. The key point of development is to use time and space table to connect all the operations, data collection actions, collected data and spatial positions at the time of collecting data. It can be queried forward and backward, and the time and space table includes ID, XM_ID, project name, time, space, audio file path, video file path, picture file path, LAS file path, model file path, etc. The thought diagram for generating the final video file is as shown in Figure 4 , and the set fields are as shown in Figure 5 .
[0085] The application also proposes a space-time linkage engineering geological surveying and mapping system, as shown in Figure 6 , which specifically includes the following modules:
[0086] The multi-view image data and echo signal collecting module is used for installing a multi-view camera and a radar on a UAV or a vehicle-mounted platform to collect a target scene and obtain multi-view image data and echo signals; wherein a wide-angle and distortion-free multi-view camera is used to record spatial positions while recording videos, and the multi-view camera is a 180° distortion-free camera module mainly used for recording distortion-free images in left, right and front directions, the horizontal field of view angle is not less than ±90° (the distortion-free image angle obtained by each camera is 60°), and the vertical field of view angle is not less than ±60°, as shown in Figure 2a and Figure 2bThe camera can rotate 360° to shoot. The radar for the unmanned aerial vehicle: needs to be installed under the rack to avoid interference from the propeller. Vehicle-mounted: generally installed on the roof or front of the vehicle to ensure the range of the field of view. Fixed mode: needs a stable support to prevent the influence of vibration on accuracy during work. Field of view requirement: consider the platform motion form to ensure that the required range of scenes can be scanned. The size and weight can be selected according to the load capacity of the platform. The radar can provide a high-speed data port to transmit radar data to the computing unit in real time. At the same time, effective heat dissipation is ensured to guarantee the working environment of the radar, and shielding measures are taken to avoid mutual interference.
[0087] A multi-target decomposition module is configured to decompose a radar echo signal into a plurality of single-target signals according to a multi-target echo model M, the multi-target echo model M comprising:
[0088]
[0089] wherein A n represents an amplitude of the nth target, V n represents a radial velocity of the nth target, c represents a speed of light, f0 represents a carrier frequency of the radar, t represents time, PRI represents a pulse repetition interval, τ n represents a delay of the nth target; μ represents a frequency modulation, represents a complex target function of the nth target, and describes a scattering characteristic of the target.
[0090] The radar echo signal model represents a comprehensive echo of a plurality of moving targets, and the echo signal is decomposed into a plurality of single-target signals according to a Doppler frequency shift caused by movement.
[0091] A target image point cloud registration module is configured to calculate a projection position of each single-target signal on a multi-view camera image plane, extract an image block of a corresponding target under each camera in the multi-view camera, calculate depth information of the image block, and register the depth information with original point cloud data to obtain a three-dimensional point cloud of the target. The target image point cloud registration module specifically comprises:
[0092] For each decomposed single-target echo signal, a projection position of the single-target echo signal on each camera image plane is calculated according to distance and direction angle information of the single-target echo signal. An image region corresponding to the single-target echo in each camera image is associated. Image blocks of the same target are screened out on each image according to parameters of the echo. Depth information of the image blocks is calculated, and the image blocks are reconstructed into a point cloud and registered with original point cloud data. A depth value is assigned to each video pixel to construct a three-dimensional point cloud representation of a video image, and the three-dimensional point cloud representation is merged with original radar point cloud data.
[0093] The target category recognition module is used to define a target recognition task, and a label is used to represent the real category of each target; different view images of the corresponding target under multiple cameras are fused and recognized to determine the category of the target; in the present application, for the target recognition task, the target can be a geological body, and geological disaster areas such as collapse, landslide and debris flow can be identified, which belongs to the semantic understanding range of the landform scene; it can also be the recognition of a road, and the diseases such as pits, cracks and oil on the surface of the road, and the ecological recognition, and the target such as vegetation can be identified to help analyze and remove non-surface objects, so as to remove the vegetation. Therefore, in the geological interpretation, the multi-target recognition can play a very important auxiliary role, and more rich environmental information can be provided, so the technology should be fully utilized to improve the intelligent level of the geological interpretation. The target category recognition module specifically comprises: according to the actual needs of the target recognition task, the target categories to be recognized are determined, and the classification accuracy and positioning accuracy of each target category of each target category are jointly detected; image or video data containing different categories of targets are acquired; target detection frame labeling and category label labeling are performed on the collected data; the multi-view image blocks of the target are extracted according to the corresponding regions of the same target in different images according to the external parameters of each camera; features are extracted for each image block or multi-scale or multi-modal features are extracted; the features under different views are spliced into longer feature vectors; or a new representation of multiple view features is learned; and a target classification model is trained using the fused features. The accuracy and other indicators are set to evaluate the labeling quality to ensure that the level reaches the usable level. The data set is arranged, and the original data, the boundary box coordinates and the category labels are arranged into the format required by the model training, such as XML, JSON and the like. Optionally, the data is enhanced, and the labeled data is enhanced through mirror image, rotation, disturbance and other operations to improve the robustness of the model.
[0094] The target recognition task includes: geological body, road, identification and ecological recognition; the joint detection parameter mAP is calculated according to the following formula for the classification accuracy and positioning accuracy of each target category:
[0095]
[0096] Wherein, mAP is a joint detection parameter used to represent the positioning accuracy and the classification accuracy, wherein l*, r* are the left and right coordinates of the target detection frame, l, r are the left and right coordinates of the ground truth frame, t*, b* are the upper and lower coordinates of the target detection frame, t, b are the upper and lower coordinates of the target truth frame; The positioning accuracy of the target detection is represented, and the calculation is the overlapping area ratio of the target detection frame and the ground truth frame, which reflects the positioning accuracy of the target detection frame to the target position; In the formula, d is the target detection frame, and dgt is the ground truth frame; is the intersection over union between a single target detection box and a ground truth box, ρ represents the center point coordinate of the ground truth box, c represents the center point coordinate of the target detection box, and |ρ-c| 2 represents the square of the distance between the center point of the target detection box and the center point of the ground truth box, represents the classification accuracy between a single target detection box and a ground truth box, wherein IoU is the overall intersection over union indicator of the detection result and the true value; the classification accuracy of a single target box is represented by the overall intersection over union indicator IoU of the detection result and the true value, the intersection over union between a single target detection box and a ground truth box, and the distance between the center points; β represents a classification accuracy adjustment coefficient, which is greater than 0; λ represents a weight factor for adjusting the positioning accuracy and the classification accuracy.
[0097] a loss function construction module for constructing a regularized empirical risk loss function, including a classification loss and a view consistency constraint; the loss function includes:
[0098]
[0099] wherein L(θ) represents the loss function of the model parameter θ; m represents the number of training samples, h θ (x i ) represents the prediction output of the model for the i th sample x i ; y i represents the true label of the i th sample x i ; (h θ (x i ), y i ) represents the classification cross-entropy loss as a whole, which measures the gap between the predicted class and the true class; α represents a regularization parameter, which is used to control the regularization strength; H(v) is a regularization term, which takes the L1 norm or L2 norm of θ, and represents the view consistency constraint term.
[0100] Optionally, wherein j and k represent two different view indices, x ij represents the i th target image under the j th view; x ik represents the i th target image under the k th view; h θ (x ij ) and h θ (x ik ) are the prediction output of the model for the i th image under the j th view and the prediction output of the model for the i th image under the k th view, respectively; Ph θ (x ij )-h θ (x ik)P represents the L2 norm of the prediction output difference of the jth and kth view pair of the same target image, that is, the square sum of the elements of the difference vector between the vectors is taken and then the square root is taken; H(v) takes the average of the sum of the prediction differences between all view pairs.
[0101] A target recognition model training module is configured to train the target recognition model, minimize the loss function, and end the training. θ (x).
[0102] The training data includes a labeled multi-view target image data set.
[0103] An optimization target is defined: the constructed loss function is used as the optimization target, for example, the cross-entropy loss + view consistency constraint term used in this embodiment. An optimizer is selected: such as SGD, Adam, etc., and the learning rate and parameters are determined. Training parameters are determined: the number of iterations, batch size, and other hyperparameters. Model training: input training data is trained in batches, and the parameters θ are updated through back propagation to minimize the defined loss function. The model training effect is evaluated on the validation set to ensure that the accuracy is improved and the view consistency constraint is effectively learned. After training, the model is saved to obtain a target recognition model that fits well on the training data and meets the set constraint conditions.
[0104] A mapping file generation module is configured to perform multi-view target recognition on newly collected target scene data through the trained target recognition model, aggregate the recognition results together with the images taken by the multi-view camera, the shooting time, and the positioning information at the time of shooting to generate a final mapping file, and output the mapping file. In the mapping file generation module, all operations, data collection actions, collected data, and spatial positions at the time of collecting data are connected in series using a time-space table. Through a time-space database, audio and video, pictures, LAS data, and GIS plane positions are established in synchronization link, and through the timeline of events, the spatial position changes and the corresponding video and three-dimensional data are obtained; or the spatial positions on the route are selected through a mouse or other pointing device.
[0105] The application further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the time-space linkage engineering geological mapping method. Since the technical features in the method embodiment correspond one by one, they will not be repeated here.
[0106] To sum up, the above is only a preferred embodiment of the application, and is not used to limit the protection scope of the application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the application shall be included in the protection scope of the application.
[0107] It is apparent for a person skilled in the art that the embodiments of the present application are not limited to the details of the above-described exemplary embodiments, but can be implemented in other concrete forms without departing from the spirit or essential characteristics of the embodiments of the present application. Therefore, the embodiments should be considered in all respects as illustrative and not restrictive, the scope of the embodiments of the present application being defined by the appended claims rather than the above description, and it is intended to include all changes falling within the meaning and scope of equivalents of the claims. Any reference signs in the claims should not be construed as limiting the claims to the figures in which the reference signs are used. Further, it is apparent that the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. Multiple units, modules or devices stated in the system, device or terminal claims can also be implemented by one unit, module or device by means of software or hardware. The words first, second, etc. are used to express names and not to indicate any particular order.
[0108] Finally, it should be noted that the above-described embodiments are merely used to illustrate the technical solutions of the present application but not limit the present application, and although the embodiments of the present application are described in detail with reference to the above preferred embodiments, it should be understood by those ordinarily skilled in the art that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. A spatiotemporal linkage engineering geological mapping method, characterized in that, The application relates to a multi-view target recognition method based on a multi-view camera and a radar. Step 1: a multi-view camera and a radar are installed on a UAV or a vehicle platform to collect target scene data and obtain multi-view image data and radar echo signals; Step 2: the radar echo signals are decomposed into multiple single-target signals according to a multi-target echo model M and Doppler shift; Step 3: the projection positions of each single-target signal in the multi-view camera image plane are calculated; the image blocks of the corresponding targets under each camera in the multi-view camera are extracted; The depth information of each image block is calculated and matched with the original point cloud data in the single-target signals to obtain the three-dimensional point cloud of each target; Step 4: a target recognition task is defined, and the real categories of each target are labeled; Based on the different view images of the corresponding targets under multiple cameras, the target categories are judged through fusion recognition; Step 5: a regularized empirical risk loss function is constructed, which contains a classification loss and a view consistency constraint; Step 6: a target recognition model is trained to minimize the loss function until the training is completed; Step 7: the trained target recognition model is used to recognize the new collected target scene data, and the recognition results are aggregated with the images taken by the multi-view camera, the shooting time and the positioning information at the shooting time to generate a final mapping file and are outputted. In step 2, the multi-target echo model M is as follows: where A n denotes the amplitude of the nth target, V n denotes the radial velocity of the nth target, c denotes the speed of light, f0denotes the carrier frequency of the radar, t denotes the radar echo time, PRI denotes the pulse repetition interval, τ n denotes the delay of the nth target; μ denotes the frequency modulation, denotes the complex target function of the nth target, which describes the scattering characteristics of the target; the radar echo signal model represents the integrated echo of multiple moving targets, and according to the Doppler shift caused by the motion of the moving targets, the echo signal is decomposed into multiple single-target signals.
2. The space-time linkage engineering geological annotation method according to claim 1, characterized in that, In step 3, the projection positions of each single-target echo signal in the image planes of the cameras are calculated according to the distance and direction angle information of the single-target echo signal; the image regions corresponding to the single-target echo in each camera image are associated; and the image blocks corresponding to the same target are screened out on the camera images according to the parameters of the radar echo.
3. The space-time linkage engineering geological delineation method according to claim 1, characterized in that, In step 4, the target categories to be recognized are determined according to the actual requirements of the target recognition task, and the classification accuracy and the positioning accuracy of each target category are jointly detected; Image or video data containing different categories of targets are obtained; target detection frame labeling and category label labeling are performed on the collected target scene data; Features are extracted for each image block, or multi-scale or multi-modal features are extracted; the image block features extracted under different views are spliced into longer feature vectors or new representations of multiple view features to obtain fusion features; and a target classification model is trained using the fusion features.
4. The space-time linkage engineering geological delineation method according to claim 3, characterized in that, The target recognition task includes geological bodies, roads, signs and ecological recognition; the classification accuracy and the positioning accuracy of each target category are jointly detected according to the following formula: wherein mAP is a joint detection parameter used to represent positioning accuracy and classification accuracy, wherein l*, r* are left and right coordinates of the target detection frame, l, r are left and right coordinates of the ground truth frame, t*, b* are upper and lower coordinates of the target detection frame, t, b are upper and lower coordinates of the target truth frame; represents the positioning accuracy of target detection, and the calculation is the proportion of the overlapping area of the target detection frame and the ground truth frame, which reflects the positioning accuracy of the target detection frame to the target position; In the formula, d is the target detection frame, and dgt is the ground truth frame. is the intersection-over-union between a single target detection frame and the ground truth frame, ρ represents the center point coordinates of the ground truth frame, c represents the center point coordinates of the target detection frame, and |ρ-c| 2 represents the square of the distance between the center point of the target detection frame and the center point of the ground truth frame, represents the classification accuracy between a single target detection frame and the ground truth frame, wherein IoU is the overall intersection-over-union index of the detection result and the truth; the classification accuracy of a single target frame is represented by the overall intersection-over-union index IoU of the detection result and the truth, the intersection-over-union between a single target detection frame and the ground truth frame, and the distance between the center points, β represents a classification accuracy adjustment coefficient, which is greater than 0; and λ represents a weight factor used to adjust the positioning accuracy and the classification accuracy.
5. The space-time linkage engineering geological annotation method according to claim 1, characterized in that, In step 5, the loss function includes: where L(0) represents the loss function of the model parameter 0; m represents the number of training samples, h θ (x i ) represents the predicted output of the model for the i-th sample x i ; y i represents the true label of the i-th sample x i ; a represents a regularization parameter, used to control the regularization strength; H(v) is a regularization term, which takes the L1 norm or L2 norm of 0.
6. The space-time linkage engineering geological annotation method according to claim 1, characterized in that, In step 7, all operations, data collection actions, collected data and spatial positions at the time of data collection are connected in time and space.
7. The space-time linkage engineering geological delineation method according to claim 6, characterized in that, In step 7, the audio and video, pictures, LAS data and GIS plane positions are synchronously linked through a time-space database, the spatial position changes and corresponding audio and video images and three-dimensional data are obtained through the time line of an event, or the space on a route is selected through a mouse or other pointing devices.
8. A space-time linkage engineering geological annotation system, characterized in that, The system comprises: a multi-view image data and echo signal acquisition module, which is used for installing a multi-view camera and a radar on a UAV or a vehicle platform to acquire a target scene and obtain multi-view image data and radar echo signals respectively; A multi-target decomposition module is used for decomposing the radar echo signals into a plurality of single-target signals according to a multi-target echo model M and Doppler shift; A target image point cloud registration module is used for calculating the projection position of each single-target signal on the image plane of the multi-view camera, extracting the image block of the corresponding target under each camera in the multi-view camera, calculating the depth information of each image block, and registering the original point cloud data in the single-target signal to obtain the three-dimensional point cloud of each target; A target category identification module is used for defining a target identification task, labeling the real category of each target, and performing fusion identification based on the different perspective images of the corresponding target under a plurality of cameras to determine the category of the target; A loss function construction module is used for constructing a regularized empirical risk loss function, which includes a classification loss and a perspective consistency constraint; A target identification model training module is used for training the target identification model to minimize the loss function until the training is completed; A mapping file generation module is used for performing multi-view target identification on newly acquired target scene data through the trained target identification model, aggregating the identification results together with the images taken by the multi-view camera, the shooting time and the positioning information at the shooting time to generate a final mapping file, and outputting the same. The multi-target echo model M of the multi-target decomposition module is as follows: where A n denotes the amplitude of the nth target, V n denotes the radial velocity of the nth target, c denotes the speed of light, f0denotes the carrier frequency of the radar, t denotes the radar echo time, PRI denotes the pulse repetition interval, τ n denotes the delay of the nth target; μ denotes the frequency modulation, denotes the complex target function of the nth target, which describes the scattering characteristics of the target; the radar echo signal model represents the integrated echo of multiple moving targets, and according to the Doppler shift caused by the motion of the moving targets, the echo signal is decomposed into multiple single-target signals.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize the spatio-temporal linkage engineering geological mapping method according to any one of claims 1 to 7.
Citation Information
Patent Citations
UNet-based radar multi-target distance and speed estimation method
CN113608193A
Multichannel ultra wide band based (UWB-based) radar life detector and positioning method thereof
WO2012055148A1