A road surface health information extraction system and extraction method based on large model analysis

By using three drones to collect data in parallel and employing laser positioning technology, the problem of low efficiency in 3D modeling during existing drone inspections has been solved, enabling the generation of high-precision 3D road surface models and health status assessments.

CN120495564BActive Publication Date: 2025-12-23CHENGDU HUAMAI COMM TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510984467.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-12-23
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

Existing drone inspection technology is inefficient in acquiring three-dimensional road surface information and struggles to achieve high-precision three-dimensional modeling, which affects the accuracy and efficiency of road surface health status assessment.

Method used

Three drones were used to collect 3D video information in parallel and synchronously. The precise relative position was obtained by combining wireless positioning and laser precision technology to construct a high-precision 3D road surface model. The model was then fused through image preprocessing and feature point matching algorithms.

Benefits of technology

It enables efficient generation and accurate analysis of 3D road surface models, improving the accuracy and efficiency of road surface health status assessment, and enhancing the accuracy of feature point matching and the spatial positioning precision of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495564B_ABST
    Figure CN120495564B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of road surface analysis, in particular to a road surface health information extraction system and method based on large model analysis. The road surface health information extraction system based on large model analysis comprises a UAV information acquisition device, a UAV control device, a three-dimensional modeling device and a model information collection and analysis device. According to the technical scheme, three UAVs are used to synchronously collect road surface three-dimensional video information in parallel, and a high-precision road surface three-dimensional modeling and dynamic analysis system is constructed by combining real-time updated relative position data. Compared with the traditional mode of reciprocating collection by a single UAV, the scheme realizes stereoscopic coverage and synchronous acquisition of road surface information by using multi-machine cooperative operation, completes rapid fusion modeling of a three-dimensional scene based on accurate relative position information, and makes the generation efficiency of the road surface three-dimensional model increase by more than 3 times and the spatial positioning accuracy reach the centimeter level.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of road surface analysis, in particular to a road surface health information extraction system and method based on large model analysis. BACKGROUND

[0002] The content of this part only provides background information related to the present application, which may not constitute prior art.

[0003] In the process of highway construction, the route is usually planned first, and then the construction is gradually carried out along the route. The construction process generally includes road leveling, tunnel excavation, bridge erection, and then foundation construction and road paving. Because the entire construction period is long, after the initial road leveling is completed, the route needs to be continuously inspected. By obtaining and analyzing the three-dimensional structural information of the road surface, the construction party can timely adjust the subsequent construction progress and plan. If the leveled road surface is buried by landslides due to rain during construction, the construction process and plan need to be reevaluated, and the adjusted cost needs to be calculated. For example, if new buildings appear on the construction route, their legality needs to be determined, and the subsequent construction plan needs to be modified according to the new ground conditions.

[0004] Therefore, periodic inspection needs to be carried out along the planned route during highway construction. The current mainstream inspection scheme uses a drone to execute: the drone flies along the planned route, collects image data along the way, and generates road surface three-dimensional information through image fusion technology, and then evaluates the health status of the road surface. The drone inspection usually uses oblique photogrammetry to obtain data. However, practice shows that the information obtained by a single shooting angle is limited, especially for tall objects on the road surface, which often need to be shot from two to three different angles to obtain complete surface information. The existing technical solution usually makes the drone repeatedly fly back and forth on the route to continuously collect data and gradually improve the three-dimensional scene, and finally build a complete three-dimensional model. This three-dimensional information extraction method is low in efficiency in actual application. In the face of complex road conditions, the drone often cannot obtain accurate three-dimensional information even after multiple flights. This leads to inaccurate road surface three-dimensional modeling, which affects the accuracy and efficiency of road surface health status evaluation. SUMMARY

[0005] Therefore, the purpose of the present application is to provide a road surface health information extraction system and method based on large model analysis to solve the technical problems mentioned in the background.

[0006] The purpose of the present application is achieved by the following technical solutions:

[0007] A road surface health information extraction system based on large model analysis, comprising:

[0008] Unmanned aerial vehicle (UAV) information acquisition device: includes three aerial photography UAVs, which are deployed parallel to the road to be measured and fly synchronously along the road to acquire three-dimensional video information. The three-dimensional video information includes: video data collected synchronously by the three UAVs; relative position information corresponding to each frame of the video data; the relative position information includes at least the real-time relative position between the three aerial photography UAVs.

[0009] Unmanned aerial vehicle (UAV) control device: used to control the flight direction and attitude of each aerial photography UAV in real time and guide it to fly along a predetermined route; 3D modeling device: receives 3D video information in real time, and generates a 3D model of the road surface based on the video data and its corresponding relative position information.

[0010] Model information collection and analysis device: Collects road surface 3D models generated at different time periods, and extracts and generates road health information that characterizes changes in road surface condition by comparing and analyzing the current road surface 3D model with historical road surface 3D models.

[0011] This technical solution utilizes three drones to synchronously and in parallel acquire 3D video information of the road surface. Combined with real-time updated relative position data, it constructs a high-precision 3D road surface modeling and dynamic analysis system. Compared to the traditional mode of reciprocating acquisition by a single drone, this solution achieves three-dimensional coverage and synchronous acquisition of road surface information through multi-drone collaborative operation. Based on accurate relative position information, it completes rapid fusion modeling of the 3D scene, increasing the efficiency of 3D road surface model generation by more than 3 times and achieving centimeter-level spatial positioning accuracy. Through temporal comparative analysis of 3D models at different time periods, it can automatically identify the development trajectory of road surface defects such as cracks, subsidence, and rutting, accurately extract quantitative indicators such as crack width changes and subsidence depth, and improve the accuracy of extracting road surface condition change information to over 95%. This effectively solves the problems of low efficiency and large subjective errors in traditional manual inspection, providing accurate data support with spatiotemporal continuity for road maintenance decisions, and significantly improving the automation level and engineering application value of road surface health status assessment.

[0012] When acquiring road surface information through high-altitude aerial photography, increasing the flight altitude is necessary to ensure drone flight safety and expand the field of view. However, this introduces significant measurement errors. On the one hand, increased flight altitude amplifies subtle angular deviations in the relative positions of drones, severely affecting the accuracy of feature point matching in the images and making accurate registration difficult when constructing 3D road surface models. On the other hand, traditional wireless positioning technology can only acquire distance information between drones and cannot effectively handle altitude errors, making it difficult to accurately determine the relative positional relationships between drones. This significantly reduces the accuracy and reliability of road surface information acquisition. To address this issue, this application provides the following technical solution:

[0013] Aerial photography drones include:

[0014] The unmanned aerial vehicle body is used to provide an aerial flight platform.

[0015] The high-speed camera is fixedly installed on the unmanned aerial vehicle body at a fixed angle and is tilted towards the lower side.

[0016] The wireless positioning module generates initial relative position parameters with the rest aerial vehicles through millimeter wave positioning.

[0017] The laser transmitter is fixedly installed on the unmanned aerial vehicle.

[0018] The laser receiver forms two laser receiving planes on both sides of the unmanned aerial vehicle body.

[0019] The information processor is in signal connection with the wireless positioning module and the laser receiver and is used to generate relative position information.

[0020] In the method, the laser transmitter of the aerial vehicle sends a laser signal to the laser receiver of the adjacent aerial vehicle, and the information processor corrects the relative position information based on the position of the laser irradiation to the laser receiving plane.

[0021] The scheme effectively solves the positioning problem of high-altitude aerial photography by constructing a relative position acquisition system of "wireless positioning preliminary measurement + laser fine correction" for the unmanned aerial vehicle. The wireless positioning module provides a preliminary position reference for the unmanned aerial vehicle to quickly establish a relative position reference. The laser transmitter cooperates with the laser receiving planes on both sides to accurately measure the subtle differences in the relative positions between the unmanned aerial vehicles based on the laser landing points, so as to realize high-precision correction at the centimeter or even millimeter level. Not only is the feature point matching deviation caused by the relative position error of the unmanned aerial vehicle significantly reduced, but also the construction accuracy and reliability of the road surface three-dimensional model are greatly improved, and the efficiency and quality of road detection are effectively improved.

[0022] In the process of collecting road surface information by multiple unmanned aerial vehicles, the error accumulation exists in the independent positioning of each unmanned aerial vehicle, and there is a lack of unified position reference, which leads to the difficulty in establishing accurate position correlation between the unmanned aerial vehicles. If each unmanned aerial vehicle only relies on its own positioning system to obtain position information, not only the positioning deviation caused by factors such as flight height and environmental interference cannot be eliminated, but also the position relationship is inaccurate in image matching, which causes problems such as feature point mismatching and three-dimensional model misalignment, seriously affecting the accuracy and integrity of the road surface three-dimensional modeling, and further reducing the reliability of the road surface health information extraction.

[0023] Further, the aerial vehicles located on both sides obtain the relative position of the aerial vehicle located in the middle, and the aerial vehicle located in the middle obtains its current geographical position.

[0024] The scheme effectively solves the positioning cooperation problem in the cooperative operation of multiple unmanned aerial vehicles by constructing a position sensing mechanism of "center positioning + two-side cooperation". The middle unmanned aerial vehicle obtains an accurate geographical position as a global reference, and the two unmanned aerial vehicles on the sides measure the relative position with the middle unmanned aerial vehicle as a reference, forming a three-in-one position relationship network. This design unifies the positioning errors of the unmanned aerial vehicles in the same coordinate system, improves the relative position accuracy to the decimeter level, greatly reduces the feature point mis-matching rate in image matching, and improves the integrity and accuracy of three-dimensional model stitching.

[0025] When the unmanned aerial vehicle performs a road information collection task at a high altitude, it is easily affected by atmospheric turbulence, air flow disturbance and other factors, resulting in unstable flight attitude, shaking, tilting and other situations. The images collected during this period have serious geometric distortion, and due to the lack of stable lens pointing parameters, the spatial correspondence between images cannot be accurately established. This makes it easy to produce feature point mis-matching during subsequent image matching, and it is difficult to achieve effective fusion during three-dimensional modeling, resulting in problems such as distortion and misplacement of the road three-dimensional model, which seriously reduces the accuracy and reliability of road health information extraction, and cannot meet the needs of high-precision road detection. To solve this problem, the application provides the following technical solutions:

[0026] The unmanned aerial vehicle information collection device further comprises:

[0027] The video information screening module is connected with the aerial unmanned aerial vehicle signal and is used to obtain the video data synchronously collected by each aerial unmanned aerial vehicle, and screen out the image frames when the aerial unmanned aerial vehicle is in a stable flight attitude to generate three-dimensional video information.

[0028] By introducing the video information screening module, the intelligent image quality control mechanism is constructed, and the negative impact of unstable unmanned aerial vehicle flight on data collection is effectively avoided. Based on the real-time monitored unmanned aerial vehicle flight attitude data, the stable flight period is accurately identified, and high-quality image frames without obvious shaking and controllable geometric distortion are automatically screened out, ensuring the quality consistency of the three-dimensional video information. The model stitching accuracy in the three-dimensional modeling process is improved to the sub-pixel level, significantly enhancing the integrity and accuracy of the road three-dimensional model.

[0029] When the unmanned aerial vehicle cooperatively collects road information, it faces double technical bottlenecks: first, the attitude data of the built-in gyroscope in the flight control system cannot be directly read due to permission restrictions, making it impossible to obtain the flight attitude of the unmanned aerial vehicle in real time; second, relying only on the attitude information of a single unmanned aerial vehicle cannot determine whether the collection array composed of three unmanned aerial vehicles is in the predetermined shooting position as a whole, making it difficult to judge the modeling value of the video frames under turbulence and other disturbances due to the lack of spatial position correlation, and causing feature point matching confusion and model stitching misplacement during three-dimensional information fusion, which seriously affects the road detection accuracy. Therefore, the application provides the following technical solutions:

[0030] The video information screening module comprises:

[0031] A video information acquisition unit is configured to acquire video data synchronously collected by each aerial unmanned vehicle, and arrange the video data acquired by the three aerial unmanned vehicles synchronously based on time labels to generate a video frame matrix;

[0032] A video information screening unit is configured to acquire a time period during which laser signals are received by the laser receivers of each aerial unmanned vehicle, and take a time period during which laser signals are received by all the laser receivers as an ideal time period;

[0033] A video information cropping unit is configured to take a portion corresponding to the ideal time period in the video frame matrix as three-dimensional video information.

[0034] The scheme breaks through the problems of flight control data permission limitation and array position determination by constructing a "laser signal-time label" dual screening mechanism. The video information screening module synchronizes the video data of the three unmanned vehicles based on time labels to form a time-sequenced video frame matrix, and accurately identifies the "ideal time period" during which the three unmanned vehicles are all in the predetermined position through the signal interaction of the laser receivers. This mechanism does not need to directly read the internal data of the flight control, but can determine the spatial position consistency of the array as a whole through the laser signal.

[0035] The road surface health information extraction system based on large model analysis further comprises an image preprocessing device configured to preprocess the three-dimensional video information; the preprocessing comprises noise denoising and lens correction.

[0036] In the technical scheme provided in the present application, the image preprocessing device can effectively increase the picture quality and increase the construction accuracy of the three-dimensional model.

[0037] There are two technical pain points in the traditional three-dimensional modeling process: first, the feature point matching is easily affected by factors such as unmanned aerial vehicle flight attitude fluctuation and image perspective change, resulting in pixel point mapping deviation, forming a feature point mismatching set, and further causing three-dimensional model geometric distortion; second, when constructing a model framework based on a single matching algorithm, there is a lack of global adjustment constraint, it is difficult to eliminate the cumulative error between multi-view images, and it cannot meet the high-precision detection demand.

[0038] The three-dimensional modeling device comprises:

[0039] A pixel point matching module is configured to extract feature points in the three-dimensional video information and map the feature points to generate a feature point mapping set; and a three-dimensional model construction module is configured to match the feature points based on a bundle adjustment to construct a model framework;

[0040] A model rendering module is configured to extract a rendering material from the three-dimensional image to render the model framework to generate a road surface three-dimensional model.

[0041] The scheme realizes high-precision reconstruction of the pavement three-dimensional model by constructing a three-level modeling system of "feature point mapping-beam net adjustment-texture rendering". The pixel point matching module combines the three-dimensional video information collected by multiple unmanned aerial vehicles synchronously, uses the relative position reference provided by laser positioning to form a high-density feature point set; the three-dimensional model construction module introduces a beam net adjustment algorithm, and based on the collinearity equation constraint of multi-view images, globally optimizes the ground coordinates of the feature points, so that the spatial positioning accuracy of the model framework reaches centimeter level, and the cumulative error is effectively eliminated; the model rendering module extracts the texture material with consistent spectrum from the original video, and realizes accurate fitting of the texture and the model framework through coordinate mapping, which can intuitively present the geometric morphology and texture characteristics of pavement cracks, rutting and other diseases, provide high-fidelity data basis for the automatic extraction of pavement health information, and significantly improve the engineering application value of three-dimensional modeling.

[0042] When constructing a three-dimensional model, enough feature points need to be extracted from three-dimensional video information, and the feature points are matched to provide enough feature points when constructing a three-dimensional model. However, feature point extraction and feature point matching require a large amount of computing power. If a neural network model is used to extract feature points, the accuracy is not high when the image feature variation in the video information is complex, which will cause a large amount of distortion in the model framework. Based on this, the application provides the following technical solutions:

[0043] The pixel point matching module extracts feature points in a three-dimensional image and maps the feature points by the following steps:

[0044] Step 1: Extract the flight speed of the aerial unmanned aerial vehicle from the three-dimensional video information, divide the three-dimensional video information into several video groups based on the flight speed, and each video group includes video streams shot by three aerial unmanned aerial vehicles at the same time period;

[0045] Step 2: Extract several sample pictures at equal intervals from the video stream, extract feature points from the sample pictures based on the SIFT algorithm, and one-to-one match the feature points based on the cosine distance to generate a sample matching set;

[0046] Step 3: Train the sample matching set to the convolutional network model and adjust the weight parameters in the convolutional network model;

[0047] Step 4: Input the remaining pictures in the video stream into the trained convolutional network model to extract feature points and the mapping relationship between the feature points;

[0048] Step 5: Take the feature points and the mapping relationship between the feature points extracted in steps 2 and 4 as a feature point mapping set.

[0049] In the technical scheme provided in the application, the SIFT algorithm is used for feature extraction on part of the pictures in the video stream, which can greatly ensure the accuracy of feature point extraction and feature point mapping. Then the convolutional network model is trained with the training data, and the convolutional network model can change its internal weight parameters, thereby having good feature recognition and feature comparison capabilities in the video group, and then the feature points and the mapping relationship between the feature points can be quickly extracted. In this way, the convolutional network and the feature extraction algorithm are combined in the application, which not only ensures the efficiency of feature point extraction, but also ensures the accuracy of feature point extraction.

[0050] In the three-dimensional model construction process, the feature point processing faces double technical bottlenecks: on the one hand, although the traditional feature point extraction and matching algorithm (such as SIFT, SURF) has high accuracy, but the calculation complexity is large, and when facing massive video data, the algorithm consumes too much, which is difficult to meet the real-time modeling demand; on the other hand, when simply using a neural network model to extract feature points, due to the complex road texture, light and perspective change law in the video image, the model generalization ability is insufficient, and feature point mismatching is easy to occur, which leads to geometric distortion of the three-dimensional model framework, and cannot meet the high-precision detection requirement.

[0051] Further, the convolutional network model comprises:

[0052] An input layer is configured to input pictures taken by three aerial unmanned aerial vehicles at the same time to generate a first feature map, a second feature map and a third feature map;

[0053] A convolutional layer is configured to perform convolutional processing on the input first feature map, second feature map and third feature map to extract a first hidden feature, a second hidden feature and a third hidden feature;

[0054] A pooling layer is connected with the convolutional layer and is configured to pool the first hidden feature, the second hidden feature and the third hidden feature;

[0055] A probability calculation layer is configured to generate probability information of each pixel point belonging to a feature point from the first hidden feature, the second hidden feature and the third hidden feature;

[0056] A first attention mechanism network is configured to input the first feature map and the probability information to generate a first transformed feature;

[0057] A second attention mechanism network is configured to input the second feature map and the probability information to generate a second transformed feature;

[0058] A third attention mechanism network is configured to input the third feature map and the probability information to generate a third transformed feature;

[0059] A mapping network is configured to map the first transformed feature, the second transformed feature and the third transformed feature to each other to generate feature points and a mapping relationship between the feature points.

[0060] The scheme breaks through the contradiction between the consumption of computing power and the matching accuracy by constructing a feature point processing system of "SIFT accurate labeling + convolution network efficient inference". First, the video is grouped based on the flight speed to ensure that the videos in the same group have similar motion characteristics. A high-precision sample matching set is generated by equidistant sampling combined with the SIFT algorithm to provide high-quality training data for the convolution network, which greatly improves the feature recognition accuracy of the network in complex scenes. The scheme not only ensures the matching accuracy of the basic feature points by using the SIFT algorithm, but also realizes the efficient processing of batch data through the neural network, which reduces the distortion rate of the three-dimensional model framework. It provides a high-density and low-error feature point set for subsequent bundle adjustment, significantly improves the efficiency and reliability of road surface three-dimensional modeling, and meets the real-time monitoring needs of the engineering site.

[0061] In the process of constructing a three-dimensional model framework of the road surface, there are two major technical problems: first, the traditional model rendering method relies on a single data source or initial parameters, which is difficult to accurately process the spatial geometric relationship between multi-view images, resulting in large calculation errors of the ground coordinates of the connection points, and distortion problems such as distortion and misplacement of the model framework; second, if a simple iterative calculation is used, continuous calculation until the required accuracy is achieved will consume a large amount of time and computing resources, while early termination of the calculation cannot guarantee the accuracy of the model, and too much intermediate calculation data will occupy a large amount of storage space, making it difficult to balance the accuracy, efficiency and storage cost.

[0062] The model rendering module generates a model framework based on the following steps:

[0063] Z1: Obtain the feature point mapping set and the relative position information corresponding to the feature point mapping set, and convert the relative position information into exterior orientation elements E. The exterior orientation elements E include the coordinates (X S , Y S , Z S ) of the photographic center S in the ground coordinate system and the rotation angles χ S represents the rotation angle of the image coordinate system around the X-axis, ψ S represents the rotation angle of the image coordinate system around the Y-axis, represents the rotation angle of the image coordinate system around the Z-axis; the feature points in the feature point set are taken as connection points i (x ij , y ij ), and j represents the serial number of the image, j = 1, 2, 3;

[0064] Z2: For each connection point i, based on the exterior orientation elements of the image j containing the connection point i and the pixel point coordinates of the connection point i in the image j, the initial ground coordinate approximation of the connection point i is solved by the collinearity equation.

[0065] Z3: The method according to any one of Z1 to Z2, wherein the exterior orientation elements are determined by a bundle adjustment method. and the exterior orientation elements build a parameter vector X k ;

[0066]

[0067] wherein k represents the iteration number, and k = 0 at the beginning, E1k represents the exterior orientation elements of the first aerial vehicle at the kth iteration, E1k represents the exterior orientation elements of the first aerial vehicle at the kth iteration, E3k represents the exterior orientation elements of the third aerial vehicle at the kth iteration, TP1k represents the initial ground coordinate approximation of the first connection point at the kth iteration, TPnk represents the initial ground coordinate approximation of the nth connection point at the kth iteration, n represents the total number of connection points, T represents the matrix transpose symbol, E1 represents the exterior orientation elements of the first aerial vehicle, E3 represents the exterior orientation elements of the third aerial vehicle; TP1 represents the initial ground coordinate approximation of the first connection point, TPn represents the initial ground coordinate approximation of the nth connection point;

[0068] Z4: For each connection point i (x ij , y ij ), extract the current exterior orientation elements of the image j and the current coordinates of the connection point i TP i k , calculate the predicted image point coordinates (x pr , y pr ); Ejk represents the exterior orientation elements of the jth image at the kth iteration;

[0069]

[0070] wherein f represents the camera focal length, Rjk represents the rotation matrix elements of the jth image, the two digits of the subscript represent the row number and column number of the rotation matrix respectively, x0, y0 represent the pixel point coordinates; TPuk represents the initial ground coordinate approximation of the connection point u at the kth iteration,

[0071] Ejk represents the exterior orientation elements of the jth aerial vehicle at the kth iteration, Δx, Δy represent the distortion correction values;

[0072] Take the observation value as a constant term I c , calculate the observation value residual l x and l y , arrange the residuals in order to generate a constant vector I im ;

[0073]

[0074] Xi represents the coordinate of connection point i measured, x pr , y pr Xi represents the coordinate of connection point i measured, x Xi represents the coordinate of connection point i measured, x x Xi represents the coordinate of connection point i measured, x y Xi represents the coordinate of connection point i measured, x Xi represents the coordinate of connection point i measured, x ct ;

[0075]

[0076] Xi represents the coordinate of connection point i measured, x Xi represents the coordinate of connection point i measured, x Xi represents the coordinate of connection point i measured, x ct ;

[0077] Xi represents the coordinate of connection point i measured, x im Xi represents the coordinate of connection point i measured, x ct Xi represents the coordinate of connection point i measured, x

[0078] Z5: Parameter vector X k Xi represents the coordinate of connection point i measured, x

[0079] Xi represents the coordinate of connection point i measured, x

[0080] Xi represents the coordinate of connection point i measured, x

[0081] Z6: Construct normal equation:

[0082] Xi represents the coordinate of connection point i measured, x T Xi represents the coordinate of connection point i measured, x T Xi represents the coordinate of connection point i measured, x

[0083] Xi represents the coordinate of connection point i measured, x

[0084] Z7: solve the method equation to generate a correction vector ΔX, input the correction vector ΔX into the parameter vector X k , update the new parameter vector X k+1 ;

[0085] Based on the new parameter vector X k+1 generate the geographical coordinates of each connection point, and render the model framework according to the geographical coordinates.

[0086] The scheme effectively breaks through the technical bottleneck of traditional model rendering by constructing an iterative optimization-based model framework generation system. The scheme converts relative position information into exterior orientation elements, combines collinear equations with forward intersection to obtain initial coordinates of connection points, and lays a precise foundation for model construction; through iterative least squares adjustment, the geographical coordinates of the connection points are continuously optimized, and the errors of the exterior orientation elements and the ground coordinates are continuously corrected, so that the spatial positioning accuracy of the model framework is improved to sub-centimeter level, and the model distortion rate is significantly reduced. At the same time, the scheme supports flexible control of rendering time according to actual needs, and terminates the calculation in time after reaching the preset accuracy standard, compared with traditional methods, which not only ensures the accuracy of model rendering, but also realizes the efficient use of computing resources, providing reliable technical support for fast and accurate road three-dimensional model construction.

[0087] A road surface health information extraction method based on large model analysis, which uses the foregoing road surface health information extraction system based on large model analysis to extract road surface health information.

[0088] The beneficial effects of the present application are:

[0089] (1) In the road three-dimensional information collection link, the traditional single unmanned aerial vehicle reciprocating flight inefficient mode is abandoned, and a three unmanned aerial vehicle synchronous collection strategy is innovatively adopted. The core advantage is:

[0090] (2) Synchronous multi-angle collection: three unmanned aerial vehicles not only synchronously collect video data during flight, but also continuously update and record real-time relative position (including attitude) information between each other.

[0091] (3) Precise spatial positioning: based on accurate real-time relative position information, the relative coordinates and viewing angles of each frame of picture in the video in three-dimensional space can be reliably determined.

[0092] (4) Efficient three-dimensional fusion modeling: using multi-view synchronous video and its accurate spatial position relationship, three-dimensional scene fusion and reconstruction can be performed more quickly and accurately, significantly improving the generation speed and modeling accuracy of road three-dimensional models.

[0093] (5) Improve evaluation efficiency: the more accurate and timely road three-dimensional model obtained directly improves the accuracy and efficiency of subsequent road health state evaluation. BRIEF DESCRIPTION OF DRAWINGS

[0094] Figure 1 A structural schematic diagram of a road surface health information extraction system based on large model analysis provided for Embodiment 1 of the present application; Figure 2 A schematic diagram of a UAV for aerial photography.

[0095] Figure 3 A flowchart of extracting feature points in a three-dimensional image in Embodiment 2.

[0096] Figure 4 A structural schematic diagram of a convolutional network model in Embodiment 2.

[0097] REFERENCE NUMERALS

[0098] 1, UAV body; 2, laser transmitter; 3, laser receiver; 4, high-speed camera. DETAILED DESCRIPTION

[0099] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below in conjunction with specific embodiments. The same reference numerals in the drawings represent the same components. It should be noted that the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the described embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without any creative effort fall within the scope of protection of the present application.

[0100] Compared with the embodiments shown in the drawings, the feasible implementation solutions within the scope of protection of the present application can have fewer components, other components not shown in the drawings, different components, differently arranged components or differently connected components, etc. In addition, two or more components in the drawings can be implemented in a single component, or a single component shown in the drawings can be implemented as multiple separate components.

[0101] Unless otherwise defined, the technical terms or scientific terms used herein should be understood as their usual meanings by those of ordinary skill in the art to which the present application belongs. The use of “first”, “second” and similar words in the specification and claims of the present application does not represent any order, quantity or importance, but is only used to distinguish different components. Similarly, “one” or “a” and similar words do not necessarily represent a quantity limitation. “Up”, “down” and the like are only used to represent a relative positional relationship, which may change accordingly when the absolute position of the described object changes.

[0102] REFERENCE Figure 1 , Embodiment 1:

[0103] The first embodiment of the present application discloses a road surface health information extraction system based on large model analysis, which comprises a UAV information collection device, a UAV control device, a three-dimensional modeling device, and a model information collection and analysis device. The UAV information collection device comprises three aerial UAVs, which are deployed in parallel to the road to be measured and fly synchronously along the road to be measured to obtain three-dimensional video information, which includes video data collected synchronously by the three UAVs, and relative position information corresponding to each frame of the video data, wherein the relative position information at least includes the real-time relative position between the three aerial UAVs. The UAV control device is used to control the flight direction and attitude of each aerial UAV in real time to guide it to fly along the predetermined route. The three-dimensional modeling device receives the three-dimensional video information in real time, generates a road surface three-dimensional model based on the video data and the corresponding relative position information, and fuses the model. The model information collection and analysis device collects road surface three-dimensional models generated at different time periods, compares and analyzes the current road surface three-dimensional model with the historical road surface three-dimensional model, extracts and generates road surface health information representing the change of road surface state. The road surface health information extraction system based on large model analysis further comprises an image preprocessing device for preprocessing the three-dimensional video information, which includes noise removal and lens correction. Noise removal and lens correction are prior art and will not be described here.

[0104] Reference Figure 2 The aerial UAV comprises a UAV body, a high-speed camera, a wireless positioning module, a laser emitter, a laser receiver, and an information processor. The UAV body carries the high-speed camera (installed at a fixed downward angle), the laser emitter is fixedly installed below the UAV body, and the laser receiver is fixedly installed on both sides of the UAV body. The laser receiver is a laser receiver with a certain area, which can identify the wavelength of the received optical signal and then find the position of the light spot on the laser receiver.

[0105] Each aerial UAV generates initial relative position parameters through the wireless positioning module using millimeter wave positioning technology to quickly establish a relative position reference. Subsequently, the laser emitter of the UAV sends laser signals to the laser receiving planes on both sides of the adjacent UAV, the laser receiver transmits the position information of the laser irradiation to the receiving plane to the information processor, and the information processor corrects the initial relative position parameters based on the laser landing position to accurately calculate the subtle differences in the relative positions between the UAVs, achieving high-precision positioning at the centimeter or even millimeter level. This process establishes a UAV relative position acquisition system of "wireless positioning preliminary measurement + laser fine correction", effectively solves the problem of relative position error of the UAV caused by the increase of flight height, significantly reduces the feature point matching deviation, and improves the construction precision of the road surface three-dimensional model and the efficiency and quality of road detection.

[0106] Further, the two side aerial unmanned vehicles obtain the relative position of the aerial unmanned vehicle in the middle, and the aerial unmanned vehicle in the middle obtains its current geographical position.

[0107] The unmanned aerial vehicle information acquisition device further comprises a video information screening module connected with the aerial unmanned vehicle signals, configured to obtain the video data synchronously collected by each aerial unmanned vehicle, and screen out the image frames when the aerial unmanned vehicle flight attitude is stable to generate three-dimensional video information.

[0108] Specifically, the video information screening module is connected with each aerial unmanned vehicle through wireless or wired signals to obtain the video data synchronously collected by the high-speed camera in real time (for example, when the aerial unmanned vehicles in formation fly over a certain section of road, the continuous video stream synchronously recorded). The video information screening module first judges whether the flight attitude of the aerial unmanned vehicle is stable through the laser receiver. For example, when the unmanned aerial vehicle shakes due to air flow, the laser receiver cannot receive the laser signal, and the video information screening module automatically filters the video frames in the corresponding period; when the unmanned aerial vehicle is in a stable flight state, the high-definition image frames in this state are screened out (for example, 10-15 stable image frames per second), and three-dimensional video information is generated.

[0109] The video information screening module comprises a video information acquisition unit, a video information screening unit, and a video information cutting unit.

[0110] The video information screening module first synchronously receives the video data collected by the high-speed cameras of the three aerial unmanned vehicles through the video information acquisition unit, arranges the video data in time sequence based on the time label of the video data, and forms a regular video frame matrix. For example, when the three unmanned vehicles are shooting the same section of road, this unit will align the video frames collected by each unmanned vehicle at the same time. Then, the video information screening unit monitors the working state of the laser receiver of each aerial unmanned vehicle in real time, obtains the laser signal receiving period, and determines the period when the laser receivers of the three unmanned vehicles all receive the laser signal as the ideal period, which means that the three unmanned vehicles are in stable flight during this period. Finally, the video information cutting unit accurately cuts out the corresponding part from the video frame matrix according to the ideal period to generate three-dimensional video information for three-dimensional modeling.

[0111] The three-dimensional modeling device comprises a pixel point matching module, a three-dimensional model construction module and a model rendering module. The pixel point matching module first analyzes the three-dimensional video information, extracts feature points in the video frame by using the SIFT algorithm, and then performs coordinate conversion on the feature points to generate a feature point mapping set containing the spatial position information of each feature point. The three-dimensional model construction module matches and optimizes the points in the feature point mapping set based on the light beam network adjustment algorithm, constructs a preliminary three-dimensional model framework by calculating the intersection relationship of light rays under different viewing angles, and ensures the accuracy of the geometric structure of the model. The model rendering module extracts rendering materials such as the rough texture of the road surface and the reflective properties of the road markings from the three-dimensional image containing information such as texture and color, assigns these materials to the model framework, and renders the model by using techniques such as light calculation and texture mapping, thereby generating a realistic and accurate three-dimensional road model and providing intuitive and effective data presentation for road detection and analysis.

[0112] The unmanned aerial vehicle control device is mainly used for controlling the flight direction and flight route of the unmanned aerial vehicle. In practice, if the interval between adjacent ideal time periods is too long, the unmanned aerial vehicle control device controls the unmanned aerial vehicle to fly in the area between adjacent ideal time periods.

[0113] The model information collection and analysis device mainly compares the current generated three-dimensional model with the historical three-dimensional model, and labels the areas where the three-dimensional model has differences. If the number of difference areas is larger and the volume of difference is larger, the corresponding road health degree is worse, and vice versa. In this way, the road health information is the difference information of the adjacent two three-dimensional models.

[0114] Reference Figure 3 Embodiment 2: Embodiment 2 provides a feature point extraction method based on embodiment 1, specifically:

[0115] The pixel point matching module extracts the feature points in the three-dimensional image by the following steps and maps the feature points:

[0116] Step 1: Extract the flight speed of the aerial unmanned aerial vehicle from the three-dimensional video information, and divide the three-dimensional video information into several video groups based on the flight speed, each video group including video streams shot by three aerial unmanned aerial vehicles at the same time period;

[0117] In actual operation, firstly, flight speed data of the aerial unmanned plane is acquired from the three-dimensional video information, for example, the speed change of the aerial unmanned plane in the flight process is recorded by a GPS module or other speed measuring equipment of the aerial unmanned plane. Based on the flight speed, the three-dimensional video information is divided into several video groups. Taking a long highway aerial video as an example, if the unmanned plane flies at a relatively stable speed, the video can be divided into multiple small segments according to the speed and time interval, and each video group contains video streams shot by three aerial unmanned planes in the same time period. The purpose of such division is to centrally process videos with similar shooting conditions and scenes, facilitating subsequent extraction of feature points and matching.

[0118] Step 2: A number of sample pictures are extracted at equal intervals from the video stream, feature points are extracted from the sample pictures based on the SIFT algorithm, and the feature points are matched one by one based on the cosine distance to generate a sample matching set.

[0119] From the video stream of each video group, a number of sample pictures are extracted at equal intervals according to fixed time or frame interval. For example, a picture is extracted as a sample every 10 frames, and these sample pictures can represent the overall picture features of the video group to a certain extent. Then, the SIFT (Scale-Invariant Feature Transform) algorithm is used to process the sample pictures. This algorithm can detect stable feature points in different scale spaces and generate unique descriptors for each feature point. Then, the feature points in different sample pictures are matched one by one based on the cosine distance. The cosine distance can measure the similarity between feature point descriptors, and feature points with high similarity are paired to generate a sample matching set. In this way, the correspondence between features in different pictures is preliminarily determined.

[0120] For example, three aerial unmanned planes shoot pictures A, B and C respectively, and feature points in pictures A, B and C are extracted and matched for similarity. If the similarity exceeds a threshold value, it means that these feature points are the same feature points, and the corresponding mapping relationship is established.

[0121] Step 3: The sample matching set is convolved with the convolutional neural network model to adjust the weight parameters in the convolutional neural network model.

[0122] The generated sample matching set is input into the convolutional network model as training data. The convolutional network model is a powerful deep learning model that can automatically learn feature patterns in images. During the training process, the model adjusts its weight parameters based on the feature points and their matching relationships in the sample matching set. For example, through the backpropagation algorithm, the error between the predicted results and the actual matching results is propagated in reverse, updating the weights in the model, so that the model can better capture the distribution and matching rules of feature points in images. After multiple iterations of training, the convolutional network model is gradually optimized, improving the accuracy of feature point extraction and matching in images.

[0123] Reference Figure 4 Specifically, the convolutional network model includes an input layer, a convolutional layer, a pooling layer, a probability calculation layer, a first attention mechanism network, a second attention mechanism network, and a mapping network.

[0124] The input layer is used to input the pictures taken by the three aerial drones at the same time to generate first, second, and third feature maps. The input layer is the entrance of the convolutional network model, and its main function is to receive the pictures taken by the three aerial drones at the same time.

[0125] The convolutional layer performs convolutional processing on the input first, second, and third feature maps to extract first, second, and third hidden features.

[0126] As a core component of the model, the convolutional layer uses multiple convolutional kernels of different sizes and quantities (such as 3x3 and 5x5 convolutional kernels) to perform convolutional processing on the input first, second, and third feature maps. Taking a 3x3 convolutional kernel as an example, by sliding the convolutional kernel over the feature map, element multiplication and accumulation operations are performed to extract local features in the picture. During the convolution process, different strides (such as stride 1 or 2) can be set to control the interval of convolutional kernel sliding to obtain features of different scales.

[0127] The pooling layer is connected to the convolutional layer and is used to pool the first, second, and third hidden features. The pooling layer is connected to the convolutional layer and mainly performs average pooling on the first, second, and third hidden features to reduce data volume and computational complexity while preserving key features.

[0128] The probability calculation layer generates probability information of each pixel point belonging to a feature point from the first, second, and third hidden features. The probability calculation layer is based on the pooled first, second, and third hidden features, and uses a fully connected layer and an activation function (such as the Softmax function) to calculate the probability information of each pixel point belonging to a feature point.

[0129] The first attention mechanism network inputs the first feature map and the probability information to generate a first transformed feature;

[0130] The second attention mechanism network inputs the second feature map and the probability information to generate a second transformed feature;

[0131] The third attention mechanism network inputs the third feature map and the probability information to generate a third transformed feature;

[0132] The first attention mechanism network, the second attention mechanism network, and the third attention mechanism network have the same structure. The attention mechanism network first fuses the input feature map and the probability information through linear transformation and an activation function to generate an attention weight matrix. This weight matrix reflects the importance of different positions in the image, and the higher the weight value, the more critical the features contained in the region. Then, the attention weight matrix is weighted with the first feature map to obtain the transformed feature.

[0133] The mapping network maps the first transformed feature, the second transformed feature, and the third transformed feature to generate feature points and mapping relationships between the feature points. The mapping network takes the first transformed feature, the second transformed feature, and the third transformed feature as input and performs nonlinear transformation through, for example, a multilayer perceptron. In this process, the network learns the spatial relationships and similarities between features under different perspectives, thereby mapping these transformed features to each other to generate feature points and mapping relationships between the feature points.

[0134] Step 4: Input the remaining pictures in the video stream into the trained convolutional network model to extract feature points and mapping relationships between the feature points;

[0135] The remaining pictures in the video stream, excluding the sample pictures, are input into the trained and optimized convolutional network model. At this time, the trained model can automatically extract feature points in the pictures and determine the mapping relationships between the feature points using the learned feature patterns.

[0136] Step 5: Take the feature points and mapping relationships between the feature points extracted in steps 2 and 4 as a feature point mapping set.

[0137] Integrate the sample picture feature points and their mapping relationships obtained by the SIFT algorithm and cosine distance matching in step 2 with the feature points and mapping relationships of the remaining pictures extracted by the convolutional network model in step 4. These feature points and mapping relationships are summarized together to form a complete feature point mapping set. This set contains feature points of each picture in the three-dimensional video information and their corresponding relationships with each other, providing an important data foundation for subsequent three-dimensional model construction, enabling the model to accurately restore the three-dimensional structure and features of the road surface.

[0138] Embodiment 3: Embodiment 3 provides a generation scheme of a model framework on the basis of Embodiment 1;

[0139] The model rendering module generates the model framework based on the following steps:

[0140] Z1: Obtain the feature point mapping set and the relative position information corresponding to the feature point mapping set, and convert the relative position information into the exterior orientation element E, the exterior orientation element E including the coordinates (X S , Y S , Z S ) of the photographic center S in the ground coordinate system and the rotation angles χ S representing the rotation angle of the image coordinate system around the X axis, ψ S representing the rotation angle of the image coordinate system around the Y axis, representing the rotation angle of the image coordinate system around the Z axis; the feature points in the feature point set are taken as the connection points i (x ij , y ij ), and j represents the serial number of the image, j = 1, 2, 3.

[0141] The photographic center S is simplified as the coordinates of the UAV body, or the coordinates of the UAV body are replaced by the photographic center. Although the UAV body is in stable flight, the high-speed camera is tilted downward to form an inclined measurement. The serial number of the image corresponds to the serial number of the UAV body.

[0142] Z2: For each connection point i, based on the exterior orientation element of the image j containing the connection point i and the pixel point coordinates of the connection point i in the image j, the initial ground coordinate approximation TP

[0143]

[0144] Z3: According to and the exterior orientation element, construct the parameter vector X k ;

[0145]

[0146] wherein k represents the iteration number, k = 0 at the beginning, represents the exterior orientation element of the first aerial UAV at the kth iteration, represents the exterior orientation element of the third aerial UAV at the kth iteration, represents the initial ground coordinate approximation of the connection point 1 at the kth iteration, represents the initial ground coordinate approximation of the connection point n at the kth iteration, n represents the total number of connection points, T represents the matrix transpose symbol, 1E represents the exterior orientation element E of the first aerial unmanned vehicle, 3E represents the exterior orientation element E of the third aerial unmanned vehicle; TP1 represents the initial ground coordinate approximation of the first connection point, TPn represents the initial ground coordinate approximation of the nth connection point;

[0147] Z4: for each connection point i (x ij , y ij ), extract the current exterior orientation element of the image j and the current coordinate TP i of the connection point i; k , calculate the predicted image point coordinate (x pr , y pr ); represents the exterior orientation element of the jth image at the kth iteration;

[0148]

[0149] wherein f represents the camera focal length, represents the rotation matrix element of the jth image, the two digits of the subscript represent the row number and column number of the rotation matrix respectively, x0, y0 represents the pixel point coordinate; represents the initial ground coordinate approximation of the connection point u at the kth iteration,

[0150] represents the exterior orientation element of the jth aerial unmanned vehicle at the kth iteration, Δx, Δy represents the distortion correction value;

[0151] , the observation value is taken as a constant term I c , the observation value residual l x and l y are calculated, the residuals are arranged in order to generate the connection point pixel point coordinate residual vector I im ;

[0152]

[0153] represents the coordinate measured and obtained by the connection point i, x pr , y pr represents the theoretical connection point coordinate value calculated by the collinearity equation, represents the corrected pixel point coordinate of the connection point i on the image j, l x represents the observation value residual component of the connection point i in the horizontal direction of the image j, l y represents the observation value residual component of the connection point i in the vertical direction of the image j; from the feature point set, the control point g is extracted, for each ground control point g, according to the known observation value Generate control point ground coordinate residual vector I ct ;

[0154]

[0155] Indicates the current coordinate approximation of the ground control point g, Indicates the observation residual of the ground control point g, Indicates the coordinate obtained by measuring the ground control point g, and the residual is arranged in order to generate a constant vector, forming a constant term vector I of the control point ct ;

[0156] Connect the pixel point coordinate residual vector I im and the control point ground coordinate residual vector I ct Vertically splice;

[0157] Z5: Parameter vector X k Carry out Taylor series expansion to obtain the linearized error equation of the observation value:

[0158] v=A△X-I;

[0159] Wherein, v represents the residual vector, which contains the residual of the coordinate observation value of all connection points and control points, △X represents the correction number vector, which contains the correction number of all exterior orientation elements E, and A represents the design matrix; Each element of the design matrix A is the predicted value of the collinear equation;

[0160] Z6: Construct the normal equation:

[0161] (A T PA)△X=A T PI;

[0162] Wherein, I represents the complete constant vector, A represents the design matrix, T represents the matrix transpose symbol, and P represents the weight matrix of the observation value;

[0163] Z7: Solve the normal equation to generate the correction number vector △X, and input the correction number vector △X into the parameter vector X k , update the new parameter vector X k+1 ;

[0164] Based on the new parameter vector X k+1 Generate the geographical position coordinates of each connection point, and render the model framework according to the geographical position coordinates. The control points in the scheme are points with specific positions known in advance, such as a ruler with known position and size placed on the road surface.

[0165] Embodiment 4: A road surface health information extraction method based on large model analysis, which extracts road surface health information by using the foregoing road surface health information extraction system based on large model analysis.

[0166] The above merely provides preferred embodiments of the present application and is not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall fall into the scope of protection of the present application.

Claims

1. A system for extracting road surface health information based on large model analysis, characterized by, The system comprises: An unmanned aerial vehicle information acquisition device: including three aerial unmanned aerial vehicles, which are deployed in parallel to the road to be measured and fly synchronously along the road to be measured to obtain three-dimensional video information, which includes video data synchronously collected by the three unmanned aerial vehicles, relative position information corresponding to each frame of the video data, and the relative position information at least including real-time relative positions among the three aerial unmanned aerial vehicles; An unmanned aerial vehicle control device: for controlling the flight direction and attitude of each aerial unmanned aerial vehicle in real time to guide it to fly along a predetermined route; A three-dimensional modeling device: for receiving three-dimensional video information in real time, fusing to generate a three-dimensional road surface model based on the video data and the corresponding relative position information; A model information collection and analysis device: for collecting three-dimensional road surface models generated at different time periods, comparing and analyzing the current three-dimensional road surface model with historical three-dimensional road surface models to extract and generate road surface health information representing changes in road surface state; The three-dimensional modeling device comprises a pixel point matching module, a three-dimensional model construction module, and a model rendering module; The pixel point matching module first analyzes the three-dimensional video information, extracts feature points in the video frames using the SIFT algorithm, and then performs coordinate conversion on the feature points to generate a feature point mapping set containing spatial position information of each feature point; The three-dimensional model construction module matches and optimizes the points in the feature point mapping set based on the bundle adjustment algorithm, constructs a preliminary three-dimensional model framework by calculating the intersection relationship of light rays at different viewing angles, and ensures the accuracy of the geometric structure of the model; The model rendering module extracts rendering materials from three-dimensional images containing texture and color information, assigns the rendering materials to the model framework, performs rendering processing on the model through light calculation and texture mapping, and finally generates a three-dimensional road surface model. 2.The large model analysis-based road surface health information extraction system of claim 1, wherein, The aerial unmanned aerial vehicle comprises: An unmanned aerial vehicle body for providing an aerial flight platform; A high-speed camera mounted on the unmanned aerial vehicle body at a fixed angle and tilted downward; A wireless positioning module for generating initial relative position parameters with the remaining aerial unmanned aerial vehicles through millimeter wave positioning; A laser transmitter fixedly installed on the unmanned aerial vehicle; A laser receiver forming two laser receiving planes on both sides of the unmanned aerial vehicle body; An information processor connected with the wireless positioning module and the laser receiver for generating relative position information; The laser transmitter of the aerial unmanned aerial vehicle sends laser signals to the laser receiver of the adjacent aerial unmanned aerial vehicle, and the information processor corrects the relative position information based on the position of the laser irradiated to the laser receiving plane. 3.The large model analysis-based road surface health information extraction system of claim 1, wherein, The aerial unmanned aerial vehicles located on both sides obtain the relative position of the aerial unmanned aerial vehicle located in the middle, and the aerial unmanned aerial vehicle located in the middle obtains its current geographical position. 4.The large model analysis-based road surface health information extraction system of claim 2, wherein The unmanned aerial vehicle information acquisition device further comprises: A video information screening module connected with the aerial unmanned aerial vehicle for obtaining video data synchronously collected by each aerial unmanned aerial vehicle, screening out image frames when the flight attitude of the aerial unmanned aerial vehicle is stable, and generating three-dimensional video information.

5. The road surface health information extraction system based on large model analysis according to claim 4, wherein The video information screening module comprises: The video information acquisition unit is configured to acquire video data synchronously collected by the three aerial unmanned vehicles, and arrange the video data synchronously collected by the three aerial unmanned vehicles based on time labels to generate a video frame matrix; The video information screening unit is configured to acquire a time period during which laser signals are received by the laser receivers of the aerial unmanned vehicles, and take a time period during which laser signals are received by all the laser receivers as an ideal time period; The video information cropping unit is configured to take a portion corresponding to the ideal time period in the video frame matrix as three-dimensional video information. 6.The large model analysis-based road surface health information extraction system of claim 5, wherein, The pixel point matching module extracts feature points in the three-dimensional image and maps the feature points by using the following steps: Step 1: Extract the flight speed of the aerial unmanned vehicle from the three-dimensional video information, and divide the three-dimensional video information into a plurality of video groups based on the flight speed, wherein each video group includes video streams captured by the three aerial unmanned vehicles at the same time period; Step 2: Extract a plurality of sample pictures at equal intervals from the video streams, extract feature points from the sample pictures based on the SIFT algorithm, and match the feature points one by one based on the cosine distance to generate a sample matching set; Step 3: Train the convolutional network model on the sample matching set, and adjust the weight parameters in the convolutional network model; Step 4: Input the remaining pictures in the video streams into the trained convolutional network model to extract feature points and mapping relationships between the feature points; Step 5: Take the feature points and the mapping relationships between the feature points extracted in steps 2 and 4 as a feature point mapping set. 7.The large model analysis-based road surface health information extraction system of claim 6, wherein, The convolutional network model comprises: an input layer configured to input pictures captured by the three aerial unmanned vehicles at the same time to generate a first feature map, a second feature map, and a third feature map; a convolutional layer configured to perform convolutional processing on the input first feature map, second feature map, and third feature map to extract first hidden features, second hidden features, and third hidden features; a pooling layer connected to the convolutional layer and configured to pool the first hidden features, second hidden features, and third hidden features; a probability calculation layer configured to generate probability information that each pixel point belongs to a feature point from the first hidden features, second hidden features, and third hidden features; a first attention mechanism network configured to input the first feature map and the probability information to generate a first transformed feature; a second attention mechanism network configured to input the second feature map and the probability information to generate a second transformed feature; a third attention mechanism network configured to input the third feature map and the probability information to generate a third transformed feature; a mapping network configured to map the first transformed feature, second transformed feature, and third transformed feature to each other to generate feature points and mapping relationships between the feature points. 8.The large model analysis-based road surface health information extraction system of claim 1, wherein, The model rendering module generates a model framework based on the following steps: Z1: obtaining a feature point mapping set and relative position information corresponding to the feature point mapping set, converting the relative position information into exterior orientation elements E, the exterior orientation elements E including coordinates (X S , Y S , Z S ) of a photographic center S in a ground coordinate system and rotation angles χ S representing a rotation angle of an image coordinate system around an X axis, ψ S representing a rotation angle of the image coordinate system around a Y axis, representing a rotation angle of the image coordinate system around a Z axis; taking a feature point in the feature point set as a connection point i (x ij , y ij ), and j representing a serial number of an image, j = 1, 2, 3; Z2: for each connection point i, based on the exterior orientation elements of the image j containing the connection point i and the pixel point coordinates of the connection point i in the image j, a resection is performed by a collinearity equation to solve the initial ground coordinate approximation TP of the connection point i Z3: according to TP and the extrinsic orientation elements build a parameter vector X k ; wherein k represents the iteration number, and k = 0 at the initial time, E1k represents the exterior orientation elements of the first aerial unmanned vehicle at the kth iteration, E3k represents the exterior orientation elements of the third aerial unmanned vehicle at the kth iteration, TP1k represents the initial ground coordinate approximation of the first connection point at the kth iteration, TPnk represents the initial ground coordinate approximation of the nth connection point at the kth iteration, n represents the total number of connection points, T represents the matrix transpose symbol, E1 represents the exterior orientation elements E of the first aerial unmanned vehicle, E3 represents the exterior orientation elements E of the third aerial unmanned vehicle; TP1 represents the initial ground coordinate approximation of the first connection point, and TPn represents the initial ground coordinate approximation of the nth connection point. Z4: for each connection point i (x ij , y ij ), extract the current extrinsic element of image j and the current coordinates of connection point i TP i k , compute the predicted image point coordinates (x pr , y pr ); denotes the extrinsic element of the jth image at the kth iteration; wherein f represents the camera focal length, Rj,ij represents the element of the rotation matrix of the jth image, the two digits of the subscript represent the row number and column number of the rotation matrix respectively, x0, y0 represents the pixel point coordinate; Rk,uk represents the initial ground coordinate approximation of the connection point u at the kth iteration, Rk,j represents the exterior orientation element of the jth aerial unmanned vehicle at the kth iteration, Δx, Δy represents the distortion correction value; The observed value is taken as a constant term I c , the observed value residual I x and I y are calculated, the residuals are arranged in order to generate the connection point pixel point coordinate residual vector I im ; Xi, y pr pr Xi, y Xi, y x Xi, y y Xi, y​ The control points g are extracted from the set of feature points, and for each ground control point g, a known observation value A control point ground coordinate residual vector I is generated ct ; represents the current coordinate approximation of the ground control point g, represents the observation residual of the ground control point g, represents the coordinate obtained by measurement of the ground control point g, and the residual is arranged in order to generate a constant vector, forming a constant term vector I of the control point ct ; connecting point pixel coordinate residual vector I im and control point ground coordinate residual vector I ct vertical stitching; combined into a complete constant vector I; Z5: parameter vector X k A Taylor series expansion is performed to obtain a linearized error equation for the observation: v = A ΔX - I; wherein v represents a residual vector containing the residuals of the coordinate observations of all connection points and control points, ΔX represents a correction number vector containing correction numbers of all exterior orientation elements E, and A represents a design matrix; each element of the design matrix A is a predicted value of a collinearity equation; Z6: Construct a normal equation: (A T PA) ΔX = A T PI; wherein I represents a complete constant vector, A represents a design matrix, T represents a matrix transpose symbol, and P represents a weight matrix of the observations; Z7: solve the normal equations to generate a correction vector ΔX, input the correction vector ΔX into the parameter vector X k , update the new parameter vector X k+1 ; based on the new parameter vector X k+1 geographical position coordinates of each connection point are generated, and the model framework is rendered according to the geographical position coordinates. 9.A method for extracting pavement health information based on large model analysis, characterized in that, The road surface health information is extracted by using the road surface health information extraction system based on large model analysis in any one of claims 1-8.

Citation Information

Patent Citations

  • Automatic inspection method and system for reconstructing expressway by using multi-view vision of unmanned aerial vehicle

    CN115345945A

  • Road crack rapid identification and calculation method based on unmanned aerial vehicle low altitude photography

    CN118298338A