Method and apparatus for constructing a three-dimensional digital large model of open-set transportation infrastructure
By combining street scene pictures and point cloud data, a three-dimensional mask is generated using open set object detection and SAM segmentation model, the problems of identification accuracy and multi-dimensional attribute acquisition in the digitization of traffic infrastructure are solved, and high adaptability and high accuracy of traffic infrastructure recognition are achieved.
Patent Information
- Application Number
- CN202510601387.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-12
AI Technical Summary
In the digitalization of transportation infrastructure, the single modal method of image and point cloud data has limitations. It is difficult to accurately obtain the three-dimensional geographical location and multi-dimensional attribute information of the target, cannot meet the needs of refined management and decision-making, and is not well adaptable to new category goals.
Combining street scene pictures, point cloud data and prompt information, target recognition is carried out through an open set object detection model, and all models are segmented using SAM to generate a two-dimensional mask, and project it into a three-dimensional space. Combining the communication coverage judgment strategy and the Otsu threshold segmentation algorithm, a high-quality three-dimensional mask is generated, and finally extracting the multi-dimensional attributes of the target through a multi-modal pre-training model.
It realizes high adaptability identification of transportation infrastructure in different urban environments, supports rich category expansion and fine-grained attribute extraction, improves identification accuracy, and improves decision-making efficiency in infrastructure management and urban planning.
Smart Images

Figure CN120125811B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision, and particularly to a method, device, storage medium, and electronic device for constructing a three-dimensional digital model of open-set transportation infrastructure. Background Art
[0002] Transportation infrastructure is a key support for urban operation, and its digital transformation is crucial for improving management efficiency, ensuring traffic safety, and promoting urban sustainable development. Governments around the world have actively invested. For example, the US government has made huge investments in the modernization of the transportation system, and China has also introduced policies to promote the digital upgrade of highway and waterway transportation infrastructure, highlighting its important position in modern society.
[0003] The core of digitalizing transportation infrastructure lies in accurate asset inventory, and this process highly depends on efficient data processing and analysis technologies. Traditional methods for investigating and identifying transportation infrastructure mainly work based on the point cloud and panoramic images of road scenes obtained by mobile mapping systems. However, these traditional methods face many difficult challenges in practical applications.
[0004] Although image-based recognition methods can capture rich texture and semantic information, they have obvious limitations in obtaining the accurate three-dimensional geographical location of targets. Due to the lack of depth information in image data itself, relying solely on images for positioning, its geographical accuracy is difficult to meet actual needs, resulting in the inability to provide accurate spatial location data for transportation infrastructure management in practical applications. In the case of point cloud-based methods, when dealing with targets with complex shapes, due to the limited ability to describe manual features, it is difficult to comprehensively and accurately depict target features, thus making its generalization and robustness poor. In addition, point cloud data cannot directly provide texture information, which hinders the understanding of the multi-dimensional attributes and deep meanings of transportation infrastructure.
[0005] To overcome the limitations of the above single-modal methods, methods combining images and point clouds have emerged. Such methods attempt to integrate the advantages of both and have made progress to a certain extent, but there are still many deficiencies. In terms of open-set detection, they have poor adaptability to new category targets and are difficult to cope with the continuously increasing and changing types of transportation infrastructure. The utilization of multi-view information is also insufficient, and the valuable spatial information contained in repeated observations has not been fully explored, resulting in a waste of data resources. Moreover, these methods can often only obtain the basic attributes of targets, such as category and location, and cannot effectively obtain richer attribute information such as the meaning of targets, shape details, and material composition. This makes the description of transportation infrastructure incomplete and in-depth, and it is difficult to meet the needs of refined management and decision-making. Summary of the Invention
[0006] This application provides a method, device, storage medium, and electronic device for constructing a three-dimensional digital large model of open-set transportation infrastructure, which has high adaptability, can accurately identify transportation infrastructure in different urban environments, supports rich category expansion and fine-grained attribute extraction, and improves the accuracy of transportation infrastructure recognition.
[0007] This application provides a method for constructing a three-dimensional digital large model of open-set transportation infrastructure, including:
[0008] Obtain street view images containing transportation infrastructure, point cloud data, and the corresponding hint information of the street view images;
[0009] Perform object detection on the street view images based on the hint information to obtain object detection boxes;
[0010] Input the object detection boxes into the Segment Anything Model to obtain two-dimensional masks;
[0011] Based on the geometric relationship between the two-dimensional masks and the point cloud data, project the two-dimensional masks into the three-dimensional space to obtain three-dimensional mask candidates;
[0012] Group the three-dimensional mask candidates based on the intersection coverage judgment strategy to obtain groups of three-dimensional mask candidates;
[0013] Calculate the scores of the points in the groups of three-dimensional mask candidates, and determine the segmentation threshold through the Otsu threshold segmentation algorithm. Take the points with scores higher than the segmentation threshold as foreground points, and the set of the foreground points as the target three-dimensional masks.
[0014] Furthermore, in the above method for constructing a three-dimensional digital large model of open-set transportation infrastructure, the hint information includes text hint information and visual hint information. The performing object detection on the street view images based on the hint information to obtain object detection boxes includes:
[0015] Input the text hint information and the street view images into an open-set object detection model to obtain object detection boxes;
[0016] Input the visual hint information and the street view images into an open-set object detection model to obtain object detection boxes.
[0017] Furthermore, in the above method for constructing a three-dimensional digital large model of open-set transportation infrastructure, the inputting the text hint information and the street view images into an open-set object detection model to obtain object detection boxes includes:
[0018] Obtain the object detection boxes through the first formula, and the first formula is:
[0019]
[0020] Among them, is an open-set object detection model, is a panoramic image of a street scene, is a text prompt describing the object category, is an object detection bounding box, is a set of bounding boxes of the object, is the corresponding semantic category;
[0021] The step of inputting the visual prompt information and the street scene picture into the open-set object detection model to obtain an object detection bounding box includes:
[0022] Obtain the object detection bounding box through a second formula, and the second formula is:
[0023]
[0024] Among them, is another open-set object detection model, is the visual prompt information, is the object detection bounding box, and are the detected bounding box and the corresponding semantic category respectively.
[0025] Furthermore, for the above open-set traffic infrastructure three-dimensional digital large model construction method, wherein, the step of inputting the object detection bounding box into the Segment Anything model to obtain a two-dimensional mask includes:
[0026] Obtain the two-dimensional mask through a third formula, and the third formula is:
[0027]
[0028] Among them, and respectively represent the number of object detection bounding boxes obtained using the text prompt information and the visual prompt information, represents the Segment Anything model, is the input image, is a set of bounding boxes of the object detection, is the output two-dimensional mask.
[0029] Furthermore, for the above open-set traffic infrastructure three-dimensional digital large model construction method, wherein, based on the geometric relationship between the two-dimensional mask and the point cloud data, projecting the two-dimensional mask into the three-dimensional space to obtain a three-dimensional mask candidate includes:
[0030] According to the extrinsic matrix and the intrinsic matrix of the camera, project the point cloud data onto the image plane to obtain a number of projection points;
[0031] Take the projection points falling within each mask range of the two-dimensional mask as three-dimensional mask candidate points;
[0032] Take the set of the three-dimensional mask candidate points as the three-dimensional mask candidate.
[0033] Furthermore, for the above method for constructing a three-dimensional digital model of open-set traffic infrastructure, wherein, grouping the three-dimensional masks based on the communication coverage judgment strategy to obtain three-dimensional mask candidate groups, includes:
[0034] Calculate the confidence of the three-dimensional mask candidate;
[0035] Calculate the three-dimensional mask candidate with the highest confidence and the coverage rate between the three-dimensional mask candidate with the highest confidence and other three-dimensional mask candidates having an intersection with the three-dimensional mask candidate with the highest confidence. If the coverage rate is greater than the coverage rate threshold, divide the two three-dimensional mask candidates into the same group of three-dimensional mask candidates.
[0036] Furthermore, for the above method for constructing a three-dimensional digital model of open-set traffic infrastructure, wherein, the calculation formula of the coverage rate is:
[0037]
[0038] wherein, represents the dilation operation on the sparse pixels of the point cloud, is the two-dimensional mask obtained by projecting the intersection point cloud of two three-dimensional mask candidates to be judged onto their respective images, is the corresponding two-dimensional mask, is the number of pixel points;
[0039] The definition of the three-dimensional mask candidate group is:
[0040]
[0041] wherein, is the three-dimensional mask candidate group, represents the coverage rate judgment threshold, is the three-dimensional mask candidate, is the convolution operation.
[0042] Furthermore, for the above method for constructing a three-dimensional digital model of open-set traffic infrastructure, wherein, calculating the scores of the points in the three-dimensional mask candidate group, includes:
[0043] Calculate the scores through the following formula:
[0044]
[0045]
[0046] Among them, represents the number of three-dimensional mask candidates containing the point , and is the number of three-dimensional mask candidates. is a three-dimensional mask candidate. is a function for calculating the absolute value. is the observation angle closest to orthogonality in multiple perspectives.
[0047] Furthermore, in the above method for constructing a large-scale three-dimensional digital model of open-set traffic infrastructure, after the step of using the set of foreground points as the target three-dimensional mask, it includes:
[0048] Determine the target detection box corresponding to the target three-dimensional mask;
[0049] Crop the target detection box to obtain a cropped image;
[0050] Input the cropped image and the preset attribute information group into the multi-modal pre-trained model to obtain the matching degree of each piece of attribute information in the cropped image and the attribute information group;
[0051] Select the piece of attribute information with the highest matching degree as the additional attribute of the corresponding traffic infrastructure.
[0052] This application also provides a device for constructing a large-scale three-dimensional digital model of open-set traffic infrastructure, including:
[0053] An acquisition module, configured to acquire a street view image containing traffic infrastructure, point cloud data, and the prompt information corresponding to the street view image;
[0054] A target detection module, configured to perform target detection on the street view image based on the prompt information to obtain a target detection box;
[0055] A two-dimensional mask generation module, configured to input the target detection box into the Segment Anything Model to obtain a two-dimensional mask;
[0056] A three-dimensional mask candidate generation module, configured to project the two-dimensional mask into three-dimensional space based on the geometric relationship between the two-dimensional mask and the point cloud data to obtain a three-dimensional mask candidate;
[0057] A grouping module, configured to group the three-dimensional mask candidates based on the communication coverage judgment strategy to obtain a group of three-dimensional mask candidates;
[0058] A target three-dimensional mask generation module is used to calculate the scores of the points in the three-dimensional mask candidate group, determine the segmentation threshold through the Otsu threshold segmentation algorithm, take the points with scores higher than the segmentation threshold as foreground points, and set the set of the foreground points as the target three-dimensional mask.
[0059] The present application also provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded by a processor to execute any one of the above open-set traffic infrastructure three-dimensional digital large model construction methods.
[0060] The present application also provides an electronic device, including a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used for the steps in any one of the above open-set traffic infrastructure three-dimensional digital large model construction methods.
[0061] For the open-set traffic infrastructure three-dimensional digital large model construction method, device, storage medium and electronic device provided by the present application, first, the present application identifies the targets in the street view image through the open-set object detection technology based on text and visual cues; secondly, using the object detection bounding box as a cue, the SAM segmentation everything model is used to generate the two-dimensional mask of the target; then, the two-dimensional mask is projected onto the 3D point cloud to obtain the three-dimensional mask candidate, and through the 2D-3D spatial projection strategy bound to the point cloud, using the multi-view repeated observation information and the observation angle, the three-dimensional mask candidate is grouped, scored and segmented, so as to generate a high-quality target three-dimensional mask; finally, a multi-modal model based on image-text contrast learning is used to extract the multi-dimensional attribute information of the target, including fine semantics, geometric shape and material characteristics. The present application has high adaptability, can accurately identify traffic infrastructure in different urban environments, supports rich category expansion and fine-grained attribute extraction, improves the accuracy of traffic infrastructure identification, and further improves the decision-making efficiency of infrastructure management and urban planning. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The following will make the technical solutions and other beneficial effects of the present application obvious by describing the specific embodiments of the present application in detail in conjunction with the drawings.
[0063] Figure 1 It is a flowchart of the open-set traffic infrastructure three-dimensional digital large model construction method provided by the embodiment of the present application.
[0064] Figure 2 It is another flowchart of the open-set traffic infrastructure three-dimensional digital large model construction method provided by the embodiment of the present application.
[0065] Figure 3The object detection results of traffic signs and traffic lights using "Traffic Sign" and "Signal Light" provided by the embodiments of this application.
[0066] Figure 4 The visual cues for roadside stone isolation piers provided by the embodiments of this application and the object detection results based on the visual cues.
[0067] Figure 5 Partial two-dimensional masks of traffic infrastructure obtained using the SAM model and object detection bounding boxes as cues provided by the embodiments of this application.
[0068] Figure 6 The recognition results of partial categories of traffic infrastructure provided by the embodiments of this application.
[0069] Figure 7 The recognition results of the attributes of partial categories of traffic infrastructure provided by the embodiments of this application.
[0070] Figure 8 The structural schematic diagram of the open-set traffic infrastructure three-dimensional digital large model construction device provided by the embodiments of this application.
[0071] Figure 9 The structural schematic diagram of the electronic device provided by the embodiments of this application. Detailed implementation manners
[0072] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.
[0073] The embodiments of this application provide an open-set traffic infrastructure three-dimensional digital large model construction method, device, storage medium, and electronic device. An open-set traffic infrastructure three-dimensional digital large model construction device provided by the embodiments of this application can be integrated in an electronic device, and the electronic device can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a micro processing box, or other devices, etc.
[0074] Please refer to Figure 1 And Figure 2 , Figure 1 which is the flowchart of the open-set traffic infrastructure three-dimensional digital large model construction method provided by the embodiments of this application, Figure 2Another flowchart of the method for constructing a three-dimensional digital model of open-set traffic infrastructure provided by the embodiments of the present application, which is applied to an electronic device. The method for constructing a three-dimensional digital model of open-set traffic infrastructure includes the following steps:
[0075] S1. Obtain street view images, point cloud data, and prompt information corresponding to the street view images that contain traffic infrastructure.
[0076] S2. Perform object detection on the street view images based on the prompt information to obtain object detection bounding boxes.
[0077] Specifically, the prompt information includes text prompt information and visual prompt information. Step S2 includes the following steps:
[0078] S21. Input the text prompt information and the street view images into an open-set object detection model to obtain object detection bounding boxes.
[0079] Specifically, use a text prompt object detection large model to implement object detection based on natural language descriptions. By inputting text prompt information (such as "traffic sign"), the model can identify various objects with the same semantic category without relying on specific appearance features. This process can be expressed as:
[0080]
[0081] Among them, represents the open-set object detection model, specifically the text prompt object detection large model, is the panoramic image of the street view, is the text prompt describing the object category, is the object detection bounding box, and the output is the set of bounding boxes of the objects, is the corresponding semantic category. Figure 3 Shows the object detection results of traffic signs and red / green signal lights using "Traffic Sign" and "SignalLight" provided by the embodiments of the present application.
[0082] S22. Input the visual prompt information and the street view images into an open-set object detection model to obtain object detection bounding boxes.
[0083] When it is difficult to accurately describe the object in language, a visual prompt object detection large model can be used to implement object detection based on visual prompts (such as a selected area or a point). In this embodiment, the visual prompts in the image are used to label the region of interest to achieve precise recognition of complex objects with minimal prompts. The object detection process is:
[0084]
[0085] Among them, is another open-set object detection model, specifically a large model for visual prompt object detection. is a visual prompt (such as a boxed area or a point). is the object detection box. and are the detected bounding box and the corresponding semantic class respectively. Figure 4 are the visual prompts provided by the embodiments of this application for roadside stone isolation piers and the object detection results based on these visual prompts.
[0086] S3. Input the object detection box into the Segment Anything Model to obtain a two-dimensional mask.
[0087] Use the Segment Anything Model (SAM) to generate a pixel-level 2D mask of the object. The SAM model takes the detected bounding box as the prompt input and generates a high-quality pixel-level object segmentation mask. The process can be expressed as:
[0088]
[0089] where and represent the number of object detection bounding boxes obtained using text prompt information and visual prompt information respectively. represents the SAM model. is the input image. is the set of bounding boxes for object detection. is the output pixel-level two-dimensional mask. This process provides high-quality 2D segmentation input for subsequent point cloud processing, which can significantly improve the accuracy and reliability of segmentation. Figure 5 are partial two-dimensional masks of traffic infrastructure obtained by the SAM model provided by the embodiments of this application and the object detection bounding boxes as prompts.
[0090] S4. Based on the geometric relationship between the two-dimensional mask and the point cloud data, project the two-dimensional mask into the three-dimensional space to obtain a three-dimensional mask candidate.
[0091] In one embodiment, step S4 includes the following steps:
[0092] S41. According to the extrinsic matrix and intrinsic matrix of the camera, project the point cloud data onto the image plane to obtain a number of projection points.
[0093] S42. Take the projection points falling within each mask range of the two-dimensional mask as three-dimensional mask candidate points.
[0094] S43. Take the set of three-dimensional mask candidate points as the three-dimensional mask candidate.
[0095] Specifically, after obtaining the two-dimensional mask , according to the extrinsic and intrinsic matrix of the camera, project the three-dimensional points in the point cloud data onto the image plane. For each mask in the two-dimensional mask , the three-dimensional mask candidate corresponding to the two-dimensional mask is composed of the points in the point cloud whose two-dimensional coordinates projected onto the image fall within . Each two-dimensional mask corresponds to a generated three-dimensional mask candidate which is expressed as:
[0096]
[0097] where is the input point cloud, is the point whose pixel row and column numbers projected onto the panoramic image according to the exposure extrinsic parameters.
[0098] By calculating the three-dimensional mask candidate points through each mask , the three-dimensional mask candidate can be obtained.
[0099] S5. Group the three-dimensional mask candidates based on the intersection coverage judgment strategy to obtain groups of three-dimensional mask candidates.
[0100] S51. Calculate the confidence of the three-dimensional mask candidates;
[0101] S52. Calculate the coverage rate between the three-dimensional mask candidate with the highest confidence and other three-dimensional mask candidates that have an intersection with the three-dimensional mask candidate with the highest confidence. If the coverage rate is greater than the coverage rate threshold, the two three-dimensional mask candidates are divided into the same group of three-dimensional mask candidates.
[0102] Specifically, in order to make full use of the spatial information embedded in the repeated observations of the same target from different angles, the present invention proposes a three-dimensional mask candidate grouping strategy based on intersection coverage judgment, which divides the three-dimensional mask candidates belonging to different targets into different clusters. The principle of this strategy is that if two three-dimensional mask candidates represent the same target, their intersection point set will basically cover the foreground points of the target. On the contrary, if they represent different targets, their intersection point set is empty, or at least does not cover the two targets simultaneously. By projecting the intersection points back to their respective two-dimensional images and calculating the coverage rate, it can be determined whether two three-dimensional masks correspond to the same target.
[0103] First, for the three-dimensional candidate mask with the highest confidence score, iteratively check other candidate masks that have an intersection with it Whether the coverage threshold is met. If the threshold condition is met, and are considered to represent the same target and are grouped into the same group. This operation is repeated for all 3D mask candidates until the grouping of all candidates is completed. The formula for calculating the coverage rate is:
[0104]
[0105] where represents the dilation operation on the sparse pixels of the point cloud, is the 2D mask obtained by projecting the intersection point cloud of two 3D candidate masks to be judged onto their respective images, is the corresponding 2D mask, is the number of pixel points.
[0106] Through the above method, all 3D mask candidates that represent the same target as can be found and a 3D mask candidate group is formed, which is defined as follows:
[0107]
[0108] where is the 3D mask candidate group, is the 3D mask candidate, represents the coverage rate judgment threshold, which is set to 0.5 in the present invention. After performing the above operations on all 3D mask candidates in each , all mask candidates in the 3D mask candidate can be divided into multiple groups, and each group corresponds to a target in the 3D world.
[0109] S6. Calculate the scores of the points in the 3D mask candidate group, and determine the segmentation threshold through the Otsu threshold segmentation algorithm. The points with scores higher than the segmentation threshold are used as foreground points, and the set of foreground points is the target 3D mask.
[0110] Specifically, after finding all 3D mask candidate groups belonging to a certain target, scores are assigned to each point in the 3D mask candidate group to measure the possibility of it belonging to the target. This strategy can simultaneously utilize the spatial information of different shapes presented by the same target from different perspectives. By considering the number of observations and the observation angles, higher scores are assigned to the points that are more likely to be foreground points.
[0111] For the target , the set of points in all its corresponding 3D mask candidates is represented as . The score of the point is calculated as follows:
[0112]
[0113]
[0114] Among them, represents the number of 3D mask candidates that contain point , is a 3D mask candidate, is a function for calculating the absolute value, is the observation angle closest to orthogonality among multiple perspectives. Assume that represents the set of line segments connecting point and the image exposure points corresponding to each mask candidate. Then represents the angle closest to 90° among the angles formed by any two line segments in
[0115] Obviously, the more times point is observed in images from different perspectives, the higher the probability that it belongs to the target. At the same time, the closer it is to 90°, the higher the confidence that the point belongs to the target. After assigning scores to each point, the Otsu threshold segmentation algorithm is used to determine the segmentation threshold between foreground points and background points. Points with scores higher than are considered to belong to the target (i.e., foreground points). The formula for generating the final target 3D mask is as follows:
[0116]
[0117] By performing the above operations on each group of 3D mask candidates, the final set of 3D masks is generated. Figure 6 is the recognition result of some categories of traffic infrastructure provided by the embodiments of this application.
[0118] Furthermore, after step S6, the following steps are further included:
[0119] S61, determining the target detection box corresponding to the target 3D mask;
[0120] S62, cropping the target detection box to obtain a cropped picture;
[0121] S63, inputting the cropped picture and the preset attribute information group into the multi-modal pre-trained model to obtain the matching degree between the cropped picture and each piece of attribute information in the attribute information group;
[0122] S64, selecting the piece of attribute information with the highest matching degree as the additional attribute of the corresponding traffic infrastructure.
[0123] Specifically, for each 3D mask, find the object detection bounding box with the highest object detection confidence corresponding to it, crop the image using the object detection bounding box enlarged by 1.5 times, input the cropped image and a preset set of attribute information into the multi-modal pre-trained model, calculate the matching degree between the image and each attribute in the set of set attribute information, and select the attribute with the highest matching degree as the additional attribute of the traffic infrastructure. Finally, obtain the target 3D mask and additional attribute of each traffic infrastructure. Figure 7 It is the recognition result of the attributes of some categories of traffic infrastructure provided by the embodiments of this application.
[0124] Among them, the multi-modal pre-trained model can be a CLIP model (Contrastive Language-Image Pre-Training).
[0125] This application combines point cloud and image data and uses a multi-modal large model to achieve efficient traffic infrastructure recognition. Specifically, first, through the open-set object detection technology based on text and visual cues, identify the objects in the street view image; second, use the object detection bounding box as a cue and use the SAM segmentation all model to generate the 2D mask of the object; then, project the 2D mask into the 3D point cloud to obtain the 3D mask candidate, and through the 2D-3D spatial projection strategy bound by the point cloud, use the multi-view repeated observation information and the viewing angle to group, score and segment the 3D mask candidate, so as to generate a high-quality target 3D mask; finally, use a multi-modal model based on image-text contrast learning to extract the multi-dimensional attribute information of the object, including fine semantics, geometric shape and material characteristics. This application has high adaptability, can identify traffic infrastructure in different urban environments, supports rich category expansion and fine-grained attribute extraction, improves the accuracy of traffic infrastructure recognition, and further improves the decision-making efficiency of infrastructure management and urban planning.
[0126] According to the method described in the above embodiments, this embodiment will be further described from the perspective of an open-set traffic infrastructure 3D digital large model construction device. The open-set traffic infrastructure 3D digital large model construction device can be specifically implemented as an independent entity, or integrated in an electronic device. The electronic device can be a terminal, a server, etc. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a micro processing box, or other devices, etc.
[0127] Please refer to Figure 8 , Figure 8Specifically described is an open-set traffic infrastructure three-dimensional digital large model construction device provided by an embodiment of the present application, which is applied to an electronic device. The open-set traffic infrastructure three-dimensional digital large model construction device may include:
[0128] Obtain street view pictures, point cloud data containing traffic infrastructure, and prompt information corresponding to the street view pictures;
[0129] Perform object detection on the street view pictures based on the prompt information to obtain object detection frames;
[0130] Input the object detection frames into a segment-anything model to obtain two-dimensional masks;
[0131] Based on the geometric relationship between the two-dimensional masks and the point cloud data, project the two-dimensional masks into three-dimensional space to obtain three-dimensional mask candidates;
[0132] Group the three-dimensional mask candidates based on an intersection coverage judgment strategy to obtain three-dimensional mask candidate groups;
[0133] Calculate the scores of the points in the three-dimensional mask candidate groups, and determine a segmentation threshold through an Otsu threshold segmentation algorithm. Take the points with scores higher than the segmentation threshold as foreground points, and set the set of the foreground points as the target three-dimensional mask.
[0134] In specific implementation, each of the above modules and / or units may be implemented as an independent entity, or may be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of each of the above modules and / or units, reference may be made to the foregoing method embodiments. For the specific beneficial effects that can be achieved, reference may also be made to the beneficial effects in the foregoing method embodiments, which will not be elaborated herein.
[0135] In addition, an embodiment of the present application further provides an electronic device, which may be a device such as a computer or a tablet computer. The electronic device may implement the steps in any embodiment of the open-set traffic infrastructure three-dimensional digital large model construction method provided by the embodiment of the present application. Therefore, it can achieve the beneficial effects that can be achieved by any of the open-set traffic infrastructure three-dimensional digital large model construction methods provided by the embodiments of the present invention. For details, reference may be made to the foregoing embodiments, which will not be elaborated herein.
[0136] Figure 9 Shows a specific structural block diagram of the electronic device provided by an embodiment of the present invention. The electronic device may be used to implement the open-set traffic infrastructure three-dimensional digital large model construction method provided in the foregoing embodiment. The electronic device 500 may be a device such as a terminal or a server. Among them, the terminal may include a tablet computer, a notebook computer, a personal computer (PC), a microprocessing box, or other devices, etc.
[0137] The RF circuit 510 is used to receive and transmit electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices. The RF circuit 510 may include various existing circuit elements for performing these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The RF circuit 510 can communicate with various networks such as the Internet, enterprise intranets, wireless networks or communicate with other devices through wireless networks. The above-mentioned wireless networks may include cellular phone networks, wireless local area networks or metropolitan area networks. The above-mentioned wireless networks can use various communication standards, protocols and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging and short messages, and any other suitable communication protocols, and may even include those protocols that have not been developed yet.
[0138] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, that is, to implement functions such as taking pictures with the front camera, processing the captured images, and switching the display colors of the display content on the display screen. The memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 520 may further include a memory remotely disposed relative to the processor 580, and these remote memories can be connected to the electronic device 500 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0139] The input unit 530 can be used to receive input digital or character information, and generate keyboards and mice related to user settings and function controls.
[0140] The display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, and these graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 540 may include a display panel 541. Optionally, the display panel 541 can be configured in the form of an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode).
[0141] The audio circuit 560, the speaker 561, and the microphone 562 can provide an audio interface between the user and the electronic device 500. The audio circuit 560 can transmit the electrical signal converted from the received audio data to the speaker 561, and the speaker 561 converts it into a sound signal for output; on the other hand, the microphone 562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 560 and then converted into audio data. After the audio data is output to the processor 580 for processing, it is sent through the RF circuit 510 to, for example, another terminal, or the audio data is output to the memory 520 for further processing. The audio circuit 560 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device 500.
[0142] The electronic device 500 can help the user receive requests, send information, etc. through the transmission module 570 (such as a Wi-Fi module), and it provides the user with wireless broadband Internet access. Although the transmission module 570 is shown in the figure, it can be understood that it does not belong to the essential components of the electronic device 500 and can be omitted completely within the scope of not changing the essence of the invention according to needs.
[0143] The processor 580 is the control center of the electronic device 500, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 520, and by calling the data stored in the memory 520, it executes various functions of the electronic device 500 and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 580 either.
[0144] The electronic device 500 also includes a power supply 590 (such as a battery) for supplying power to each component. In some embodiments, the power supply can be logically connected to the processor 580 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 590 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0145] Although not shown, the electronic device 500 also includes a camera (such as a front camera, a rear camera), a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. One or more programs include instructions for performing the following operations:
[0146] An acquisition module, configured to acquire a street view image, point cloud data including traffic infrastructure, and prompt information corresponding to the street view image;
[0147] A target detection module, configured to perform target detection on the street view image based on the prompt information to obtain a target detection frame;
[0148] A two-dimensional mask generation module, configured to input the target detection frame into a segment-anything model to obtain a two-dimensional mask;
[0149] A three-dimensional mask candidate generation module, configured to project the two-dimensional mask into three-dimensional space based on the geometric relationship between the two-dimensional mask and the point cloud data to obtain a three-dimensional mask candidate;
[0150] A grouping module, configured to group the three-dimensional mask candidates based on a communication coverage judgment strategy to obtain a group of three-dimensional mask candidates;
[0151] A target three-dimensional mask generation module is configured to calculate the scores of the points in the three-dimensional mask candidate group, determine a segmentation threshold through an Otsu threshold segmentation algorithm, use the points with scores higher than the segmentation threshold as foreground points, and set the set of the foreground points as the target three-dimensional mask.
[0152] In specific implementation, each of the above modules can be implemented as an independent entity, or can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of each of the above modules, reference can be made to the foregoing method embodiments, which will not be elaborated herein.
[0153] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the foregoing embodiments can be completed by instructions, or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this purpose, an embodiment of the present invention provides a storage medium, in which multiple instructions are stored, and the instructions can be loaded by a processor to execute the steps of any one of the embodiments of the method for constructing an open-set traffic infrastructure three-dimensional digital model provided by the embodiments of the present invention.
[0154] Wherein, the computer-readable storage medium may include: a read-only memory (ROM, Read Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disc, etc.
[0155] Since the instructions stored in the storage medium can execute the steps of any one of the embodiments of the method for constructing an open-set traffic infrastructure three-dimensional digital model provided by the embodiments of the present invention, the beneficial effects achievable by any of the methods for constructing an open-set traffic infrastructure three-dimensional digital model provided by the embodiments of the present invention can be achieved. For details, reference can be made to the foregoing embodiments, which will not be elaborated herein.
[0156] The foregoing has introduced in detail a method, an apparatus, a storage medium, and an electronic device for constructing an open-set traffic infrastructure three-dimensional digital model provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for constructing a three-dimensional digital large model of open-set transportation infrastructure, characterized in that, Including: Obtain street view images, point cloud data including traffic infrastructure, and hint information corresponding to the street view images; Perform object detection on the street view images based on the hint information to obtain object detection boxes; Input the object detection boxes into a segment - anything model to obtain 2D masks; Based on the geometric relationship between the 2D masks and the point cloud data, project the 2D masks into 3D space to obtain 3D mask candidates; Group the 3D mask candidates based on a communication coverage judgment strategy to obtain a 3D mask candidate group; Calculate the scores of the points in the 3D mask candidate group, and determine a segmentation threshold through the Otsu threshold segmentation algorithm. Consider the points with scores higher than the segmentation threshold as foreground points, and set the set of the foreground points as the target 3D mask; where calculating the scores of the points in the 3D mask candidate group includes: Calculate the scores through the following formula: Among them, represents the number of three-dimensional mask candidates including point is a three-dimensional mask candidate, is a function for calculating the absolute value, is the observation angle closest to orthogonality among multiple perspectives.
2. The method for constructing a three-dimensional digital large model of open-set transportation infrastructure according to claim 1, wherein The hint information includes text hint information and visual hint information. Performing object detection on the street view images based on the hint information to obtain object detection boxes includes: Input the text hint information and the street view images into an open - set object detection model to obtain object detection boxes; Input the visual hint information and the street view images into an open - set object detection model to obtain object detection boxes.
3. The method for constructing a three-dimensional digital model of open-set transportation infrastructure according to claim 2, wherein Inputting the text hint information and the street view images into an open - set object detection model to obtain object detection boxes includes: Obtain the object detection boxes through a first formula, and the first formula is: Among them, is an open-set object detection model, is a panoramic image of a street scene, is a text prompt describing the object category, is an object detection box, is a set of bounding boxes of the object, is the corresponding semantic category; Inputting the visual hint information and the street view images into an open - set object detection model to obtain object detection boxes includes: Obtain the object detection boxes through a second formula, and the second formula is: Among them, is another open-set object detection model, is visual cue information, is the object detection box, and are the detected bounding box and the corresponding semantic class respectively.
4. The method for constructing a three-dimensional digital large model of open-set transportation infrastructure according to claim 1, wherein Inputting the object detection boxes into a segment - anything model to obtain 2D masks includes: Obtain the 2D masks through a third formula, and the third formula is: Among them, and respectively represent the number of object detection bounding boxes obtained using text prompt information and visual prompt information. represents the segmentation all model. is the input image. is the set of bounding boxes for object detection. is the output two-dimensional mask.
5. The method for constructing a three-dimensional digital model of an open-set transportation infrastructure according to claim 1, wherein Based on the geometric relationship between the 2D masks and the point cloud data, projecting the 2D masks into 3D space to obtain 3D mask candidates includes: According to the external parameter matrix and internal parameter matrix of the camera, project the point cloud data onto the image plane to obtain a number of projection points; Take the projection points falling within each mask range of the 2D masks as 3D mask candidate points; Set the set of the 3D mask candidate points as the 3D mask candidates.
6. The method for constructing a three-dimensional digital large model of open-set transportation infrastructure according to claim 1, wherein Grouping the 3D masks based on a communication coverage judgment strategy to obtain a 3D mask candidate group includes: Calculate the confidence of the 3D mask candidates; Calculate the 3D mask candidate with the highest confidence and other 3D mask candidates that intersect with the 3D mask candidate with the highest confidence to calculate the coverage rate therebetween. If the coverage rate is greater than the coverage rate threshold, the two 3D mask candidates are classified into the same group of 3D mask candidates.
7. The method for constructing a three-dimensional digital model of open-set transportation infrastructure according to claim 6, characterized in that, The calculation formula for the coverage rate is: Among them, represents the dilation operation on the sparse pixels of the point cloud, is the two-dimensional mask obtained by projecting the intersection point cloud of two three-dimensional mask candidates to be judged onto their respective images, is the corresponding two-dimensional mask, is the number of pixel points; The definition of the 3D mask candidate group is: Among them, is a three-dimensional mask candidate group, represents the coverage judgment threshold, is a three-dimensional mask candidate, is a convolution operation.
8. The method for constructing a three-dimensional digital model of open-set transportation infrastructure according to claim 1, characterized in that, After the step of setting the set of the foreground points as the target 3D mask, it includes: Determine the object detection box corresponding to the target 3D mask; Crop the object detection box to obtain a cropped image; Input the cropped image and a preset group of attribute information into a multi - modal pre - trained model to obtain the matching degree between the cropped image and each piece of attribute information in the group of attribute information; Select the piece of attribute information with the highest matching degree as the additional attribute of the corresponding traffic infrastructure.
9. An apparatus for constructing a three-dimensional digital large model of open-set transportation infrastructure, which is used to implement the method for constructing a three-dimensional digital large model of open-set transportation infrastructure described in claim 1, characterized in that, Including: An acquisition module, configured to acquire a street view image containing traffic infrastructure, point cloud data, and hint information corresponding to the street view image; A target detection module, configured to perform target detection on the street view image based on the hint information to obtain a target detection box; A two-dimensional mask generation module, configured to input the target detection box into a Segment Anything Model to obtain a two-dimensional mask; A three-dimensional mask candidate generation module, configured to project the two-dimensional mask into three-dimensional space based on the geometric relationship between the two-dimensional mask and the point cloud data to obtain a three-dimensional mask candidate; A grouping module, configured to group the three-dimensional mask candidates based on a communication coverage judgment strategy to obtain a three-dimensional mask candidate group; A target three-dimensional mask generation module, configured to calculate the scores of the points in the three-dimensional mask candidate group, determine a segmentation threshold through an Otsu threshold segmentation algorithm, use the points with scores higher than the segmentation threshold as foreground points, and use the set of the foreground points as a target three-dimensional mask.
Citation Information
Patent Citations
Interactive bridge point cloud semantic segmentation method based on visual large model
CN118864850A
Processing method and system for full-scene ground object segmentation based on visual large model
CN119693823A