A method and system for rapidly constructing a large-scale indoor and outdoor high-precision scene
Through the composite recognition model and dynamic request mechanism, combined with remote sensing images and video data, the problem of identifying temporary equipment and messy scenes in the prior art is solved, and the rapid construction of ultra-large-scale indoor and outdoor high-precision scenes is realized.
Patent Information
- Application Number
- CN202510286901.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-12
AI Technical Summary
When the existing technology builds ultra-large-scale indoor and outdoor high-precision scenarios, it is difficult to accurately identify temporary equipment and infrastructure in messy scenarios, resulting in identification errors and affecting the accuracy of scene construction.
The composite recognition model is used to initially identify the remote sensing image, confirm the suspected infrastructure twice, and obtain indoor video data through the dynamic request mechanism, coordinate mapping and position alignment are carried out to form a super-large-scale indoor and outdoor high-precision scenario.
It realizes accurate identification of indoor and outdoor infrastructure, quickly creates high-precision scenarios, and improves the accuracy and efficiency of scene construction.
Smart Images

Figure CN119784971B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of scene construction, and in particular to a method and system for quickly constructing ultra-large-scale indoor and outdoor high-precision scenes. Background Art
[0002] With the rapid development of science and technology, the demand for large-scale high-precision indoor and outdoor scene construction is growing, mainly used in the field of traffic navigation. High-precision scene models are the key to achieving efficient decision-making and precise control of navigation driving. This kind of navigation application includes both conventional vehicle navigation and navigation of various robots.
[0003] In the process of high-precision scene construction, it is necessary to identify and extract various types of infrastructure and their contours, location distribution information, etc. from the images of these indoor and outdoor scenes, so as to construct outdoor high-precision scenes and indoor high-precision scenes, and then align and integrate them to obtain the required ultra-large-scale indoor and outdoor high-precision scenes. At present, outdoor scene construction often relies on images taken by drones or satellite remote sensing, while indoor scene construction mostly uses surveillance cameras to obtain video images. However, there are many problems with the existing construction methods, mainly reflected in: 1) When identifying infrastructure, it is difficult to distinguish temporary equipment that does not belong to infrastructure in the image. For example, in a road scene, a car is carrying a statue and parked on the side of the road (for example, it is parked at the entrance and exit of the main road and the parallel auxiliary road along the road). At this time, it is easy to identify it as part of the road based on remote sensing images, that is, it is determined that there is no entrance and exit between the main road and the auxiliary road, but the area corresponding to the entrance and exit is a statue. These temporary objects similar to infrastructure will seriously interfere with the identification of infrastructure, resulting in many errors in the constructed high-precision scenes.
[0004] 2) In some indoor areas where scene construction is required, the actual boundary distribution of the passable area cannot be accurately identified because the temporary stacking objects in the scene are too messy.
[0005] Therefore, there is an urgent need for a new method that can quickly and accurately construct ultra-large-scale indoor and outdoor high-precision scenes. Summary of the invention
[0006] In response to the above technical problems, the present invention provides a method, system, electronic device, computer storage medium and computer program product for quickly constructing ultra-large-scale indoor and outdoor high-precision scenes.
[0007] The present invention discloses a method for rapidly constructing a large-scale indoor and outdoor high-precision scene. The method includes the following steps: receiving a target area corresponding to a high-precision scene input by a user, and dividing the target area into an outdoor area and an indoor area based on map information corresponding to the target area; obtaining multiple groups of remote sensing image datasets of the outdoor area corresponding to different acquisition periods, identifying a first infrastructure dataset and a suspected infrastructure dataset from the multiple groups of remote sensing image datasets, performing secondary confirmation on each suspected infrastructure in the suspected infrastructure dataset, and updating the first infrastructure dataset based on the secondary confirmation result to obtain a second infrastructure dataset; using a dynamic request mechanism to request and obtain video data corresponding to the indoor area, and identifying a third infrastructure dataset from the video data; performing coordinate mapping and position alignment on the second infrastructure dataset and the third infrastructure dataset to obtain a large-scale indoor and outdoor high-precision scene.
[0008] Optionally, the identifying a first infrastructure dataset and a suspected infrastructure dataset from the multiple groups of remote sensing image datasets includes: using a composite recognition model to identify objects in any group of the remote sensing image datasets, obtaining several objects and their recognition confidence levels, and identifying objects with a recognition confidence level higher than a confidence threshold as infrastructure and classifying them into the first infrastructure dataset; wherein, the composite recognition model is used for identifying several specified types of infrastructure; and classifying objects with a recognition confidence level not higher than the confidence threshold into the suspected infrastructure dataset.
[0009] Optionally, the performing secondary confirmation on each suspected infrastructure in the suspected infrastructure dataset includes: determining the recognition confidence level of each suspected infrastructure, and determining a first target quantity and a target spanning duration according to the high or low recognition confidence level; taking the shooting moment of the remote sensing image dataset corresponding to the recognition confidence level as a reference, screening out multiple groups of other remote sensing image datasets of the first target quantity according to the target spanning duration, and performing secondary confirmation on the suspected infrastructure based on multiple remote sensing images corresponding to each other remote sensing image dataset to obtain a secondary confirmation result, where the secondary confirmation result includes confirmed as infrastructure and confirmed as non-infrastructure.
[0010] Optionally, the performing secondary confirmation on the suspected infrastructure based on multiple remote sensing images corresponding to each other remote sensing image dataset to obtain a secondary confirmation result includes: determining several associated sub-recognition models corresponding to the sub-recognition model of the suspected infrastructure; each of the associated sub-recognition models refers to having a similarity higher than a similarity threshold with the type of infrastructure for which the sub-recognition model is used for recognition; using each of the associated sub-recognition models to respectively identify multiple remote sensing images, and fusing each recognition result to obtain the secondary confirmation result.
[0011] Optionally, the step of requesting to obtain video data corresponding to the indoor area by using the dynamic request mechanism includes: requesting to obtain a set of video data corresponding to the indoor area, analyzing the change range of the item positions from the first set of video data, determining a second target quantity according to the change range level to which the item position change range belongs; and requesting to obtain multiple sets of video data corresponding to the second target quantity and corresponding to the indoor area.
[0012] The present invention also discloses a system for rapidly constructing a large-scale indoor and outdoor high-precision scene. The system includes a receiving and preprocessing module, an outdoor analysis and recognition module, an indoor analysis and recognition module, and a scene integration and generation module. The receiving and preprocessing module receives a target area corresponding to a high-precision scene input by a user, and divides the target area into an outdoor area and an indoor area based on map information corresponding to the target area. The outdoor analysis and recognition module obtains multiple sets of remote sensing image data sets of the outdoor area corresponding to different acquisition time periods, identifies a first infrastructure data set and a suspected infrastructure data set from the multiple sets of remote sensing image data sets, performs secondary confirmation on each suspected infrastructure in the suspected infrastructure data set, and updates the first infrastructure data set based on the secondary confirmation result to obtain a second infrastructure data set. The indoor analysis and recognition module requests to obtain video data corresponding to the indoor area by using a dynamic request mechanism, and identifies a third infrastructure data set from the video data. The scene integration and generation module performs coordinate mapping and position alignment on the second infrastructure data set and the third infrastructure data set to obtain a large-scale indoor and outdoor high-precision scene.
[0013] Optionally, the outdoor analysis and recognition module is configured to: use a composite recognition model to identify objects in any set of the remote sensing image data sets, obtain several objects and their recognition confidence levels, and identify the objects with recognition confidence levels higher than a confidence threshold as infrastructure and classify them into the first infrastructure data set; wherein the composite recognition model includes multiple sub-recognition models for respectively identifying several specified types of infrastructure; and classify the objects with recognition confidence levels not higher than the confidence threshold into the suspected infrastructure data set.
[0014] Optionally, the outdoor analysis and recognition module is further configured to: determine the recognition confidence of each of the suspected infrastructures, and determine a first target quantity and a target spanning duration according to the level of the recognition confidence; based on the shooting moment of the remote sensing image dataset corresponding to the recognition confidence, screen out a first target quantity of other remote sensing image datasets according to the target spanning duration, and perform secondary confirmation on the suspected infrastructure based on multiple remote sensing images corresponding to each other remote sensing image dataset to obtain a secondary confirmation result, where the secondary confirmation result includes confirmed as infrastructure and confirmed as non-infrastructure.
[0015] Optionally, the outdoor analysis and recognition module is further configured to: determine a number of associated sub-recognition models corresponding to the sub-recognition model of the suspected infrastructure; where the similarity between each of the associated sub-recognition models and the type of infrastructure used for recognition by the sub-recognition model is higher than a similarity threshold; use each of the associated sub-recognition models to respectively recognize multiple remote sensing images, and fuse the recognition results of each to obtain the secondary confirmation result.
[0016] Optionally, the indoor analysis and recognition module is configured to: request to obtain a set of video data corresponding to the indoor area, analyze the change amplitude of the item positions from the set of video data, and determine a second target quantity according to the change amplitude level to which the item position change amplitude belongs; request to obtain multiple sets of video data corresponding to the second target quantity of the indoor area.
[0017] The present invention also discloses an electronic device, including: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, where the processor executes the computer program to implement the method as described in any one of the foregoing.
[0018] The present invention also discloses a computer storage medium, where the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method as described in any one of the foregoing.
[0019] The present invention also discloses a computer program product, where the computer program product contains computer code, and when the computer code is executed by a processor of an electronic device, the method as described in any one of the foregoing is implemented.
[0020] The above solution of the present invention performs secondary recognition on outdoor targets of suspected infrastructures, and uses a dynamic request mechanism to obtain video data of indoor areas, realizing accurate recognition of indoor and outdoor infrastructures, and further realizing rapid creation of a super-large-scale and high-precision scene. Description of the Drawings
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.
[0022] Figure 1 It is a schematic flowchart of a method for quickly constructing a large-scale indoor and outdoor high-precision scene disclosed in an embodiment of the present invention.
[0023] Figure 2 It is a schematic structural diagram of a composite recognition model disclosed in an embodiment of the present invention.
[0024] Figure 3 It is a schematic structural diagram of a system for quickly constructing a large-scale indoor and outdoor high-precision scene disclosed in an embodiment of the present invention. Specific Embodiments
[0025] The following specific embodiments illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.
[0026] In addition, the technical features involved in different implementation manners of the present application described below can be combined with each other as long as they do not conflict with each other.
[0027] Regarding the above technical problems, as Figure 1 shown, an embodiment of the present invention discloses a method for quickly constructing a large-scale indoor and outdoor high-precision scene. The method includes the following steps: S101, receiving a target area corresponding to a high-precision scene input by a user, and dividing the target area into an outdoor area and an indoor area based on the map information corresponding to the target area.
[0028] For example, if a user wants to construct a high-precision scene of a large commercial area and its surrounding office buildings in a certain city, the user inputs the name or coordinate range of this area on the cloud platform at this time. The cloud then obtains the electronic map information of this area, and based on the building outlines, roads and other markings in the electronic map, divides the open-air parts such as streets and squares in the commercial area into outdoor areas, and divides the interiors of shopping malls and office buildings in the commercial area into indoor areas.
[0029] S102. Obtain multiple groups of remote sensing image datasets of the outdoor area corresponding to different acquisition time periods, identify a first infrastructure dataset and a suspected infrastructure dataset from the multiple groups of remote sensing image datasets, conduct secondary confirmation on each suspected infrastructure in the suspected infrastructure dataset, and update the first infrastructure dataset based on the secondary confirmation result to obtain a second infrastructure dataset.
[0030] For the above urban block, the cloud obtains remote sensing images of this area taken at different time periods (satellite or drone images), such as images taken during weekday daytime, weekend daytime, and at night. Through image recognition algorithms, roads, buildings, etc. are initially identified from these images to form a first infrastructure dataset. However, during the recognition process, some temporary objects (such as a stage temporarily built with a color similar to the road surface, a large truck temporarily parked by the roadside for loading and unloading a sculpture) may be misidentified as infrastructure, and these misjudged objects constitute the suspected infrastructure dataset.
[0031] Then, the cloud conducts secondary confirmation on the suspected infrastructure by analyzing the changes in the position, status, etc. of the objects in the images at different time periods (such as the position of the truck changing at different time periods while the position of the real road infrastructure remains unchanged). Finally, the objects in the suspected infrastructure dataset that are confirmed to belong to the infrastructure are added to the first infrastructure dataset to obtain a more accurate second infrastructure dataset.
[0032] S103. Use a dynamic request mechanism to request and obtain video data corresponding to the indoor area, and identify a third infrastructure dataset from the video data.
[0033] For the interior of shopping malls and office buildings in the business district, the cloud uses a dynamic request mechanism to communicate and connect with relevant terminal devices in these indoor areas to request them to provide corresponding video data. From the obtained video data, using image recognition technology, infrastructure such as counters, aisles, elevators in the shopping mall, and office area divisions, corridors, stairs, etc. in the office building are identified to form a third infrastructure dataset.
[0034] Due to the adoption of the dynamic request mechanism, different video image acquisition strategies can be decided according to the actual situations of different indoor areas, thereby achieving high-precision scene construction of the indoor scene.
[0035] S104. Perform coordinate mapping and position alignment on the second infrastructure dataset and the third infrastructure dataset to obtain an ultra-large-scale high-precision indoor and outdoor scene.
[0036] The cloud maps the coordinates of the second infrastructure dataset containing information such as the three-dimensional contours and positions of outdoor roads and buildings, and the third infrastructure dataset containing information such as the layout of indoor shopping mall counters and the internal structure of office buildings based on a unified geographic coordinate system, and then performs alignment processing. For example, the geographical coordinates of the mall entrance outdoors are associated with the coordinates at the indoor entrance of the mall, and the outdoor position of the office building is aligned with the corresponding positions on each floor inside the office building, so as to construct a super-large-scale indoor and outdoor high-precision scene of the entire urban block. Users can obtain the complete high-precision scene of the target area through the cloud for navigation of vehicles, robots (such as food delivery robots), etc.
[0037] In the above solution of the present invention, secondary identification is performed on outdoor targets suspected of being infrastructure, and a dynamic request mechanism is used to obtain video data of indoor areas, realizing the accurate identification of indoor and outdoor infrastructure, and further realizing the rapid creation of a super-large-scale high-precision scene.
[0038] Optionally, the identifying the first infrastructure dataset and the suspected infrastructure dataset from multiple groups of the remote sensing image datasets includes: using a composite identification model to identify objects in any group of the remote sensing image datasets, obtaining several objects and their identification confidence levels, and identifying the objects with the identification confidence level higher than the confidence threshold as infrastructure and classifying them into the first infrastructure dataset; wherein, the composite identification model includes multiple sub-identification models for respectively identifying several specified types of infrastructure; and classifying the objects with the identification confidence level not higher than the confidence threshold into the suspected infrastructure dataset.
[0039] In this embodiment, as Figure 2 shown, the composite identification model is designed to identify multiple specified types of infrastructure, which may include several sub-identification models. Each sub-identification model is used to identify one type of infrastructure. For example, sub-identification model A is used to identify roads, and sub-identification model B is used to identify buildings. When processing the remote sensing images of the urban block, the sub-identification models only identify these pre-set specified types of infrastructure.
[0040] When processing the remote sensing images of the above urban blocks, for any set of acquired remote sensing image datasets (such as the set of images taken during the day on weekdays), input them into the composite recognition model. The composite recognition model analyzes and identifies each object in the image (such as roads, buildings, vehicles, temporary structures, etc.) and assigns a recognition confidence level to each identified object. This recognition confidence level reflects the degree of certainty of the model that the object is correctly identified as a certain object. For example, it is represented by a numerical value, and the higher the value, the greater the possibility that the model believes the object is correctly identified. For example, for an office building in the image, the composite recognition model gives a corresponding recognition confidence level by analyzing its shape, texture, relationship with the surrounding environment, and other features.
[0041] When the recognition confidence level of a certain object given by the composite recognition model is higher than this set confidence threshold, it can be more reliably considered that the object is an infrastructure. For example, in the remote sensing image of an urban block, after a certain road is analyzed by the composite recognition model and its recognition confidence level is high, exceeding the confidence threshold, then this road is identified as an infrastructure, and its relevant information (such as location, shape, etc.) is included in the first infrastructure dataset. In this way, the first infrastructure dataset initially contains the information of those objects that the model believes are very likely to be infrastructures.
[0042] If the recognition confidence level given by the composite recognition model for a certain object is not higher than the set confidence threshold, it means that the model is not sure whether the object is an infrastructure. For example, in the urban block image, a temporarily built stage with a color close to the road surface, or a large truck temporarily parked by the roadside for loading and unloading sculptures. After these objects are recognized by the model, their recognition confidence levels may be low and do not reach the confidence threshold. At this time, these objects are included in the suspected infrastructure dataset. Subsequently, the objects in the suspected infrastructure dataset will be reconfirmed to further determine whether they really belong to the infrastructure.
[0043] Optionally, the reconfirmation of each suspected infrastructure in the suspected infrastructure dataset includes: determining the recognition confidence level of each suspected infrastructure, and determining the first target quantity and the target spanning duration according to the level of the recognition confidence level; based on the shooting time of the remote sensing image dataset corresponding to the recognition confidence level as a reference, screening out the first target quantity of other remote sensing image datasets according to the target spanning duration, and reconfirming the suspected infrastructure based on the corresponding multiple remote sensing images in each other remote sensing image dataset to obtain the reconfirmation result, and the reconfirmation result includes being confirmed as an infrastructure and being confirmed as not an infrastructure.
[0044] In this embodiment, for each object that has been classified into the suspected infrastructure dataset (such as a temporary stage in an urban block image, a truck loading and unloading a sculpture, etc.), two key parameters are first determined based on the previously obtained recognition confidence levels: the first target quantity and the target crossing duration.
[0045] Generally speaking, the lower the recognition confidence level, the more uncertain the model's judgment of the object. In this case, a larger first target quantity (i.e., more groups of other remote sensing image datasets need to be referenced, such as 10 or 20) and a longer target crossing duration (i.e., a larger time span between the shooting times of the referenced remote sensing images, such as 1 hour or 2 hours) are set to more accurately determine the true nature of the object. On the contrary, for suspected infrastructure with a relatively high recognition confidence level (but still lower than the confidence threshold), a smaller first target quantity and a shorter target crossing duration are set.
[0046] After determining the first target quantity and the target crossing duration corresponding to each suspected infrastructure, the shooting moment of the remote sensing image corresponding to when the suspected infrastructure was initially recognized is used as the reference time point. Then, according to the determined target crossing duration, the first target quantity of other remote sensing image datasets is selected from the existing remote sensing image datasets taken at different time periods.
[0047] After selecting the first target quantity of other remote sensing image datasets, multiple remote sensing images related to the suspected infrastructure in each dataset are analyzed. For example, for the truck loading and unloading the sculpture, the remote sensing images 1 hour, 2 hours... before and 1 hour, 2 hours... after the initial remote sensing image (the target crossing duration is 1 hour) are viewed, and the position, state, and relationship with the surrounding environment of the truck at these different times are observed. If it is found that the position of the truck changes significantly at different time points and does not conform to the static characteristics that infrastructure should have, it can be confirmed that it is non-infrastructure; if it is found that the position, shape, and other characteristics of the object remain unchanged in multiple groups of images and conform to the characteristics of a certain infrastructure, it can be confirmed that it is infrastructure. Through such comprehensive analysis based on multiple remote sensing images at different times, a secondary confirmation result for each suspected infrastructure is obtained, and it is clearly classified as infrastructure or non-infrastructure, thereby further improving the accurate recognition of outdoor infrastructure.
[0048] Among them, the remote sensing images corresponding to the same suspected infrastructure in different remote sensing image datasets, in addition to having different shooting times, will also have slight differences in shooting angles, which can be used to analyze whether the suspected infrastructure is infrastructure or a temporary object.
[0049] Optionally, perform secondary confirmation on the suspected infrastructure based on multiple corresponding remote sensing images in each other remote sensing image dataset to obtain a secondary confirmation result, including: determining a number of associated sub-identification models corresponding to the sub-identification model of the suspected infrastructure; wherein, the similarity between each of the associated sub-identification models and the type of infrastructure to be identified by the sub-identification model is higher than a similarity threshold; using each of the associated sub-identification models to identify multiple remote sensing images respectively, and fusing the identification results of each to obtain the secondary confirmation result.
[0050] In this embodiment, although the multiple sub-identification models included in the composite identification model are for different types of infrastructure, some infrastructure has a certain visual similarity in remote sensing images. For example, a water tower and a signal tower set on a road usually both have a relatively high columnar or tower-like shape and have a high visual similarity in remote sensing images.
[0051] In view of this, the present invention determines these sub-identification models as the associated sub-identification models corresponding to this sub-identification model, and then schedules these associated sub-identification models and this sub-identification model to identify the multiple obtained remote sensing images respectively, and fuses the identification results of the associated sub-identification models and the sub-identification model to obtain the final secondary confirmation result. For example, the identification result of the sub-identification model is: infrastructure - sculpture, identification confidence 90%; the identification result of the associated sub-identification model 1 is: infrastructure - A, identification confidence 65%; the identification result of the associated sub-identification model 2 is: infrastructure - B, identification confidence 73%; the identification result of the associated sub-identification model 3 is: infrastructure - C, identification confidence 86%. Integrate the identification confidences of each associated sub-identification model, for example, calculate the mean, median or maximum value, and the integrated identification confidence obtained by fusion is, for example, 86% (maximum value). At this time, according to the preset comparison relationship, the auxiliary coefficient corresponding to 86% is determined to be 0.9 (the auxiliary coefficient is positively correlated with the integrated value). Correspondingly, the identification confidence of this suspected infrastructure is 90% * 0.9 = 81%, and 81% is higher than the confidence threshold of 80%. At this time, it is determined that the secondary confirmation result of this suspected infrastructure is confirmed as infrastructure.
[0052] Therefore, while using multiple remote sensing images to perform secondary confirmation on the suspected infrastructure, the present invention further refers to the associated sub-identification models for identifying similar infrastructure to assist in analyzing whether this suspected infrastructure belongs to the corresponding type of infrastructure, thereby improving the analysis accuracy.
[0053] Optionally, the step of requesting to obtain video data corresponding to the indoor area by using the dynamic request mechanism includes: requesting to obtain a set of video data corresponding to the indoor area, analyzing the change range of the item positions from this set of video data, and determining a second target quantity according to the change range level to which the item position change range belongs; requesting to obtain multiple sets of video data corresponding to the indoor area and corresponding to the second target quantity.
[0054] In this embodiment, when it is necessary to obtain video data of an indoor area (such as the interior of a shopping mall) to construct a high-precision scene, the cloud first requests a set of video data from relevant video acquisition devices (such as surveillance cameras) in this indoor area. Then, using computer vision and image analysis technologies, various items in the video (such as goods, shopping carts, pedestrians, etc. in the shopping mall) are identified, and the position changes of these items in the video frames are tracked to calculate the overall position change range of the items in the video frame. For example, most of the goods in the video frame are located in area A in the first period and mainly located in area B in the second period, which indicates that there are obvious situations of goods entering, stacking, and being transported away in this video frame, such as a storage area. In this case, since the goods handling is relatively frequent, it is easier to quickly determine the passable areas in this indoor area. At this time, fewer sets of video data are requested.
[0055] On the contrary, if the goods in the video frame have always been concentrated in area A, it indicates that the goods handling is not frequent. At this time, it is very difficult to determine the passable areas in this indoor area from this set of video data because the goods always block and occupy the passable areas. At this time, more sets of video data are requested to promote obtaining video frames in which the goods in a certain passable area are removed, and then the boundaries of this passable area are extracted.
[0056] The boundaries of the passable areas obtained by the above method can be used for the passage navigation of various robots, for example.
[0057] Such as Figure 3As shown in the figure, an embodiment of the present invention also discloses a system for rapidly constructing a large-scale indoor and outdoor high-precision scene. The system includes a receiving and preprocessing module, an outdoor analysis and recognition module, an indoor analysis and recognition module, and a scene integration and generation module. The receiving and preprocessing module receives the target area corresponding to the high-precision scene input by the user, and divides the target area into an outdoor area and an indoor area based on the map information corresponding to the target area. The outdoor analysis and recognition module acquires multiple groups of remote sensing image datasets of the outdoor area corresponding to different acquisition time periods, identifies a first infrastructure dataset and a suspected infrastructure dataset from the multiple groups of remote sensing image datasets, performs secondary confirmation on each suspected infrastructure in the suspected infrastructure dataset, and updates the first infrastructure dataset based on the secondary confirmation result to obtain a second infrastructure dataset. The indoor analysis and recognition module requests and acquires video data corresponding to the indoor area by using a dynamic request mechanism, and identifies a third infrastructure dataset from the video data. The scene integration and generation module performs coordinate mapping and position alignment on the second infrastructure dataset and the third infrastructure dataset to obtain a large-scale indoor and outdoor high-precision scene.
[0058] An embodiment of the present invention also discloses an electronic device, including: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, where the processor executes the computer program to implement the method as described in the foregoing embodiment.
[0059] An embodiment of the present invention also discloses a computer storage medium, where the computer storage medium stores a computer program, and the computer program is executed by a processor to implement the method as described in the foregoing embodiment.
[0060] The above-mentioned computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the above. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0061] An embodiment of the present invention also discloses a computer program product, where the computer program product contains computer code, and when the computer code is executed by a processor of an electronic device, it implements the method as described in the foregoing embodiment.
[0062] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0063] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for rapidly constructing a high-precision indoor and outdoor scene on a super-large scale, characterized in that: The method includes the following steps: receiving a target area corresponding to a high-precision scenario input by a user, and dividing the target area into an outdoor area and an indoor area based on map information corresponding to the target area; acquiring multiple groups of remote sensing image datasets of the outdoor area corresponding to different acquisition time periods, identifying a first infrastructure dataset and a suspected infrastructure dataset from the multiple groups of remote sensing image datasets, performing secondary confirmation on each suspected infrastructure in the suspected infrastructure dataset, and updating the first infrastructure dataset based on the secondary confirmation result to obtain a second infrastructure dataset; using a dynamic request mechanism to request and obtain video data corresponding to the indoor area, and identifying a third infrastructure dataset from the video data; performing coordinate mapping and position alignment on the second infrastructure dataset and the third infrastructure dataset to obtain an ultra-large-scale indoor and outdoor high-precision scenario; performing secondary confirmation on each suspected infrastructure in the suspected infrastructure dataset, including: determining the recognition confidence of each suspected infrastructure, and determining a first target quantity and a target spanning duration according to the level of the recognition confidence; taking the shooting moment of the remote sensing image dataset corresponding to the recognition confidence as a reference, screening out multiple groups of other remote sensing image datasets of the first target quantity according to the target spanning duration, and performing secondary confirmation on the suspected infrastructure based on multiple remote sensing images corresponding to each other remote sensing image dataset to obtain a secondary confirmation result, where the secondary confirmation result includes being confirmed as infrastructure and being confirmed as non-infrastructure; 2. The rapid construction method for a very large-scale indoor and outdoor high-precision scene according to claim 1, wherein: identifying a first infrastructure dataset and a suspected infrastructure dataset from multiple groups of remote sensing image datasets, including: using a composite recognition model to identify objects in any group of remote sensing image datasets, obtaining several objects and their recognition confidences, determining the objects with recognition confidences higher than a confidence threshold as infrastructure and classifying them into the first infrastructure dataset; where the composite recognition model includes multiple sub-recognition models for respectively identifying several specified types of infrastructure; classifying the objects with recognition confidences not higher than the confidence threshold into the suspected infrastructure dataset; 3. A method for rapidly constructing a very large-scale indoor and outdoor high-precision scene according to claim 2, characterized in that: performing secondary confirmation on the suspected infrastructure based on multiple remote sensing images corresponding to each other remote sensing image dataset to obtain a secondary confirmation result, including: determining several associated sub-recognition models corresponding to the sub-recognition model of the suspected infrastructure; where the similarity between each associated sub-recognition model and the type of infrastructure used for recognition by the sub-recognition model is higher than a similarity threshold; using each associated sub-recognition model to respectively identify multiple remote sensing images, and fusing each recognition result to obtain the secondary confirmation result; 4. A method for rapidly constructing a large-scale indoor and outdoor high-precision scene according to claim 1, characterized in that: Using a dynamic request mechanism to request and obtain video data corresponding to the indoor area, including: requesting to obtain a set of video data corresponding to the indoor area, analyzing the magnitude of item position changes from this set of video data, determining a second target quantity according to the change magnitude level to which the item position change magnitude belongs; requesting to obtain multiple sets of video data corresponding to the second target quantity for the indoor area.
5. A system for rapidly constructing a large-scale indoor and outdoor high-precision scene, characterized in that: The system includes a receiving and preprocessing module, an outdoor analysis and recognition module, an indoor analysis and recognition module, and a scene integration and generation module; The receiving and preprocessing module receives the target area corresponding to the high-precision scene input by the user, and divides the target area into an outdoor area and an indoor area based on the map information corresponding to the target area; The outdoor analysis and recognition module obtains multiple sets of remote sensing image data sets of the outdoor area corresponding to different acquisition time periods, identifies a first infrastructure data set and a suspected infrastructure data set from the multiple sets of remote sensing image data sets, performs secondary confirmation on each suspected infrastructure in the suspected infrastructure data set, and updates the first infrastructure data set based on the secondary confirmation result to obtain a second infrastructure data set; The indoor analysis and recognition module uses a dynamic request mechanism to request and obtain video data corresponding to the indoor area, and identifies a third infrastructure data set from the video data; The scene integration and generation module performs coordinate mapping and position alignment on the second infrastructure data set and the third infrastructure data set to obtain an ultra-large-scale indoor and outdoor high-precision scene; The outdoor analysis and recognition module is further used for: determining the recognition confidence of each of the suspected infrastructures, and determining a first target quantity and a target crossing duration according to the high or low of the recognition confidence; Taking the shooting moment of the remote sensing image data set corresponding to the recognition confidence as a reference, screening out multiple sets of other remote sensing image data sets of the first target quantity according to the target crossing duration, and performing secondary confirmation on the suspected infrastructure based on multiple remote sensing images corresponding to each other remote sensing image data set to obtain a secondary confirmation result, where the secondary confirmation result includes confirmed as infrastructure and confirmed as non-infrastructure.
6. The rapid construction system for a very large-scale indoor and outdoor high-precision scene according to claim 5, characterized in that: The outdoor analysis and recognition module is used for: using a composite recognition model to recognize object items in any set of the remote sensing image data sets, obtaining several object items and their recognition confidences, identifying the object items with recognition confidences higher than the confidence threshold as infrastructures and classifying them into the first infrastructure data set; where the composite recognition model includes multiple sub-recognition models for respectively recognizing several specified types of infrastructures; classifying the object items with recognition confidences not higher than the confidence threshold into the suspected infrastructure data set.
7. An electronic device, comprising: At least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, characterized in that: the processor executes the computer program to implement the method according to any one of claims 1-4.
8. A computer storage medium storing a computer program, characterized in that: The computer program is executed by the processor to implement the method according to any one of claims 1-4.
9. A computer program product, characterized in that: The computer program product contains computer code which, when executed by a processor of an electronic device, implements the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Air-ground integrated city ecological civilization managing system and method based on Beidou positioning
CN104021586A
Indoor and outdoor integrated map construction system based on vision and construction method thereof
CN110702078A