A global map construction method and device
Through the multi-object collaborative mapping system, visual feature fusion and full convolutional network extract image features are used to solve the problem of unstable mapping of bicycles in large-scale construction scenarios of construction machinery, and efficient and economical global map construction is achieved.
Patent Information
- Application Number
- CN202210190958.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-28
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-02-28
AI Technical Summary
In the prior art, bicycle environment modeling in the field of engineering machinery is difficult to achieve stable mapping effect in large scenarios, and the collaborative mapping method based on lidar is expensive and is not suitable for every engineering vehicle to be equipped.
The global map construction method is adopted, and the collaborative mapping system is used to fusion of local local maps through multiple target objects (such as engineering vehicles), and image features are extracted in combination with a full convolutional network to reduce dependence on lidar and improve mapping efficiency and accuracy.
It realizes that while reducing the cost of building global maps, it improves the efficiency and accuracy of map construction, and reduces the accumulated error of bicycles, and is suitable for multi-vehicle collaborative environment modeling scenarios.
Smart Images

Figure CN114648598B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method and device for constructing a global map. Background Art
[0002] With the development of the construction machinery industry in recent years, the demand for intelligence, unmanned operation, and digitization has become increasingly urgent. Intelligence and unmanned operation first require a comprehensive perception and modeling of the operating environment of construction machinery to facilitate tasks such as autonomous driving and operation path planning of subsequent engineering vehicles. The operating and driving environments of construction machinery often have a large global scope and geographical coherence and stability, which pose challenges to the full-scenario environmental modeling.
[0003] Currently, the modeling of the environment focuses on the research and use of single-vehicle mapping. By installing multiple environmental perception devices on a single vehicle, the surrounding environment is modeled and located during the continuous driving of the vehicle. Currently, the relatively mature SLAM (simultaneous localization and mapping) system based on lidar has been successfully applied to various scenarios. Unmanned aerial vehicles or indoor scene robots can perform autonomous mapping and path planning through a visual SLAM system. In the field of construction machinery, single-vehicle environmental modeling will continuously accumulate errors as the vehicle operates, and in the large-scale scenarios of construction machinery, single-vehicle environmental modeling cannot obtain a good stable mapping effect, and the stability of the system will be tested. The lidar-based collaborative mapping method is only applicable to the direct collaborative mapping of multiple vehicles where the environmental data collected by multiple vehicles are all lidar point clouds. The expensive lidar is not suitable for being installed on each engineering vehicle, and applying it to the engineering operation scenario will greatly increase the input cost of engineering vehicles. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method and device for constructing a global map to overcome the problem that it is difficult to balance mapping accuracy and economy in the panoramic map construction method in the prior art.
[0005] According to a first aspect, an embodiment of the present invention provides a method for constructing a global map, which is applied to any one of the target objects in a collaborative mapping system. The collaborative mapping system includes multiple target objects, and the method includes:
[0006] Obtain a first local partial map constructed by a first target object and a first environmental image collected by the first target object, and extract first visual features from the first environmental image;
[0007] Receive a second local partial map and second visual features sent by the remaining target objects, where the second visual features are visual features extracted by the current target object from a second environmental image collected by the current target object;
[0008] Based on the relationship between the first visual feature and the second visual features corresponding to the remaining target objects, the first local local map and the second local local maps corresponding to each of the remaining target objects are respectively fused to obtain the global map corresponding to the first target object.
[0009] Optionally, the fusing the first local local map and the second local local maps corresponding to each of the remaining target objects respectively based on the relationship between the first visual feature and the second visual features corresponding to the remaining target objects includes:
[0010] Performing feature point matching on the first visual feature and the second visual feature corresponding to the current target object;
[0011] Processing the first local local map and the second local local map corresponding to the current target object based on the feature point matching result to obtain the current global map of the first target object and the current target object;
[0012] Based on the current global maps of the first target object and all the remaining target objects, the global map corresponding to the first target object is obtained.
[0013] Optionally, the processing the first local local map and the second local local map based on the feature point matching result to obtain the current global map of the first target object and the current target object includes:
[0014] Judging whether the feature point matching result meets a preset feature point matching requirement;
[0015] When the feature point matching result meets the preset feature point matching requirement, performing map fusion at the overlapping positions of the first local local map and the second local local map to obtain the global map corresponding to the current target object.
[0016] Optionally, the method further includes:
[0017] When the feature point matching result does not meet the preset feature point matching requirement, deleting the second visual feature and the second local local map corresponding to the current target object, and continuing to perform feature point matching on the first visual feature and the second visual feature corresponding to the next target object.
[0018] Optionally, before performing feature point matching on the first visual feature and the second visual feature corresponding to the current target object, the method further includes:
[0019] Obtaining a first visual word bag model corresponding to the first visual feature and a second visual word bag model corresponding to the second visual feature corresponding to the current target object;
[0020] Calculate a first distance between the second visual word bag model and the first visual word bag model;
[0021] Determine whether the first distance is less than a distance threshold;
[0022] When the first distance is less than the distance threshold, perform feature point matching on the first visual feature and the second visual feature corresponding to the current target object.
[0023] Optionally, the method further includes:
[0024] When the first distance is not less than the distance threshold, splice maps at different positions of the first local local map and the second local local map corresponding to the current target object to obtain the current global map of the first target object and the current target object.
[0025] Optionally, the extracting the first visual feature from the first environmental image includes:
[0026] Extract the first visual feature from the first environmental image by using a fully convolutional network.
[0027] According to a second aspect, an embodiment of the present invention provides a global map construction device, which is applied to any target object in a collaborative mapping system. The collaborative mapping system includes multiple target objects. The device includes:
[0028] An acquisition module, configured to acquire a first local local map constructed by a first target object and a first environmental image collected, and extract a first visual feature from the first environmental image;
[0029] A first processing module, configured to receive a second local local map and a second visual feature sent by the remaining target objects. The second visual feature is a visual feature extracted by the current target object from a second environmental image collected by it;
[0030] A second processing module, configured to fuse the first local local map and the second local local map corresponding to each of the remaining target objects respectively based on the relationship between the first visual feature and the second visual features corresponding to the remaining target objects, to obtain the global map corresponding to the first target object.
[0031] According to a third aspect, an embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method described in the first aspect of the present invention and any optional manner thereof is implemented.
[0032] According to a fourth aspect, an embodiment of the present invention provides an electronic device, including:
[0033] A memory and a processor, which are communicatively connected to each other. Computer instructions are stored in the memory, and the processor executes the computer instructions to execute the method according to the first aspect of the present invention and any of its optional modes.
[0034] The technical solution of the present invention has the following advantages:
[0035] An embodiment of the present invention provides a global map construction method and apparatus, which are applied to any target object in a collaborative mapping system. The collaborative mapping system includes multiple target objects. By obtaining a first local partial map constructed by a first target object and a first environmental image collected, and extracting first visual features from the first environmental image; receiving second local partial maps and second visual features sent by the remaining target objects; based on the relationship between the first visual features and the second visual features corresponding to the remaining target objects, respectively fusing the first local partial map with the second local partial maps corresponding to each of the remaining target objects to obtain a global map corresponding to the first target object. Thus, by using the relationship between the visual features in the environmental images collected by each target object, the positional relationship between the local partial maps constructed by the remaining target objects and the local partial map of the first target object is determined, and the collaborative construction of the global map of the first target object is realized through the fusion of the local partial maps of multiple target objects and the local partial map of the first target object. While improving the mapping efficiency, it is not necessary to load a lidar, reducing the cost of global map construction, and further ensuring the accuracy of the final global map by using the visual features in the original environmental images collected by each target object. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0037] Figure 1 It is a flowchart of the global map construction method in the embodiment of the present invention;
[0038] Figure 2 It is a schematic structural diagram of the global map construction apparatus in the embodiment of the present invention;
[0039] Figure 3 It is a schematic structural diagram of the electronic device in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.
[0042] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and may also be the internal communication of two elements. It may be a wireless connection or a wired connection. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0043] The technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0044] Currently, the environmental modeling focuses on the research and use of single-vehicle mapping. By installing multiple environmental perception devices on a single vehicle, the surrounding environment is modeled and located during the continuous driving of the vehicle. Currently, the relatively mature SLAM system based on lidar has been successfully applied to various scenarios. UAVs or indoor scene robots can perform autonomous mapping and path planning through visual SLAM systems. In the field of construction machinery, the single-vehicle environmental modeling will continuously accumulate errors as the vehicle runs, and in the large-scale construction machinery scenarios, the single-vehicle environmental modeling cannot obtain a stable mapping effect well, and the stability of the system will be tested; while the SLAM system based on lidar is often costly and has poor economy.
[0045] Based on the above problems, an embodiment of the present invention provides a global map construction method, which is applied to any target object in a collaborative mapping system. The collaborative mapping system includes multiple target objects. In the embodiment of the present invention, the collaborative mapping system is composed of multiple engineering vehicles in the same large scene as target objects. It can be understood that the engineering vehicles mentioned here include both traditional engineering vehicles such as forklifts, dump trucks, container trucks, flatbed trucks, and tractor trucks, and engineering machinery and equipment such as excavators, pavers, mixers, and graders. Each vehicle serves as one of the nodes, and a complete single-node system runs inside the node: including an independent VSLAM module for local environmental modeling of a single vehicle, a communication link for inter-node communication, a module for collaborative mapping of adjacent nodes, etc. The vehicle is equipped with binocular or depth vision sensors for obtaining image information of the vehicle's surrounding environment, etc. In practical applications, the collaborative mapping system can also be composed of other vehicles or intelligent devices with collaborative mapping requirements, such as intelligent devices such as edge servers set at the field end. The present invention is not limited thereto.
[0046] As Figure 1 shown, this global map construction method is applied to the controller of any target object in the collaborative mapping system, and specifically includes the following steps:
[0047] Step S101: Obtain the first local partial map constructed by the first target object and the first environmental image collected, and extract the first visual features from the first environmental image.
[0048] Among them, the first target object is an engineering vehicle with a global map construction requirement, and the first environmental image is a key frame of the environmental image collected by the visual sensor carried by the engineering vehicle. The number of key frames can be flexibly set according to the modeling accuracy. The present invention is not limited thereto. Exemplarily, the engineering vehicle can perform local mapping and positioning using this key frame by running its carried VSLAM module to obtain the above first local partial map. In addition, in practical applications, the construction method of the local partial map of each engineering vehicle can also be obtained by using other existing single-vehicle-based map construction methods. Exemplarily, if a certain engineering vehicle is equipped with a lidar device, it can also construct a local partial map based on the lidar point cloud data collected by the lidar device. In the embodiment of the present invention, different engineering vehicles can construct local partial maps in different ways according to the sensor devices they carry. The present invention is not limited thereto.
[0049] Specifically, SLAM that only uses a camera as an external perception sensor is called visual SLAM, i.e., vSLAM. The camera has the advantages of rich visual information and low hardware cost. A classic vSLAM system generally includes four main parts: a front-end visual odometry, a back-end optimization, a loop closing detection, and a mapping. The main implementation process is as follows: Visual Odometry: Pose estimation with only visual input; Optimization: The back-end receives the camera poses measured by the visual odometry at different times, as well as the information of the loop closing detection, and optimizes them to obtain a globally consistent trajectory and map; Loop Closing: It means that during the map construction process of the robot, it detects whether a trajectory loop occurs through sensor information such as vision, that is, determines whether it enters the same historical location; Mapping: According to the estimated trajectory, a map corresponding to the task requirements is established. For the detailed process of how to use the VSLAM module to construct a local local map, reference can be made to the relevant descriptions in the prior art, and details will not be elaborated here.
[0050] Specifically, since the traditional ORB (Oriented Fast and Rotated Brief) feature extraction method cannot extract the deep semantic features of vision well, during the loop closing detection, that is, the above-mentioned loop closing detection, the image features based on ORB cannot be well matched. Especially in the scenario of multi-vehicle collaborative environment modeling, due to different environmental factors such as the viewing angle and distance of different vehicles for the same location, the loop closing detection error based on ORB feature matching is relatively large, which in turn affects the accuracy of the final global map. In the embodiments of the present invention, a fully convolutional network is used to extract the first visual features from the first environmental image. By using the fully convolutional network, during the image feature extraction process, it adapts to image rotation invariance, translation invariance, and scale invariance, better extracts the deep semantic features of the image, reduces the ineffective extraction of unnecessary low-level image features and the interference to the subsequent matching calculation, and further improves the accuracy of the finally constructed global map.
[0051] Step S102: Receive the second local local map and the second visual features sent by the other target objects.
[0052] Among them, the other target objects are other engineering vehicles in the collaborative mapping system, and the second visual features are the visual features extracted by the current target object from the second environmental image collected by its visual sensor.
[0053] Specifically, in the collaborative mapping system, each construction vehicle runs its onboard VSLAM module to perform local mapping and positioning using the key frames of the collected environmental images, obtaining a local partial map. The visual features corresponding to the collected environmental images are extracted using the visual feature extraction method of the above process, and the local partial map and visual features are sent to other construction vehicles through a communication link, so that each construction vehicle can obtain the local partial mapping and visual features of other construction vehicles. Through the construction of the local partial map by the construction vehicle locally and the extraction of the depth image in the key frame of the environmental image, the global map modeling of any construction vehicle can be realized. Moreover, this decentralized data processing method greatly reduces the computational processing volume, reduces the unnecessary consumption of computing resources, and further improves the efficiency of global map construction.
[0054] Step S103: Based on the relationship between the first visual feature and the second visual features corresponding to the remaining target objects, fuse the first local partial map and the second local partial maps corresponding to each of the remaining target objects respectively to obtain the global map corresponding to the first target object.
[0055] Specifically, since the visual features are extracted from the key frames of the environmental images collected by the construction vehicle, the relationship between the visual features corresponding to different construction vehicles reflects the positional relationship between different construction vehicles. Based on this positional relationship, the local partial maps corresponding to each construction vehicle can be fused to obtain a global map centered on a certain construction vehicle.
[0056] By performing the above steps, the global map construction method provided by the embodiment of the present invention determines the positional relationship between the local partial map constructed by the remaining target object and the local partial map of the first target object by using the relationship between the visual features in the environmental images collected by each target object. Through the fusion of the local partial maps of multiple target objects and the local partial map of the first target object, the collaborative construction of the global map of the first target object is realized. While improving the mapping efficiency, it does not require the installation of a lidar, reducing the cost of global map construction, and further ensuring the accuracy of the final global map by using the visual features in the original environmental images collected by each target object.
[0057] Specifically, in one embodiment, the above step S103 specifically includes the following steps:
[0058] Step S201: Perform feature point matching on the first visual feature and the second visual feature corresponding to the current target object.
[0059] Step S202: Process the first local partial map and the second local partial map corresponding to the current target object based on the feature point matching result to obtain the current global map of the first target object and the current target object.
[0060] Specifically, in one embodiment, step S202 specifically includes the following steps:
[0061] Step S301: Determine whether the feature point matching result meets the preset feature point matching requirements.
[0062] Specifically, the preset feature point matching requirement refers to the number of feature point matches. The specific quantity requirement can be flexibly set according to the global map construction accuracy requirement and speed requirement, and the present invention is not limited thereto.
[0063] Step S302: When the feature point matching result meets the preset feature point matching requirements, perform map fusion on the overlapping positions of the first local local map and the second local local map to obtain the global map corresponding to the current target object.
[0064] Specifically, by obtaining the coordinates of the successfully matched feature points in their corresponding environmental images, and using the conversion relationship between the visual sensor coordinate system and the world coordinate system on the corresponding engineering vehicle, the overlapping coordinates of the feature points in the world coordinate system are obtained. Then, the overlapping coordinates of the successfully matched feature points of two construction machinery in the world coordinate system are used to perform map fusion on the overlapping positions of the two local local maps.
[0065] Step S303: When the feature point matching result does not meet the preset feature point matching requirements, delete the second visual feature and the second local local map corresponding to the current target object, and continue to perform feature point matching on the first visual feature and the second visual feature corresponding to the next target object.
[0066] Specifically, only when the number of feature point matches reaches a certain value does it indicate that there is a certain overlapping area between the local local maps of two engineering vehicles. If the number of feature point matches is not met, it means that the overlapping area is too small, and the small number of matched feature points is prone to fusion errors, making it difficult to ensure the accuracy of local local map fusion. In the embodiment of the present invention, by deleting the data sent by other engineering vehicles that do not meet the preset feature point matching requirements, while avoiding fusion errors, it reduces the occupation of unnecessary computing resources, improves the map construction efficiency, and ensures the accuracy of the fused map by using a certain number of successfully matched feature points to fuse the two local local maps.
[0067] Step S203: Based on the current global maps of the first target object and all the remaining target objects, obtain the global map corresponding to the first target object.
[0068] Specifically, since all current global maps are established based on the local partial map of the first target object, that is, all current global maps contain the local partial map of the first target object, directly taking the union of all current global maps can obtain the global map corresponding to the first target object. Similarly, by adopting the above process, the global map of any target object, that is, the position of any engineering vehicle, in the collaborative mapping system can be obtained.
[0069] Specifically, in one embodiment, before performing the above step S201, the global map construction method provided by the embodiment of the present invention further includes the following steps:
[0070] Step S401: Obtain the first visual word bag model corresponding to the first visual feature and the second visual word bag model corresponding to the second visual feature of the current target object.
[0071] Specifically, the visual word bag model is used to represent the visual features of the key frames of the environmental images. In practical applications, the acquisition method of the visual word bag model of each target object can be processed by each target object and then sent to the target object that is performing global map modeling through a communication link, or the target object that is performing global map modeling can use the received visual features to obtain the corresponding visual word bag model. The present invention is not limited thereto.
[0072] Step S402: Calculate the first distance between the second visual word bag model and the first visual word bag model.
[0073] Specifically, the above first distance is obtained by calculating the Euclidean distance between the two visual word bag models.
[0074] Step S403: Determine whether the first distance is less than the distance threshold.
[0075] Among them, the distance threshold can be dynamically configured. The distance threshold adopts a normalization operation. During the test, by adjusting the size of the distance threshold, the matching strictness can be dynamically adjusted to obtain the best mapping effect. The initial threshold can be set between 0.3 and 0.5. The present invention only takes this as an example and is not limited thereto.
[0076] Specifically, when the first distance is less than the distance threshold, perform the above step S201. When the first distance is not less than the distance threshold, splice the maps of different positions of the first local partial map and the second local partial map corresponding to the current target object to obtain the current global map of the first target object and the current target object.
[0077] In practical applications, if the distance between two visual bag-of-words models is larger, it indicates that the possibility of overlap between the corresponding local local maps is smaller. Since continuous feature point matching calculations are required, a large amount of computing resources are consumed. By setting a distance threshold, two local local maps with no overlap can be directly stitched according to their respective positions without feature point matching. Thus, by using the visual bag-of-words model for pre-computation of feature matching of different target objects, unnecessary specific matching calculations can be reduced, unnecessary consumption of computing resources can be reduced, and the global map modeling efficiency can be further improved.
[0078] Next, a specific application example will be used to explain in detail the global map construction method provided by the embodiments of the present invention.
[0079] The controller of each engineering vehicle in the collaborative mapping system is represented by an Agent.
[0080] Each Agent sender accesses a visual sensor and runs the VSLAM system. In the VSLAM system, the internal storage consumption is reduced by restricting the number of key frames. After obtaining the image input, the visual odometer is run, feature extraction is performed on the key frames, the feature extraction process uses a fully convolutional network to replace the ORB feature extraction, and at the same time, the visual bag-of-words model is used to represent the features of the key frames, and the visual features and bag-of-words vectors of the key frames are locally cached.
[0081] The Agent performs mapping and positioning of the local local map based on the above single-vehicle VSLAM system at the sender, and saves the local local map and the global map after coordinate transformation. Before effective communication with other Agents, the local local map and the global map data are the same.
[0082] The Agent establishes communication with neighboring Agents through the communication module. Through the communication link, the Agent sends the key frame bag-of-words vector to neighboring other Agents for pre-computation of the visual features of the key frame images, and at the same time sends the corresponding key frame visual feature point data and local local map information.
[0083] The target Agent at the receiver receives the key frame bag-of-words vector of other Agents through the communication link, and performs pre-matching calculation of the visual features of the key frames by calculating the Euclidean distance between the received key frame bag-of-words vector and the locally cached key frame bag-of-words vector to obtain the result of whether the set distance threshold is satisfied. If the set threshold is satisfied, further feature point matching calculation is performed on the key frame visual feature data sent by other Agents and the key frame visual feature data cached at this node, otherwise it is determined to directly perform map stitching at different positions of the global map.
[0084] The target Agent is at the receiving end. By performing feature point matching, if the set threshold of the number of feature point matches is not met, the visual features of other Agent key frames cached in the previous step and the corresponding local partial maps are destroyed. If the set threshold is met, loop closure detection is triggered to perform map fusion at the overlapping positions of the global map. Loop closure detection is a prior art, and the specific implementation process will not be elaborated here.
[0085] Each Agent dynamically maintains and updates the local global map. When a single Agent performs a rescan of a repeated location or multiple Agents fuse maps at the same location, the loop closure detection mechanism is triggered.
[0086] The following uses a specific embodiment to illustrate how multiple construction machinery uses the technical solution of the present invention to perform collaborative environment modeling in a large scene. In the scenario of an automated forklift transporting goods, multiple automated forklifts need to complete the task of loading goods at point A, passing through an intermediate driving stage, unloading goods at point B, and then returning to point A to repeat the above steps in a relatively fixed park; each forklift is an independent Agent. By installing a visual sensor, it obtains the surrounding environment image information. The vehicle-mounted controller provides a computing unit and a storage unit, and the 5G module of the controller provides a communication function. Each forklift runs a collaborative mapping system internally, and the system includes a sending end and a receiving end; at the sending end, the VSLAM system based on the fully convolutional network in the above technical solution runs independently to perform local map modeling and global map modeling locally, continuously model the surrounding environment where the forklift travels, and continuously update the local partial map and the global map; at the receiving end, the environmental collaborative modeling of different forklift nodes runs synchronously. When a forklift has other adjacent forklift nodes, a connection is established through the communication module to perform mutual communication between different forklift nodes. Each forklift node performs pre-computation of key frame feature matching through the visual bag-of-words vectors of mutual communication, and then judges whether further image visual feature matching calculation is required according to the pre-computation result; if not, directly perform map stitching of the multi-forklift global map at different position points at the receiving end of each forklift. If required, use the received image visual feature data to perform specific visual feature matching calculation. According to the set threshold, if the set threshold is met, perform map fusion of the multi-forklift global map at the same position point; if the set threshold is not met, delete the received image visual features and the local partial map; through the collaborative environment modeling of different forklifts, a single forklift can obtain the map information of the unpassed area, and at the same time, multi-vehicle collaboration reduces the cumulative error of a single vehicle over a long time and long distance. Through the vision-based VSLAM system and multi-object collaborative environment modeling, the configuration cost is reduced, the global mapping efficiency is improved, the cumulative error of a single vehicle is reduced, and the ultra-long-range global map information is quickly constructed inside different target objects.
[0087] By performing the above steps, the global map construction method provided by the embodiments of the present invention determines the positional relationship between the local partial maps constructed by the remaining target objects and the local partial map of the first target object by utilizing the relationship between the visual features in the environmental images collected by each target object. Through the fusion of the local partial maps of multiple target objects and the local partial map of the first target object, the collaborative construction of the global map of the first target object is achieved. While improving the mapping efficiency, it is not necessary to load a lidar, reducing the global map construction cost. Moreover, by utilizing the visual features in the original environmental images collected by each target object, the accuracy of the final global map is further ensured.
[0088] Embodiments of the present invention also provide a global map construction device, as Figure 2 shown. The global map construction device specifically includes:
[0089] An acquisition module 101, configured to acquire a first local partial map constructed by a first target object and a first environmental image collected, and extract first visual features from the first environmental image. For detailed content, refer to the detailed description of step S101 above, and details will not be elaborated here.
[0090] A first processing module 102, configured to receive a second local partial map and second visual features sent by the remaining target objects, where the second visual features are visual features extracted by the current target object from a second environmental image collected by it. For detailed content, refer to the detailed description of step S102 above, and details will not be elaborated here.
[0091] A second processing module 103, configured to fuse the first local partial map and the second local partial maps corresponding to the remaining target objects respectively based on the relationship between the first visual features and the second visual features corresponding to each of the remaining target objects, to obtain a global map corresponding to the first target object. For detailed content, refer to the detailed description of step S103 above, and details will not be elaborated here.
[0092] Through the collaborative cooperation of the above various components, the global map construction device provided by the embodiments of the present invention determines the positional relationship between the local partial maps constructed by the remaining target objects and the local partial map of the first target object by utilizing the relationship between the visual features in the environmental images collected by each target object. Through the fusion of the local partial maps of multiple target objects and the local partial map of the first target object, the collaborative construction of the global map of the first target object is achieved. While improving the mapping efficiency, it is not necessary to load a lidar, reducing the global map construction cost. Moreover, by utilizing the visual features in the original environmental images collected by each target object, the accuracy of the final global map is further ensured.
[0093] As Figure 3As shown in the figure, an embodiment of the present invention further provides an electronic device, which may include a processor 901 and a memory 902. The processor 901 and the memory 902 may be connected by a bus or other means. Figure 3 Taking the connection by bus as an example.
[0094] The processor 901 may be a central processing unit (CPU). The processor 901 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or combinations of the above types of chips.
[0095] As a non-transitory computer-readable storage medium, the memory 902 can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the embodiments of the present invention. The processor 901 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory 902, that is, the above methods are implemented.
[0096] The memory 902 may include a program storage area and a data storage area. Among them, the program storage area may store an operating device and application programs required for at least one function; the data storage area may store data created by the processor 901, etc. In addition, the memory 902 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 902 may optionally include a memory remotely provided with respect to the processor 901, and these remote memories may be connected to the processor 901 through a network. Examples of the above networks include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0097] One or more modules are stored in the memory 902 and, when executed by the processor 901, execute the above methods.
[0098] For specific details of the above server, reference may be made to the corresponding relevant descriptions and effects in the above method embodiments for understanding, and details are not described herein again.
[0099] Those skilled in the art can understand that to implement all or part of the processes in the above-described embodiment methods, it can be completed by instructing relevant hardware through a computer program. The implemented program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above various methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.
[0100] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A global map construction method, which is applied to any target object in a visual SLAM collaborative mapping system, and the collaborative mapping system includes multiple target objects, characterized in that, The method includes: Obtaining a first local partial map constructed by a first target object and a first environmental image collected, and extracting first visual features from the first environmental image; Receiving a second local partial map and second visual features sent by the remaining target objects, where the second visual features are visual features extracted by the current target object from a second environmental image collected by it; Based on the relationship between the first visual features and the second visual features corresponding to the remaining target objects, respectively fusing the first local partial map and the second local partial maps corresponding to each of the remaining target objects to obtain a global map corresponding to the first target object; The respectively fusing the first local partial map and the second local partial maps corresponding to each of the remaining target objects based on the relationship between the first visual features and the second visual features corresponding to the remaining target objects includes: Performing feature point matching on the first visual features and the second visual features corresponding to the current target object; Processing the first local partial map and the second local partial map corresponding to the current target object based on the feature point matching result to obtain a current global map of the first target object and the current target object; Based on the current global maps of the first target object and all the remaining target objects, obtaining a global map corresponding to the first target object; The processing the first local partial map and the second local partial map based on the feature point matching result to obtain a current global map of the first target object and the current target object includes: Judging whether the feature point matching result meets a preset feature point matching requirement; When the feature point matching result meets the preset feature point matching requirement, performing map fusion at the overlapping positions of the first local partial map and the second local partial map to obtain a global map corresponding to the current target object.
2. The method according to claim 1, wherein It further includes: When the feature point matching result does not meet the preset feature point matching requirement, deleting the second visual features and the second local partial map corresponding to the current target object, and continuing to perform feature point matching on the first visual features and the second visual features corresponding to the next target object.
3. The method according to claim 1, wherein Before performing feature point matching on the first visual features and the second visual features corresponding to the current target object, the method further includes: Obtaining a first visual word bag model corresponding to the first visual features and a second visual word bag model corresponding to the second visual features corresponding to the current target object; Calculating a first distance between the second visual word bag model and the first visual word bag model; Judging whether the first distance is less than a distance threshold; When the first distance is less than the distance threshold, performing feature point matching on the first visual features and the second visual features corresponding to the current target object.
4. The method according to claim 3, wherein It further includes: When the first distance is not less than the distance threshold, performing map stitching at different positions on the first local partial map and the second local partial map corresponding to the current target object to obtain a current global map of the first target object and the current target object.
5. The method according to claim 1, wherein The extracting first visual features from the first environmental image includes: Extract first visual features from the first environmental image using a fully convolutional network.
6. A global map construction device is applied to any one of the target objects in a visual SLAM collaborative mapping system, and the collaborative mapping system includes multiple target objects, and is characterized in that, The device includes: An acquisition module, configured to acquire a first local local map constructed by a first target object and a first environmental image collected, and extract first visual features from the first environmental image; A first processing module, configured to receive a second local local map and second visual features sent by other target objects, where the second visual features are visual features extracted by a current target object from a second environmental image collected by the current target object; A second processing module, configured to fuse the first local local map with the second local local map corresponding to each other target object respectively based on the relationship between the first visual features and the second visual features corresponding to other target objects, to obtain a global map corresponding to the first target object; the fusing the first local local map with the second local local map corresponding to each other target object respectively based on the relationship between the first visual features and the second visual features corresponding to other target objects includes: performing feature point matching on the first visual features and the second visual features corresponding to the current target object; processing the first local local map and the second local local map corresponding to the current target object based on the feature point matching result, to obtain a current global map of the first target object and the current target object; obtaining a global map corresponding to the first target object based on the current global maps of the first target object and all other target objects; the processing the first local local map and the second local local map based on the feature point matching result, to obtain a current global map of the first target object and the current target object includes: determining whether the feature point matching result meets a preset feature point matching requirement; when the feature point matching result meets the preset feature point matching requirement, performing map fusion at the overlapping positions of the first local local map and the second local local map, to obtain a global map corresponding to the current target object.
7. An electronic device, characterized in that, Including: A memory and a processor, where the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Method for collaborative mapping and locating of multiple robots for large-scale environment
CN106272423A
Map construction method and device, SLAM system and storage medium
CN112446845A