Model training method, visual positioning method, device and apparatus
By training deep neural network models in different environments and using image and sensor data to determine anchor samples, positive samples, and negative samples, the accuracy problem of visual positioning algorithms in extreme environments is solved, and accurate visual positioning in variable environments is achieved.
Patent Information
- Application Number
- CN202110518372.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-12
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2041-05-12
AI Technical Summary
Visual positioning algorithms become less accurate under extreme environmental changes, leading to reduced control accuracy and safety of autonomous vehicles.
By acquiring image and sensor data from multiple different environments, anchor samples, positive samples, and negative samples are identified, and a deep neural network model is trained to obtain a model capable of accurate visual localization in different environments.
It improves the robustness of visual positioning, ensures positioning accuracy under different lighting, weather, seasons and scenarios, solves the problem of visual positioning failure or instability, and enhances the practicality of the method.
Smart Images

Figure HDA0003062795310000011 
Figure HDA0003062795310000021 
Figure HDA0003062795310000022
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the positioning technical field, and in particular to a model training method, a visual positioning method, a device and equipment. BACKGROUND
[0002] With the rapid development of science and technology, the application field of visual positioning technology is more and more extensive, for example: unmanned driving field, robot navigation field, augmented reality field, etc. Taking the unmanned driving field as an example, the visual positioning technology can realize real-time positioning of the position of the unmanned vehicle itself through the image captured by the camera, thereby providing the position information and attitude information of the vehicle itself for the vehicle planning and control of the unmanned vehicle.
[0003] In the visual positioning operation, the positioning accuracy has a crucial influence on the control accuracy of the unmanned vehicle and the safety of the vehicle. However, in the real unmanned vehicle operation scene, the visual positioning algorithm also faces great challenges caused by various environmental changes, for example: extreme environmental changes such as light, weather, season and scene changes will reduce the accuracy of visual positioning, and thus reduce the control accuracy of the unmanned vehicle and the safety of the vehicle operation. SUMMARY
[0004] Therefore, the embodiments of the present application provide a model training method, a visual positioning method, a device and equipment, which can solve the positioning failure problem caused by extreme environmental changes in the visual positioning process, and improve the robustness of the positioning operation under different light, weather, season and different scenes.
[0005] In a first aspect, the embodiments of the present application provide a model training method, comprising:
[0006] obtaining a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image;
[0007] determining an anchor sample, a positive sample corresponding to the anchor sample and a negative sample corresponding to the anchor sample from the plurality of images according to the sensor data;
[0008] learning and training the anchor sample, the positive sample and the negative sample to obtain a deep neural network model, the deep neural network model being used for visual positioning operation based on the image.
[0009] In a second aspect, the embodiments of the present application provide a model training device, comprising:
[0010] a first obtaining module, configured to obtain a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image;
[0011] The first determining module is configured to determine, according to the sensor data, an anchor sample, a positive sample corresponding to the anchor sample, and a negative sample corresponding to the anchor sample from the plurality of images.
[0012] The first training module is configured to perform learning training on the anchor sample, the positive sample, and the negative sample to obtain a deep neural network model, the deep neural network model being used for visual positioning operation based on images.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, the memory being used for storing one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the model training method in the first aspect.
[0014] In a fourth aspect, an embodiment of the present application provides a computer storage medium for storing a computer program, the computer program causing a computer to implement the model training method in the first aspect when the computer program is executed.
[0015] In a fifth aspect, an embodiment of the present application provides a visual positioning method, including:
[0016] An image to be analyzed collected by a vehicle in a space is acquired;
[0017] A simultaneous localization and mapping (SLAM) database used for analyzing and processing the image to be analyzed is determined, the SLAM database including a plurality of standard images and standard location information corresponding to each standard image;
[0018] The image to be analyzed and the SLAM database are input into a deep neural network model to obtain positioning information of the vehicle in the space, wherein the deep neural network model is trained to be used for visual positioning operation based on images.
[0019] In a sixth aspect, an embodiment of the present application provides a visual positioning device, including:
[0020] A second acquiring module is configured to acquire an image to be analyzed collected by a vehicle in a space;
[0021] A second determining module is configured to determine a simultaneous localization and mapping (SLAM) database used for analyzing and processing the image to be analyzed, the SLAM database including a plurality of standard images and standard location information corresponding to each standard image;
[0022] A second positioning module is configured to input the image to be analyzed and the SLAM database into a deep neural network model to obtain positioning information of the vehicle in the space, wherein the deep neural network model is trained to be used for visual positioning operation based on images.
[0023] In a seventh aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, the memory being configured to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the visual positioning method in the fifth aspect.
[0024] In an eighth aspect, an embodiment of the present application provides a computer storage medium configured to store a computer program, the computer program causing a computer to implement the visual positioning method in the fifth aspect when executed.
[0025] In a ninth aspect, an embodiment of the present application provides a visual positioning method, comprising:
[0026] obtaining a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image;
[0027] determining, according to the sensor data, an anchor sample, a positive sample corresponding to the anchor sample, and a negative sample corresponding to the anchor sample from the plurality of images;
[0028] taking the anchor sample, the positive sample, and the negative sample as input parameters of a to-be-trained model, and executing the to-be-trained model to obtain a loss function;
[0029] optimizing the to-be-trained model according to the loss function to obtain an optimized model, the optimized model being configured to perform visual positioning on a vehicle.
[0030] In a tenth aspect, an embodiment of the present application provides a visual positioning device, comprising:
[0031] a third obtaining module configured to obtain a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image;
[0032] a third determining module configured to determine, according to the sensor data, an anchor sample, a positive sample corresponding to the anchor sample, and a negative sample corresponding to the anchor sample from the plurality of images;
[0033] a third processing module configured to take the anchor sample, the positive sample, and the negative sample as input parameters of a to-be-trained model, and execute the to-be-trained model to obtain a loss function;
[0034] a third optimization module configured to optimize the to-be-trained model according to the loss function to obtain an optimized model, the optimized model being configured to perform visual positioning on a vehicle.
[0035] In an eleventh aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, the memory being configured to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the visual positioning method in the ninth aspect.
[0036] In a twelfth aspect, an embodiment of the present application provides a computer storage medium for storing a computer program, and the computer program causes a computer to execute the visual positioning method in the ninth aspect when executed.
[0037] The technical scheme provided by the embodiment determines an anchor sample, a positive sample corresponding to the anchor sample, and a negative sample corresponding to the anchor sample from the plurality of images according to the sensor data, and then learns and trains the anchor sample, the positive sample, and the negative sample, so as to obtain a deep neural network model capable of realizing accurate visual positioning operation. Specifically, the deep neural network model can accurately analyze and identify a plurality of different images corresponding to a plurality of different environments, and then realize visual positioning operation based on the identification result, and ensures the accuracy of the visual positioning operation, effectively solves the problem of visual positioning failure or instability caused by various environmental changes, and effectively improves the practicability of the method. BRIEF DESCRIPTION OF DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0039] Figure 1 A scene flow chart of a model training method provided by the embodiment of the present application;
[0040] Figure 2 A flowchart of a model training method provided by the embodiment of the present application;
[0041] Figure 3 A flowchart of obtaining a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image provided by the embodiment of the present application;
[0042] Figure 4 A flowchart of another model training method provided by the embodiment of the present application;
[0043] Figure 5 A flowchart of another model training method provided by the embodiment of the present application;
[0044] Figure 6 A flowchart of another model training method provided by the embodiment of the present application;
[0045] Figure 7A process schematic diagram of a deep neural network model obtained by learning and training the anchor sample, the positive sample and the negative sample provided by the embodiment of the present application is provided.
[0046] Figure 8 A process schematic diagram of another model training method provided by the embodiment of the present application is provided.
[0047] Figure 9 A scene schematic diagram of a visual positioning method provided by the embodiment of the present application is provided.
[0048] Figure 10 A process schematic diagram of a visual positioning method provided by the embodiment of the present application is provided.
[0049] Figure 11 A process schematic diagram of another visual positioning method provided by the embodiment of the present application is provided.
[0050] Figure 12 A process schematic diagram of a visual positioning method provided by the application embodiment of the present application is provided.
[0051] Figure 13 A structure schematic diagram of a model training device provided by the embodiment of the present application is provided.
[0052] Figure 14 A structure schematic diagram of a model training device provided by the embodiment of the present application is provided. Figure 13
[0053] A structure schematic diagram of a visual positioning device provided by the embodiment of the present application is provided. Figure 15
[0054] A structure schematic diagram of a visual positioning device provided by the embodiment of the present application is provided. Figure 16 Figure 15 A process schematic diagram of another visual positioning method provided by the embodiment of the present application is provided.
[0055] Figure 17 A structure schematic diagram of another visual positioning device provided by the embodiment of the present application is provided.
[0056] Figure 18 A structure schematic diagram of a visual positioning device provided by the embodiment of the present application is provided.
[0057] Figure 19 Figure 18 A structure schematic diagram of a visual positioning device provided by the embodiment of the present application is provided. DETAILED DESCRIPTION
[0058] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some but not all of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts should fall into the scope of the present application.
[0059] The terms used in the embodiments of the present application are only for the purpose of describing particular embodiments and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Plural" generally includes at least two, but does not exclude the case of including at least one.
[0060] It should be understood that the term "and / or" used herein is only to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " herein generally represents an "or" relationship between the front and rear associated objects.
[0061] Depending on the context, the word "if" as used herein can be interpreted as "when" or "upon" or "in response to determining" or "in response to identifying". Similarly, depending on the context, the phrase "if it is determined" or "if it is identified (a stated condition or event)" can be interpreted as "when it is determined" or "in response to determining" or "when it is identified (a stated condition or event)" or "in response to identifying (a stated condition or event)".
[0062] It should also be noted that the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that a product or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such product or system. Without more limitations, the element defined by the sentence "including a" does not exclude the presence of other identical elements in the product or system including the element.
[0063] In order to facilitate the understanding of the technical solutions of the embodiments, the following will make relevant description on the prior art:
[0064] Visual positioning algorithm has important applications in the field of unmanned driving, robot navigation and augmented reality. For example, in the field of unmanned driving, the position of the unmanned vehicle is located in real time through the image captured by the camera, so as to provide the position information and attitude information of the vehicle itself for vehicle planning and control.
[0065] Specifically, the visual positioning algorithm in the prior art includes the following steps: extracting manually designed image key point features, such as Scale-invariant feature transform (SIFT), Speeded Up Robust Features (SURF), and Oriented Fast and Rotated Brief (ORB); then matching the key point features of the image with the key point features in the map, and finally solving the pose of the camera in the world coordinate system through the feature point matching relationship, so as to obtain the position and attitude of the vehicle body.
[0066] When performing visual positioning operation, the accuracy of positioning has a crucial influence on the control accuracy of unmanned vehicle and vehicle safety. However, in real unmanned vehicle operation scenarios, the visual positioning algorithm also faces great challenges caused by various environmental changes, such as light change (caused by camera light change, such as camera backlight and overexposure), weather change (overcast, rain, snow and other weather), seasonal change (environmental change caused by spring, summer, autumn and winter), scene change (operation scene change, such as open road and closed community) and other extreme environmental changes, which will reduce the accuracy of visual positioning, and further reduce the control accuracy of unmanned vehicle and the safety degree of vehicle operation.
[0067] Specifically, in order to solve the problem of visual positioning failure or instability caused by various environmental changes, the embodiment provides a model training method, a visual positioning method, a device and equipment. The execution subject of the model training method is a model training device, and the execution subject of the visual positioning method is a visual positioning device. The model training device and the visual positioning device are in communication connection with the controller of the target to be positioned (unmanned vehicle, unmanned ship, unmanned aerial vehicle, mobile robot, etc.), so as to realize the visual positioning operation for the target to be positioned.
[0068] As Figure 1As shown, the to-be-positioned target (vehicle) can be provided with an image acquisition device (for example: a camera, a video camera, other devices with image acquisition capabilities, etc.) for acquiring or generating a plurality of images corresponding to a plurality of different environments. The plurality of images can include a plurality of different environmental features corresponding to a plurality of different environments. After obtaining the plurality of images, the image acquisition device can transmit the plurality of images to the model training device. In addition, the to-be-positioned target can also be provided with a plurality of sensors. The above-mentioned sensors are used to acquire corresponding sensor data during the movement of the to-be-positioned target. The sensor data can correspond to the above-mentioned images. In specific implementation, the sensor data can include positioning data, acceleration data, angular velocity data, etc. After obtaining the sensor data, the plurality of sensors can transmit the plurality of sensor data to the model training device.
[0069] The model training device can be a device with model training capability. In specific implementation, it can be implemented as an electronic device, a server, etc. The server usually refers to a server that plans information using a network. In physical implementation, the model training device can be any device that can provide computing services, respond to service requests, and process, such as a conventional server, a cloud server, a cloud host, a virtual center, etc. The constitution of the model training device mainly includes a processor, a hard disk, a memory, a system bus, etc., and is similar to the general computer architecture.
[0070] Specifically, the model training device can obtain a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image. Then, according to the sensor data, the anchor sample, the positive sample corresponding to the anchor sample, and the negative sample corresponding to the anchor sample are determined from the plurality of images, and the anchor sample, the positive sample and the negative sample are trained, so as to obtain a deep neural network model capable of realizing accurate visual positioning operation.
[0071] It should be noted that the above model training method and visual positioning method can be used on an unmanned vehicle, and at this time, the execution subject for implementing the model training method and visual positioning method, i.e., the model training device and visual positioning device, can be arranged on the vehicle to be identified. Specifically, for the model training device and visual positioning device, the model training device and visual positioning device can be adjusted according to different types of vehicles, that is, the algorithm modules included in the model training device and visual positioning device will also be different according to the different types of vehicles, at this time, the model training device can not only implement the model training operation, but also implement other operations; the visual positioning device can not only implement the visual positioning operation, but also implement other operations. For example, for logistics vehicles, public service vehicles, medical service vehicles, and terminal service vehicles, different model training devices and visual positioning devices are involved. The algorithm modules included in the model training device and visual positioning device for the four types of autonomous vehicles will be described below:
[0072] Among them, the logistics vehicle refers to the vehicle used in the logistics scene, for example: it can be a logistics vehicle with automatic sorting function, a logistics vehicle with cold storage function, and a logistics vehicle with measurement function. These logistics vehicles involve different algorithm modules.
[0073] For example, for a logistics vehicle, it can be equipped with an automatic sorting device that can automatically take out, transport, sort, and store goods after the logistics vehicle arrives at the destination. This involves an algorithm module for goods sorting, which mainly implements logic control of goods taking out, transporting, sorting, and storing.
[0074] For example, for a cold chain logistics scene, the logistics vehicle can also be equipped with a cold storage device that can achieve cold storage or preservation of fruits, vegetables, aquatic products, frozen foods, and other perishable foods during transportation, so that they are in a suitable temperature environment, solving the problem of long-distance transportation of perishable foods. This involves an algorithm module for cold storage control, which is mainly used to dynamically and adaptively calculate the appropriate temperature for cold storage or preservation according to the properties, perishability, transportation time, current season, climate, and other information of the food (or goods), and automatically adjust the cold storage device according to the appropriate temperature. In this way, the transportation personnel do not need to manually adjust the temperature when transporting different foods or goods, which liberates the transportation personnel from tedious temperature regulation and improves the efficiency of cold storage transportation.
[0075] For example, in most logistics scenarios, the charges are based on the volume and / or weight of the package, and the number of logistics packages is very large, so simply relying on the courier to measure the volume and / or weight of the package is very inefficient and has high labor costs. Therefore, in some logistics vehicles, a measuring device is added to automatically measure the volume and / or weight of the logistics package and calculate the cost of the logistics package. This involves an algorithm module for measuring logistics packages, which is mainly used to identify the type of logistics package, determine the measurement method of the logistics package, such as volume measurement, weight measurement, or combined volume and weight measurement, and can complete the volume and / or weight measurement according to the determined measurement method, and complete the cost calculation according to the measurement result.
[0076] Among them, the public service vehicle refers to a vehicle that provides a certain public service, for example: it can be a fire truck, a deicing truck, a water truck, a snowplow truck, a garbage disposal vehicle, a traffic command vehicle, etc. These public service vehicles will involve different algorithm modules.
[0077] For example, for an automatic driving fire truck, its main task is to perform a reasonable fire extinguishing task on the fire scene, which involves an algorithm module for the fire extinguishing task, which at least needs to implement the identification of the fire situation, the planning of the fire extinguishing scheme, and the automatic control of the fire extinguishing device, etc.
[0078] For example, for a deicing truck, its main task is to remove ice and snow on the road surface, which involves an algorithm module for deicing, which at least needs to implement the identification of the ice and snow situation on the road surface, the formulation of the deicing scheme according to the ice and snow situation, such as which road section needs to be deiced, which road section does not need to be deiced, whether to use a salt method, the salt dosage, etc., and the automatic control of the deicing device under the condition of determining the deicing scheme, etc.
[0079] Among them, the medical service vehicle refers to an automatic driving vehicle that can provide one or more medical services, which can provide disinfection, temperature measurement, medicine dispensing, isolation, etc. This involves algorithm modules for providing various self-service medical services, which mainly implement the identification of disinfection needs and the control of disinfection devices to disinfect patients, or the identification of patient positions, the control of temperature measurement devices to automatically approach the positions such as the forehead of the patient to measure the temperature of the patient, or the judgment of the disease, the prescription according to the judgment result, and the identification of medicines / medicine containers, and the control of the medicine grabbing manipulator to grab medicines according to the prescription for the patient, etc.
[0080] Among them, the terminal service vehicle refers to a self-service type of automatic driving vehicle that can replace some terminal devices to provide certain convenient services to users, for example, these vehicles can provide printing, attendance, scanning, unlocking, payment, retail, etc. services for users.
[0081] For example, in some application scenarios, users often need to go to a specific location to print or scan documents, which is time-consuming and laborious. Therefore, a terminal service vehicle that can provide printing / scanning services for users appears. These service vehicles can interconnect with user terminal devices, and users can issue printing instructions through terminal devices. The service vehicles respond to the printing instructions, automatically print the documents required by the users, and can automatically deliver the printed documents to the user location. Users do not need to queue at the printer, which can greatly improve the printing efficiency. Alternatively, the service vehicles can respond to the scanning instructions issued by the users through the terminal devices, move to the user location, and the users can place the documents to be scanned on the scanning tools of the service vehicles to complete the scanning, without the need to queue at the printer / scanner, saving time and effort. This involves an algorithm module for providing printing / scanning services, which at least needs to identify the interconnection with the user terminal device, the response to the printing / scanning instruction, the positioning of the user location, and the travel control, etc.
[0082] For another example, with the development of new retail scenarios, more and more e-commerce companies use self-service vending machines to deliver goods to major office buildings and public areas, but these self-service vending machines are placed in fixed positions and cannot be moved. Users need to be close to the self-service vending machine to purchase the required goods, which is still not very convenient. Therefore, self-driving vehicles that can provide retail services appear. These service vehicles can carry goods and automatically move, and can provide corresponding self-service shopping APPs or shopping portals. Users can use mobile phones and other terminal devices to place orders through the APPs or shopping portals. The order includes the name and quantity of the goods to be purchased and the user location. After receiving the order request, the vehicle can determine whether the remaining goods have the goods to be purchased by the user and whether the quantity is sufficient. If the goods to be purchased by the user are available and the quantity is sufficient, the vehicle can automatically move to the user location and provide the goods to the user, further improving the convenience of user shopping, saving user time, and allowing users to use their time for more important things. This involves algorithm modules for providing retail services, which mainly implement the logic of responding to user order requests, order processing, goods information maintenance, user location positioning, and payment management, etc.
[0083] The technical scheme provided by the embodiment comprises the following steps: acquiring a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image; determining an anchor sample, a positive sample corresponding to the anchor sample and a negative sample corresponding to the anchor sample from the plurality of images according to the sensor data; and then performing learning training on the anchor sample, the positive sample and the negative sample, so as to obtain a deep neural network model capable of realizing accurate visual positioning operation. Specifically, the deep neural network model can accurately analyze and identify a plurality of different images corresponding to a plurality of different environments, and then realize visual positioning operation based on the identification result, thereby ensuring the accuracy of the visual positioning operation and effectively solving the problem of visual positioning failure or instability caused by various environmental changes, thereby effectively improving the practicability of the method.
[0084] In addition, the step sequence in each method embodiment described below is only an example and is not strictly limited.
[0085] Figure 1 A scene flowchart of a model training method provided by the embodiment of the application is shown in FIG. 1. Figure 2 A flowchart of a model training method provided by the embodiment of the application is shown in FIG. 1. Figures 1-2 The model training method provided by the embodiment comprises the following steps:
[0086] In step S201, a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image are acquired.
[0087] The plurality of different environments comprise at least one of the following: a plurality of different time periods, a plurality of different weathers, and a plurality of different operation scenarios. Specifically, the plurality of different time periods can comprise at least one of the following: any time point in the morning time period, any time point in the noon time period, any time point in the afternoon time period, any time point in the evening time period, and the like. The plurality of different weathers can comprise at least one of the following: spring weather, summer weather, spring weather, winter weather, rainy and snowy weather, and windy and sandy weather. The plurality of different operation scenarios can comprise at least one of the following: an open road scenario, an indoor scenario, and a closed community scenario. It can be understood that the plurality of different environments are not limited to the scenarios described above, and a person skilled in the art can set the plurality of different environments according to specific application scenarios and application requirements, which will not be described here.
[0088] Specifically, in order to realize the visual positioning operation on the object to be positioned, the image acquisition device and the sensor can be arranged on the object to be positioned. The image acquisition device is used to acquire images of the object to be positioned in different environments, so as to obtain a plurality of images corresponding to a plurality of different environments. The sensor is used to detect the movement state corresponding to the object to be positioned, which can include positioning information, attitude information and the like. In some examples, the sensor data includes at least one of the detection data of the inertial measurement unit and the detection data of the global positioning system. It can be understood that the detection data of the inertial measurement unit can include acceleration information, angular velocity information, attitude information and the like corresponding to the object to be positioned, and the detection data of the global positioning system can include positioning data of the global positioning system.
[0089] When the plurality of images corresponding to the plurality of different environments are acquired by the image acquisition device, the sensor data corresponding to each image can be acquired by the sensor. It can be understood that there is a corresponding relationship between the sensor data and the image. After acquiring the plurality of images corresponding to the plurality of different environments and the sensor data corresponding to each image, the plurality of images corresponding to the plurality of different environments and the sensor data corresponding to each image can be actively or passively transmitted to the model training device, so that the model training device can accurately and effectively acquire the plurality of images and the sensor data.
[0090] Step S202: determining an anchor sample, a positive sample corresponding to the anchor sample and a negative sample corresponding to the anchor sample from the plurality of images according to the sensor data.
[0091] After acquiring the sensor data and the plurality of images, the plurality of images can be analyzed and processed based on the sensor data to determine the anchor sample, the positive sample corresponding to the anchor sample and the negative sample corresponding to the anchor sample. The anchor sample can be any one of the plurality of images. The positive sample can refer to an image having a set relationship with the anchor sample. The negative sample can refer to an image not having a set relationship with the anchor sample.
[0092] In some examples, the set relationship can include a co-view relationship with the anchor sample and a distance less than a preset range from the anchor sample. The co-view relationship with the anchor sample can mean that the shooting angle corresponding to the positive sample and the shooting angle corresponding to the anchor sample are the same, and the distance between the shooting location corresponding to the positive sample and the shooting location corresponding to the anchor sample is not far.
[0093] In some examples, determining the anchor sample, the positive sample corresponding to the anchor sample, and the negative sample corresponding to the anchor sample from the plurality of images according to the sensor data can include: selecting one image as the anchor sample in the plurality of images; determining, according to the sensor information, an image having a set relationship with the anchor sample in the plurality of images as the positive sample; and determining, according to the sensor information, an image not having the set relationship with the anchor sample in the plurality of images as the negative sample.
[0094] For example, the plurality of images include image 1, image 2, image 3, image 4, and image 5, and the sensor data can include sensor data 1 corresponding to image 1, sensor data 2 corresponding to image 2, sensor data 3 corresponding to image 3, sensor data 4 corresponding to image 4, and sensor data 5 corresponding to image 5.
[0095] Taking image 1 included in the plurality of images as an anchor sample, after obtaining the plurality of sensor data, whether the anchor sample and image 2 have a set relationship can be detected based on sensor data 1 and sensor data 2, whether the anchor sample and image 3 have a set relationship can be detected based on sensor data 1 and sensor data 3, whether the anchor sample and image 4 have a set relationship can be detected based on sensor data 1 and sensor data 4, and whether the anchor sample and image 5 have a set relationship can be detected based on sensor data 1 and sensor data 5.
[0096] When the anchor sample and image 2 have a set relationship, and the anchor sample and image 3 have a set relationship, image 2 and image 3 can be determined as positive samples corresponding to the anchor sample. When the anchor sample and image 4 do not have a set relationship, and the anchor sample and image 5 do not have a set relationship, image 4 and image 5 can be determined as negative samples corresponding to the anchor sample, thereby effectively achieving the accuracy and reliability of determining the anchor sample, the positive sample corresponding to the anchor sample, and the negative sample corresponding to the anchor sample.
[0097] Step S203: learning and training the anchor sample, the positive sample, and the negative sample to obtain a deep neural network model, the deep neural network model being used for visual positioning operation based on images.
[0098] After obtaining the anchor sample, the positive sample, and the negative sample, the anchor sample, the positive sample, and the negative sample can be learned and trained, thereby obtaining a deep neural network model. The obtained deep neural network model is used for visual positioning operation based on images. Since the deep neural network model contains a large number of learnable parameters, the trained deep neural network model can better solve the problem of inaccurate positioning caused by changes in light (such as overexposure), extreme weather changes, and seasonal and scene changes, and changes in perspective.
[0099] The model training method provided by the embodiment can obtain multiple images corresponding to multiple different environments and sensor data corresponding to each image, determine anchor samples, positive samples corresponding to the anchor samples and negative samples corresponding to the anchor samples from the multiple images according to the sensor data, and then perform learning training on the anchor samples, the positive samples and the negative samples, so as to obtain a deep neural network model capable of realizing accurate visual positioning operation. Specifically, the deep neural network model can accurately analyze and identify multiple different images corresponding to multiple different environments, and then realize visual positioning operation based on the identification result, thereby ensuring the accuracy of the visual positioning operation and effectively solving the problem of visual positioning failure or instability caused by various environmental changes, thereby effectively improving the practicality of the method.
[0100] Figure 3 The flowchart for obtaining multiple images corresponding to multiple different environments and sensor data corresponding to each image is provided for the embodiment of the application. Based on the above embodiment, with reference to the accompanying drawings, the embodiment provides an implementation manner for obtaining multiple images corresponding to multiple different environments and sensor data corresponding to each image. Figure 3 The implementation manner for obtaining multiple images corresponding to multiple different environments and sensor data corresponding to each image in the embodiment can include the following steps.
[0101] Step S301: Obtain all images corresponding to multiple different environments and all sensor data corresponding to each image.
[0102] In order to improve the quality of learning training on the deep neural network model, all images corresponding to multiple different environments and all sensor data corresponding to each image can be obtained. It can be understood that different images in all images can correspond to the same or different image quality, and different sensor data in all sensor data can correspond to the same or different data accuracy. For example, the image definition of some images in all images is high (high image quality), and the image definition of some images in all images is low (low image quality); the data accuracy of some sensor data in all sensor data is high, and the data accuracy of some sensor data in all sensor data is low.
[0103] Step S302: Identify invalid data included in all sensor data.
[0104] When all the sensor data is acquired by using the sensor, the data accuracy of some sensor data is higher, and the data accuracy of some sensor data is lower. Therefore, in order to improve the quality and efficiency of learning and training of the deep neural network model, invalid data included in all the sensors can be identified, and the invalid data can include at least one of abnormal data and repeated data, wherein the abnormal data can refer to sensor data that is out of a preset range.
[0105] In some examples, identifying the invalid data included in all the sensor data includes: identifying whether there is repeated data in all the sensors, and if there is repeated data, the repeated data is determined as invalid data; then, a preset range for all the sensor data is acquired, all the sensor data is analyzed and compared with the preset range, when the sensor data is out of the preset range, the sensor data is determined as abnormal data; and when the sensor data is within the preset range, the sensor data is determined as valid data.
[0106] Step S303: removing the invalid data included in all the sensor data to obtain target sensor data.
[0107] After the invalid data included in all the sensor data is acquired, the invalid data included in all the sensor data can be removed, so as to obtain target sensor data.
[0108] Step S304: determining an image corresponding to the target sensor data as a target image in all the images.
[0109] Since there is a corresponding relationship between the sensor data and the image, after the target sensor data is acquired, an image corresponding to the target sensor data can be determined as a target image in all the images, and the target image is used for learning and training to generate a deep neural network model.
[0110] In the embodiment, by acquiring all the images corresponding to multiple different environments and all the sensor data corresponding to each image, the invalid data included in all the sensor data is identified and removed, and the target sensor data is obtained, and then an image corresponding to the target sensor data is determined as a target image in all the images, so that the image used for learning and training does not include invalid data, and the quality and efficiency of learning and training of the deep neural network model are effectively improved when learning and training is performed based on the image not including invalid data.
[0111] Figure 4 Another model training method provided by the embodiment of the present application is provided. On the basis of the above-mentioned embodiment, reference is made to the accompanying drawings Figure 4As shown, before identifying the invalid data included in all sensor data, the method in this embodiment can further include:
[0112] Step S401: Obtain a localization confidence corresponding to each image and sensor data.
[0113] Step S402: Based on the localization confidence, determine a candidate image and sensor data corresponding to the candidate image from all images and all sensor data corresponding to each image.
[0114] To further improve the quality of learning and training of the deep neural network model, before identifying the invalid data included in all sensor data, a localization confidence corresponding to each image and sensor data can be obtained. Specifically, a preset algorithm or a pre-trained machine learning model can be used to analyze and process all images and sensor data, so as to obtain a localization confidence corresponding to each image and sensor data. The localization confidence is used to identify the accuracy when visual localization is performed using each image and sensor data. It can be understood that the higher the localization confidence corresponding to each image and sensor data, the higher the positioning accuracy of the deep neural network model obtained when learning and training using the above-mentioned each image and sensor data.
[0115] After obtaining the localization confidence corresponding to each image and sensor data, a candidate image and sensor data corresponding to the candidate image can be determined from all images and all sensor data corresponding to each image based on the localization confidence. In some examples, determining a candidate image and sensor data corresponding to the candidate image from all images and all sensor data corresponding to each image based on the localization confidence can include: when the localization confidence is less than or equal to a preset threshold, removing the image and sensor data corresponding to the localization confidence to obtain a candidate image and candidate sensor data corresponding to the candidate image.
[0116] Specifically, after obtaining the localization confidence, the localization confidence can be compared with a preset threshold. When the localization confidence is less than or equal to the preset threshold, it indicates that the positioning accuracy of the image and sensor data corresponding to the localization confidence is low, and then the image and sensor data corresponding to the localization confidence can be removed to obtain a candidate image and candidate sensor data corresponding to the candidate image.
[0117] In this embodiment, by acquiring the positioning confidence corresponding to each image and sensor data, and then determining the candidate image and the candidate sensor data corresponding to the candidate image in all images and all sensor data corresponding to each image based on the positioning confidence, the candidate image with higher positioning confidence and the candidate sensor data corresponding to the candidate image can be effectively obtained, so that the deep neural network model with higher visual positioning accuracy can be obtained when learning and training are performed using the candidate image and the candidate sensor data, and the practicability of the visual positioning operation is further improved.
[0118] Figure 5 A flowchart of another model training method provided by an embodiment of the present application is shown in FIG. 6. Based on the above embodiment, the method of the present embodiment can further include the following steps: Figure 5 As shown in FIG. 6, before identifying the invalid data included in all sensor data, the method of the present embodiment can further include the following steps:
[0119] Step S501: acquire model requirement information.
[0120] Step S502: determine, in all images and all sensor data, a candidate image corresponding to the model requirement information and candidate sensor data corresponding to the candidate image.
[0121] In order to ensure the range and accuracy of the deep neural network model, the model requirement information can be acquired, which can include at least one of the following: image and sensor data with strong light, image and sensor data with weak light, image and sensor data in daytime, image and sensor data in nighttime, etc. It can be understood that a person skilled in the art can set different model requirement information according to different application scenarios and application requirements. After the model requirement information is configured, the configured model requirement information can be stored in a preset area, and the model requirement information can be acquired by accessing the preset area.
[0122] After obtaining the model requirement information, the corresponding candidate image and the corresponding candidate sensor data of the candidate image can be determined in all images and all sensor data. In some examples, determining the corresponding candidate image and the corresponding candidate sensor data of the candidate image in all images and all sensor data can include: obtaining a requirement feature corresponding to the model requirement information; determining an image feature corresponding to all images, when the image feature matches the requirement feature, the image corresponding to the image feature and the sensor data are determined as the corresponding candidate image and the corresponding candidate sensor data of the candidate image corresponding to the model requirement information. When the image feature does not match the requirement feature, the image corresponding to the image feature is determined as an invalid image, and the invalid image in all images can be deleted.
[0123] In the embodiment, by obtaining the model requirement information, and then determining the corresponding candidate image and the corresponding candidate sensor data of the candidate image in all images and all sensor data, the corresponding candidate image and the corresponding candidate sensor data of the candidate image satisfying the model requirement information can be effectively obtained, so that when learning and training are performed by using the candidate image and the candidate sensor data, the deep neural network model satisfying the scene requirement and the design requirement can be obtained, and the practicability of the deep neural network model is further improved.
[0124] Figure 6 Another flowchart of a model training method provided by the embodiment is shown in FIG. 6. Based on the above embodiment, the method provided by the embodiment can further include the following steps. Figure 6 As shown in FIG. 6, after selecting one image as an anchor sample in the plurality of images, the method provided by the embodiment can further include the following steps.
[0125] Step S601: In the plurality of images, all similar sample images corresponding to the anchor sample are obtained.
[0126] In the plurality of images, any one image can be determined as the anchor sample, and then the anchor sample and other images are analyzed and processed to obtain all similar sample images corresponding to the anchor sample. It can be understood that the number of similar sample images corresponding to the anchor sample can be one or more, and the similarity between the similar sample image and the anchor sample is greater than or equal to a preset threshold.
[0127] In some examples, in the plurality of images, all similar sample images corresponding to the anchor sample can include: in the plurality of images, any one image is determined as the anchor sample; the similarity between the anchor sample and other images is determined, and when the similarity is greater than or equal to a preset threshold, the image corresponding to the similarity is determined as the similar sample image corresponding to the anchor sample.
[0128] Step S602: According to the sensor information, the relative pose information and the feature point matching information between the similar sample image and the anchor sample are obtained.
[0129] Wherein, after the similar sample image and the anchor sample are obtained, the similar sample image and the anchor sample can be analyzed and processed based on the sensor information to determine the relative pose information and the feature point matching information between the similar sample image and the anchor sample. It can be understood that the relative pose information between the similar sample image and the anchor sample can refer to the relative pose deviation between the first camera pose information (the first target pose information) and the second camera pose information (the second target pose information), wherein the first camera pose information refers to the pose information corresponding to the acquisition of the similar sample image, and the second camera pose information refers to the pose information corresponding to the acquisition of the anchor sample; or the relative pose information between the similar sample image and the anchor sample can refer to the relative pose deviation between the first vehicle pose information corresponding to the acquisition of the similar sample image and the second vehicle pose information corresponding to the acquisition of the anchor sample.
[0130] In some examples, after the sensor information is obtained, the scene can be mapped based on the sensor information by using a Simultaneous Localization And Mapping (SLAM) algorithm, and the relative pose information and the feature point matching information between the similar sample image and the anchor sample can be obtained in the mapping process.
[0131] In other examples, according to the sensor information, the relative pose information between the similar sample image and the anchor sample can include: obtaining the first sensor information corresponding to the similar sample image and the second sensor information corresponding to the anchor sample; determining the relative pose information between the similar sample image and the anchor sample based on the first sensor information and the second sensor information.
[0132] Specifically, the first sensor information can include the first pose information corresponding to the similar sample image, and the second sensor information can include the second pose information corresponding to the anchor sample. After the first pose information and the second pose information are obtained, the first pose information and the second pose information can be analyzed and processed, so that the pose offset information can be obtained, and the pose offset information can be determined as the relative pose information between the similar sample image and the anchor sample.
[0133] Of course, those skilled in the art can also use other ways to obtain the relative pose information and the feature point matching information between the similar sample image and the anchor sample, as long as the relative pose information and the feature point matching information between the similar sample image and the anchor sample can be stably obtained, which will not be described here.
[0134] Step S603: Based on the relative pose information and the feature point matching information, it is detected whether the similar sample image and the anchor sample exist a set relationship.
[0135] After obtaining the relative pose information and the feature point matching information, the relative pose information and the feature point matching information can be analyzed and processed to detect whether the similar sample image and the anchor sample exist a set relationship. In some examples, based on the relative pose information and the feature point matching information, it is detected whether the similar sample image and the anchor sample exist a set relationship can include: when the relative pose information is less than or equal to a preset threshold, and the feature matching information is greater than or equal to a preset threshold, it is determined that the similar sample image and the anchor sample exist a set relationship; when the relative pose information is greater than the preset threshold, or the feature matching information is less than the preset threshold, it is determined that the similar sample image and the anchor sample do not exist a set relationship.
[0136] Specifically, after obtaining the relative pose information and the feature matching information, the relative pose information and the preset threshold, and the feature matching information and the preset threshold can be analyzed and compared. It should be noted that the preset threshold used for analyzing and comparing the relative pose information and the preset threshold used for analyzing and processing the feature matching information can be the same or different.
[0137] When the relative pose information is less than or equal to the preset threshold, it indicates that the vehicle pose or the camera pose when obtaining the similar sample image is basically the same as the vehicle pose or the camera pose when obtaining the anchor sample; when the relative pose information is greater than the preset threshold, it indicates that the vehicle pose or the camera pose when obtaining the similar sample image is different from the vehicle pose or the camera pose when obtaining the anchor sample. When the feature matching information is greater than or equal to the preset threshold, it indicates that the similarity between the similar sample image and the anchor sample is high; when the feature matching information is less than the preset threshold, it indicates that the similarity between the similar sample image and the anchor sample is low.
[0138] When the relative pose information is less than or equal to the preset threshold, and the feature matching information is greater than or equal to the preset threshold, it is determined that the similar sample image and the anchor sample exist a set relationship; when the relative pose information is greater than the preset threshold, or the feature matching information is less than the preset threshold, it is determined that the similar sample image and the anchor sample do not exist a set relationship, thereby effectively realizing that whether the similar sample image and the anchor sample exist a set relationship can be detected based on the relative pose information and the feature point matching information.
[0139] In the embodiment, in the plurality of images, all similar sample images corresponding to the anchor sample are obtained, the relative pose information and the feature point matching information between the similar sample images and the anchor sample are obtained according to the sensor information, and then whether the similar sample images and the anchor sample exist the set relationship is detected according to the relative pose information and the feature point matching information, thereby effectively realizing the accuracy and reliability of detecting whether the similar sample images and the anchor sample exist the set relationship, and further ensuring the quality and efficiency of training the deep neural network model.
[0140] Figure 7 The flowchart of another model training method provided by the embodiment of the present application is shown in FIG. 8. On the basis of any one of the above embodiments, with continued reference to FIG. 8, the method provided by the embodiment of the present application can further include: Figure 7 The embodiment provides an implementation manner of learning and training the anchor sample, the positive sample and the negative sample, and specifically, the learning and training of the anchor sample, the positive sample and the negative sample to obtain the deep neural network model in the embodiment can include:
[0141] Step S701: obtaining a loss function used for analyzing and processing the anchor sample, the positive sample and the negative sample.
[0142] Step S702: learning and training the anchor sample, the positive sample and the negative sample by using the loss function to obtain the deep neural network model.
[0143] After the anchor sample, the positive sample and the negative sample are obtained, a loss function corresponding to the anchor sample, the positive sample and the negative sample can be obtained, and the loss function is used for minimizing the matching distance between the anchor sample and the positive sample and maximizing the matching distance between the anchor sample and the negative sample. In some examples, the loss function can be a triplet loss function.
[0144] After the loss function is obtained, the anchor sample, the positive sample and the negative sample can be learned and trained based on the loss function, so that the deep neural network model can be obtained, thereby effectively ensuring the quality and efficiency of training the deep neural network model.
[0145] Figure 8 The flowchart of another model training method provided by the embodiment of the present application is shown in FIG. 8. On the basis of any one of the above embodiments, with continued reference to FIG. 8, the method provided by the embodiment of the present application can further include: Figure 8 After the deep neural network model is obtained, the method in the embodiment can further include:
[0146] Step S801: obtaining an image to be analyzed collected by a target to be positioned in space.
[0147] Step S802: determine a plurality of standard images for analyzing the to-be-analyzed image, each standard image corresponding to standard location information.
[0148] Step S803: input the to-be-analyzed image and the plurality of standard images into the deep neural network model to obtain positioning information of the to-be-positioned target in the space.
[0149] Wherein, after obtaining the deep neural network model, the visual positioning operation can be performed by using the deep neural network model. Specifically, when the user has a visual positioning demand, the to-be-positioned target in the space can be acquired. The to-be-positioned target can include at least one of the following: a vehicle, a ship, an airplane, an unmanned vehicle, an unmanned ship, an unmanned aerial vehicle, a mobile robot, and the like.
[0150] In order to accurately implement the visual positioning operation, a plurality of standard images for analyzing the to-be-analyzed image can be determined, each standard image corresponding to standard location information. It can be understood that the plurality of standard images correspond to the space where the to-be-positioned target is located. When the space where the to-be-positioned target is located is different, the plurality of corresponding standard images are also different.
[0151] After obtaining the to-be-analyzed image and the plurality of standard images, the to-be-analyzed image and the plurality of standard images can be input into the deep neural network model. The deep neural network model can identify a target standard image matching the to-be-analyzed image in the plurality of standard images. Then, the standard location information corresponding to the target standard image can be determined as the positioning information of the to-be-positioned target in the space. Thus, the visual positioning operation is effectively implemented, and the practicality of the deep neural network model is further improved.
[0152] Figure 9 A scene schematic diagram of a visual positioning method provided by an embodiment of the present application; Figure 10 A flowchart of a visual positioning method provided by an embodiment of the present application; referring to FIG. 8B, Figures 9-10 The embodiment provides a visual positioning method. The execution subject of the visual positioning method is a visual positioning device. It can be understood that the visual positioning device can be implemented as software or a combination of software and hardware. Specifically, the method can include the following steps:
[0153] Step S1001: acquiring a to-be-analyzed image collected by a vehicle in a space.
[0154] Step S1002: determining a simultaneous localization and mapping (SLAM) database for analyzing the to-be-analyzed image. The SLAM database includes a plurality of standard images and standard location information corresponding to each standard image.
[0155] Step S1003: input the to-be-analyzed image and the SLAM database into a deep neural network model to obtain the positioning information of the vehicle in the space, wherein the deep neural network model is trained for visual positioning operation based on images.
[0156] The specific implementation process of each step is described in detail as follows:
[0157] Step S1001: acquire the to-be-analyzed image collected by the vehicle in the space.
[0158] The image acquisition device can be arranged on the vehicle, and the image acquisition device is configured to acquire images of the vehicle in different environments, so that the to-be-analyzed images corresponding to the different environments can be obtained. When there is a demand for visual positioning, the to-be-analyzed images corresponding to the different environments can be actively or passively transmitted to the visual positioning device, so that the visual positioning device can accurately and effectively acquire the to-be-analyzed image collected by the vehicle in the space.
[0159] Step S1002: determine a simultaneous localization and mapping (SLAM) database for analyzing and processing the to-be-analyzed image, wherein the SLAM database includes a plurality of standard images and standard location information corresponding to each standard image.
[0160] In order to accurately implement the visual positioning operation of the vehicle, a SLAM database for analyzing and processing the to-be-analyzed image can be determined, and the SLAM database can include a plurality of standard images and standard location information corresponding to each standard image. It can be understood that the SLAM database corresponds to the space in which the vehicle is located, and when the space in which the vehicle is located is different, the SLAM database corresponding to the space is also different.
[0161] Step S1003: input the to-be-analyzed image and the SLAM database into a deep neural network model to obtain the positioning information of the vehicle in the space, wherein the deep neural network model is trained for visual positioning operation based on images.
[0162] After the to-be-analyzed image and the SLAM database are acquired, the to-be-analyzed image and the SLAM database can be input into a deep neural network model. The deep neural network model can identify a target standard image matching the to-be-analyzed image from the plurality of standard images included in the SLAM database, and then determine the standard location information corresponding to the target standard image as the positioning information of the vehicle in the space, thereby effectively implementing the visual positioning operation and further improving the practicability of the visual positioning method.
[0163] The visual positioning method provided by the embodiment can accurately identify the positioning information of the vehicle in the space by acquiring the to-be-analyzed image collected by the vehicle in the space, determining a simultaneous localization and mapping (SLAM) database used for analyzing and processing the to-be-analyzed image, and inputting the to-be-analyzed image and the SLAM database into a deep neural network model. In addition, the deep neural network model can analyze and identify the images corresponding to various environmental features, thereby effectively solving the positioning failure problem caused by extreme environmental changes in the visual positioning process, improving the robustness of positioning operation under different illumination, weather, seasons and different scenes, and further ensuring the accuracy and reliability of the method.
[0164] Figure 11 The flowchart of another visual positioning method provided by the embodiment is shown in FIG. 11. Figure 11 The vehicle can be provided with an inertial measurement unit. After obtaining the positioning information of the vehicle in the space, the method in the embodiment can further include:
[0165] Step S1101: acquiring detection data of the inertial measurement unit.
[0166] Step S1102: determining the attitude information of the vehicle based on the detection data.
[0167] Step S1103: generating navigation information corresponding to the vehicle according to the attitude information and the positioning information.
[0168] The vehicle can be provided with an inertial measurement unit, which can detect the running state of the vehicle in real time during the running of the vehicle, so as to obtain detection data. The detection data can include the acceleration of the vehicle and the angular velocity of the vehicle. After obtaining the detection data, the detection data can be analyzed and processed to determine the attitude information of the vehicle. The attitude information is the pose information of the vehicle in the world coordinate system. Therefore, the attitude information of the vehicle can include the heading direction of the vehicle.
[0169] After obtaining the attitude information of the vehicle, the navigation information corresponding to the vehicle can be generated based on the attitude information and the positioning information. The generated navigation information can include at least one of road information and movement speed information, so as to control the vehicle to move based on the navigation information, and further ensure the safety and reliability of the vehicle running.
[0170] In a specific application, the application embodiment provides a visual positioning method. The visual positioning method can obtain a deep neural network model by continuously training and updating visual features on image data under different illuminations, weathers, seasons and different scenes on a large scale, and then can perform visual positioning operation by using the deep neural network model, thereby solving the positioning failure problem caused by extreme environmental changes in the visual positioning process, and improving the feature stability and robustness of positioning effect under different environmental conditions. Specifically, taking a vehicle as an example, as shown in the accompanying drawings, the method can include the following steps: Figure 12
[0171] Step 1: Regression test.
[0172] The vehicle is provided with a camera, a global positioning system (GPS) and an inertial measurement unit (IMU), and the vehicle is controlled to move in different operation scenes, different time periods and different weathers, so that images in different time periods, different weathers and different operation scenes can be obtained by the camera, positioning information corresponding to the images in different time periods, different weathers and different operation scenes can be obtained by the GPS, and attitude information corresponding to the images in different time periods, different weathers and different operation scenes can be obtained by the IMU, that is, the vehicle is continuously tested. In general, the image, sensor data and vehicle state data are obtained by the camera, GPS and IMU, so that a large-scale data set under different illuminations, different weathers and different scenes can be continuously established.
[0173] Step 2: Data mining.
[0174] After obtaining the data set corresponding to different illuminations, different weathers and different scenes, the data in the data set can be processed by data mining. Specifically, the data mining operation can include passive mining operation and active mining operation.
[0175] The passive mining operation is mainly used to identify data with poor positioning effect. Specifically, the visual positioning confidence corresponding to the data can be obtained, and then the data with poor positioning effect can be mined based on the visual positioning confidence. These data with poor positioning quality are often caused by various environmental changes. In order to ensure the quality of learning and training of the deep neural network model, the above-mentioned data with poor positioning effect can be deleted.
[0176] In addition, the active mining operation is mainly used to identify data useful or effective for the model training operation. Specifically, model requirement information can be obtained, which can include data with excessively strong light, data with excessively dark light, and the like. Based on the model requirement information and the mining algorithm, it is determined whether the data in the data set is valid data. In this way, valid data required by the model training operation is obtained, and the generated deep neural network model based on the valid data can analyze and process data corresponding to different illuminations, different weathers, and different scenes, thereby further improving the application range of the method.
[0177] Step 3: Data screening.
[0178] The data after data mining is further screened to eliminate noise data (for example, sensor abnormal data) or repeated data included in the data after data mining, and finally valid data for model training can be screened out.
[0179] Step 4: Model training.
[0180] The screened valid data is used for learning and training, so that a deep neural network model can be obtained. When the deep neural network model is used for visual positioning operation, the stability of the visual positioning operation can be improved.
[0181] The deep neural network model mainly consists of a shared backbone network and two head branches: 1) The backbone network can be any backbone network, for example, Visual Geometry Group Network (VGG), Deep residual network (ResNet), and the like. Considering the real-time performance of the visual positioning operation, the preferred backbone network of the present application embodiment can be a MobileNet network. 2) The head branch includes a feature point detection branch and a feature description branch. The feature point detection branch is output by the features output by the backbone network through a convolution layer and a normalization layer, and the feature description branch is obtained by the features output by the backbone network through multiple convolution layers.
[0182] When the model training operation is performed, the following steps can be included:
[0183] Step 41: Relationship identification of image association.
[0184] According to the effective data screened and the recorded sensor information (image, IMU, GPS, etc.), a scene is mapped by using a SLAM mapping algorithm, and in the mapping process, a relative pose relationship between two similar images and a feature point matching relationship can be obtained, wherein the relative pose relationship refers to the relative pose information corresponding to the vehicle poses corresponding to the two similar images.
[0185] In some other examples, the above relative pose information can also be obtained by a laser sensor. Specifically, first pose information corresponding to an image and second pose information corresponding to another image are obtained by the laser sensor, and based on the first pose information and the second pose information, the pose corresponding relationship corresponding to the two similar images can be determined.
[0186] Step 42: training sample.
[0187] After all the effective images are obtained, any one image can be determined as an anchor sample (Anchor example), and then a positive sample (Positive example) and a negative sample (Negative example) corresponding to the anchor sample are determined, that is, each set of training samples contains three images: anchor sample, positive sample, and negative sample.
[0188] For each anchor sample A screened, all images having a co-view relationship with the anchor sample A are determined based on the relative pose information and the feature point matching relationship. Specifically, when the relative pose information is less than a preset deviation and the feature point matching relationship is greater than a preset threshold, it can be determined that the anchor sample and the image have a co-view relationship; when the relative pose information is greater than or equal to the preset deviation and the feature point matching relationship is less than or equal to the preset threshold, it can be determined that the anchor sample and the image do not have a co-view relationship.
[0189] After determining all images having a co-view relationship with the anchor sample A, the distance between the position information corresponding to the above image and the position information corresponding to the anchor sample A is compared with a preset range. When the above position information is less than the preset range (for example, 5m, 6m, 7m, or 10m, etc.), the image having a co-view relationship with the anchor sample A and having a distance less than the preset range from the anchor sample A is determined as a positive sample P. If the image does not have a co-view relationship with the anchor sample A or has a distance greater than or equal to the preset range from the anchor sample A, it is determined as a negative sample N.
[0190] Step 43: model training.
[0191] Based on the anchor sample, the positive sample, and the negative sample, learning training is performed, so that a deep neural network model can be obtained.
[0192] In some examples, the anchor sample, the positive sample and the negative sample are input into a (pre-set) deep neural network model, and a triplet loss function is used to minimize the matching distance between the anchor sample and the positive sample, while maximizing the matching distance between the anchor sample and the negative sample, and a back propagation algorithm is used to update the model parameters, so as to obtain a deep neural network model which can analyze and identify the images corresponding to various environments.
[0193] Step 5: visual positioning.
[0194] After obtaining the deep neural network model, the deep neural network model can be used for real-time visual positioning of the vehicle. In addition, the deep neural network model can be optimized and updated based on the visual positioning result, which is conducive to improving the identification accuracy of the deep neural network model.
[0195] The visual positioning method provided by the application embodiment can continuously collect large-scale data, mine and screen effective training data, and then obtain a deep neural network model based on the training data, or update a configured deep neural network model, thereby improving the stability of the deep neural network model, enabling it to robustly position in different light, different weather, different seasons and different scenes, and thereby solving the problem of positioning failure or instability caused by various environmental changes, and further improving the accuracy and reliability of the method.
[0196] Figure 13 A structural schematic diagram of a model training device provided by the application embodiment is shown in the figure. The application embodiment provides a model training device for executing the above-mentioned Figure 13 corresponding model training method, which can include a first acquisition module 11, a first determination module 12 and a first training module 13. Specifically, Figure 2 The first acquisition module 11 is configured to acquire a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image.
[0197] The first determination module 12 is configured to determine an anchor sample, a positive sample corresponding to the anchor sample and a negative sample corresponding to the anchor sample from the plurality of images according to the sensor data.
[0198] The first training module 13 is configured to learn and train the anchor sample, the positive sample and the negative sample to obtain a deep neural network model, and the deep neural network model is configured to perform visual positioning based on the images.
[0199]
[0200] In some examples, the plurality of different environments comprises at least one of: a plurality of different time periods, a plurality of different weathers, a plurality of different operation scenarios.
[0201] In some examples, the sensor data comprises at least one of: detection data of an inertial measurement unit, detection data of a global positioning system.
[0202] In some examples, when the first acquisition module 11 acquires the plurality of images corresponding to the plurality of different environments and the sensor data corresponding to each image, the first acquisition module 11 is configured to: acquire all the images corresponding to the plurality of different environments and all the sensor data corresponding to each image; identify invalid data included in all the sensor data; remove the invalid data included in all the sensor data to obtain target sensor data; and determine, among all the images, an image corresponding to the target sensor data as a target image.
[0203] In some examples, the invalid data comprises at least one of: abnormal data, repeated data.
[0204] In some examples, before identifying the invalid data included in all the sensor data, the first acquisition module 11 and the first determination module 12 in this embodiment are configured to perform the following steps:
[0205] The first acquisition module 11 is configured to acquire a positioning confidence corresponding to each image and sensor data.
[0206] The first determination module 12 is configured to determine, based on the positioning confidence, among all the images and all the sensor data corresponding to each image, a candidate image and sensor data corresponding to the candidate image.
[0207] In some examples, when the first determination module 12 determines, based on the positioning confidence, among all the images and all the sensor data corresponding to each image, a candidate image and sensor data corresponding to the candidate image, the first determination module 12 is configured to perform: when the positioning confidence is less than or equal to a preset threshold, removing the image and sensor data corresponding to the positioning confidence to obtain a candidate image and candidate sensor data corresponding to the candidate image.
[0208] In some examples, before identifying the invalid data included in all the sensor data, the first acquisition module 11 and the first determination module 12 in this embodiment are configured to perform the following steps:
[0209] The first acquisition module 11 is configured to acquire model requirement information.
[0210] The first determining module 12 is configured to determine, from all the images and all the sensor data, a candidate image corresponding to the model requirement information and candidate sensor data corresponding to the candidate image.
[0211] In some examples, when the first determining module 12 determines, from the sensor data, the anchor sample, the positive sample corresponding to the anchor sample and the negative sample corresponding to the anchor sample from the plurality of images, the first determining module 12 is configured to perform the following steps: selecting one image as the anchor sample from the plurality of images; determining, according to the sensor information, an image having a set relationship with the anchor sample as the positive sample from the plurality of images; and determining, according to the sensor information, an image not having the set relationship with the anchor sample as the negative sample from the plurality of images.
[0212] In some examples, the set relationship includes: having a common view relationship with the anchor sample and having a distance less than a preset range with the anchor sample.
[0213] In some examples, after selecting one image as the anchor sample from the plurality of images, the first obtaining module 11 and the first determining module 12 in the embodiment are configured to perform the following steps:
[0214] The first obtaining module 11 is configured to obtain, from the plurality of images, all the similar sample images corresponding to the anchor sample;
[0215] The first obtaining module 11 is configured to obtain, according to the sensor information, relative pose information and feature point matching information between the similar sample images and the anchor sample;
[0216] The first determining module 12 is configured to detect, based on the relative pose information and the feature point matching information, whether the similar sample images have the set relationship with the anchor sample.
[0217] In some examples, when the first obtaining module 11 obtains, according to the sensor information, the relative pose information between the similar sample images and the anchor sample, the first obtaining module 11 can be configured to perform the following steps: obtaining first sensor information corresponding to the similar sample images and second sensor information corresponding to the anchor sample; and determining, based on the first sensor information and the second sensor information, the relative pose information between the similar sample images and the anchor sample.
[0218] In some examples, when the first determining module 12 detects, based on the relative pose information and the feature point matching information, whether the similar sample images have the set relationship with the anchor sample, the first determining module 12 is configured to perform the following steps: when the relative pose information is less than or equal to a preset threshold value and the feature matching information is greater than or equal to a preset threshold value, it is determined that the similar sample images have the set relationship with the anchor sample; and when the relative pose information is greater than the preset threshold value or the feature matching information is less than the preset threshold value, it is determined that the similar sample images do not have the set relationship with the anchor sample.
[0219] In some examples, when the first training module 13 learns and trains the anchor sample, the positive sample and the negative sample to obtain the deep neural network model, the first training module 13 is configured to perform: obtaining a loss function used for analyzing and processing the anchor sample, the positive sample and the negative sample; learning and training the anchor sample, the positive sample and the negative sample by using the loss function to obtain the deep neural network model.
[0220] In some examples, the loss function is configured to minimize the matching distance between the anchor sample and the positive sample, and maximize the matching distance between the anchor sample and the negative sample.
[0221] In some examples, after obtaining the deep neural network model, the first obtaining module 11 and the first determining module 12 in the embodiment are configured to perform the following steps:
[0222] The first obtaining module 11 is configured to obtain a to-be-analyzed image collected by the to-be-positioned target in the space.
[0223] The first determining module 12 is configured to determine a plurality of standard images used for analyzing and processing the to-be-analyzed image, each standard image corresponding to standard location information; input the to-be-analyzed image and the plurality of standard images into the deep neural network model to obtain positioning information of the to-be-positioned target in the space.
[0224] Figure 13 The apparatus can execute the method of the embodiments described in Figures 1 to 8 、 Figure 12 The embodiments not described in detail in the embodiments described in Figures 1 to 8 、 Figure 12 The execution process and technical effects of the technical solutions can refer to the descriptions of the embodiments described in Figures 1 to 8 、 Figure 12 Therefore, the detailed description is omitted here.
[0225] In one possible design, the structure of the model training apparatus described in Figure 13 may be implemented as an electronic device, which can be a mobile phone, a tablet computer, a server, or various devices. As shown in Figure 14 , the electronic device can include a first processor 21 and a first memory 22. The first memory 22 is configured to store a program supporting the electronic device to execute at least part of the embodiments described in Figures 1-8 、 Figure 12 The first processor 21 is configured to execute the program stored in the first memory 22.
[0226] The program comprises one or more computer instructions, wherein the one or more computer instructions are executed by the first processor 21 to implement the following steps:
[0227] obtaining a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image;
[0228] determining, according to the sensor data, an anchor sample, a positive sample corresponding to the anchor sample, and a negative sample corresponding to the anchor sample from the plurality of images;
[0229] learning and training the anchor sample, the positive sample, and the negative sample to obtain a deep neural network model, the deep neural network model being used for visual positioning operation based on images.
[0230] Optionally, the first processor 21 is further configured to execute all or part of the steps in at least some of the embodiments shown in Figures 1-8 、 Figure 12 .
[0231] The structure of the electronic device can further include a first communication interface 23 for communication between the electronic device and other devices or communication networks.
[0232] In addition, the embodiment of the present application provides a computer storage medium for storing computer software instructions for the electronic device, which includes a program for executing the model training method in at least some of the embodiments shown in Figures 1-8 、 Figure 12 .
[0233] Figure 15 A structural schematic diagram of a visual positioning device provided by the embodiment of the present application; as shown in the accompanying drawings, Figure 15 the embodiment provides a visual positioning device for executing the visual positioning method corresponding to the above-mentioned Figure 10 , the visual positioning device can include a second acquisition module 31, a second determination module 32, and a second positioning module 33. Specifically,
[0234] The second acquisition module 31 is configured to acquire a to-be-analyzed image collected by a vehicle in a space;
[0235] The second determination module 32 is configured to determine a simultaneous localization and mapping (SLAM) database for analyzing and processing the to-be-analyzed image, the SLAM database including a plurality of standard images and standard location information corresponding to each standard image;
[0236] The second positioning module 33 is configured to input the to-be-analyzed image and the SLAM database into a deep neural network model to obtain positioning information of the vehicle in the space, wherein the deep neural network model is trained to be used for visual positioning operation based on images.
[0237] In some instances, the vehicle is equipped with an inertial measurement unit; after obtaining the vehicle's positioning information in space, the second acquisition module 31 and the second determination module 32 in this embodiment can be used to perform the following steps:
[0238] The second acquisition module 31 is used to acquire the detection data of the inertial measurement unit;
[0239] The second determining module 32 is used to determine the vehicle's attitude information based on the detection data; and to generate navigation information corresponding to the vehicle based on the attitude information and positioning information.
[0240] Figure 15 The device shown can perform Figures 9 to 12 For the methods in the embodiments shown, the parts not described in detail in this embodiment can be referred to the [examples / descriptions]. Figures 9 to 12 The embodiments shown are described in detail below. For the implementation process and technical effects of this technical solution, please refer to... Figures 9 to 12 The descriptions in the illustrated embodiments will not be repeated here.
[0241] In one possible design, Figure 15 The structure of the visual positioning device shown can be implemented as an electronic device, which can be various devices such as mobile phones, tablets, and servers. Figure 16 As shown, the electronic device may include a second processor 41 and a second memory 42. The second memory 42 is used to store data supporting the electronic device in performing the above-described actions. Figures 9 to 12 In at least some of the embodiments shown, the visual positioning method program is provided, and the second processor 41 is configured to execute the program stored in the second memory 42.
[0242] The program includes one or more computer instructions, wherein the one or more computer instructions, when executed by the second processor 41, can perform the following steps:
[0243] Acquire images of the vehicle in space that are to be analyzed;
[0244] A synchronous localization and mapping (SLAM) database is determined for analyzing and processing the images to be analyzed. The SLAM database includes multiple standard images and standard location information corresponding to each standard image.
[0245] The images to be analyzed and the SLAM database are input into a deep neural network model to obtain the vehicle's spatial positioning information. The deep neural network model is trained to perform visual positioning operations based on images.
[0246] Optionally, the second processor 41 is also used to perform the aforementioned Figures 9 to 12All or part of the steps in at least some of the embodiments shown.
[0247] In the structure of the electronic device, a second communication interface 43 can also be included, which is used for communication between the electronic device and other devices or communication networks.
[0248] In addition, the embodiments of the present application provide a computer storage medium for storing computer software instructions for electronic devices, which include computer software instructions for executing the above Figures 9 to 12 The program involved in the visual positioning method in at least some of the embodiments shown.
[0249] Figure 17 Another flowchart of the visual positioning method provided by the embodiments of the present application is shown in FIG. 17. Figure 17 As shown, the embodiments provide another visual positioning method, and the execution subject of the visual positioning method is a visual positioning device. It can be understood that the visual positioning device can be implemented as software or a combination of software and hardware. Specifically, the method can include:
[0250] Step S1701: Obtain a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image.
[0251] Step S1702: According to the sensor data, determine an anchor sample, a positive sample corresponding to the anchor sample, and a negative sample corresponding to the anchor sample from the plurality of images.
[0252] In the embodiments, the specific implementation process and implementation effect of steps S1701-S1702 are similar to those of steps S201-S202 in the above embodiments, and specific reference can be made to the above statements, which will not be repeated here.
[0253] Step S1703: The anchor sample, the positive sample, and the negative sample are used as input parameters of a to-be-trained model, and the to-be-trained model is executed to obtain a loss function.
[0254] The to-be-trained model is pre-set for visual positioning operation of the vehicle, and the positioning accuracy corresponding to the to-be-trained model does not meet the preset requirements, especially the positioning accuracy of the visual positioning operation in a plurality of different environments is low. Therefore, in order to improve the accuracy of the visual positioning operation, the anchor sample, the positive sample, and the negative sample obtained can be used as input parameters of the to-be-trained model, and the to-be-trained model is controlled to execute the operation, so that a loss function corresponding to the to-be-trained model can be obtained. The loss function is used to minimize the matching distance between the anchor sample and the positive sample, and maximize the matching distance between the anchor sample and the negative sample.
[0255] Step S1704: optimizing the to-be-trained model according to the loss function, to obtain an optimized model, the optimized model being used for visual positioning of the vehicle.
[0256] After the loss function is obtained, the to-be-trained model can be optimized according to the loss function, so that an optimized model is obtained, which is used for visual positioning of the vehicle, thereby effectively improving the robustness of positioning operation in different light, weather, seasons and different scenes.
[0257] The visual positioning method provided in the embodiment, by obtaining a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image, determining an anchor sample, a positive sample corresponding to the anchor sample and a negative sample corresponding to the anchor sample from the plurality of images according to the sensor data, taking the anchor sample, the positive sample and the negative sample as input parameters of a to-be-trained model, executing the to-be-trained model to obtain a loss function, and then optimizing the to-be-trained model according to the loss function to obtain an optimized model for visual positioning of the vehicle, effectively solves the positioning failure problem caused by extreme environmental changes in the visual positioning process, improves the robustness of positioning operation in different light, weather, seasons and different scenes, and further improves the accuracy and reliability of the visual positioning method.
[0258] Figure 18 Another structural schematic diagram of a visual positioning device provided in the embodiment of the application is shown in FIG. 5. Figure 18 The embodiment provides another visual positioning device, which is used to execute the above-mentioned Figure 17 corresponding visual positioning method, the visual positioning device can include a third acquisition module 51, a third determination module 52, a third processing module 53 and a third optimization module 54. Specifically,
[0259] The third acquisition module 51 is configured to acquire a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image.
[0260] The third determination module 52 is configured to determine an anchor sample, a positive sample corresponding to the anchor sample and a negative sample corresponding to the anchor sample from the plurality of images according to the sensor data.
[0261] The third processing module 53 is configured to take the anchor sample, the positive sample and the negative sample as input parameters of a to-be-trained model, execute the to-be-trained model to obtain a loss function.
[0262] The third optimization module 54 is configured to optimize the to-be-trained model according to the loss function, to obtain an optimized model, the optimized model being used for visual positioning of the vehicle.
[0263] Figure 18 The device shown can perform Figure 17 For the methods in the embodiments shown, the parts not described in detail in this embodiment can be referred to the [examples / descriptions]. Figure 17 The embodiments shown are described in detail below. For the implementation process and technical effects of this technical solution, please refer to... Figure 17 The descriptions in the illustrated embodiments will not be repeated here.
[0264] In one possible design, Figure 18 The structure of the visual positioning device shown can be implemented as an electronic device, which can be various devices such as mobile phones, tablets, and servers. Figure 19 As shown, the electronic device may include a third processor 61 and a third memory 62. The third memory 62 is used to store data supporting the electronic device in performing the above-described actions. Figure 17 In at least some of the embodiments shown, the program of the visual positioning method is provided, and the third processor 61 is configured to execute the program stored in the third memory 62.
[0265] The program includes one or more computer instructions, wherein the one or more computer instructions, when executed by the third processor 61, can perform the following steps:
[0266] Acquire multiple images corresponding to multiple different environments, as well as sensor data corresponding to each image;
[0267] Based on the sensor data, anchor samples, positive samples corresponding to anchor samples, and negative samples corresponding to anchor samples are determined from the plurality of images;
[0268] The anchor sample, the positive sample, and the negative sample are used as input parameters of the model to be trained, and the model to be trained is executed to obtain the loss function;
[0269] The model to be trained is optimized according to the loss function to obtain an optimized model, which is used for visual localization of vehicles.
[0270] Optionally, the third processor 61 is also used to perform the aforementioned Figure 17 All or part of the steps in at least some of the embodiments shown.
[0271] The structure of the electronic device may also include a third communication interface 63 for communication between the electronic device and other devices or communication networks.
[0272] In addition, embodiments of the present invention provide a computer storage medium for storing computer software instructions used by an electronic device, which includes instructions for executing the above-described... Figure 17The program involved in the visual positioning method in at least part of the embodiments shown.
[0273] The device embodiments described above are merely illustrative, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0274] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of a general hardware platform as necessary, and of course can also be realized by means of combination of hardware and software. Based on such understanding, the above technical solutions can be embodied in the form of computer products, and the present application can be in the form of computer program products implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program codes.
[0275] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be realized by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices produce a device for realizing the functions specified in the flowcharts and / or block diagrams. Figure One The functions specified in one flow or multiple flows and / or blocks Figure One The device for realizing the functions specified in one block or multiple blocks.
[0276] These computer program instructions can also be stored in a computer readable storage medium to guide the computer or other programmable data processing devices to work in a specific way, so that the instructions stored in the computer readable storage medium produce a product including instruction devices, which realize the functions specified in the flowcharts and / or block diagrams. Figure One The functions specified in one flow or multiple flows and / or blocks Figure One The device for realizing the functions specified in one block or multiple blocks.
[0277] These computer program instructions can also be loaded into a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure One Figure One
[0278] In one typical arrangement, the computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0279] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) and / or cache memory, non-volatile memory, such as read-only memory (ROM), EPROM, and / or flash memory, etc. The memory is an example of computer readable media.
[0280] Computer readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer readable media does not include transitory media, such as modulated data signals and carrier waves.
[0281] Finally, it should be noted that the above-mentioned embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A model training method comprising: obtaining a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image; determining, from the sensor data, an anchor sample, a positive sample corresponding to the anchor sample, and a negative sample corresponding to the anchor sample from the plurality of images; the positive sample is an image having a set relationship with the anchor sample, and the negative sample is an image not having the set relationship with the anchor sample; the set relationship includes: having a co-view relationship with the anchor sample, and a shooting location corresponding to the positive sample and a shooting location corresponding to the anchor sample being less than a preset range; learning and training the anchor sample, the positive sample, and the negative sample to obtain a deep neural network model, the deep neural network model being used for visual positioning operation based on an image.
2. The method of claim 1, wherein, Obtaining a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image comprises: obtaining all images corresponding to the plurality of different environments and all sensor data corresponding to each image; identifying invalid data included in all sensor data; removing invalid data included in all sensor data to obtain target sensor data; in all images, determining an image corresponding to the target sensor data as a target image.
3. The method of claim 2, wherein, Before identifying the invalid data included in all sensor data, the method further comprises: obtaining a positioning confidence corresponding to each image and sensor data; based on the positioning confidence, determining a candidate image and sensor data corresponding to the candidate image from all images and all sensor data corresponding to each image.
4. The method of claim 2, wherein, Before identifying the invalid data included in all sensor data, the method further comprises: obtaining model requirement information; determining, from all images and all sensor data, a candidate image corresponding to the model requirement information and candidate sensor data corresponding to the candidate image.
5. The method of claim 1, wherein, According to the sensor data, determining an anchor sample, a positive sample corresponding to the anchor sample, and a negative sample corresponding to the anchor sample from the plurality of images comprises: selecting one image as an anchor sample from the plurality of images; determining, from the sensor data, an image having a set relationship with the anchor sample as a positive sample from the plurality of images; determining, from the sensor data, an image not having a set relationship with the anchor sample as a negative sample from the plurality of images.
6. The method of claim 5, wherein, After selecting one image as an anchor sample from the plurality of images, the method further comprises: obtaining all similar sample images corresponding to the anchor sample from the plurality of images; obtaining relative pose information and feature point matching information between the similar sample images and the anchor sample according to the sensor data; based on the relative pose information and the feature point matching information, detecting whether the similar sample images have a set relationship with the anchor sample.
7. The method of claim 6, wherein, According to the sensor data, obtaining relative pose information between the similar sample images and the anchor sample comprises: obtaining first sensor information corresponding to the similar sample image and second sensor information corresponding to the anchor sample; determining relative pose information between the similar sample image and the anchor sample based on the first sensor information and the second sensor information.
8. The method of claim 6, wherein, detecting whether the similar sample image and the anchor sample have a set relationship based on the relative pose information and feature point matching information, including: when the relative pose information is less than or equal to a preset threshold, and the feature point matching information is greater than or equal to a preset threshold, it is determined that the similar sample image and the anchor sample have a set relationship; when the relative pose information is greater than a preset threshold, or the feature point matching information is less than a preset threshold, it is determined that the similar sample image and the anchor sample do not have a set relationship.
9. The method of any of claims 1-8, wherein, learning and training the anchor sample, the positive sample and the negative sample to obtain a deep neural network model, including: obtaining a loss function for analyzing and processing the anchor sample, the positive sample and the negative sample; learning and training the anchor sample, the positive sample and the negative sample using the loss function to obtain the deep neural network model.
10. The method of any of claims 1-8, wherein, After obtaining the deep neural network model, the method further includes: obtaining a to-be-analyzed image collected by a to-be-positioned target in space; determining a plurality of standard images for analyzing and processing the to-be-analyzed image, each standard image corresponding to standard location information; inputting the to-be-analyzed image and the plurality of standard images into the deep neural network model to obtain positioning information of the to-be-positioned target in space.
11. A visual positioning method, comprising: obtaining a to-be-analyzed image collected by a vehicle in space; determining a simultaneous localization and mapping (SLAM) database for analyzing and processing the to-be-analyzed image, the SLAM database including a plurality of standard images and standard location information corresponding to each standard image; inputting the to-be-analyzed image and the SLAM database into a deep neural network model to obtain positioning information of the vehicle in space, wherein the deep neural network model is trained for visual positioning operation based on images, and is obtained based on learning and training of an anchor sample, a positive sample and a negative sample; the positive sample is an image having a set relationship with the anchor sample, and the negative sample is an image not having a set relationship with the anchor sample; the set relationship includes having a co-view relationship with the anchor sample, and a distance between a shooting location corresponding to the positive sample and a shooting location corresponding to the anchor sample being less than a preset range.
12. A visual positioning method, comprising: obtaining a plurality of images corresponding to a plurality of different environments and sensor data corresponding to each image; determining an anchor sample, a positive sample corresponding to the anchor sample, and a negative sample corresponding to the anchor sample from the plurality of images according to the sensor data; The positive sample is an image having a set relationship with the anchor sample, and the negative sample is an image not having the set relationship with the anchor sample; the set relationship includes: having a co-view relationship with the anchor sample, a shooting location corresponding to the positive sample being less than a preset range from a shooting location corresponding to the anchor sample; The anchor sample, the positive sample and the negative sample are taken as input parameters of a to-be-trained model, and a loss function is obtained by executing the to-be-trained model; The to-be-trained model is optimized according to the loss function, and an optimized model is obtained, which is used for visual positioning of a vehicle.
Citation Information
Patent Citations
Visual positioning method and device, storage medium and electronic equipment
CN110211181A
Image recognition model training method, image recognition method and device
CN112329826A