Positioning method and device and storage medium

By enhancing the initial visual data and fusing it with GNSS and IMU data, the problem of insufficient positioning accuracy during the take-off and landing stage of the drone is solved, and the full utilization of visual information and the improvement of positioning accuracy is achieved.

CN120468894APending Publication Date: 2025-08-12CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510686245.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing GNSS/INS/visual fusion positioning algorithms lack positioning accuracy during the take-off and landing stage of the drone and do not fully utilize visual sensor information.

Method used

By obtaining the initial visual data of the carrier to be tested and inputting the pre-trained visual enhancement model, the number of image features of the visual data is enhanced, and then input with the GNSS data and IMU data into the fusion positioning model to perform data fusion to obtain the positioning result.

Benefits of technology

It improves the utilization rate and positioning accuracy of visual information, solves the problem of insufficient utilization of visual sensor information, and improves the positioning accuracy of the drone's take-off and landing stage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120468894A_ABST
    Figure CN120468894A_ABST
Patent Text Reader

Abstract

The invention provides a positioning method, equipment and a storage medium, relates to the technical field of navigation and positioning services, and aims to solve the technical problems of low visual information utilization rate and insufficient positioning precision of a positioning method in a general technology. The positioning method comprises the following steps: acquiring initial visual data, global navigation satellite system (GNSS) data and inertial measurement unit (IMU) data of a carrier to be measured; the initial visual data comprises multiple frames of associated visual images; inputting the initial visual data into a pre-trained visual enhancement model to obtain enhanced visual data; the number of image features in the enhanced visual data is greater than the number of image features in the initial visual data; and inputting the enhanced vision data, the GNSS data and the IMU data into the fusion positioning model to obtain a positioning result of the to-be-detected carrier. According to the invention, the utilization rate of visual information is improved, and the positioning precision of the positioning result of the to-be-detected carrier is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of navigation and positioning services, and in particular to a positioning method, device, and storage medium. Background Art

[0002] With the development of location services and drone technology, the Global Navigation Satellite System (GNSS) is now able to provide high-precision location information for outdoor environments. For urban canyons and semi-occluded urban environments, the introduction of an inertial navigation system (Inertial Navigation System) can solve the problem of reduced accuracy of GNSS positioning and navigation due to signal interference in the above-mentioned environments. However, during the take-off and landing phase of the drone, the occlusion environment is complex, the GNSS error increases, and the INS positioning error accumulates too quickly, which will cause the positioning accuracy of the GNSS / INS combined positioning structure to decrease rapidly during the take-off and landing phase of the drone. Among the existing algorithm models, GNSS / INS / visual information has become the optimal choice for seamless positioning of drones in all scenarios, but the current algorithm model has not fully utilized the role of visual sensor information in fusion positioning results.

[0003] Therefore, the GNSS / INS / vision fusion positioning and navigation algorithm model for seamless positioning of drones in all scenarios needs to be further improved. Summary of the Invention

[0004] The embodiments of the present disclosure provide a positioning method, device, and storage medium, which aim to solve the technical problems of low utilization of visual information and insufficient positioning accuracy in positioning methods in general technologies.

[0005] To achieve the above objectives, this application adopts the following technical solutions:

[0006] In a first aspect, a positioning method is provided, comprising: obtaining initial visual data, global navigation satellite system (GNSS) data, and inertial measurement unit (IMU) data of a carrier to be measured; the initial visual data comprises multiple frames of associated visual images; the initial visual data is input into a pre-trained visual enhancement model to obtain enhanced visual data; the number of image features in the enhanced visual data is greater than the number of image features in the initial visual data; and the enhanced visual data, GNSS data, and IMU data are input into a fusion positioning model to obtain a positioning result of the carrier to be measured.

[0007] Optionally, the visual enhancement model is trained as follows:

[0008] A training sample set is obtained; the training sample set includes training samples and labels of the training samples; the training samples include: multiple frames of sample visual images collected by the sample carrier; the labels of the training samples include: sample enhanced images after visual enhancement of the visual images collected by the sample carrier; the original visual enhancement model is trained based on the training sample set to obtain a trained visual enhancement model.

[0009] Optionally, the multiple frames of sample visual images include a first sample data image set and a second sample data image set; and training the original visual enhancement model based on the training sample set to obtain a trained visual enhancement model includes:

[0010] The sample data images in the first sample data image set are merged to obtain multidimensional image data; the multidimensional image data are input into the first network structure in the visual enhancement model to obtain a first processing result; the first network structure includes multiple convolution layers and / or normalized pooling layers; the first processing result is input into the spatial attention mechanism structure and the channel attention mechanism structure in the visual enhancement model respectively, and the output result of the spatial attention mechanism structure and the output result of the channel attention mechanism structure are merged to obtain a second processing result; the sample data images in the second sample data image set are convolutionally processed to obtain a third processing result; the second processing result and the third processing result are merged, and the merged processing result is input into the long short-term memory network model structure in the visual enhancement model to obtain a prediction result of the training sample; the loss value of the prediction result of the training sample and the label of the training sample is determined according to the loss function, and the visual enhancement model is trained according to the loss value until the visual enhancement model meets the convergence condition to obtain a trained visual enhancement model.

[0011] Optionally, the enhanced visual data, GNSS data, and IMU data are input into a fusion positioning model to obtain the positioning results of the vehicle under test, including:

[0012] GNSS data and IMU data are input into the first local filter in the fusion positioning model to obtain output information of the first local filter; the output information of the first local filter includes a first fusion result of the GNSS data and the IMU data; the first local filter is used to perform data fusion through the Kalman filtering algorithm; the enhanced visual data and IMU data are input into the second local filter in the fusion positioning model to obtain output information of the second local filter; the output information of the second local filter includes a second fusion result of the enhanced visual data and the IMU data; the second local filter is used to perform data fusion through the Kalman filtering algorithm; the first fusion result and the second fusion result are input into the main filter in the fusion positioning model to obtain output information of the main filter; the output information of the main filter includes the positioning result of the carrier to be measured.

[0013] Optionally, the output information of the first local filter also includes: weight information of the first local filter; the weight information of the first local filter is the first covariance matrix obtained by the first local filter according to the Kalman filter algorithm to fuse the GNSS data and the IMU data, which is used to measure the accuracy of the first fusion result; the output information of the second local filter also includes: weight information of the second local filter; the weight information of the second local filter is the second covariance matrix obtained by the second local filter according to the Kalman filter algorithm to fuse the enhanced visual data and the IMU data, which is used to measure the accuracy of the second fusion result; the output information of the main filter also includes: comprehensive weight information and allocation factor of the main filter; the comprehensive weight information of the main filter is the global covariance matrix determined by the main filter according to the first covariance matrix and the second covariance matrix, which is used to measure the accuracy of the positioning result of the carrier to be tested; the allocation factor is used to distribute information to the first local filter and the second local filter; the information in the information distribution includes the global covariance matrix and / or the system noise covariance matrix.

[0014] Optionally, obtain initial visual data of the carrier to be tested, including:

[0015] Acquire state data of a carrier to be tested and binocular vision information data of the carrier to be tested; the binocular vision information data includes: at least one of image data, parallax information, depth information, and three-dimensional coordinates collected by the carrier to be tested; determine a motion state of the carrier to be tested based on the state data of the carrier to be tested, and determine an image frame extraction rule for the binocular vision information data based on the motion state; perform image frame extraction on the binocular vision information data based on the image frame extraction rule for the binocular vision information data to obtain initial visual data of the carrier to be tested.

[0016] Optionally, determining the motion state of the carrier to be tested according to the state data of the carrier to be tested includes:

[0017] The state data of the carrier to be tested is input into the pre-trained particle swarm optimization-back propagation deep learning network model to obtain the motion state of the carrier to be tested.

[0018] Optionally, the motion state of the carrier to be measured includes the motion speed of the carrier to be measured; and the image frame extraction rule of the binocular vision information data includes:

[0019] The number of image frame extraction intervals of binocular vision information data; the number of image frame extraction intervals is positively correlated with the movement speed of the carrier to be measured.

[0020] In a second aspect, a positioning device is provided, which includes: a communication unit and a processing unit.

[0021] The communication unit is used to obtain initial visual data, global navigation satellite system (GNSS) data, and inertial measurement unit (IMU) data of the vehicle under test. The initial visual data includes multiple frames of associated visual images.

[0022] The processing unit is configured to input the initial visual data into a pre-trained visual enhancement model to obtain enhanced visual data. The number of image features in the enhanced visual data is greater than the number of image features in the initial visual data.

[0023] The processing unit is also used to input the enhanced visual data, GNSS data and IMU data into the fusion positioning model to obtain the positioning result of the carrier to be tested.

[0024] In a third aspect, a positioning device is provided, comprising a memory and a processor; the memory is used to store computer-executable instructions, and the processor is connected to the memory via a bus; when the positioning device is running, the processor executes the computer-executable instructions stored in the memory, so that the positioning device performs the positioning method of the first aspect.

[0025] The positioning device may be an electronic device or a portion of an electronic device, such as a chip system within the electronic device. The chip system is configured to support the electronic device in implementing the functions described in the first aspect and any possible implementation thereof, such as acquiring and determining the data and / or information involved in the positioning method. The chip system includes a chip and may also include other discrete components or circuit structures.

[0026] In a fourth aspect, a computer-readable storage medium is provided, the computer-readable storage medium including computer-executable instructions. When the computer-executable instructions are executed on a computer, the computer is caused to execute the positioning method described in the first aspect.

[0027] In a fifth aspect, a computer program product is also provided, which includes a computer program or instructions. When the computer instructions are run on a positioning device, the positioning device executes the positioning method as described in the first aspect above.

[0028] It should be noted that the above-mentioned computer instructions may be stored in whole or in part on a computer-readable storage medium. The computer-readable storage medium may be packaged together with the processor of the positioning device or separately from the processor of the positioning device, and this embodiment of the application does not limit this.

[0029] The description of the second, third, fourth and fifth aspects of this application can refer to the detailed description of the first aspect.

[0030] In the embodiments of this application, the names of the positioning devices described above do not limit the devices or functional modules themselves. In actual implementation, these devices or functional modules may appear by other names. For example, the processing unit may also be referred to as a processing module, a processor, etc. As long as the functions of the various devices or functional modules are similar to those of this application, they are within the scope of the claims of this application and their equivalents.

[0031] The technical solution provided by this application brings at least the following beneficial effects:

[0032] Based on any of the above aspects, an embodiment of the present application provides a positioning method, including: first, obtaining the initial visual data, GNSS data and IMU data of the carrier to be tested (the initial visual data includes multiple frames of associated visual images); secondly, inputting the initial visual data into a pre-trained visual enhancement model to obtain enhanced visual data (the number of image features in the enhanced visual data is greater than the number of image features in the initial visual data); then inputting the enhanced visual data, GNSS data and IMU data into the fusion positioning model to obtain the positioning result of the carrier to be tested.

[0033] From the above, it can be seen that this application obtains enhanced visual data by inputting the initial visual data into a pre-trained visual enhancement model, thereby enhancing the image features of the visual data, improving the utilization rate of visual information data, and solving the problem of insufficient utilization of visual information data collected by visual sensors.

[0034] Secondly, this application will enhance the input of visual data, GNSS data and IMU data into the fusion positioning model to improve the positioning accuracy of the fusion positioning method and give full play to the role of visual information data in the positioning method.

[0035] The beneficial effects of the first, second, third, fourth and fifth aspects of this application can all be referred to the analysis of the above beneficial effects, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A schematic diagram of the structure of a positioning system provided in an embodiment of the present application;

[0037] Figure 2 A schematic diagram of the hardware structure of a positioning device provided in an embodiment of the present application;

[0038] Figure 3 A flowchart of a positioning method provided in an embodiment of the present application;

[0039] Figure 4 A flowchart of another positioning method provided in an embodiment of the present application;

[0040] Figure 5A flowchart of another positioning method provided in an embodiment of the present application;

[0041] Figure 6 A schematic diagram of the training process of a visual enhancement model provided in an embodiment of the present application;

[0042] Figure 7 A flowchart of another positioning method provided in an embodiment of the present application;

[0043] Figure 8 A flowchart of another positioning method provided in an embodiment of the present application;

[0044] Figure 9 A schematic diagram of the overall process of a positioning method provided in an embodiment of the present application;

[0045] Figure 10 A schematic diagram of the overall process of another positioning method provided in an embodiment of the present application;

[0046] Figure 11 A schematic diagram of the structure of a fusion positioning model provided in an embodiment of the present application;

[0047] Figure 12 A flowchart of another positioning method provided in an embodiment of the present application;

[0048] Figure 13 A schematic structural diagram of a positioning device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0050] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being more preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0051] In order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order.

[0052] Before giving a detailed introduction to the positioning method provided by this application, a brief introduction to the application scenarios and implementation environment involved in this application is first given.

[0053] First, a brief introduction to the application scenarios involved in this application is given.

[0054] As described in the background technology, in existing positioning methods, during the take-off and landing phase of the UAV, due to the complex occlusion environment, the GNSS error increases, and then the INS positioning error accumulates too quickly, resulting in a decrease in the positioning accuracy of the existing GNSS / INS positioning structure during the take-off and landing phase of the UAV.

[0055] What you need to know is that lidar has high positioning accuracy, is less affected by the environment, and has high construction and maintenance costs; visual cameras have high accuracy within the field of view, high accuracy in areas with rich feature textures, and low cost; UWB positioning has high accuracy, and long-sequence positioning requires the deployment of a series of auxiliary base stations, which invisibly makes construction and maintenance costs high.

[0056] To address these challenges, existing positioning methods have proposed a GNSS / INS / visual information algorithm model, which has become the optimal choice for seamless, full-scene positioning of drones. However, these existing GNSS / INS / visual information algorithm models fail to fully utilize the role of visual sensor information in fusion positioning results. Therefore, how to fully utilize visual sensor information in positioning methods and improve their accuracy is an urgent issue.

[0057] In response to the above problems, an embodiment of the present application provides a positioning method, including: first, obtaining the initial visual data, GNSS data and IMU data of the carrier to be tested (the initial visual data includes multiple frames of associated visual images); secondly, inputting the initial visual data into a pre-trained visual enhancement model to obtain enhanced visual data (the number of image features in the enhanced visual data is greater than the number of image features in the initial visual data); then inputting the enhanced visual data, GNSS data and IMU data into a fusion positioning model to obtain the positioning result of the carrier to be tested.

[0058] From the above, it can be seen that this application obtains enhanced visual data by inputting the initial visual data into a pre-trained visual enhancement model, thereby enhancing the image features of the visual data, improving the utilization rate of visual information data, and solving the problem of insufficient utilization of visual information data collected by visual sensors.

[0059] Secondly, this application will enhance the input of visual data, GNSS data and IMU data into the fusion positioning model to improve the positioning accuracy of the fusion positioning method and give full play to the role of visual information data in the positioning method.

[0060] The implementation environment of the above positioning method can be the positioning system provided in the embodiment of the present application.

[0061] Figure 1 FIG. 1 shows a schematic diagram of the structure of a positioning system provided by an embodiment of the present application. Figure 1 As shown, the positioning system includes: a positioning device 101 and a data acquisition device 102.

[0062] The positioning device 101 and the data acquisition device 102 are in communication connection with each other.

[0063] In practical applications, the positioning device 101 can be connected to any number of data acquisition devices 102. Figure 1 An example of a positioning device 101 connected to a data acquisition device 102 is used for description.

[0064] In an embodiment of the present application, the data acquisition device 102 is used to collect relevant data of the carrier to be tested (including status data, GNSS data, IMU data, binocular vision information data, etc.) and provide it to the positioning device 101, so that the positioning device 101 determines the positioning result of the carrier to be tested based on the relevant data of the carrier to be tested provided by the data acquisition device 102.

[0065] Optionally, the physical device of the positioning device 101 may be a server, a terminal, or other types of electronic devices, which is not limited in the embodiment of the present application.

[0066] Optionally, the terminal may be a device that provides voice and / or data connectivity to a user, a handheld device with wireless connection capabilities, or other processing devices connected to a wireless modem. A wireless terminal may communicate with one or more core networks via a radio access network (RAN). A wireless terminal may be a mobile terminal, such as a mobile phone (or "cellular" phone) and a computer with a mobile terminal, or a portable, pocket-sized, handheld, computer-built-in, or vehicle-mounted mobile device that exchanges voice and / or data with a radio access network, such as a mobile phone, tablet computer, laptop computer, netbook, or personal digital assistant (PDA).

[0067] Optionally, the above-mentioned server can be a server in a server cluster (consisting of multiple servers), or a chip in the server, or a system on a chip in the server, or can be implemented through a virtual machine (VM) deployed on a physical machine. This embodiment of the present application does not limit this.

[0068] Optionally, the physical device of the data acquisition device 102 can be an airborne sensor for collecting relevant data of the carrier to be tested, or it can be other types of acquisition devices for collecting relevant data of the carrier to be tested. The data acquisition device 102 can also include at least one of an IMU sensor, a GNSS receiver, and a binocular camera, which is not limited here.

[0069] Optionally, the positioning device 101 and the data acquisition device 102 can be two independent devices or integrated into the same device. When the positioning device 101 and the data acquisition device 102 are integrated into the same device, the data acquisition device 102 can be an acquisition module (e.g., a data collector, etc.) of the positioning device 101.

[0070] It is easy to understand that when the positioning device 101 and the data acquisition device 102 are integrated into the same device, the communication between the positioning device 101 and the data acquisition device 102 is carried out by means of communication between the internal modules of the device. In this case, the communication process between the two is the same as when the positioning device 101 and the data acquisition device 102 are independent of each other.

[0071] For ease of understanding, this application is described by taking the positioning device 101 and the data acquisition device 102 as an example in which they are independent of each other.

[0072] The positioning device 101 in the positioning system includes: Figure 2 The following are the components included. Figure 2 Taking the positioning device shown as an example, the hardware structure of the positioning device 101 is introduced.

[0073] Figure 2 Schematic diagram of the hardware structure of a positioning device provided in an embodiment of the present application. The positioning device includes a processor 21, a memory 22, a communication interface 23, and a bus 24. The processor 21, the memory 22, and the communication interface 23 can be connected via the bus 24.

[0074] The processor 21 is the control center of the positioning device and can be a single processor or a collective term for multiple processing elements. For example, the processor 21 can be a general-purpose central processing unit (CPU) or other general-purpose processor. The general-purpose processor can be a microprocessor or any conventional processor.

[0075] As an embodiment, the processor 21 may include one or more CPUs, such as Figure 2 CPU0 and CPU1 are shown in the figure.

[0076] The memory 22 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0077] In one possible implementation, the memory 22 can exist independently of the processor 21 and can be connected to the processor 21 via a bus 24 to store instructions or program codes. When the processor 21 calls and executes the instructions or program codes stored in the memory 22, the positioning method provided in the following embodiments of the present application can be implemented.

[0078] In the embodiment of the present application, for the positioning device, the software programs stored in the memory 22 are different, so the positioning device implements different functions. The functions performed by each device will be described in conjunction with the following flowchart.

[0079] In another possible implementation, the memory 22 may also be integrated with the processor 21 .

[0080] The communication interface 23 is used to connect the positioning device to other devices via a communication network, which may be Ethernet, wireless access network, wireless local area network (WLAN), etc. The communication interface 23 may include a receiving unit for receiving data and a sending unit for sending data.

[0081] The bus 24 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, and the like. For ease of presentation, Figure 2 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0082] It should be pointed out that Figure 2 The structure shown in the figure does not constitute a limitation on the positioning device, except Figure 2 In addition to the components shown, the positioning device may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0083] The positioning method provided in the embodiments of the present application is described in detail below with reference to the accompanying drawings.

[0084] The positioning method provided in the embodiment of the present application is applied to Figure 1 The positioning device 101 in the positioning system shown is as follows: Figure 3 As shown, the positioning method provided in the embodiment of the present application includes:

[0085] S301: The positioning device obtains initial visual data, GNSS data and IMU data of the carrier to be measured.

[0086] The initial visual data includes multiple frames of associated visual images.

[0087] Specifically, the positioning device needs to obtain the positioning result of the carrier to be measured based on the enhanced visual data, GNSS data and IMU data of the carrier to be measured. Therefore, it is necessary to obtain the initial visual data, GNSS data and IMU data of the carrier to be measured.

[0088] Optionally, the initial visual data is a continuous visual image frame data.

[0089] Illustratively, the GNSS data of the carrier to be measured may include satellite signal information measured by a GNSS receiver, such as pseudorange, carrier phase, and other data.

[0090] Exemplarily, the IMU includes an accelerometer sensor and a gyroscope sensor, and the IMU data of the carrier to be measured includes acceleration data measured by the accelerometer and angular velocity data measured by the gyroscope, as well as attitude, position and velocity data obtained by integrating the acceleration data and angular velocity data.

[0091] S302: The positioning device inputs the initial visual data into a pre-trained visual enhancement model to obtain enhanced visual data.

[0092] The number of image features in the enhanced visual data is greater than the number of image features in the initial visual data.

[0093] Specifically, when determining the positioning result of the carrier to be measured, the positioning device needs to use enhanced visual data to improve the positioning accuracy of the positioning result of the carrier to be measured, improve the utilization rate of visual data, and enhance visual features. Therefore, the positioning device needs to obtain enhanced visual data.

[0094] Optionally, the vision enhancement model may also be referred to as a parallel multi-attention mechanism-convolutional neural network-long short-term memory network (PMA-CNN-LSTM) model.

[0095] S303: The positioning device inputs the enhanced visual data, GNSS data, and IMU data into a fusion positioning model to obtain a positioning result of the carrier to be measured.

[0096] Specifically, the positioning device inputs the enhanced visual data, GNSS data and IMU data into the fusion positioning model to achieve the fusion of the three data, and finally obtains the positioning result of the carrier to be tested.

[0097] Optionally, the positioning result of the carrier to be measured is the fusion positioning information of enhanced visual data, GNSS data and IMU data. The positioning result of the carrier to be measured may include the position, speed and attitude information of the carrier to be measured.

[0098] In some embodiments, as Figure 4 As shown in Figure 2, the visual enhancement model is trained in the following way:

[0099] S401: The positioning device obtains a training sample set.

[0100] The training sample set includes training samples and labels of training samples; the training samples include: multiple frames of sample visual images collected by the sample carrier; the labels of training samples include: sample enhanced images after visual enhancement of the visual images collected by the sample carrier.

[0101] Optionally, the label of the training sample can be a single enhanced image.

[0102] For example, a time series of camera video data can be collected and extracted frame by frame into image data. The first six frames of the image data are clustered together to form a sample carrier, which is used to collect multiple frames of sample visual images. Starting from the seventh frame, the images are manually enhanced. The enhanced image results, also called enhanced single images, serve as labels for the training samples.

[0103] S402: The positioning device trains the original vision enhancement model based on the training sample set to obtain a trained vision enhancement model.

[0104] Specifically, the positioning device trains the original visual enhancement model using a training sample set. During the training process, the loss function is used to determine the loss value between the prediction results of the visual enhancement model and the training sample labels until the visual enhancement model meets the convergence conditions, thereby obtaining a trained visual enhancement model.

[0105] In some embodiments, combined Figure 4 ,like Figure 5 As shown, in the above S402, the multiple frames of sample visual images include a first sample data image set and a second sample data image set.

[0106] Optionally, the first sample data image set and the second sample data image set are a section of continuous image frame data.

[0107] For example, the first five frames in a cluster of data (including six frames of data) can be divided into a first sample data image set, and the sixth frame can be divided into a second sample data image set.

[0108] The positioning device trains the original visual enhancement model based on the training sample set to obtain a trained visual enhancement model, which specifically includes:

[0109] S501: The positioning device merges sample data images in a first sample data image set to obtain multi-dimensional image data.

[0110] Specifically, the positioning device merges the sample data images in the first sample data image set to form multi-dimensional image data, in preparation for subsequent convolution processing of the image data.

[0111] For example, the positioning device may stack the first five frames of data in a cluster of data to form multi-dimensional data.

[0112] S502: The positioning device inputs the multi-dimensional image data into a first network structure in the vision enhancement model to obtain a first processing result.

[0113] The first network structure includes multiple convolutional layers and / or normalized pooling layers.

[0114] Optionally, the positioning device processes the multi-dimensional image data through multiple convolution layers, extracts and converts the features of the image to obtain multiple feature maps, and then performs feature selection and dimensionality reduction on the feature maps through normalized pooling layers.

[0115] It can be seen that by processing multi-dimensional image data through a network structure composed of multiple convolutional layers and normalized pooling layers, the accuracy of the visual enhancement model can be improved through operations such as image feature extraction and optimization, and the generalization ability of the visual enhancement model can be enhanced through normalized pooling layer processing, as well as the computational efficiency of the visual enhancement model.

[0116] S503. The positioning device inputs the first processing result into the spatial attention mechanism structure and the channel attention mechanism structure in the visual enhancement model respectively, and merges the output result of the spatial attention mechanism structure and the output result of the channel attention mechanism structure to obtain a second processing result.

[0117] For example, the results of multi-layer convolution processing are fed into a spatial attention mechanism (SAM) and a channel attention mechanism (CAM) respectively, and then the two results are stacked to obtain a second processing result. The SAM structure combines the spatial information in the feature map to enhance the contribution of task-related areas. The CAM structure combines the importance of each channel in the feature map to enhance the contribution of task-related channels and suppress redundant information.

[0118] It can be seen that the positioning device can improve the quality of image feature extraction and enhance the adaptability of the visual enhancement model to spatial changes through SAM structure processing. Through CAM structure processing, the impact of redundant information can be reduced, enabling the visual enhancement model to efficiently and concisely perform tasks based on key image features. By processing the first processing result through the SAM structure and CAM structure, the image feature texture is enhanced.

[0119] S504: The positioning device performs convolution processing on the sample data images in the second sample data image set to obtain a third processing result.

[0120] Specifically, the positioning device extracts features of the sample data images by performing convolution processing on the sample data images in the second sample data image set, so as to facilitate other subsequent operations based on the image features.

[0121] Exemplarily, the sample data image in the second sample data image set (the sixth frame of data in a cluster of data) is subjected to convolution processing to obtain a third processing result.

[0122] S505: The positioning device merges the second processing result and the third processing result, and inputs the merged processing result into the long short-term memory network model structure in the vision enhancement model to obtain a prediction result of the training sample.

[0123] Specifically, the positioning device merges the second processing result and the third processing result to provide continuous image frame data for subsequent processing; the merged processing result is processed through the long short-term memory network model structure in the visual enhancement model to extract image features and retain key image information of the image frame data.

[0124] Exemplarily, the second processing result in S503 and the third processing result in S504 are superimposed and put into a long short-term memory network (LSTM) structure to obtain the final result (the third processing result).

[0125] It can be understood that the positioning device can fuse the information of multiple frames of images and integrate image features under different perspectives or states by dividing multiple frames of sample visual images into a first sample data set and a second sample data set, performing different operations on them respectively (processing multiple frames of image data separately) and then superimposing the results of the two. This can enhance the visual enhancement model's ability to understand images, help to mine the temporal information in the image data (6 frames of continuous image data), namely, the movement changes and state evolution of objects over time, and improve the visual enhancement model's ability to process image data.

[0126] S506. The positioning device determines the loss value of the prediction result of the training sample and the label of the training sample according to the loss function, and trains the visual enhancement model according to the loss value until the visual enhancement model meets the convergence condition, thereby obtaining a trained visual enhancement model.

[0127] For example, during the training of the visual enhancement model, the loss value between the predicted result of the visual enhancement model and the label of the training sample (the enhanced single image can also be called the true result) is determined by the loss function until the loss value meets the convergence condition (the convergence condition can be less than a preset threshold), thereby obtaining a trained visual enhancement model.

[0128] It can be understood that by judging the loss value between the prediction result of the visual model and the label of the training sample through the loss function, the prediction result error of the visual enhancement model can be minimized, that is, the accuracy of the visual enhancement model can be improved.

[0129] For example, Figure 6 This is a diagram of a training process of a visual enhancement model provided in an embodiment of the present application. Figure 6As shown in the figure, the visual enhancement model mainly includes: LSTM structure, SAM structure and CAM structure. The training process of this visual enhancement model includes:

[0130] First, the k-5th to k-1th frame images (5 frames of data in the historical record) are acquired, and the above frame images are stacked to obtain stacked data.

[0131] Then, the stacked data is input into multiple convolutional layers and normalized pooling layers to obtain the backbone network processing results.

[0132] That is, the stacked data is processed by a backbone network structure constructed by multiple convolutional layers and normalized pooling layers to obtain the backbone network processing result.

[0133] Next, the processing results of the backbone network are input into the spatial attention mechanism structure and the channel attention mechanism structure respectively, and the processing results of the channel attention mechanism structure and the spatial attention mechanism structure are obtained respectively. The processing results of the channel attention mechanism structure and the processing results of the spatial attention mechanism structure are stacked to form a fusion result.

[0134] For example, Figure 6 As shown in Figure 2, the process of inputting the backbone network processing results into the channel attention mechanism structure (i.e., CAM structure) to obtain the processing results of the channel attention mechanism structure is as follows:

[0135] First, the input features are obtained and the input features are subjected to maximum pooling and average pooling respectively to obtain the maximum pooling results and the average pooling results.

[0136] For example, the input feature can be an input feature map of size H × W × C. By performing maximum pooling and average pooling in the spatial dimension on an input feature map of size H × W × C, two 1 × 1 × C feature maps can be obtained (pooling in the spatial dimension compresses the spatial size to facilitate subsequent learning of channel features), namely the maximum pooling result and the average pooling result.

[0137] Then, the maximum pooling result and the average pooling result are respectively input into a shared multi-layer perceptron (also called a shared MLP) for learning to obtain the MLP first processing result and the MLP second processing result after processing by the multi-layer perceptron.

[0138] Exemplarily, the maximum pooling result and the average pooling result are respectively sent to the multi-layer perceptron (MLP) for learning to obtain two 1×1×C feature maps (based on the features of the MLP learning channel dimension and the importance of each channel), namely the first processing result of the MLP and the second processing result of the MLP after processing by the multi-layer perceptron.

[0139] Subsequently, the first processing result of MLP and the second processing result of MLP are added, and then mapped by Sigmoid activation function to obtain the channel attention weight matrix M C .

[0140] It can be seen that in the above-mentioned visual enhancement model, when processing image frame data, the stacked data is processed by convolution layers and normalized pooling layers. The processing of convolution layers and normalized pooling layers is based on convolution operations. It extracts image features by fusing spatial and channel information through local perception. This method usually assigns the same importance to each channel feature, but in fact some channels are more important than others. For this reason, the CAM structure is introduced into the visual enhancement model to focus on important channels in the feature map. The CAM structure can assign different weights to each channel, emphasizing channels that contribute more to the task while suppressing channels that contribute less to the task, thereby improving the overall performance of the visual enhancement model.

[0141] For example, Figure 6 As shown in Figure 2, the process of inputting the backbone network processing results into the spatial attention mechanism structure (i.e., SAM structure) to obtain the processing results of the spatial attention mechanism structure is as follows:

[0142] First, obtain the channel feature F, and perform maximum pooling and average pooling on the channel feature F.

[0143] For example, the channel feature F can be an input feature map of size H×W×C. Maximum pooling and average pooling are performed on the feature map in the channel dimension to obtain two H×W×1 feature maps (pooling is performed in the channel dimension to compress the channel size to facilitate the subsequent learning of spatial features). Then, the results of maximum pooling and average pooling are concatenated according to the channel dimension to obtain a feature map of size H×W×2.

[0144] Then, convolution is performed through the convolution layer.

[0145] For example, a convolution operation may be performed on the concatenated result according to a 7×7 convolutional layer to obtain a feature map with a size of H×W×1.

[0146] Subsequently, the convolution processing result is processed by the Sigmoid activation function to obtain the spatial attention weight matrix M S .

[0147] It can be seen that in the above-mentioned visual enhancement model, when processing image frame data, the stacked data is processed by the convolution layer and the normalized pooling layer. When processing images, the convolution layer and the normalized pooling layer often process the entire image uniformly without distinguishing the importance of different regions. This method may ignore the important information contained in certain areas of the image, thereby limiting the performance of the model when processing complex images. For this reason, the SAM structure is introduced into the visual enhancement model. Similar to the CAM structure, the SAM structure focuses more on the spatial information in each channel. The SAM structure emphasizes the areas that contribute more to the task by weighting specific spatial positions in the feature map, while suppressing the areas that contribute less to the task, thereby improving the performance of the visual enhancement model.

[0148] At the same time, the k-th frame image (positioning frame) is obtained, and the k-th frame image (positioning frame) is convolved through the convolution layer to obtain the convolution result.

[0149] Secondly, the convolution result and the fusion result are superimposed to obtain the superimposed result, and the superimposed result is input into the long short-term memory network model structure to obtain the prediction result.

[0150] For example, Figure 6 As shown, the superimposed results are input into the Long Short-Term Memory (LSTM) model structure to obtain the prediction result. The LSTM structure mainly consists of four substructures: cell state, input gate, forget gate, and output gate. The core of the LSTM structure is the cell state. These gate control mechanisms jointly control the inflow, retention, and outflow of information. The calculation formula for the superimposed results of the Long Short-Term Memory (LSTM) model structure is as follows:

[0151] For example, it is assumed that the current moment is t (six frames of images are processed sequentially) and the previous moment is t-1.

[0152] Forget Gate:

[0153] f t =σ f (W xf *X t +W hf *H t-1 +W cf *C t-1 +b f );

[0154] Among them, f t is the output of the forget gate at time t, σ f is the activation function of the forget gate (usually the Sigmoid function), W xf For the current input Xt To the weight matrix of the forget gate, X t is the input vector at time t, W hf H is the hidden state at the previous moment t-1 To the weight matrix of the forget gate, H t-1 is the hidden state vector of the previous moment, W cf is the cell state C at the previous moment t-1 To the weight matrix of the forget gate, C t-1 is the cell state vector at the previous moment, b f is the bias vector of the forget gate.

[0155] For example, the forget gate determines which image information is forgotten from the cell state at the previous moment. t-1 The feature vector X corresponding to the current image frame t , and then through the weight matrix (W xf 、W hf 、W cf ) and bias (b f ) are linearly combined and then pass through the activation function (σ f ) to get f t , f t The value range is between [0,1].

[0156] Input Gate:

[0157] i t =σ in (W xi *X t +W hi *H t-1 +W ci *C t-1 +b i );

[0158] Among them, i t is the output of the input gate at time t, σ in is the activation function of the input gate (usually Sigmoid function), W xi For the current input X t The weight matrix to the input gate, X t is the input vector at time t, W hi H is the hidden state at the previous moment t-1 The weight matrix to the input gate, H t-1 is the hidden state vector at the previous moment, W ci is the cell state C at the previous moment t-1 The weight matrix to the input gate, C t-1 is the cell state vector at the previous moment, b iis the bias vector of the input gate.

[0159] For example, the input gate determines how much image information is incorporated into the cell state C at the current moment. t Combined with the hidden state H of the previous moment t-1 and cell state C t-1 The image information accumulated in the image is converted into xi 、W hi 、W ci ) and bias (b i ) are linearly combined and then pass through the activation function (σ in ) get i t (range is between [0,1]), i t Each element in represents the proportion of the corresponding image information integrated into the cell state (i t The closer the element value is to 1, the more important the corresponding image information is), which is used to determine how much image information needs to be added to the cell state at the current moment.

[0160] Cell status update:

[0161] C t =f t ·C t-1 +i t tanh(W xc *X t +W hc *H t-1 +b c );

[0162] Among them, C t is the cell state vector at time t, C t-1 is the cell state vector at the previous moment, i t is the output of the input gate at time t, tanh is the hyperbolic tangent activation function, W xc For the current input X t to the weight matrix of the candidate cell state, X t is the input vector at time t, W hc H is the hidden state at the previous moment t-1 To the weight matrix of candidate cell states, H t-1 is the hidden state vector of the previous moment, b c is the bias vector of the candidate cell state.

[0163] For example, according to the output of the forget gate and the input gate, the cell state C at the current moment is updated t First, according to the output f of the forget gate t For the cell state C at the previous moment t-1 Selectively forget, then pass through the input gate it Determine how much new image information to incorporate (by tanh(W xc *X t +W hc *H t-1 +b c ) is generated), the tanh function maps the output value to between -1 and 1, and obtains a processed candidate cell state, which is then compared with i t After multiplication, the cell state C at the current moment is obtained by adding the previously forgotten state. t .

[0164] Output gate:

[0165] o t =σ out (W xo *X t +W ho *H t-1 +W co *C t +b o );

[0166] Among them, t is the output of the output gate at time t, σ out is the activation function of the output gate (usually Sigmoid function), W xo For the current input X t The weight matrix to the output gate, X t is the input vector at time t, W ho H is the hidden state at the previous moment t-1 The weight matrix to the output gate, H t-1 is the hidden state vector at the previous moment, W co is the cell state C at time t t The weight matrix to the output gate, C t is the cell state vector at time t, b o is the bias vector of the output gate.

[0167] For example, the output gate determines the current cell state C t Which image information can be output to the hidden state H at the current moment? t According to the current cell state C t , current image feature vector X t And the hidden state H at the previous moment t-1 etc., through the weight matrix (W xo 、W ho 、W co ) and bias (b o ) are linearly combined and then pass through the activation function (σ out ) get ot , o t The value range is between [0,1], and the importance of the output image information is the same as o t The corresponding element values in are positively correlated.

[0168] Hide status update:

[0169] H t =o t tanh(C t );

[0170] Among them, H t is the hidden state vector at time t, o t is the output of the output gate at time t, tanh is the hyperbolic tangent activation function, C t is the cell state vector at time t.

[0171] For example, from the above formula, we can know that based on the current cell state C t (accumulated image information), after being processed by the tanh function, the image information is mapped to the appropriate range (-1 to 1), and then passes through the output gate o t The control of the current image determines the more important image information in the cell state and the proportion of these image information that needs to be output to the hidden state, generating the hidden state H at the current moment. t . H t It contains the most valuable image information for the current image frame after screening and integration based on image features.

[0172] It is understandable that the visual enhancement model uses an LSTM structure to process image frames, and the dynamic temporal relationship between image frames is modeled through the gate control mechanism (forget gate, input gate, and output gate) in the LSTM structure. According to the temporal relationship between image frames, the image features are used to determine which information in the image is forgotten and which information is retained, which can better identify the target's motion trajectory. By analyzing the position offset of the target in multiple frames, its motion pattern is implicitly learned and the future position of the target is predicted. Through temporal modeling, spatial features (image feature information extracted from the image) and dynamic changes (changes in the state, position, etc. of the target at consecutive time points) are combined, giving the visual enhancement model "memory" and "prediction" capabilities.

[0173] Subsequently, the model parameters in the above structures can be trained according to the loss values of the predicted results and the actual results until the visual enhancement model meets the convergence conditions and the trained visual enhancement model (i.e. Figure 6 Model training is performed based on the prediction results as shown).

[0174] In some embodiments, combined Figure 3 ,like Figure 7 As shown, in the above S303, the positioning device inputs the enhanced visual data, GNSS data and IMU data into the fusion positioning model to obtain the positioning result of the carrier to be measured, which specifically includes:

[0175] Optionally, the fusion positioning model can also be called a GNSS / IMU / vision-enhanced coupling structure.

[0176] S701: The positioning device inputs GNSS data and IMU data into a first local filter in a fusion positioning model to obtain output information of the first local filter.

[0177] The output information of the first local filter includes a first fusion result of the GNSS data and the IMU data; the first local filter is used to perform data fusion through a Kalman filter algorithm.

[0178] Optionally, the first local filter is responsible for fusing IMU data and GNSS data, and its input information is information such as the position, speed or attitude of the carrier to be measured formed by GNSS and IMU, and outputs the fusion result and weight information of GNSS data and IMU data.

[0179] Optionally, the IMU data fused by the first local filter includes acceleration and angular velocity data of the carrier to be measured measured by the IMU. The GNSS data fused by the first local filter includes pseudorange and carrier phase data of the carrier to be measured measured by a BeiDou or GNSS receiver.

[0180] In some embodiments, the output information of the first local filter further includes: weight information of the first local filter.

[0181] The weight information of the first local filter is a first covariance matrix obtained by fusing the GNSS data and the IMU data by the first local filter according to the Kalman filter algorithm, which is used to measure the accuracy of the first fusion result.

[0182] Optionally, the weight information of the first local filter is the first covariance matrix of the first local filter, which is used to evaluate the accuracy or credibility of the first fusion result of the first local filter. The smaller the value of the first covariance matrix, the higher the accuracy or credibility of the first fusion result.

[0183] It is important to note that the Kalman filter algorithm estimates the system state in a noisy environment by fusing the predicted values of the system dynamic model with the observed data. The core idea of the Kalman filter algorithm is to build a system state estimation framework based on the system state equation and the observation equation, and to make a real-time, optimal estimate of the system state. It continuously cycles through the two stages of prediction and update, using the optimal estimate at the previous moment to predict the state at the current moment, and then correcting this prediction value based on the observed value at the current moment, thereby gradually obtaining a more accurate estimate of the system state. The relevant formula is as follows:

[0184] System state equation:

[0185]

[0186] Where X(t) is the state vector of the system at time t, including all key state information of the carriers that need to be estimated in the system; is the time derivative of the system state vector X(t), representing the rate of change of the system state over time; F(t) is the state transfer matrix, describing the transfer relationship of the system state vector X(t) from time t to the next time; G(t) is the control input matrix, describing the mode and speed of influence of external control input on the system state vector X(t); W(t) is the process noise vector, which is the interference term that causes random changes in the system state due to various unpredictable factors during the operation of the system.

[0187] Observation equation:

[0188] Z(t)=H(t)X(t)+V(t);

[0189] Among them, Z(t) is the observation vector of the system at time t, which is a vector composed of data actually measured by the sensor; H(t) is the measurement matrix, which is the mapping relationship between each state variable in the system state vector X(t) and the observation vector Z(t); V(t) is the measurement noise vector, which reflects the error of the sensor itself during the measurement process. It is also usually assumed to be Gaussian white noise with zero mean and a certain covariance.

[0190] Optionally, the above system state equation and observation equation are discretized and the optimal estimate at the current moment (assuming the current moment is moment k) is used to predict the state at the next moment (assuming the next moment is moment k+1). The specific formula is as follows:

[0191] System state prediction equation:

[0192]

[0193] in, It is the state prediction value at time k+1, which contains the relevant state information of the carrier to be estimated in the system (such as the position, speed, attitude, etc. of the carrier in the positioning system); is the optimal estimate at the current moment; F is the state transfer matrix, which describes the transition relationship of the system state from time k to time k+1; G is the control input matrix; is the process noise vector, which represents the uncertainty caused by unpredictable factors in system operation. It is generally assumed to be Gaussian white noise with zero mean and a certain covariance (represented by Q).

[0194] Covariance matrix prediction equation:

[0195]

[0196] in, is the covariance prediction matrix at time k+1; is the covariance prediction matrix at time k; F T is the transposed matrix of the state transfer matrix; Q is the noise covariance matrix.

[0197] Optionally, it can be seen from the above formula that the above system state prediction equation and covariance matrix prediction equation are equations of the prediction stage.

[0198] Kalman filter gain equation:

[0199]

[0200] Among them, K k+1 is the Kalman gain, which is the key weight for fusing the predicted value and the observed value; H is the measurement matrix, which reflects the mapping relationship between the system state vector and the observation vector; H T is the transposed matrix of the measurement matrix H; R is the measurement noise covariance matrix.

[0201] Observation state update equation:

[0202]

[0203] Among them, Z k+1 is the observation vector at time k+1 (a vector composed of data actually measured by the sensor), which contains the observation information related to the system state that can be obtained at time k+1.

[0204] Covariance matrix update equation:

[0205]

[0206] Among them, P k+1 is the state covariance matrix at time k+1; I is the identity matrix.

[0207] Optionally, it can be seen from the above formula that the above Kalman filter gain equation, observation state update equation and covariance matrix update equation are equations of the update stage.

[0208] Exemplarily, the first local filter fuses the GNSS data and the IMU data according to the Kalman filter algorithm, specifically including:

[0209] First, the state equation of the first local filter and the observation equation of the first local filter are constructed as follows:

[0210] The overall state equation of the fusion positioning model:

[0211]

[0212] Among them, X g is the system state vector, X g =[δp x ,δp y ,δp z ,δv x ,δv y ,δv z ,δψ x ,δψ y ,δψ z , ε x , ε y , ε z ] T ,δp x ,δp y ,δp z is the three-dimensional position change of the carrier, δv x ,δv y ,δv z is the velocity variation of the carrier in three directions, δψ x ,δψ y ,δψ z is the three-dimensional attitude change of the carrier, is the constant bias of the carrier's accelerometer, ε x , ε y , ε z is the constant drift of the carrier's gyroscope; is the derivative form of the state vector; F is the system state transfer matrix; G is the system noise input matrix; W is the system white noise vector, its mean is 0, and its variance is Q, W(t)=[W ax , W ay , W az , W gx , W gy , W gz] T , W ax , W ay , W az is the accelerometer white noise, W gx , W gy , W gz is the gyroscope white noise.

[0213] Optionally, it can be seen from the above formula that because the state vector X1 of the first local filter and the state vector X2 of the second local filter are the same as the overall state vector X of the fusion positioning model g Consistent, so X g =X1=X2. Therefore, the state equation of the first local filter and the state equation of the second local filter are consistent with the overall state equation of the fusion positioning model.

[0214] The state equation of the first local filter is:

[0215]

[0216] Where X1 is the state vector of the first local filter; is the derivative form of the state vector.

[0217] Optionally, for the first local filter, the position change (also called three-way position change) and attitude angle change of the IMU and GNSS are selected as observation quantities.

[0218] The observation equation of the first local filter is:

[0219] Z1(t)=H1(t)X1(t)+V1(t);

[0220] Where Z1 is the observation information, including the measurement information of GNSS and IMU, Z1(t)=[δp x ,δp y ,δp z , H imu -H gnss , P imu -P gnss , R imu -R gnss ] T ,δp x ,δp y ,δp z is the three-dimensional position change of the carrier, H imu , P imu , R imu Calculate the azimuth, pitch and roll angles of the output carrier for IMU, H gnss , P gnss , R gnssis the observation noise of the azimuth, pitch and roll angles of the GNSS measurement output carrier; V1(t) is the observation noise vector, The observation noise of the three-way position of the carrier output by GNSS measurement; is the observation noise of the azimuth, pitch and roll angles of the carrier output by GNSS measurement; H1(t) is the measurement matrix.

[0221] Secondly, the overall state equation of the fusion positioning model and the observation equation of the first local filter are discretized. The discretized formula is as follows:

[0222]

[0223] Among them, X1(k) is the system state vector of the first local filter at time k; Φ(k+1,k) is the state transfer matrix, which describes the transition relationship of the system state from time k to time k+1; Γ(k+1,k) is the system noise input matrix; W(k) is the system white noise vector, which represents the uncertainty caused by unpredictable factors in the operation of the carrier; Z1(k) is the observation vector of the first local filter at time k (i.e., the input GNSS data and IMU data); H1(k) is the observation matrix; V1(k) is the observation noise vector, which reflects the error in the fusion process of GNSS data and IMU data.

[0224] Then, based on the state equation of the first local filter in discrete form and the observation equation of the first local filter, the time update and measurement update of the first local filter are calculated. The specific calculation process and formula are as follows:

[0225] Optionally, the time update of the first local filter includes state prediction and covariance prediction.

[0226] The state prediction equation of the first local filter:

[0227] X1(k+1)=Φ(k+1,k)X1(k)+W(k+1,k);

[0228] Among them, X1(k) is the state estimate of the first local filter at time k; X1(k+1) is the state prediction value of the first local filter at time k+1; Φ(k+1,k) is the state transfer matrix, which describes the change law of the first local filter state from time k to time k+1.

[0229] Optionally, it can be seen from the above formula that the main content of the state prediction calculation of the first local filter is to map the current state X1(k) to the predicted value X1(k+1) at the next moment through the state transfer matrix Φ(k+1,k), and the system white noise W(k+1,k) reflects the model prediction uncertainty (such as the impact of air resistance on carrier motion).

[0230] The covariance prediction equation of the first local filter is:

[0231] P1(k+1,k)=Φ(k+1,k)P1(k)Φ T (k+1,k)+Γ(k+1,k)Q1(k)Γ T (k+1,k);

[0232] Among them, P1(k+1,k) is the error covariance matrix of the predicted k+1 moment; P1(k) is the error covariance matrix of the k moment, which reflects the uncertainty of the state estimation; Φ T (k+1,k) is the transposed matrix of the state transfer matrix Φ(k+1,k); Γ(k+1,k) is the noise input matrix; Γ T (k+1,k) is the transposed matrix of the noise input matrix Γ(k+1,k); Q1(k) is the noise covariance matrix.

[0233] Optionally, as can be seen from the above formula, the main content of the covariance prediction calculation of the first local filter is calculated by Φ(k+1,k)P1(k)Φ T (k+1,k) transmits the uncertainty P1(k) of the current state estimation, reflecting the impact of the model linear transformation on the error. Among them, the credibility of the prediction result is negatively correlated with the prediction covariance P1(k+1,k).

[0234] Optionally, the measurement update of the first local filter includes calculating the Kalman gain, state update and covariance update.

[0235] The Kalman gain equation of the first local filter is:

[0236] K1(k+1)=P1(k+1,k)H1 T (k+1)[H1(k+1)·P1(k+1,k)H1 T (k+1)+R1(k+1)] -1 ;

[0237] Among them, K1(k+1) is the Kalman gain of the first local filter at time k+1; H1(k+1) is the measurement matrix; H1 T(k+1) is the transposed matrix of the measurement matrix H1(k+1); R1(k+1) is the measurement noise covariance matrix, which describes the statistical characteristics of the noise generated by the carrier during the measurement data process.

[0238] Optionally, as can be seen from the above formula, the main content of the Kalman gain calculation of the first local filter is to achieve the optimal fusion of the state prediction value of the first local filter and the actual observation value (IMU data and enhanced visual data). Specifically, when the observation noise R1(k+1) (also known as measurement noise) is low, the Kalman gain increases, giving priority to trusting the actual observation value; when the prediction covariance P1(k+1,k) is small, it means that the prediction result X1(k+1) is highly credible.

[0239] The state update equation of the first local filter is:

[0240] X1(k+1)=X1(k+1,k)+K1(k+1)[Z1(k+1)-H1(k+1)X1(k+1,k)];

[0241] Among them, X1(k+1) is the state update value of the first local filter (also known as the optimal estimate value of the first local filter at time k+1, representing the fusion result of GNSS data and IMU data); Z1(k+1) is the observation value of the first local filter at time k+1.

[0242] Optionally, it can be seen from the above formula that the main content of the state update calculation of the first local filter is to correct the predicted state (the predicted state is X1(k+1,k)) through the measurement residual (the measurement residual is Z1(k+1)-H1(k+1)X1(k+1,k)). The measurement residual reflects the deviation between the predicted value and the actual observation value, and the Kalman gain K1(k+1) is used to dynamically adjust the correction weight.

[0243] The covariance update equation of the first local filter is:

[0244] P1(k+1)=[I-K1(k+1)H1(k+1)]P1(k+1,k);

[0245] Among them, P1(k+1) is the state covariance matrix at time k+1, which is used to reflect the uncertainty of the updated state estimation; I is the identity matrix.

[0246] Optionally, it can be seen from the above formula that the main content of the covariance update calculation of the first local filter is to update the predicted covariance through the Kalman gain K1(k+1), and the updated covariance matrix P1(k+1) is negatively correlated with the credibility of the optimal estimate X1(k+1) (also called the fusion result).

[0247] It can be seen that the first local filter fuses GNSS data and IMU data according to the Kalman filter algorithm. First, the state equation and observation equation of the first local filter are determined, the equations are discretized, and the state value at time k+1 is predicted based on the state value at time k. This includes the state prediction and covariance prediction process of the first local filter, which enables the prediction of the carrier state at the next moment from the carrier state at the current moment. Then, the prediction results are corrected and corrected based on the measured GNSS data and IMU data, including the three processes of Kalman gain calculation, state update, and covariance update of the first local filter, to obtain the optimal estimate of the fusion of GNSS data and IMU data and the covariance matrix of the first local filter.

[0248] It is understandable that by fusing GNSS data and IMU data, the error of a single data source can be compensated. By fusing data from two different sources and comprehensively considering information such as the position, speed and attitude of the carrier to be measured, a more accurate position estimate can be obtained, thereby improving the positioning accuracy of the positioning results. The fusion of GNSS data and IMU data can cope with situations where the signal is interrupted. For example, in some complex environments, such as canyon cities, tunnels, indoors, etc., the GNSS signal may be temporarily interrupted and unable to provide continuous and reliable position information. The IMU, with its own inertial measurement characteristics, can continue to infer the position, speed and attitude changes of the carrier based on previous data in a short time. After fusing with the GNSS data, it can ensure that the positioning information remains continuous during periods of poor GNSS signals, avoiding blank periods in positioning and maintaining stable operation of the carrier.

[0249] S702: The positioning device inputs the enhanced visual data and the IMU data into the second local filter in the fusion positioning model to obtain output information of the second local filter.

[0250] The output information of the second local filter includes a second fusion result of enhanced visual data and IMU data; the second local filter is used to perform data fusion through a Kalman filter algorithm.

[0251] Optionally, the second local filter is responsible for fusing IMU data and enhanced visual data (also called binocular visual information data), whose input information is the observation information of the carrier formed by the enhanced visual data and IMU data, and outputs the fusion result and weight information of the enhanced visual data and IMU data.

[0252] Optionally, the IMU data fused by the second local filter includes acceleration and angular velocity data of the carrier to be measured measured by the IMU. The enhanced visual data fused by the second local filter is obtained by applying the binocular visual information of the carrier to be measured obtained by the binocular vision camera to the visual enhancement model to obtain enhanced visual data, and then analyzing and calculating the successfully matched feature points in the image in the enhanced visual data (including the left and right images obtained by the binocular vision camera) to obtain the distance information of the binocular image matching feature points.

[0253] In some embodiments, the output information of the second local filter further includes: weight information of the second local filter.

[0254] Among them, the weight information of the second local filter is the second covariance matrix obtained by the second local filter fusing the enhanced visual data and the IMU data according to the Kalman filter algorithm, which is used to measure the accuracy of the second fusion result.

[0255] Optionally, the weight information of the second local filter is the second covariance matrix of the second local filter, which is used to evaluate the accuracy or credibility of the second fusion result of the second local filter. The smaller the second covariance matrix value, the higher the accuracy or credibility of the second fusion result.

[0256] Exemplarily, the second local filter fuses the enhanced visual data and the IMU data according to the Kalman filter algorithm, specifically including:

[0257] First, the state equation of the second local filter and the observation equation of the second local filter are constructed as follows:

[0258] The overall state equation of the fusion positioning model:

[0259]

[0260] Among them, X g is the system state vector, X g =[δp x ,δp y ,δp z ,δv x ,δv y ,δv z ,δψ x ,δψ y ,δψ z , ε x , ε y , ε z ] T ,δp x ,δp y ,δp zis the three-dimensional position change of the carrier, δv x ,δv y ,δv z is the velocity variation of the carrier in three directions, δψ x ,δψ y ,δψ z is the three-dimensional attitude change of the carrier, is the constant bias of the carrier's accelerometer, ε x , ε y , ε z is the constant drift of the carrier's gyroscope; is the derivative form of the state vector; F is the system state transfer matrix; G is the system noise input matrix; W is the system white noise vector, its mean is 0, and its variance is Q, W(t)=[W ax , W ay , W az , W gx , W gy , W gz ] T , W ax , W ay , W az is the accelerometer white noise, W gx , W gy , W gz is gyro white noise.

[0261] Optionally, it can be seen from the above formula that because the state vector X1 of the first local filter and the state vector X2 of the second local filter are the same as the overall state vector X of the fusion positioning model g Consistent, so X g =X1=X2. Therefore, the state equation of the first local filter and the state equation of the second local filter are consistent with the overall state equation of the fusion positioning model.

[0262] The state equation of the second local filter is:

[0263]

[0264] Where X2 is the state vector of the second local filter; is the derivative form of the state vector.

[0265] Optionally, for the second local filter, the difference between the pseudorange values of the IMU and the enhanced visual data and the pseudorange rate are selected as the observation values.

[0266] The observation equation of the second local filter is:

[0267] Z2(t)=H2(t)X2(t)+V2(t);

[0268] Where Z2 is the measurement information, Z2(t)=[ρ imu -ρ visual , ρ imu , are the pseudorange and pseudorange rate of IMU respectively, ρ visual , are the pseudorange and pseudorange rate of the binocular visual odometry respectively; V2(t) is the observation noise vector, ε ρ is pseudorange random noise, The random noise of pseudorange rate; H2(t) is the observation matrix.

[0269] Secondly, the overall state equation of the fusion positioning model and the observation equation of the second local filter are discretized. The discretized formula is as follows:

[0270]

[0271] Among them, X2(k) is the system state vector of the second local filter at time k; Φ(k+1,k) is the state transfer matrix, which describes the transition relationship of the system state from time k to time k+1; Γ(k+1,k) is the system noise input matrix; W(k) is the system white noise vector, which represents the uncertainty caused by unpredictable factors in the operation of the carrier; Z2(k) is the observation vector of the second local filter at time k, that is, the input enhanced visual data and IMU data; H2(k) is the observation matrix; V2(k) is the observation noise vector, which reflects the error in the fusion process of enhanced visual data and IMU data.

[0272] Then, based on the state equation of the second local filter and the observation equation of the second local filter in discrete form, the time update and measurement update of the second local filter are calculated. The specific calculation process and formula are as follows:

[0273] Optionally, the time update of the second local filter includes state prediction and covariance prediction.

[0274] The state prediction equation of the second local filter is:

[0275] X2(k+1)=Φ(k+1,k)X2(k)+W(k+1,k);

[0276] Among them, X2(k) is the state estimate of the second local filter at time k; X2(k+1) is the state prediction value of the second local filter at time k+1; Φ(k+1,k) is the state transfer matrix, which describes the change law of the system state from time k to time k+1, and W(k+1,k) is the system white noise.

[0277] Optionally, as can be seen from the above formula, the main content of the state prediction calculation of the second local filter is to map the current state X2(k) to the predicted value X2(k+1) at the next moment through the state transfer matrix Φ(k+1,k), and the system white noise W(k+1,k) reflects the uncertainty of the model prediction (such as the impact of air resistance on the movement of the carrier).

[0278] The covariance prediction equation of the second local filter is:

[0279] P2(k+1,k)=Φ(k+1,k)P2(k)Φ T (k+1,k)+Γ(k+1,k)Q2(k)Γ T (k+1,k);

[0280] Among them, P2(k+1,k) is the error covariance matrix of the predicted k+1 moment; P2(k) is the error covariance matrix of the k moment, which reflects the uncertainty of the state estimation; Φ T (k+1,k) is the transposed matrix of the state transfer matrix Φ(k+1,k); Γ(k+1,k) is the noise input matrix; Γ T (k+1,k) is the transposed matrix of the noise input matrix Γ(k+1,k); Q2(k) is the system noise covariance matrix.

[0281] Optionally, as can be seen from the above formula, the main content of the covariance prediction calculation of the second local filter is calculated by Φ(k+1,k)P2(k)Φ T (k+1,k) transmits the uncertainty P2(k) of the current state estimation, reflecting the impact of the model linear transformation on the error. Among them, the credibility of the prediction result is negatively correlated with the prediction covariance P2(k+1,k).

[0282] Optionally, the measurement update of the second local filter includes calculating the Kalman gain, state update and covariance update.

[0283] The Kalman gain equation of the second local filter is:

[0284] K2(k+1)=P2(k+1,k)H2 T (k+1)[H2(k+1)·P2(k+1,k)H2 T (k+1)+R2(k+1)] -1 ;

[0285] Among them, K2(k+1) is the Kalman gain of the second local filter at time k+1, H2(k+1) is the measurement matrix; H2 T(k+1) is the transposed matrix of the measurement matrix H2(k+1); R2(k+1) is the measurement noise covariance matrix, which describes the statistical characteristics of the noise generated by the carrier in the measurement data process.

[0286] Optionally, as can be seen from the above formula, the main content of the Kalman gain calculation of the second local filter is to achieve the optimal fusion of the state prediction value of the second local filter and the actual observation value (IMU data and enhanced visual data). Specifically, when the observation noise R2(k+1) (also known as measurement noise) is low, the Kalman gain increases, giving priority to trusting the actual observation value; when the prediction covariance P2(k+1,k) is small, it means that the prediction result X2(k+1) is highly credible.

[0287] The state update equation of the second local filter is:

[0288] X2(k+1)=X2(k+1,k)+K2(k+1)[Z2(k+1)-H2(k+1)X2(k+1,k)];

[0289] Among them, X2(k+1) is the state update value of the second local filter (also known as the optimal estimate value at time k+1, representing the fusion result of enhanced visual data and IMU data); Z2(k+1) is the observation value of the second local filter at time k+1.

[0290] Optionally, it can be seen from the above formula that the main content of the state update calculation of the second local filter is to correct the predicted state (the predicted state is X2(k+1,k)) through the measurement residual (the measurement residual is Z2(k+1)-H2(k+1)X2(k+1,k)). The measurement residual reflects the deviation between the predicted value and the actual observation value, and the Kalman gain K2(k+1) is used to dynamically adjust the correction weight.

[0291] The covariance update equation of the second local filter is:

[0292] P2(k+1)=[I-K2(k+1)H2(k+1)]P2(k+1,k);

[0293] Among them, P2(k+1) is the covariance matrix at time k+1, which is used to reflect the uncertainty of the updated state estimation; I is the identity matrix.

[0294] Optionally, it can be seen from the above formula that the main content of the covariance update calculation of the second local filter is to update the predicted covariance through the Kalman gain K2(k+1), and the updated covariance matrix P2(k+1) is negatively correlated with the credibility of the optimal estimate X2(k+1) (also called the fusion result).

[0295] It can be seen that the second local filter fuses the enhanced visual data and IMU data according to the Kalman filter algorithm, and outputs the optimal estimate X2(k+1) at time k+1 and the covariance matrix P2(k+1) of the second local filter.

[0296] It can be seen that the second local filter fuses the enhanced visual data and IMU data according to the Kalman filter algorithm. First, the state equation and observation equation of the second local filter are determined, the equations are discretized, and the state value at time k+1 is predicted based on the state value at time k. This includes the state prediction and covariance prediction process of the second local filter, which enables the prediction of the carrier state at the next moment by the carrier state at the current moment. Then, the prediction results are corrected and corrected based on the measured enhanced visual data and IMU data. This includes three processes: Kalman gain calculation, state update, and covariance update of the second local filter, to obtain the optimal estimate of the fusion of the enhanced visual data and IMU data and the covariance matrix of the second local filter.

[0297] It can be understood that by fusing enhanced visual data with IMU data, combining IMU motion vividness with visual feature point matching, outputting environmental perception information such as relative displacement, obstacle distance, scene depth, etc., the ability to understand local scenes is enhanced.

[0298] S703: The positioning device inputs the first fusion result and the second fusion result into a main filter in the fusion positioning model to obtain output information of the main filter.

[0299] The output information of the main filter includes the positioning result of the carrier to be measured.

[0300] Optionally, the main filter integrates and distributes information of the first local filter and the second local filter on the one hand, and feeds back the estimated value of the system state error to the IMU to correct its cumulative error on the other hand; the input of the main filter is the output information of the first local filter and the second local filter, and the output result is the fused positioning information and distribution factor and comprehensive weight information of the IMU data, GNSS data and enhanced visual data.

[0301] In some embodiments, the output information of the main filter further includes: comprehensive weight information and allocation factors of the main filter.

[0302] The comprehensive weight information of the main filter is a global covariance matrix determined by the main filter based on the first covariance matrix and the second covariance matrix, which is used to measure the accuracy of the positioning result of the measured carrier. The allocation factor is used to allocate information between the first local filter and the second local filter; the information in the information allocation includes the global covariance matrix and / or the system noise covariance matrix.

[0303] Optionally, the comprehensive weight information of the main filter is a global covariance matrix obtained by weighted summation of the first covariance matrix value and the second covariance matrix value, which is used to evaluate the accuracy and robustness of the positioning result of the carrier to be measured (also called the global optimal estimate), and is negatively correlated with the global optimal estimate. The allocation factor of the main filter is that after the main filter completes the overall optimal synthesis of the fusion positioning model, the information feedback amount is formed based on the information allocation principle in the federal filter, and information is allocated to the first local filter and the second local filter to obtain the information allocation coefficient of the first local filter and the second local filter, which is the allocation factor, which is used to manage the information flow of the first local filter and the second local filter to ensure reasonable distribution and conservation of information. Because the state vectors of the first local filter and the second local filter are consistent with the system state vector, the main filter allocates the global covariance matrix and the system noise covariance matrix according to the allocation factor.

[0304] Exemplarily, the main filter determines output information of the main filter through the first fusion result and the second fusion result, specifically including:

[0305] First, determine the overall optimal synthesis of the fusion positioning model. Fuse the output information of the first local filter and the second local filter to obtain the global optimal estimate. The relevant formula is as follows:

[0306] State fusion equation:

[0307] X g (k+1)=P g (k+1)[P1 -1 (k+1)X1(k+1)+P2 -1 (k+1)X2(k+1)];

[0308] Among them, X g is the positioning result of the carrier to be tested, including the position information, speed information and attitude information of the carrier, P g is the global covariance matrix, P1 -1 is the inverse covariance matrix of the first local filter, P2 -1 is the inverse covariance matrix of the second local filter, X1 is the first fusion result of the first local filter, and X2 is the second fusion result of the second local filter.

[0309] Optionally, as can be seen from the above formula, the above state fusion equation is a weighted average method based on information fusion, which achieves the optimal linear unbiased estimation by maximizing the amount of information and thus solves the global optimal estimation value. g is the global covariance matrix, which reflects the reliability and accuracy of the predicted positioning results of the carrier to be measured, and is negatively correlated with the global optimal estimate. Therefore, when the global optimal estimate is solved, P gShould be minimum.

[0310] It's important to note that the formula uses the information matrix, which is the inverse of the covariance. While the covariance represents the uncertainty of the result, the information matrix represents the amount of information. A large covariance indicates poor system reliability, meaning the system contains little information, resulting in a smaller information matrix.

[0311] The inverse covariance matrix fusion equation:

[0312] P g -1 (k+1)=P1 -1 (k+1)+P2 -1 (k+1);

[0313] Among them, P g -1 Inverse of the global covariance matrix.

[0314] The process noise covariance inverse matrix fusion equation:

[0315] Q g -1 (k+1)=Q1 -1 (k+1)+Q2 -1 (k+1);

[0316] Among them, Q g -1 is the inverse matrix of the global noise covariance matrix, Q1 -1 is the inverse matrix of the noise covariance matrix of the first local filter, Q2 -1 is the inverse matrix of the noise covariance matrix of the second local filter.

[0317] Secondly, the main filter performs information allocation. After the main filter completes the optimal synthesis of the overall state, it forms an information feedback amount based on the information allocation principle in the federated filter, and allocates information to the first local filter and the second local filter (the information allocation coefficients of the first local filter and the second local filter are β1 and β2 respectively). The relevant formula is as follows:

[0318] P i -1 (k+1)=β i P g -1 (k+1);

[0319] Q i -1 (k+1)=β i Q g -1 (k+1);

[0320] X i (k+1)=X g (k+1);

[0321] Where i is 1 or 2, P i -1 (k+1) is the inverse covariance matrix of the i-th local filter, Q i -1 (k+1) is the inverse covariance matrix of the system noise of the i-th local filter, X i (k+1) is the state update value of the i-th local filter, X g The overall state value of the system, the information distribution coefficient β1 and β2 should satisfy β1+β2=1, and 0<β i <1.

[0322] For example, according to federated filtering theory, the values of the information allocation coefficients β1 and β2 determine the performance of the fused positioning model. By selecting different information allocation coefficients, the performance of the filter can be modified to meet different needs. For the combined navigation fusion positioning model designed above, the values of β1 and β2 will directly affect the performance of the fusion positioning model. Based on this, the values of β1 and β2 are adjusted according to the operating conditions of the GNSS and binocular camera, that is, whether a fault has occurred, thereby realizing an adaptive federated Kalman filter for intelligent GNSS / IMU / binocular vision combined navigation, enabling the fusion positioning model to isolate faulty sensors.

[0323] It is understandable that in order to reduce the computational complexity and increase the speed of the fusion positioning model, the time update and measurement update of the navigation information are only performed in the local filter, and the main filter is only responsible for information fusion and distribution to the local filter.

[0324] For example, according to the information distribution principle, when β1 is small, β2 is large. The first local filter output X1(k+1) is proportional to the overall state of the system X g (k+1) has little influence on the state output X2(k+1), while the state estimate value X2(k+1) of the second local filter has a greater influence on the overall state output X g (k+1) has a greater impact; similarly, when β2 is small, the overall state of the system X g (k+1) is primarily determined by the output state of the first local filter. Therefore, when the GNSS in the first local filter fails, the overall performance of the fusion positioning model should be primarily determined by the performance of the second local filter, i.e., β2 should be maximized as much as possible. When the binocular camera fails, the fusion positioning model should primarily rely on the first local filter where the GNSS is located, i.e., β1 should be maximized as much as possible. When all parts of the system are operating normally, the overall performance of the fusion positioning model can be optimized by properly setting the values of β1 and β2.

[0325] It can be understood that based on the state estimation and covariance matrix of the first local filter and the second local filter, weighted fusion is used to generate the global optimal state vector (the position, velocity, attitude angle and environmental characteristics of the carrier to be measured), so as to solve the inconsistency problem between the positioning information measured by different data (GNSS data, IMU data and enhanced visual data), and avoid the influence of abnormal positioning data on the positioning results of the carrier to be measured.

[0326] In some embodiments, combined Figure 3 ,like Figure 8 As shown, in the above S301, obtaining the initial visual data of the carrier to be tested specifically includes:

[0327] S801. The positioning device obtains status data of the carrier to be measured and binocular vision information data of the carrier to be measured.

[0328] The binocular vision information data includes at least one of image data, parallax information, depth information and three-dimensional coordinates collected by the carrier to be tested.

[0329] Optionally, the image data can be the original two-dimensional grayscale or color image synchronously captured by the left and right lenses of the binocular camera, which contains the texture, color and brightness information of the scene, and is the basic data for subsequent visual data processing. The disparity information can be the horizontal displacement difference (in pixels) of the pixels of the same object in the left and right images, reflecting the geometric relationship of the objects in space. The depth information can be the actual distance from the scene point to the camera calculated based on the focal length, baseline (binocular camera spacing) and disparity, which can be used as a physical quantity to quantify the distance between scene objects. The three-dimensional coordinates can be the pixel coordinates combined with the depth information, and the point cloud coordinates that are back-projected into the three-dimensional space through the camera model. They can be used as a spatial structure expression of the environment for positioning and navigation.

[0330] S802: The positioning device determines the motion state of the carrier to be measured according to the state data of the carrier to be measured, and determines an image frame extraction rule of the binocular vision information data according to the motion state.

[0331] Understandably, when a carrier moves quickly, the relative displacement of objects between adjacent frames in the captured image varies significantly. Extracting frames at a fixed, low frame rate can easily lead to image blur. Therefore, determining the image frame extraction rules based on the carrier's motion state ensures the effectiveness of image features.

[0332] S803: The positioning device extracts image frames from the binocular vision information data according to an image frame extraction rule of the binocular vision information data to obtain initial visual data of the carrier to be tested.

[0333] Specifically, the current positioning frame data needs to be extracted before obtaining the enhanced visual data. Therefore, the positioning device needs to extract the image frame of the binocular vision information data to obtain the initial visual data of the carrier to be measured.

[0334] In some embodiments, determining the motion state of the carrier to be tested according to the state data of the carrier to be tested includes:

[0335] The state data of the carrier to be tested is input into a pre-trained particle swarm optimization-back propagation (PSO-BP) deep learning network model to obtain the motion state of the carrier to be tested.

[0336] In some embodiments, the motion state of the carrier to be measured includes the motion speed of the carrier to be measured; and the image frame extraction rules of the binocular vision information data include:

[0337] The number of image frame extraction intervals for binocular vision information data.

[0338] The number of image frame extraction intervals is positively correlated with the movement speed of the carrier to be measured.

[0339] Optionally, the image frame extraction rule of the binocular vision information data may be to determine the number of image frame extraction intervals according to the movement speed of the carrier to be measured, thereby achieving image frame extraction of the binocular vision information data.

[0340] Exemplarily, the movement speed of the carrier to be measured may include three modes. Mode 1 represents a low speed, and the transmission rule is continuous frame data reading (the image frames of the binocular vision information data are extracted in a manner of continuous reading of image frames); Mode 2 represents a medium speed, and the transmission rule is to extract the corresponding frame data in the form of interval 1 (the image frames of the binocular vision information data are extracted in a manner of reading image frames at intervals of 1 frame); Mode 3 represents a high speed, and the transmission rule is to extract the corresponding frame data in the form of interval 2 (the image frames of the binocular vision information data are extracted in a manner of reading image frames at intervals of 2 frames).

[0341] For example, Figure 9 FIG. 1 shows a schematic diagram of the overall process of a positioning method provided by an embodiment of the present application. Figure 9 As shown, the positioning method includes:

[0342] Optionally, the positioning method mainly includes three processes: process 1, process 2 and process 3.

[0343] Optional, such as Figure 9 As shown, the steps of process 1 are as follows:

[0344] First, the onboard sensors are used to measure the motion carrier parameters (including speed and direction), the number of satellites and other body status data.

[0345] Next, the vehicle state data is identified using the vehicle identification mechanism, and the vehicle's motion mode category is output. This includes three modes: Mode 1, Mode 2, and Mode 3. Mode 1 represents low speed, with continuous frame data reading; Mode 2 represents medium speed, with frame data extracted at intervals of 1; and Mode 3 represents high speed, with frame data extracted at intervals of 2.

[0346] Then, the mode is selected according to the motion mode of the carrier to determine the optimal state mode.

[0347] Optional, such as Figure 9 As shown, the steps of process 2 are as follows:

[0348] First, the onboard sensor measures the motion timing data (stereo camera).

[0349] Next, the time series camera data is optimized. The time series camera data is optimized according to the optimal state mode of process 1.

[0350] Then, the optimized data is input into the parallel multi-attention mechanism-convolutional neural network-long short-term memory network model for model training.

[0351] Subsequently, the binocular image data at the positioning frame is enhanced. The binocular image data obtained by positioning is input into the trained model to achieve binocular image data enhancement at the positioning frame.

[0352] Optional, such as Figure 9 As shown, the steps of process 3 are as follows:

[0353] First, onboard sensors measure satellite data and inertial navigation data.

[0354] Secondly, pseudorange, carrier phase, acceleration, angular velocity, and distance information of feature points in binocular image matching are used to obtain the positioning frame data after image enhancement and the three positioning data of GNSS and IMU.

[0355] Then, the adaptive federated Kalman filter algorithm is used to fuse the above three positioning data.

[0356] Subsequently, GNSS / IMU / binocular camera fusion positioning is performed, and the positioning result of the vehicle under test obtained by GNSS / IMU / binocular camera fusion is output.

[0357] Another example, Figure 10FIG. 1 shows a schematic diagram of the overall process of another positioning method provided by an embodiment of the present application, such as Figure 10 As shown, the positioning method includes:

[0358] First, the BeiDou / GNSS receiver uses serial port transmission to obtain the GNSS data of the carrier under test, including pseudorange and carrier phase data.

[0359] Secondly, the inertial measurement unit (IMU) is used to obtain the IMU data of the carrier to be tested, including acceleration and acceleration data.

[0360] Thirdly, the binocular pixel metadata of the carrier to be tested is obtained through the serial port transmission method of the binocular camera, and the binocular pixel metadata is input into the parallel multi-attention mechanism-convolutional neural network-long short-term memory network model to obtain the binocular pixel enhanced matching feature point distance data.

[0361] Then, the GNSS data and IMU data of the carrier to be tested are respectively input into the first local filter of the federated filter through serial port transmission; the IMU data and binocular pixel enhanced matching feature point distance data are input into the second local filter of the federated filter.

[0362] Subsequently, the output results of the first local filter and the second local filter are input into the main filter for fusion to obtain the positioning result of the carrier to be measured.

[0363] For example, Figure 11 A structural diagram of a fusion positioning model provided in an embodiment of the present application is shown in FIG. Figure 11 As shown, the fusion positioning model includes:

[0364] First, the IMU data and GNSS data are input into the first local filter. The IMU outputs data at 100 Hz.

[0365] Secondly, the K-5 to K-1 frame images and the K frame image (positioning frame) are input into the binocular vision enhancement model for frame image fusion to obtain enhanced image data.

[0366] Again, the enhanced image data and IMU data are input into the second local filter.

[0367] Then, the fusion result of the first local filter (X1 and P1) and the fusion result of the second local filter (X2 and P2) are fused through the main filter to obtain the global optimal estimate (the positioning result of the carrier to be measured). The fusion result of the main filter is X g and P g .

[0368] Subsequently, the main filter performs information distribution. On the one hand, the main filter distributes information to the two local filters (the first local filter and the second local filter) (the information distribution coefficients of the two local filters are β1 and β2 respectively), and the global optimal result, the global covariance matrix and the distribution coefficients (β1 -1 and β2 -1 ) is fed back to the corresponding local filter, and on the other hand, the estimated value of the system state error is fed back to the IMU at a data output frequency of 1 Hz to correct the accumulated error.

[0369] Optionally, because the IMU data frequency is high and the data continuity is good, the estimated value of the system state error is fed back to the IMU to correct the accumulated error.

[0370] It can be seen that the fusion positioning model uses the IMU to output high-frequency data through local filters to fuse GNSS data, IMU data and enhanced visual data. Therefore, the use of this fusion positioning model not only has high-frequency navigation output capability, but also has strong fault tolerance, which facilitates fault diagnosis and isolation of each sensor.

[0371] In some embodiments, the present application provides a GNSS / INS / binocular vision fusion positioning algorithm that considers temporal and spatial characteristics, which mainly includes the following three processes:

[0372] Process 1: Carrier state judgment, using onboard sensors to measure the body state to determine the frame data reading rules of binocular vision data.

[0373] Process 2: Visual Information Enhancement: Using the rules for optimal state pattern transfer in process 1, the corresponding time series image data is extracted, including the five frames of data in the historical record and the current positioning frame data. The extracted data is then input into the hybrid network model to achieve binocular visual information fusion.

[0374] Process 3: GNSS / IMU / binocular vision adaptive federated Kalman filter structure.

[0375] The embodiment of the present application provides a deep learning model of parallel multi-attention mechanism-deep convolutional network-long short-term memory, which can realize the fusion of time series image data to enhance the data at the positioning frame and improve the positioning accuracy of the visual odometry.

[0376] The embodiment of the present application provides a combined navigation algorithm model of GNSS / IMU / binocular image data, which can achieve deep fusion of GNSS, IMU, and time-series image data to reduce the positioning limitations of a single positioning technology in practical applications.

[0377] For example, Figure 12A flow chart of another positioning method provided in an embodiment of the present application is shown. Figure 12 As shown, the positioning method includes:

[0378] First, the status information of the carrier to be classified is obtained.

[0379] Then, the carrier state information to be classified is input into the trained carrier state classification network. The classification network determines which mode the carrier state belongs to and outputs the mode category.

[0380] Optionally, a trained carrier state classification network is obtained by:

[0381] First, a carrier status classification dataset is constructed.

[0382] Secondly, a classification network based on the recurrent convolutional neural network structure is designed.

[0383] Again, train the carrier state classification neural network.

[0384] For example, the specific construction and training process of the above-mentioned carrier state classification network is as follows:

[0385] First, a time series of motion state data is collected, and its state mode categories are defined according to the expert database. Thus, a data set is formed one by one, which includes speed information, speed change rate information and state mode information. Among them, speed information and speed change rate information are model training input parameters, and state mode information is model training output parameter.

[0386] Then, a PSO-BP deep learning network structure is established, which specifically includes: constructing an initial BP neural network and using the weights and biases between each layer as the population particles in the PSO algorithm; initializing the position and velocity of the population particles and determining the population size; calculating the particle fitness value according to the fitness function and determining the optimal solution for individuals and groups; updating the velocity and position of the particles; using the optimal population particles as the weights and biases of the BP neural network and establishing a PSO-BP neural network model.

[0387] Secondly, the constructed data set is used to train the PSO-BP network structure.

[0388] Subsequently, the motion state of the carrier to be classified is input into the trained PSO-BP network to implement specific classification and output the mode category (motion state of the carrier to be tested).

[0389] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of method. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily appreciate that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0390] In the embodiments of the present application, the functional modules of the positioning device can be divided according to the above-mentioned method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated modules can be implemented in the form of hardware or software functional modules. Optionally, the division of modules in the embodiments of the present application is schematic and is only a logical functional division. In actual implementation, other division methods may be used.

[0391] Figure 13 FIG. 1 shows a schematic diagram of the structure of a positioning device provided in an embodiment of the present application. Figure 13 As shown, the positioning device includes: a communication unit 1301 and a processing unit 1302.

[0392] The communication unit 1301 is used to obtain initial visual data, global navigation satellite system GNSS data and inertial measurement unit IMU data of the carrier to be tested.

[0393] The initial visual data includes multiple frames of associated visual images.

[0394] The processing unit 1302 is configured to input the initial visual data into a pre-trained visual enhancement model to obtain enhanced visual data.

[0395] The number of image features in the enhanced visual data is greater than the number of image features in the initial visual data.

[0396] The processing unit 1302 is further configured to input the enhanced visual data, GNSS data, and IMU data into a fusion positioning model to obtain a positioning result of the carrier to be measured.

[0397] Optionally, the visual enhancement model is trained as follows:

[0398] The processing unit 1302 is further configured to obtain a training sample set.

[0399] The training sample set includes training samples and labels of training samples; the training samples include: multiple frames of sample visual images collected by the sample carrier; the labels of training samples include: sample enhanced images after visual enhancement of the visual images collected by the sample carrier.

[0400] The processing unit 1302 is further configured to train the original vision enhancement model based on the training sample set to obtain a trained vision enhancement model.

[0401] Optionally, the multiple frames of sample visual images include a first sample data image set and a second sample data image set; the processing unit 1302 is specifically configured to:

[0402] The sample data images in the first sample data image set are merged to obtain multi-dimensional image data.

[0403] The multidimensional image data is input into a first network structure in a visual enhancement model to obtain a first processing result; the first network structure includes multiple convolutional layers and / or normalized pooling layers.

[0404] The first processing result is respectively input into the spatial attention mechanism structure and the channel attention mechanism structure in the visual enhancement model, and the output result of the spatial attention mechanism structure and the output result of the channel attention mechanism structure are merged to obtain the second processing result.

[0405] The sample data images in the second sample data image set are subjected to convolution processing to obtain a third processing result.

[0406] The second processing result and the third processing result are merged, and the merged processing result is input into the long short-term memory network model structure in the visual enhancement model to obtain the prediction result of the training sample.

[0407] The loss value between the prediction result of the training sample and the label of the training sample is determined according to the loss function, and the visual enhancement model is trained according to the loss value until the visual enhancement model meets the convergence condition, thereby obtaining a trained visual enhancement model.

[0408] The processing unit 1302 is specifically configured to:

[0409] The GNSS data and the IMU data are input into a first local filter in a fusion positioning model to obtain output information of the first local filter; the output information of the first local filter includes a first fusion result of the GNSS data and the IMU data; the first local filter is used to perform data fusion through a Kalman filter algorithm.

[0410] The enhanced visual data and the IMU data are input into the second local filter in the fusion positioning model to obtain output information of the second local filter; the output information of the second local filter includes a second fusion result of the enhanced visual data and the IMU data; the second local filter is used to perform data fusion through the Kalman filter algorithm.

[0411] The first fusion result and the second fusion result are input into a main filter in the fusion positioning model to obtain output information of the main filter; the output information of the main filter includes the positioning result of the carrier to be measured.

[0412] Optionally, the output information of the first local filter also includes: weight information of the first local filter; the weight information of the first local filter is the first covariance matrix obtained by the first local filter by fusing GNSS data and IMU data according to the Kalman filtering algorithm, which is used to measure the accuracy of the first fusion result.

[0413] The output information of the second local filter also includes: the weight information of the second local filter; the weight information of the second local filter is the second covariance matrix obtained by the second local filter by fusing the enhanced visual data and the IMU data according to the Kalman filtering algorithm, which is used to measure the accuracy of the second fusion result.

[0414] The output information of the main filter also includes: the comprehensive weight information and allocation factor of the main filter; the comprehensive weight information of the main filter is the global covariance matrix determined by the main filter based on the first covariance matrix and the second covariance matrix, which is used to measure the accuracy of the positioning result of the carrier to be measured; the allocation factor is used to allocate information to the first local filter and the second local filter; the information in the information allocation includes the global covariance matrix and / or the system noise covariance matrix.

[0415] The communication unit 1301 is specifically configured to:

[0416] Acquire status data of the carrier to be tested and binocular vision information data of the carrier to be tested; the binocular vision information data includes: at least one of image data, parallax information, depth information and three-dimensional coordinates collected by the carrier to be tested.

[0417] The motion state of the carrier to be tested is determined according to the state data of the carrier to be tested, and the image frame extraction rule of the binocular vision information data is determined according to the motion state.

[0418] Image frame extraction is performed on the binocular vision information data according to an image frame extraction rule of the binocular vision information data to obtain initial visual data of the carrier to be tested.

[0419] The communication unit 1301 is specifically configured to:

[0420] The state data of the carrier to be tested is input into the pre-trained particle swarm optimization-back propagation deep learning network model to obtain the motion state of the carrier to be tested.

[0421] Optionally, the motion state of the carrier to be measured includes the motion speed of the carrier to be measured; and the image frame extraction rule of the binocular vision information data includes:

[0422] The number of image frame extraction intervals for binocular vision information data.

[0423] The number of image frame extraction intervals is positively correlated with the movement speed of the carrier to be measured.

[0424] An embodiment of the present application further provides a computer-readable storage medium, which includes computer-executable instructions. When the computer-executable instructions are executed on a computer, the computer executes the positioning method provided in the above embodiment.

[0425] An embodiment of the present application further provides a computer program, which can be directly loaded into a memory and contains software code. After being loaded and executed by a computer, the computer program can implement the positioning method provided in the above embodiment.

[0426] Those skilled in the art will appreciate that, in one or more of the examples above, the functions described herein can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer-readable storage media and communication media, wherein communication media include any medium that facilitates the transmission of a computer program from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0427] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0428] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place, or they may be distributed in multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0429] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or in other words, the part that contributes to the general technology or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for making a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.

[0430] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A positioning method, characterized in that: include: Obtain initial visual data, global navigation satellite system (GNSS) data, and inertial measurement unit (IMU) data of the vehicle to be tested; The initial visual data includes multiple frames of associated visual images; Inputting the initial visual data into a pre-trained visual enhancement model to obtain enhanced visual data; the number of image features in the enhanced visual data is greater than the number of image features in the initial visual data; The enhanced visual data, the GNSS data and the IMU data are input into a fusion positioning model to obtain a positioning result of the carrier to be measured.

2. The method according to claim 1, characterized in that The visual enhancement model is trained in the following way: Obtaining a training sample set; the training sample set includes training samples and labels of the training samples; The training samples include: multiple frames of sample visual images collected by the sample carrier; the labels of the training samples include: sample enhanced images obtained by visually enhancing the visual images collected by the sample carrier; The original visual enhancement model is trained based on the training sample set to obtain the trained visual enhancement model.

3. The method according to claim 2, characterized in that The multiple frames of sample visual images include a first sample data image set and a second sample data image set; and the original visual enhancement model is trained based on the training sample set to obtain the trained visual enhancement model, including: merging the sample data images in the first sample data image set to obtain multidimensional image data; Inputting the multidimensional image data into a first network structure in the visual enhancement model to obtain a first processing result; the first network structure includes a plurality of convolutional layers and / or normalized pooling layers; Inputting the first processing result into a spatial attention mechanism structure and a channel attention mechanism structure in the visual enhancement model respectively, and merging an output result of the spatial attention mechanism structure and an output result of the channel attention mechanism structure to obtain a second processing result; performing convolution processing on the sample data images in the second sample data image set to obtain a third processing result; Merging the second processing result and the third processing result, and inputting the merged processing result into a long short-term memory network model structure in the visual enhancement model to obtain a prediction result of the training sample; The loss value of the prediction result of the training sample and the label of the training sample is determined according to the loss function, and the visual enhancement model is trained according to the loss value until the visual enhancement model meets the convergence condition, thereby obtaining the trained visual enhancement model.

4. The method according to claim 1, wherein The step of inputting the enhanced visual data, the GNSS data, and the IMU data into a fusion positioning model to obtain a positioning result of the carrier to be measured includes: Inputting the GNSS data and the IMU data into a first local filter in the fusion positioning model to obtain output information of the first local filter; the output information of the first local filter includes a first fusion result of the GNSS data and the IMU data; the first local filter is used to perform data fusion through a Kalman filter algorithm; Inputting the enhanced visual data and the IMU data into a second local filter in the fusion positioning model to obtain output information of the second local filter; the output information of the second local filter includes a second fusion result of the enhanced visual data and the IMU data; the second local filter is used to perform data fusion through the Kalman filter algorithm; The first fusion result and the second fusion result are input into a main filter in the fusion positioning model to obtain output information of the main filter; the output information of the main filter includes the positioning result of the carrier to be measured.

5. The method according to claim 4, characterized in that Also includes: The output information of the first local filter further includes: weight information of the first local filter; the weight information of the first local filter is a first covariance matrix obtained by the first local filter fusing the GNSS data and the IMU data according to the Kalman filter algorithm, which is used to measure the accuracy of the first fusion result; The output information of the second local filter also includes: weight information of the second local filter; the weight information of the second local filter is a second covariance matrix obtained by the second local filter fusing the enhanced visual data and the IMU data according to the Kalman filter algorithm, which is used to measure the accuracy of the second fusion result; The output information of the main filter also includes: the comprehensive weight information and allocation factor of the main filter; the comprehensive weight information of the main filter is the global covariance matrix determined by the main filter based on the first covariance matrix and the second covariance matrix, which is used to measure the accuracy of the positioning result of the carrier to be measured; the allocation factor is used to allocate information to the first local filter and the second local filter; the information in the information allocation includes the global covariance matrix and / or the system noise covariance matrix.

6. The method according to claim 1, characterized in that Obtain the initial visual data of the carrier to be tested, including: Acquiring state data of the carrier to be tested and binocular vision information data of the carrier to be tested; the binocular vision information data includes: at least one of image data, parallax information, depth information and three-dimensional coordinates collected by the carrier to be tested; Determining a motion state of the carrier to be tested according to the state data of the carrier to be tested, and determining an image frame extraction rule of the binocular vision information data according to the motion state; Image frame extraction is performed on the binocular vision information data according to an image frame extraction rule of the binocular vision information data to obtain initial visual data of the carrier to be tested.

7. The method according to claim 6, characterized in that The determining the motion state of the carrier to be tested according to the state data of the carrier to be tested includes: The state data of the carrier to be tested is input into a pre-trained particle swarm optimization-back propagation deep learning network model to obtain the motion state of the carrier to be tested.

8. The method according to claim 6, characterized in that The motion state of the carrier to be tested includes the motion speed of the carrier to be tested; the image frame extraction rule of the binocular vision information data includes: The number of image frame extraction intervals of the binocular vision information data; The number of image frame extraction intervals is positively correlated with the movement speed of the carrier to be measured.

9. A positioning device, characterized in that: include: A processor and a memory; wherein the memory is used to store one or more programs, and the one or more programs include computer-executable instructions. When the positioning device is running, the processor executes the computer-executable instructions stored in the memory to enable the positioning device to perform the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that When the computer-executable instructions stored in the computer-readable storage medium are executed by a processor of a positioning device, the positioning device can perform the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Unmanned aerial vehicle positioning method and system

    CN121067877A

  • A method and system for locating unmanned aerial vehicles (UAVs)

    CN121067877B

  • Self-adaptive inspection device and method for deep and far sea wind power energy transmission medium

    CN121417495A