Image processing method and device
By matching the connection points between satellite stereo imagery and camera images, the camera imaging model is updated, solving the problem of inaccurate camera pose and achieving higher-precision camera positioning and target area measurement, thereby improving the accuracy of disaster monitoring and dynamic target tracking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-11-14
- Publication Date
- 2026-05-15
AI Technical Summary
The lack of accurate imaging model parameters in existing cameras makes it impossible to accurately determine the camera's pose, affecting its application in fields such as rapid positioning, disaster monitoring, autonomous driving, augmented reality, and robot navigation.
By acquiring satellite stereo imagery and camera images, a two-step matching method using salient feature points and dense template fusion is employed to determine the connection points between the camera and satellite images, update the camera's imaging model to improve accuracy, and thus determine the camera pose.
It achieves accurate positioning of camera pose, improves the accuracy of camera positioning in three-dimensional space and target area measurement, and supports the accuracy of applications such as road hazard forecasting, disaster emergency response, and dynamic target tracking.
Smart Images

Figure CN122053785A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of media technology, and more particularly to image processing methods and apparatus. Background Technology
[0002] An imaging model is a mathematical model describing how a point in three-dimensional space is projected onto a two-dimensional image plane. The camera pose can be obtained through the camera imaging model. Camera pose refers to the position and orientation of the camera in three-dimensional space. Camera pose includes the camera's position (i.e., its coordinates in the world coordinate system) and orientation (i.e., the camera's facing direction, usually represented by a rotation matrix or quaternion). Camera pose plays a crucial role in many fields (such as rapid localization, disaster monitoring, autonomous driving, augmented reality, and robot navigation).
[0003] Currently, many cameras (such as those in cities) lack accurate imaging model parameters, making it impossible to obtain accurate camera poses and preventing cameras from playing their due role in many fields. Summary of the Invention
[0004] This application provides an image processing method and apparatus for obtaining a more accurate camera imaging model. To achieve the above objectives, this application adopts the following technical solution:
[0005] In a first aspect, embodiments of this application provide an image processing method, the method comprising: acquiring satellite stereo images and camera images, the satellite stereo images comprising multiple satellite images captured by a satellite; determining connection points between the multiple satellite images and the camera images; and determining a second imaging model of the camera based on the connection points between the multiple satellite images and the camera images, a first imaging model of the camera, and an imaging model of the satellite.
[0006] The solution provided in this application embodiment can obtain the connection point between the satellite image and the camera image by matching the satellite image with the camera image. The obtained connection point is combined with the satellite imaging model to update the camera's initial imaging model (i.e., the camera's first imaging model) to obtain a more accurate camera imaging model (i.e., the camera's second imaging model).
[0007] In one possible implementation, a first feature point of the multiple satellite images and a second feature point of the camera image can be determined based on the multiple satellite images and the camera image; a first descriptor of the first feature point and a second descriptor of the second feature point can be determined based on the first feature point and the second feature point; a connection region between the multiple satellite images and the camera image can be determined based on the first descriptor and the second descriptor; and a connection point between the multiple satellite images and the camera image can be determined based on the connection region.
[0008] Understandably, a two-step matching method using salient features (first and second feature points) and dense template fusion (the connection region between the aforementioned multiple satellite images and the aforementioned camera images) is used for cross-modal matching between space and ground. This method addresses the significant nonlinear radiation and geometric distortion differences in cross-modal images (the aforementioned multiple satellite images and the aforementioned camera images), improves the accuracy and reliability of connection point matching, accurately obtains the connection points between satellite images and camera images, and provides a reliable foundation for solving the camera imaging model.
[0009] In one possible implementation, the camera pose can be determined based on the second imaging model described above.
[0010] Understandably, determining the camera pose based on a more accurate camera imaging model can improve the accuracy of the determined camera pose compared to using an inaccurate camera.
[0011] In one possible implementation, the position and / or area of the target region in the camera image in real three-dimensional space can be determined based on the camera pose described above.
[0012] Understandably, accurate camera pose can determine the precise location and area of the target region in the camera image in real three-dimensional space, thereby enabling accurate road hazard forecasting (such as accurately locating and measuring urban road collapses and avoiding road risks), accurate disaster emergency response (such as accurately locating and imaging the location of disasters such as fires, floods, and building collapses), and accurate dynamic target tracking (such as accurately locating vehicles).
[0013] In one possible implementation, multiple satellite images include a first satellite image and a second satellite image, wherein the intersection angle between the first satellite image and the second satellite image is greater than a first threshold.
[0014] For example, the first threshold could be 30 degrees.
[0015] Understandably, satellite stereo imagery is used to provide three-dimensional spatial information, but a single satellite image and multiple satellite images with small intersection angles (i.e., intersection angles less than a first threshold) cannot provide sufficient geometric constraints in the elevation direction. Therefore, multiple satellite images, including a first satellite image and a second satellite image with intersection angles greater than the first threshold, are needed to provide sufficient elevation accuracy.
[0016] In one possible implementation, the connection point includes a first connection point and a second connection point, wherein the first connection point is a one-degree overlap connection point between the camera image and the single-view satellite image, and the second connection point is a multi-degree overlap connection point between the camera image and the multi-view stereo satellite image.
[0017] Understandably, due to the shooting angle, some feature points in a camera image may only be connected to feature points in one satellite image, while others may be connected to feature points in multiple satellite images. Therefore, the connection points in the camera image and satellite images can be categorized as first connection points and second connection points.
[0018] Secondly, embodiments of this application provide an image processing apparatus, comprising a transceiver unit and a processing unit. The transceiver unit is configured to acquire satellite stereo images and camera images, the satellite stereo images comprising multiple satellite images captured by a satellite; the processing unit is configured to determine connection points between the multiple satellite images and the camera images; the processing unit is further configured to determine a second imaging model of the camera based on the connection points, a first imaging model of the camera, and an imaging model of the satellite.
[0019] In one possible implementation, the processing unit is specifically configured to: determine a first feature point of the multiple satellite images and a second feature point of the camera image based on the multiple satellite images and the camera image; determine a first descriptor of the first feature point and a second descriptor of the second feature point based on the first feature point and the second feature point; determine a connection region between the multiple satellite images and the camera image based on the first descriptor and the second descriptor; and determine a connection point between the multiple satellite images and the camera image based on the connection region.
[0020] In one possible implementation, the processing unit is further configured to: determine the camera pose based on the second imaging model.
[0021] In one possible implementation, the processing unit is further configured to: determine the position and / or area of the target region of the camera image in real three-dimensional space based on the camera pose.
[0022] In one possible implementation, the aforementioned multiple satellite images include a first satellite image and a second satellite image, wherein the intersection angle between the first satellite image and the second satellite image is greater than a first threshold.
[0023] In one possible implementation, the connection point includes a first connection point and a second connection point, wherein the first connection point is a one-degree overlap connection point between the camera image and the single-view satellite image, and the second connection point is a multi-degree overlap connection point between the camera image and the multi-view stereo satellite image.
[0024] Thirdly, embodiments of this application also provide an image processing apparatus, which includes: at least one processor, which, when executing program code or instructions, implements the method described in the first aspect or any possible implementation thereof.
[0025] Optionally, the device may further include at least one memory for storing the program code or instructions.
[0026] Fourthly, embodiments of this application also provide a chip, including: an input interface, an output interface, at least one processor, and at least one memory. The at least one processor is used to execute code in the at least one memory, and when the at least one processor executes the code, the chip implements the method described in the first aspect or any possible implementation thereof.
[0027] Alternatively, the chip described above can also be an integrated circuit.
[0028] Fifthly, embodiments of this application also provide a computer-readable storage medium for storing a computer program, the computer program including methods for implementing the first aspect or any possible implementation thereof.
[0029] Sixthly, embodiments of this application also provide a computer program product containing instructions that, when run on a computer, cause the computer to implement the method described in the first aspect or any possible implementation thereof.
[0030] The image processing apparatus, computer storage medium, computer program product, and chip provided in this embodiment are all used to execute the image processing method provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects in the image processing method provided above, and will not be repeated here. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This is a schematic diagram of the structure of an image processing system provided in an embodiment of this application;
[0033] Figure 2 This is a schematic diagram of another image processing system provided in an embodiment of this application;
[0034] Figure 3A schematic flowchart of an image processing method provided in an embodiment of this application;
[0035] Figure 4 A satellite stereo image and camera image provided for embodiments of this application;
[0036] Figure 5 A schematic diagram of an image connection point provided in an embodiment of this application;
[0037] Figure 6 A schematic diagram of another image connection point provided in an embodiment of this application;
[0038] Figure 7 A flowchart illustrating another image processing method provided in an embodiment of this application;
[0039] Figure 8 A schematic flowchart illustrating another image processing method provided in an embodiment of this application;
[0040] Figure 9 A flowchart illustrating an image connection point determination method provided in an embodiment of this application;
[0041] Figure 10 A schematic flowchart illustrating another image processing method provided in an embodiment of this application;
[0042] Figure 11 A schematic diagram illustrating an image pre-positioning method provided in an embodiment of this application;
[0043] Figure 12 A schematic diagram illustrating an image calibration method provided in an embodiment of this application;
[0044] Figure 13 A schematic diagram illustrating the application scenario of the solution provided in the embodiments of this application;
[0045] Figure 14 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this application;
[0046] Figure 15 This is a schematic diagram of the structure of a chip provided in an embodiment of this application;
[0047] Figure 16 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0048] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the embodiments of this application.
[0049] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0050] The terms "first" and "second," etc., in the specification and drawings of the embodiments of this application are used to distinguish different objects or to distinguish different treatments of the same object, rather than to describe a specific order of objects.
[0051] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of the embodiments of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0052] It should be noted that in the description of the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0053] Before introducing the solutions provided in the embodiments of this application, the terms involved in the embodiments of this application will be explained.
[0054] Pose: This refers to both position and attitude information. Position information is typically represented using Euclidean coordinates X, Y, Z, while attitude information is represented using rotational coordinates such as pitch, yaw, and roll. Therefore, position information (X, Y, Z) and attitude information (yaw, pitch, roll) can be collectively referred to as 6 degrees of freedom (6DOF) position and attitude information.
[0055] Satellite stereo imagery: In satellite photogrammetry and remote sensing, it usually refers to satellite images that form a certain stereo intersection geometry, with two views (front and back) or three views (front, bottom, and back).
[0056] Space-to-ground cross-modal image matching: Space-to-ground cross-modal imagery refers to space-based satellite imagery and ground-based camera images. Due to the differences between space-based and ground-based remote sensing platforms, space-based and ground-based images have significant cross-modal characteristics such as differences in viewpoint, resolution, geometric distortion, and radiation distortion. Image matching refers to obtaining corresponding image points on the image through techniques such as image signal correlation.
[0057] Camera self-calibration positioning: In photogrammetry or computer vision, a technique for accurately recovering the geometric parameters (position, orientation, and intrinsic parameters) of a camera image.
[0058] Rational function model: In satellite photogrammetry and remote sensing, a mathematical fitting model is used to directly express the image point coordinates of satellite images and the corresponding three-dimensional geodetic coordinates of ground points. Its model parameters are obtained by fitting a rigorous imaging model of the satellite image.
[0059] Feature points: Composed of keypoints and descriptors. Keypoints are the location of the feature point in the image, and some also include information such as orientation and size. Descriptors are a method of describing the pixels around the keypoint. Keypoints with similar appearances should have similar descriptors. To determine whether two keypoints in different locations are similar, the distance between their descriptors can be calculated.
[0060] Scale-invariant feature transform (SIFT) is a feature point extraction algorithm based on scale space that can extract stable feature points under different scales, rotations, and illumination variations. The SIFT algorithm includes steps such as scale space extremum detection, keypoint localization, orientation assignment, keypoint description, and feature point matching.
[0061] Speeded-up RobustFeatures: This is an accelerated version of SIFT, which improves computational speed while maintaining high accuracy. The SURF algorithm uses the Hessian matrix to detect local features in the image and calculates the Haar wavelet response to achieve feature description.
[0062] Oriented Fast and Rotated Brief (ORB) is a feature point extraction algorithm based on fast corner detection and brief descriptors, offering high speed and good performance. The ORB algorithm uses the fast corner detection algorithm to detect image corners and the brief algorithm to describe feature points.
[0063] Harris Corner Detection Algorithm: Harris corner detection is a feature point extraction algorithm based on image grayscale changes. It extracts corner features by calculating the corner response function of each point in the image. This method is well-adapted to image rotation and scaling.
[0064] An imaging model uses mathematical formulas to describe the entire imaging process, that is, the geometric transformation relationship between the spatial point of the photographed object and the imaging point of the photograph. The core parameters of the imaging model include camera intrinsic parameters and camera extrinsic parameters. Camera intrinsic parameters describe the internal optical characteristics of the camera, such as focal length and optical center; camera extrinsic parameters describe the camera's position and orientation in the external world.
[0065] Intrinsic parameters, also known as intrinsic orientation parameters, include: focal length, principal point, and distortion parameters.
[0066] Focal length (f): The distance from the camera lens to the imaging plane.
[0067] Principal point (optical center): A point on the image plane, usually the center of the image.
[0068] Distortion parameters: used to correct image distortion caused by lens defects.
[0069] Extrinsic parameters, also known as extrinsic orientation parameters, include: rotation matrix and translation vector.
[0070] Rotation matrix (R): Describes the rotation relationship between the camera coordinate system and the world coordinate system.
[0071] Translation vector (t): describes the translation relationship between the camera coordinate system and the world coordinate system.
[0072] World coordinate system: The absolute coordinates that describe the objective world.
[0073] Camera coordinate system: A coordinate system with the camera optical center as the origin and the optical axis as the Z-axis.
[0074] Image coordinate system: describes the position of pixels on the image plane, usually in millimeters.
[0075] Pixel coordinate system: describes the digital location of image data.
[0076] Image overlap region: refers to the image regions in two images that contain the same object. The proportion of the image overlap region to the whole image is called the overlap degree.
[0077] One-dimensional overlap: refers to superposition on a single layer, such as the overlap of two objects or concepts in one-dimensional space or time. This type of overlap is relatively simple and usually only involves overlap in one dimension.
[0078] Multi-dimensional overlap: This involves overlap at multiple levels or in multiple aspects, such as overlap in two-dimensional space or across multiple dimensions. This type of overlap is more complex, encompassing more dimensions and levels.
[0079] Image matching: From two images with overlapping information, feature points and their corresponding feature vectors are extracted from each image as descriptors. The feature vectors are then used to identify corresponding image points of the same object in the two images. Multiple image points of the same object in different images are called homonymous image points.
[0080] Epipolar transform is an image transformation method based on epipolar geometry, primarily used in 3D reconstruction, stereo vision, and camera calibration. Epipolar transform transforms an image so that corresponding points satisfy epipolar geometric constraints, thereby achieving image correction and processing. The core of epipolar transform lies in utilizing the properties of epipolar geometry to map points in an image to new locations through a transformation, ensuring the transformed image satisfies specific geometric relationships. This transformation can help solve matching problems in stereo vision, improve the accuracy of 3D reconstruction, or provide more constraints during camera calibration.
[0081] The technical solutions provided in this application can be applied to image processing systems.
[0082] Figure 1 A schematic diagram of a possible, non-limiting image processing system described above is shown. For example... Figure 1 As shown, the image processing system 10 includes a satellite 100, a camera 200, and an image processing device 300.
[0083] Satellite 100 is used to capture satellite stereo images.
[0084] For example, the aforementioned satellite 100 may be an optical remote sensing satellite.
[0085] Camera 200 is used to capture camera images.
[0086] For example, the camera 100 described above can be a city camera.
[0087] For example, the camera 200 can be a mobile phone camera, a city camera, a security lens, or other monocular camera.
[0088] The image processing apparatus 300 is used to execute any of the image processing methods provided in the embodiments of this application.
[0089] For example, the image processing device 300 can be distributed computing general-purpose hardware, servers, computers, laptop computers, mobile phones, tablets, mice, remote controls, styluses, set-top boxes, routers, cameras, screens, smart screens, wireless data cards, personal digital assistant computers (PDAs), smartwatches, smart bracelets, wireless headphones, electronic whiteboards, virtual reality (VR) terminals, augmented reality (AR) terminals, smart home devices (e.g., refrigerators, televisions, air conditioners, washing machines, rice cookers, table lamps, electricity meters, etc.), smart robots, robotic arms, workshop equipment, or other electronic devices.
[0090] Distributed computing general-purpose hardware refers to hardware devices capable of supporting distributed computing tasks. Distributed computing refers to a distributed system formed by multiple computer nodes connected through a network, where each node has its own processor and storage resources.
[0091] In one possible implementation, the image processing apparatus 300 described above may include a cloud-side device and an edge-side device. The cloud-side device and the edge-side device may cooperate to execute any of the image processing methods provided in the embodiments of this application. Alternatively, the cloud-side device and the edge-side device may execute any of the image processing methods provided in the embodiments of this application independently.
[0092] Figure 2 A schematic diagram of another possible, non-limiting image processing system described above is shown. For example... Figure 2 As shown, the image processing system includes a hardware platform, a cloud-side processing module, and an edge-side processing module.
[0093] like Figure 2 As shown, the aforementioned hardware platform includes optical remote sensing satellites, urban cameras, and general-purpose distributed computing hardware.
[0094] The cloud-side processing module is used to perform satellite image pre-positioning, cross-modal image matching, self-calibration adjustment modeling, and self-calibration adjustment calculation based on satellite stereo imagery, RPC imaging model, camera images, and camera initial imaging model to obtain principal points, principal distances, optical distortion coefficients, installation positions, and installation attitudes.
[0095] The edge processing module is used for building accurate imaging models for the camera and for precise target localization for the camera.
[0096] The hardware foundation of the solution provided in this application embodiment is distributed computing general-purpose hardware, cameras, and optical remote sensing satellites. Its hardware drivers and data reading and writing are carried out through the hardware abstraction layer data interface, and data and control interaction is performed in accordance with the standard system interface and the upper-layer positioning software service program.
[0097] The data source for the solution provided in this application is satellite stereo imagery and images captured by ground-based urban cameras. Based on the satellite imagery, photogrammetry and remote sensing techniques are used to obtain the connection points between the satellite stereo imagery and the ground-based urban camera images, creating a cross-modal imagery. Accurate calculation of the camera imaging model parameters is achieved through self-calibration adjustment.
[0098] The architecture provided in this application is an edge-cloud combined model. On the cloud side, photogrammetry and remote sensing technologies are used to acquire cross-modal image connection points, perform camera self-calibration and adjustment calculations, and manage the storage of urban camera model parameters. On the edge side, high-precision imaging models of urban cameras are constructed and targets are located.
[0099] like Figure 2 The application scenarios of the image processing system shown include, but are not limited to, disaster emergency response, road hazard forecasting, dynamic target tracking, and urban planning and management.
[0100] Among them, disaster emergency support refers to using cameras to detect various disasters in the city (fires, floods, building collapses, etc.), and quickly locating the disaster location through the camera's target positioning to provide real-time location information for disaster emergency support.
[0101] Road hazard forecasting refers to the method provided in this application to complete the self-calibration and positioning of urban road cameras, and combine camera target detection to realize the location forecast of road hazards, supporting early warning and emergency response to road hazards.
[0102] Dynamic target tracking refers to the real-time relay tracking and positioning of dynamic targets (people, vehicles, etc.) in the city by stitching together video images from precise line-of-sight cameras in urban areas, thereby improving the accuracy and reliability of dynamic target tracking.
[0103] Urban planning and management refers to the use of urban cameras to locate the positions of mobile facilities in the city, thereby assisting in urban planning and management.
[0104] Figure 3 This application illustrates an image processing method provided by an embodiment of the present application. This method can be executed by the aforementioned image processing apparatus, such as... Figure 3 As shown, the method includes:
[0105] S301. Acquire satellite stereo images and camera images.
[0106] Among them, satellite stereo imagery includes multiple satellite images taken by the satellite.
[0107] For example, such as Figure 4 As shown, it can acquire rear-view satellite images, forward-view satellite images, and camera images.
[0108] For example, satellite imagery can be acquired using one or more satellites deployed in the high atmosphere, and camera imagery can be acquired using cameras. One or more satellites and cameras can take pictures of the same area on the ground, acquiring multiple satellite and camera images, and then transmitting the acquired satellite and camera images to an image processing device.
[0109] For example, it is possible to acquire camera images taken by city cameras, and based on the location of the city cameras, to obtain sub-meter resolution satellite stereo images (such as 0.5-meter or 0.7-meter resolution satellite images) through open-source data download or commercial procurement.
[0110] The acquisition of camera images can be achieved through a mobile phone, camera, security lens, or other device based on monocular camera technology. All or part of the frames in the video can be extracted as camera images. In some exemplary embodiments, a video frame extraction algorithm can be used to convert the video image into a frame-based image.
[0111] The aforementioned satellite stereo imagery includes multiple satellite images taken by a satellite. For example, satellite stereo imagery may include forward-looking satellite images and rear-looking satellite images. As another example, satellite stereo imagery may include forward-looking satellite images, rear-looking satellite images, and downward-looking satellite images.
[0112] In one possible implementation, multiple satellite images include a first satellite image and a second satellite image, wherein the intersection angle between the first satellite image and the second satellite image is greater than a first threshold.
[0113] Optionally, the first threshold can be 30 degrees.
[0114] For example, multiple satellite images include forward-looking and backward-looking satellite images. The intersection angle between the forward-looking and backward-looking satellite images is 30 degrees.
[0115] For example, forward-looking satellite images, rear-looking satellite images, and downward-looking satellite images. Among them, the intersection angle between forward-looking and rear-looking satellite images is 30 degrees.
[0116] Understandably, satellite stereo imagery is used to provide three-dimensional spatial information, but a single satellite image and multiple satellite images with small intersection angles (i.e., intersection angles less than a first threshold) cannot provide sufficient geometric constraints in the elevation direction. Therefore, multiple satellite images, including a first satellite image and a second satellite image with intersection angles greater than the first threshold, are needed to provide sufficient elevation accuracy.
[0117] In one possible implementation, the camera images can also be preprocessed.
[0118] Preprocessing the acquired camera images involves operations such as exposure correction, blur restoration, and rain / fog removal to optimize image quality and improve image clarity, which is beneficial for subsequent processing. Camera image preprocessing can also include operations such as using exposure detection to exclude overexposed and underexposed images, using blur detection to exclude blurred images, and using raindrop detection algorithms to exclude images containing raindrops. It should be understood that camera image preprocessing can be performed on the local device acquiring the camera images, such as a camera, security camera, or mobile phone. This allows for preprocessing of the acquired camera images at the acquisition end, reducing the complexity of subsequent operations and saving resources and improving efficiency.
[0119] S302. Determine the connection points between multiple satellite images and camera images.
[0120] like Figure 5 As shown, more than 100 robust and reliable connection points can be identified between satellite images and camera images.
[0121] In one possible implementation, the connection points between multiple satellite images and camera images include a first connection point and a second connection point.
[0122] The aforementioned first tie point, also known as the estimated tie point (ETP), is a tie point where the camera image and the single-view satellite image overlap by one degree. In other words, it is a tie point in the camera image that has a connection relationship with only one satellite image, and a tie point in the satellite image that has a connection relationship with the first tie point in the camera image.
[0123] like Figure 6 As shown, feature point A1 in the camera image is only connected to feature point B1 in the forward-looking satellite image, so feature point A1 and feature point B1 are the first connection points.
[0124] The aforementioned second connection point, namely the stereo control tie point (CTP), is a multi-degree overlapping connection point between the camera image and the multi-view stereo satellite image. In other words, it is a connection point in the camera image that has a connection relationship with multiple satellite images, and a connection point in the satellite image that has a connection relationship with the second connection point of the camera image.
[0125] like Figure 6 As shown, feature point A2 in the camera image is connected to feature point B2 in the forward-looking satellite image and feature point C1 in the rear-looking satellite image. Therefore, feature point A2, feature point B2 and C1 are the second connection points.
[0126] like Figure 7 As shown, image matching can be performed based on satellite stereo images and camera images to obtain connection points, stereo control connection points, and encrypted connection points between multiple satellite images and camera images.
[0127] Understandably, due to the shooting angle, some feature points in a camera image may only be connected to feature points in one satellite image, while others may be connected to feature points in multiple satellite images. Therefore, the connection points in the camera image and satellite images can be categorized as first connection points and second connection points.
[0128] In one possible implementation, a first feature point in the satellite images and a second feature point in the camera images can be determined based on the aforementioned multiple satellite images and camera images. A first descriptor for the first feature point and a second descriptor for the second feature point are then determined based on these two feature points. A connection region between the multiple satellite images and the camera images is then determined based on the first and second descriptors. Finally, a connection point between the multiple satellite images and the camera images is determined based on this connection region.
[0129] Feature points, in this context, refer to representative points in an image, that is, points with characteristic properties. Feature points can represent unique and salient features in an image. Specifically, feature points can represent the target in a similar or identical form in other similar images containing the same object. Simply put, for the same object or scene, if multiple images are captured from different angles and the same areas can be identified as identical, then these scale-invariant points or patches are called feature points. Feature points are points identified through algorithmic analysis, containing rich local information, and typically appear at corners or locations with drastic texture changes in an image.
[0130] Because satellite images and camera images are captured from different angles and positions, there are significant differences in resolution and illumination between them. Therefore, image processing devices can treat the matching problem between satellite and camera images as a multimodal image matching problem, using logarithmic polar coordinates for feature description, including but not limited to Gradient Location Orientation Histogram (GLOH). GLOH is an extension of the SIFT feature descriptor, designed to increase its robustness and uniqueness. SIFT is a widely used image feature descriptor. Using GLOH to match feature points between satellite and camera images ensures that all matched feature points are located on buildings. Due to the rigidity and stability of buildings, high alignment accuracy between satellite and camera images can be guaranteed.
[0131] Understandably, a two-step matching method using salient features (first and second feature points) and dense template fusion (the connection region between the aforementioned multiple satellite images and the aforementioned camera images) is used for cross-modal matching between space and ground. This method addresses the significant nonlinear radiation and geometric distortion differences in cross-modal images (the aforementioned multiple satellite images and the aforementioned camera images), improves the accuracy and reliability of connection point matching, accurately obtains the connection points between satellite images and camera images, and provides a reliable foundation for solving the camera imaging model.
[0132] The specific implementation of determining the first feature point of the satellite image and the second feature point of the camera image based on the satellite image and the camera image can be any method that can be conceived by those skilled in the art, and the embodiments of this application do not limit this.
[0133] For example, SIFT, SURF, ORB, Harris corner detection, or other feature point detection algorithms can be used to determine the first feature point of the satellite images and the second feature point of the camera images based on the satellite images and camera images.
[0134] For example, feature point detection guided by a Gaussian directional adjustable filter can be used to determine the first feature points of the satellite images and the second feature points of the camera images based on the aforementioned multiple satellite images and camera images. Figure 8 and Figure 9 As shown, the basis filter of a first-order Gaussian Directionally Tunable Filter (FoGSF) can be constructed in a nonlinear scale space based on satellite stereo imagery and camera images. They are first-order difference operators of a two-dimensional circularly symmetric Gaussian function in the x and y axes. According to the definition of a direction-tunable filter, a FoGSF convolutional image in any direction can be obtained by a linear combination of the convolutional images of the basis filters, and convolutional images of different scales can be obtained by adjusting the standard deviation of the Gaussian function. Therefore, multi-directional and multi-scale convolutional images can achieve adaptive control. Here, we choose to perform directional control by interpolation at equal intervals and scale control by interpolation at equal proportions, thus forming a multi-scale, multi-directional convolutional image sequence. Furthermore, a multi-scale feature response aggregation model is introduced to synthesize convolutional images of different scales in a fixed direction, forming a feature map that reflects the saliency of local features and fusing multi-scale feature information. Then, the feature response aggregation results of multiple directions are synthesized according to the moment analysis equation, the direction Φ of the principal moment is calculated, and the corresponding maximum moment value map M is obtained. Φ Finally, in M Φ Feature detection is completed using FAST features.
[0135] It is understandable that the purpose of determining the first feature point of the satellite image and the second feature point of the camera image based on the above satellite image and camera image is to provide a data basis for connection point matching. The final result of the matching step is the pair of points with the same name, while the feature points (salient feature points) are the data input for feature point matching.
[0136] The specific implementation of determining the first descriptor of the first feature point and the second descriptor of the second feature point based on the first feature point and the second feature point can be any method that can be conceived by those skilled in the art, and the embodiments of this application do not limit this.
[0137] For example, a first descriptor for the first feature point and a second descriptor for the second feature point can be determined based on the first feature point and the second feature point using scale-invariant feature transform (SIFT), speeded-up robust features (SURF), or other methods.
[0138] For example, the local self-similarity (LSS) feature description information extraction method can be used to determine the first descriptor of the first feature point and the second descriptor of the second feature point based on the first feature point and the second feature point mentioned above. Figure 8 and Figure 9 As shown, a correlation surface S can be generated based on the statistical co-occurrence rate of surrounding small image patches within an image region using a local self-similarity model. q (x,y), and set S qThe image is converted to logarithmic polar representation and divided into multiple tangential and radial partitions, serving as the basis for image feature description. Then, based on the partitioned images, description vectors are calculated for the extracted feature points. Specifically, the average value of the radial partition images at a fixed angle is calculated, generating statistical histograms of pixel values at different tangential angles to obtain the angle number corresponding to the peak value at each pixel position. Next, the principal direction of the feature points is calculated to obtain rotation invariance. A circular region is defined on the statistical histogram centered on the target feature, and a distribution histogram is calculated to obtain the region corresponding to the peak value of the histogram. The principal direction of the feature is determined using the coordinates of the centroid of this region. Subsequently, the values of the circular region are normalized based on the histogram peak value. Finally, the positive x-axis of the sampling coordinate system is rotated to the principal direction, sub-regions are divided according to a logarithmic polar grid, and a distribution histogram is generated for each sub-region. The final feature vector (descriptor) is generated by connecting the histograms and performing L2 regularization.
[0139] It is understandable that the purpose of determining the first descriptor of the first feature point and the second descriptor of the second feature point based on the first feature point and the second feature point is to construct a unique descriptor (feature vector) for each feature point (salient feature point), which is used to identify corresponding points by calculating distance measures during the matching process.
[0140] The specific implementation of determining the connection area between the multiple satellite images and the camera images based on the first descriptor and the second descriptor can be any method that can be conceived by those skilled in the art, and the embodiments of this application do not limit this.
[0141] For example, a method of constructing dense image templates with locally self-similar features can be used to determine the connection regions between the multiple satellite images and the camera images based on the first descriptor and the second descriptor. Figure 8 and Figure 9 As shown, a template window of a certain size can be set on the target image (satellite image and camera image), and the template window can be divided into several sub-regions. Then, an LSS template is constructed for each sub-region, and the LSS templates of each sub-region are connected to form a dense image template. Finally, L2 regularization is performed to achieve normalization processing, realizing the construction of a dense image template with local self-similar features, so as to obtain the connection region between the above-mentioned multiple satellite images and camera images. This method captures the image structure characteristics by integrating multiple LSS templates in a dense sampling grid, and further improves the accuracy through template matching with the support of feature matching results.
[0142] The specific implementation method for determining the connection point between the multiple satellite images and the camera images based on the above-mentioned connection area can be any method that can be conceived by those skilled in the art, and the embodiments of this application do not limit it.
[0143] For example, brute-force matching and mismatch elimination can be used to determine the connection points between the multiple satellite images and the camera images based on the aforementioned connection regions, such as... Figure 8 and Figure 9 As shown, the obvious geometric distortion between images can be eliminated based on the feature matching results, and a search window is determined on the reference images (satellite images and camera images). To achieve template matching, normalized cross-correlation is chosen as the similarity metric to evaluate the similarity between templates. Finally, the Random Sample Consensus (RANSAC) algorithm is used for gross error elimination to obtain the connection points between the aforementioned satellite images and camera images.
[0144] S303. Determine the second imaging model of the camera based on the connection points between multiple satellite images and camera images, the first imaging model of the camera, and the imaging model of the satellite.
[0145] For example, a second imaging model of the camera can be determined based on the connection points between multiple satellite images and camera images, a first imaging model of the camera, and an RPC imaging model.
[0146] Here, the first imaging model of the camera refers to the initial imaging model of the camera, and the second imaging model of the camera refers to the precise imaging model of the camera.
[0147] The RPC imaging model is a mathematical model that expresses the one-to-one correspondence between the coordinates of satellite image points and the corresponding ground point coordinates. It is calculated based on parameters such as the satellite's attitude, orbit, and time during satellite data processing and is a standard model product that comes with satellite imagery.
[0148] For example, such as Figure 7 and Figure 8 As shown, firstly, based on the camera's factory design parameters (the first imaging model mentioned above), the camera's principal point (x0, y0), principal distance f, lens distortion coefficient (k1, p1), and other intrinsic parameters are initialized. Then, based on the camera's initial installation position and angle, the position parameter (X) is initialized. s ,Y s Z s ) and attitude parameters Furthermore, based on the measured image point (u,v) and the corresponding object point (X,Y,Z), the true image point (x,y) after correcting for lens distortion can be obtained. This allows for the further construction of the intrinsic parameter model's gaze vector (x-x0,y-y0,-f) and the extrinsic parameter model's gaze vector (XX) in the camera imaging model.s YY s ZZ s The system first constructs a collinear imaging model for the cameras using the rotation matrix R. Then, using the RPC model parameters (the imaging model of the satellite) corresponding to the satellite images, it constructs a rational function imaging model for the image points and object points on the satellite images at the connection points. The object coordinates of the connection points are initialized based on the forward intersection, i.e., the three-dimensional coordinates (X, Y, Z) of the object points corresponding to the connection points are initialized based on the spatial forward intersection of the connection points. The system then constructs a self-calibration adjustment model for the cameras using the geometric constraints of the same object points corresponding to the image points of the satellite images and the image points of the cameras at the connection points. Subsequently, the unknown parameters to be solved in the above model (including calibration parameters such as the camera principal point, principal distance, lens distortion coefficient, position, and attitude, and the object coordinates of the densified connection points) are linearized. Based on this, a weighted constraint equation is introduced for the object coordinates of the connection points considering the connection point type, and then a self-calibration adjustment model is constructed.
[0149] like Figure 7 and Figure 8 As shown, based on the self-calibration adjustment model constructed above, the camera imaging model parameters can be solved using least squares adjustment iteration. First, weighting is performed using multi-source tie point observations. Specifically, the first and second tie points are weighted according to their observations. To leverage the constraint effect of satellite stereo imagery, the adjustment equations constructed from tie points on satellite imagery are assigned a larger weight, while those on camera imagery are assigned a smaller weight. The weighted constraint equations for the object-space coordinates of the stereo control tie points are assigned a larger weight, while those for the object-space coordinates of the densified tie points are assigned a smaller weight. Then, the error equations for self-calibration adjustment are constructed point-by-point, and the self-calibration adjustment parameters are solved based on least squares estimation. If the self-calibration adjustment parameter solution results do not converge (diverge), gross error detection and elimination (i.e., during the least squares estimation of the self-calibration adjustment parameter solution, statistically analyze the post-verification relative accuracy of the connection points and iteratively eliminate connection point gross errors with a limit of three times the standard error) are performed, and the above steps (initialization of the object coordinates of the connection points based on forward intersection, construction of the self-calibration adjustment model, weighting of multi-source connection point observations, and least squares estimation of the self-calibration adjustment parameter solution) are repeated to construct the camera imaging model. The estimation is iteratively performed according to the above method until the least squares estimation of the self-calibration adjustment parameter solution results converge to obtain the accurate camera imaging model (i.e., the second imaging model). The accurate camera imaging model includes the camera position (installation position), camera attitude (installation attitude), and interior orientation parameters.
[0150] The solution provided in this application embodiment can obtain the connection point between the satellite image and the camera image by matching the satellite image with the camera image. The obtained connection point is combined with the satellite imaging model to update the camera's initial imaging model (i.e., the camera's first imaging model) to obtain a more accurate camera imaging model (i.e., the camera's second imaging model).
[0151] In one possible implementation, the satellite's imaging model can be corrected based on satellite imagery.
[0152] For example, based on the city location corresponding to the camera image, sub-meter resolution satellite stereo imagery (such as 0.5-meter or 0.7-meter resolution satellite imagery) can be obtained by downloading open-source data or purchasing commercially. Based on the prior positioning accuracy of the satellite imagery, sparse control points or laser points (3-5) can be used to correct the RPC imaging model of the satellite imagery using bundle adjustment techniques to perform pre-positioning of the satellite stereo imagery, thereby ensuring the geometric positioning accuracy of the satellite imagery.
[0153] In this system, control points are ground points with accurate geographic information (longitude, latitude, and elevation), while laser elevation points are ground points with only accurate elevation information. The bundle adjustment technique involves adding an error correction model (which can be a translational or affine transformation model) to the image side of the RPC model. Then, equations are established using sparse control points or laser elevation points to calculate the error correction model parameters. These parameters are then fitted back into the RPC model to obtain a new model, thus completing the pre-positioning process.
[0154] For example, high-resolution multi-mode satellite stereo imagery of the campus area at 0.5 meters resolution can be acquired through commercial procurement. This stereo imagery includes both forward and backward views, with an intersection angle greater than 30°. Considering that the initial positioning accuracy of high-resolution multi-mode satellite imagery is 5-10 meters, to ensure the accuracy of subsequent camera calibration, such as... Figure 10 As shown, satellite imagery can be pre-positioned using three ground control points within the image coverage area, correcting the initial geometric errors in its imaging RPC model. Simultaneously, natural scene images captured by a camera installed on a building on campus are collected, and this data (such as camera images and the camera's initial position) is aggregated to a cloud device. The cloud device uses the acquired initial video location information, such as the city and district of XX, and then employs a deep learning-based cross-view geolocation method to perform location retrieval, obtaining more precise location information from the camera images (e.g., reference data). Figure 11The school shown has 45 buildings. A deep learning-based cross-view geolocation method can be used to locate the target of the camera image (building 23 of the school), identifying connection points between multiple satellite images and camera images. Then, the camera imaging parameters are calculated. Forward intersection is performed using the corrected satellite image's RPC imaging model and the camera's first imaging model (initial imaging model) to initialize the 3D coordinates of the object points corresponding to the connection points. A self-calibration adjustment model is constructed point-by-point using the connection point observations, and linearization is performed to construct error equations. Then, the unknowns in the object coordinates of the connection points are eliminated using elimination methods, constructing a modified method equation that only includes camera imaging parameters (principal point, principal distance, distortion coefficient, position, and attitude). By integrating observations from all connections, the camera imaging parameters are calculated and iteratively optimized using least squares adjustment until the optimal result is obtained, yielding the calibrated camera imaging parameters. Figure 12 As shown, the camera imaging parameters (principal point, principal distance, distortion coefficient, position, and attitude) obtained from the calibration can be used to construct an accurate imaging model of the camera (i.e., the second imaging model of the camera), and the accurate principal point, principal distance, distortion coefficient, position, and attitude at the moment of camera imaging can be recovered, thereby realizing the full utilization of urban cameras (such as achieving precise positioning based on the accurate imaging model).
[0155] In one possible implementation, the camera pose can also be determined based on the second imaging model described above.
[0156] Pose is short for position and orientation; pose can be represented by six variables, three of which indicate position and the other three indicate orientation.
[0157] For example, the position of the second feature point in the camera image can be input into the second imaging model mentioned above to obtain the camera pose.
[0158] Understandably, determining the camera pose based on a more accurate camera imaging model can improve the accuracy of the determined camera pose compared to using an inaccurate camera.
[0159] In one possible implementation, the position and / or area of the target region in the camera image in real three-dimensional space can also be determined based on the aforementioned camera pose.
[0160] For example, the position (longitude, latitude, and elevation) of the target in real three-dimensional space can be obtained based on the camera imaging parameters (principal point, principal distance, distortion coefficient, position, and attitude) and the target coordinates (row and column numbers) on the camera image.
[0161] For example, such as Figure 13As shown in Figure a, the location of the fire point in the camera image can be determined based on the camera's pose, so as to realize the alarm and location of the fire point in the city and provide real-time location information for fire emergency response.
[0162] For example, such as Figure 13 As shown in b, the location and area of the road collapse in the camera image can be determined based on the camera's pose, so as to achieve accurate calculation of the road collapse area and provide accurate road location and condition for road maintenance work.
[0163] For example, such as Figure 13 As shown in Figure c, multiple camera images can be stitched together based on the camera's pose to obtain a stitched image. This stitched image provides a wider field of view, eliminates blind spots, and achieves panoramic monitoring. In this way, operators can view images of multiple areas on the same screen, saving time and effort and improving work efficiency.
[0164] For example, such as Figure 13 As shown in d, the location of urban mobile facilities in the camera image can be determined based on the camera's pose, enabling urban management personnel to accurately and promptly obtain the location of urban mobile facilities and assist in urban planning and management.
[0165] For example, such as Figure 13 As shown in e, the position of a vehicle in a camera image can be determined based on the camera's pose, enabling vehicle management personnel to accurately and promptly obtain the vehicle's location and driving trajectory, thus assisting in vehicle management.
[0166] For example, such as Figure 13 As shown in f, the position of dynamic targets (such as animals, vehicles, aircraft, etc.) in the camera image can be determined based on the camera's pose, enabling city management personnel to accurately and promptly obtain the position of dynamic targets and assist them in managing and tracking them.
[0167] Understandably, accurate camera pose can determine the precise location and area of the target region in the camera image in real three-dimensional space, thereby enabling accurate road hazard forecasting (such as accurately locating and measuring urban road collapses and avoiding road risks), accurate disaster emergency response (such as accurately locating and imaging the location of disasters such as fires, floods, and building collapses), and accurate dynamic target tracking (such as accurately locating vehicles).
[0168] The image processing apparatus used to perform the above image processing method is described below.
[0169] It is understood that, in order to achieve the above-mentioned functions, the image processing apparatus includes hardware and / or software modules that perform the respective functions. Based on the algorithm steps of the examples described in conjunction with the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in conjunction with the embodiments, but such implementation should not be considered beyond the scope of the embodiments of this application.
[0170] This application embodiment can divide the image processing device into functional modules according to the above method example. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0171] When dividing each function into modules according to its corresponding function. Figure 14 The diagram illustrates a possible configuration of the image processing apparatus involved in the above embodiments. This apparatus can be an image processing device, a module (e.g., a chip or chip system) applied to an image processing device, or a logic node, logic module, or software capable of implementing all or part of the functions of the image processing device. Figure 14 As shown, the image processing device 1400 may include a transceiver unit 1401 and a processing unit 1402.
[0172] The transceiver unit 1401 is used to acquire satellite stereo images and camera images, wherein the satellite stereo images include multiple satellite images taken by the satellite.
[0173] The processing unit 1402 is used to determine the connection points between the aforementioned multiple satellite images and the aforementioned camera images.
[0174] The processing unit 1402 is also configured to determine the second imaging model of the camera based on the connection point, the first imaging model of the camera, and the imaging model of the satellite.
[0175] In one possible implementation, the processing unit 1402 is specifically configured to: determine a first feature point of the multiple satellite images and a second feature point of the camera image based on the multiple satellite images and the camera image; determine a first descriptor of the first feature point and a second descriptor of the second feature point based on the first feature point and the second feature point; determine a connection region between the multiple satellite images and the camera image based on the first descriptor and the second descriptor; and determine a connection point between the multiple satellite images and the camera image based on the connection region.
[0176] In one possible implementation, the processing unit 1402 is further configured to: determine the camera pose based on the second imaging model.
[0177] In one possible implementation, the processing unit 1402 is further configured to: determine the position and / or area of the target region of the camera image in real three-dimensional space based on the camera pose.
[0178] In one possible implementation, the aforementioned multiple satellite images include a first satellite image and a second satellite image, wherein the intersection angle between the first satellite image and the second satellite image is greater than a first threshold.
[0179] In one possible implementation, the connection point includes a first connection point and a second connection point, wherein the first connection point is a one-degree overlap connection point between the camera image and the single-view satellite image, and the second connection point is a multi-degree overlap connection point between the camera image and the multi-view stereo satellite image.
[0180] This application also provides a chip, which can be the chip of the above-described image processing device. Figure 15 A schematic diagram of a chip 1500 is shown. The chip 1500 includes one or more processors 1501 and interface circuitry 1502. Optionally, the chip 1500 may also include a bus 1503.
[0181] Processor 1501 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above image processing method can be completed through the integrated logic circuitry in the hardware of processor 1501 or through software instructions.
[0182] Optionally, the processor 1501 described above may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods and steps disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0183] The interface circuit 1502 can be used to send or receive data, instructions or information. The processor 1501 can use the data, instructions or other information received by the interface circuit 1502 to process the data, instructions or other information, and can send the processed information out through the interface circuit 1502.
[0184] Optionally, the chip may also include memory, which may include read-only memory and random access memory, providing operation instructions and data to the processor. A portion of the memory may also include non-volatile random access memory (NVRAM).
[0185] Optionally, the memory stores executable software modules or data structures, and the processor can execute corresponding operations by calling the operation instructions stored in the memory (which may be stored in the operating system).
[0186] Optionally, the chip can be used in the image processing apparatus or image processing device involved in the embodiments of this application. Optionally, the interface circuit 1502 can be used to output the execution result of the processor 1501. For the image processing methods provided by one or more embodiments of the present application, please refer to the foregoing embodiments, which will not be repeated here.
[0187] It should be noted that the functions of processor 1501 and interface circuit 1502 can be implemented through hardware design, software design, or a combination of hardware and software; no restrictions are imposed here.
[0188] Figure 16 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device can be an image processing device, a chip within the image processing device, or a functional module. For example... Figure 16 As shown, the electronic device 1600 includes a processor 1601, a transceiver 1602, and a communication line 1603.
[0189] The processor 1601 is used to execute any step of the image processing method provided in the embodiments of this application, and during the execution of any step of the image processing method provided in the embodiments of this application, it may choose to call the transceiver 1602 and the communication line 1603 to complete the corresponding operation.
[0190] Furthermore, the electronic device 1600 may also include a memory 1604. The processor 1601, the memory 1604, and the transceiver 1602 can be connected via a communication line 1603.
[0191] Processor 1601 can be a processor, a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. Processor 1601 can also be other devices with processing capabilities, such as circuits, devices, or software modules, without limitation.
[0192] Transceiver 1602 is used to communicate with other devices or other communication networks, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), etc. Transceiver 1602 can be a module, circuit, transceiver, or any device capable of enabling communication.
[0193] The transceiver 1602 is mainly used for sending and receiving commands and information, and may include a transmitter and a receiver to send and receive commands and information, respectively; operations other than sending and receiving commands and information are implemented by the processor.
[0194] Communication line 1603 is used to transmit information between the various components included in electronic device 1600.
[0195] In one design, the processor can be viewed as a logic circuit, and the transceiver as an interface circuit.
[0196] Memory 1604 is used to store instructions. These instructions can be computer programs.
[0197] The memory 1604 can be volatile memory or non-volatile memory, or both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM). Memory 1604 can also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed discs, laser discs, optical discs, digital universal discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices. It should be noted that the memory in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0198] It should be noted that the memory 1604 can exist independently of the processor 1601, or it can be integrated with the processor 1601. The memory 1604 can be used to store instructions, program code, or some data, etc. The memory 1604 can be located inside or outside the electronic device 1600, without limitation. The processor 1601 is used to execute the instructions stored in the memory 1604 to implement the methods provided in the above embodiments of this application.
[0199] In one example, processor 1601 may include one or more processor cores, for example Figure 16 The processor cores are 0 and 1.
[0200] As an optional implementation, the electronic device 1600 includes multiple processors, for example, besides Figure 16 In addition to processor 1601, it may also include processor 1607.
[0201] As an optional implementation, the electronic device 1600 also includes an output device 1605 and an input device 1606. For example, the input device 1606 is a device such as a keyboard, mouse, microphone, or joystick, and the output device 1605 is a device such as a display screen or speaker.
[0202] It should be noted that the electronic device 1600 can be a chip system or... Figure 16 Devices with similar structures. The chip system can be composed of chips or include chips and other discrete components. Actions, terminology, etc., involved in the various embodiments of this application can be referenced interchangeably without limitation. The message names or parameter names in the messages used for interaction between devices in the embodiments of this application are merely examples; other names can be used in specific implementations without limitation. Furthermore, Figure 16 The structural composition shown does not constitute a limitation on the electronic device 1600, except... Figure 16 In addition to the components shown, the electronic device 1600 may include more than Figure 16 This may indicate more or fewer components, or combinations of certain components, or different component arrangements.
[0203] The processor and transceiver described in this application can be implemented on integrated circuits (ICs), analog ICs, radio frequency integrated circuits, mixed-signal ICs, application-specific integrated circuits (ASICs), printed circuit boards (PCBs), electronic devices, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal-oxide semiconductors (CMOS), n-metal-oxide-semiconductor (NMOS), positive-channel metal-oxide semiconductors (PMOS), bipolar junction transistors (BJTs), bipolar CMOS (BiCMOS), silicon germanium (SiGe), gallium arsenide (GaAs), etc.
[0204] This application also provides an image processing apparatus, which includes at least one processor. When the at least one processor executes program code or instructions, it implements the image processing method described above.
[0205] Optionally, the device may further include at least one memory for storing the program code or instructions.
[0206] This application also provides a computer storage medium storing computer instructions. When the computer instructions are executed on an image processing device, the image processing device performs the aforementioned related method steps to implement the image processing method described above.
[0207] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the image processing method described above.
[0208] This application also provides an image processing apparatus, which may specifically be a chip, integrated circuit, component, or module. Specifically, the apparatus may include a connected processor and a memory for storing instructions, or the apparatus may include at least one processor for fetching instructions from external memory. When the apparatus is running, the processor can execute the instructions to cause the chip to perform the image processing methods described in the above-described method embodiments.
[0209] It should be understood that in various embodiments of this application, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of this application.
[0210] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.
[0211] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0212] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0213] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0214] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0215] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of this application embodiment, essentially, or the parts that contribute to the prior art, or parts of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0216] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the embodiments of this application should be included within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.
Claims
1. An image processing method, characterized in that, include: Acquire satellite stereo images and camera images, wherein the satellite stereo images include multiple satellite images captured by the satellite; Determine the connection points between the multiple satellite images and the camera images; The second imaging model of the camera is determined based on the connection point, the first imaging model of the camera, and the imaging model of the satellite.
2. The method according to claim 1, characterized in that, Determining the connection points between the multiple satellite images and the camera images includes: Based on the multiple satellite images and the camera images, determine the first feature points of the multiple satellite images and the second feature points of the camera images; Determine a first descriptor for the first feature point and a second descriptor for the second feature point based on the first feature point and the second feature point; The connection region between the plurality of satellite images and the camera images is determined based on the first descriptor and the second descriptor; The connection points between the multiple satellite images and the camera images are determined based on the connection area.
3. The method according to claim 1 or 2, characterized in that, The method further includes: The camera pose is determined based on the second imaging model.
4. The method according to claim 3, characterized in that, The method further includes: The position and / or area of the target region in the camera image in real three-dimensional space are determined based on the camera pose.
5. The method according to any one of claims 1 to 4, characterized in that, The multiple satellite images include a first satellite image and a second satellite image, and the intersection angle between the first satellite image and the second satellite image is greater than a first threshold.
6. The method according to any one of claims 1 to 5, characterized in that, The connection points include a first connection point and a second connection point. The first connection point is a one-degree overlap connection point between the camera image and the single-view satellite image, and the second connection point is a multi-degree overlap connection point between the camera image and the multi-view stereo satellite image.
7. An image processing apparatus, characterized in that, include: Transceiver unit and processing unit; The transceiver unit is used to acquire satellite stereo images and camera images, wherein the satellite stereo images include multiple satellite images captured by the satellite; The processing unit is used to determine the connection points between the multiple satellite images and the camera images; The processing unit is further configured to determine the second imaging model of the camera based on the connection point, the first imaging model of the camera, and the imaging model of the satellite.
8. The apparatus according to claim 7, characterized in that, The processing unit is specifically used for: Based on the multiple satellite images and the camera images, determine the first feature points of the multiple satellite images and the second feature points of the camera images; Determine a first descriptor for the first feature point and a second descriptor for the second feature point based on the first feature point and the second feature point; The connection region between the plurality of satellite images and the camera images is determined based on the first descriptor and the second descriptor; The connection points between the multiple satellite images and the camera images are determined based on the connection area.
9. The apparatus according to claim 7 or 8, characterized in that, The processing unit is also used for: The camera pose is determined based on the second imaging model.
10. The apparatus according to claim 9, characterized in that, The processing unit is also used for: The position and / or area of the target region in the camera image in real three-dimensional space are determined based on the camera pose.
11. The apparatus according to any one of claims 7 to 10, characterized in that, The multiple satellite images include a first satellite image and a second satellite image, and the intersection angle between the first satellite image and the second satellite image is greater than a first threshold.
12. The apparatus according to any one of claims 7 to 11, characterized in that, The connection points include a first connection point and a second connection point. The first connection point is a one-degree overlap connection point between the camera image and the single-view satellite image, and the second connection point is a multi-degree overlap connection point between the camera image and the multi-view stereo satellite image.
13. An image processing apparatus, comprising at least one processor and a memory, characterized in that, The at least one processor executes a program or instructions stored in a memory to cause the image processing apparatus to implement the method of any one of claims 1 to 6.
14. A computer-readable storage medium for storing a computer program, characterized in that, When the computer program is run on a computer or processor, it causes the computer or processor to perform the method of any one of claims 1 to 6.
15. A computer program product, the computer program product comprising instructions, characterized in that, When the instructions are executed on a computer or processor, the computer or processor causes the computer or processor to perform the method of any one of claims 1 to 6.