An image processing method and apparatus

By splicing multiple frames of fisheye images under spherical coordinate system, the distortion problem during fisheye camera stitching is solved, and more accurate panoramic image display is achieved, improving the safety and efficiency of autonomous driving and assisted driving.

CN114240769BActive Publication Date: 2025-07-22HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111369299.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-18
Publication Date
2025-07-22
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

In the prior art, the panoramic images obtained by stitching of multiple cameras have severe distortions, especially when the viewing angle of the fisheye camera is greater than 180 degrees, resulting in loss of viewing angle and loss of image content, and unable to effectively display the surrounding environment of the vehicle.

Method used

The image stitching method based on the spherical coordinate system is adopted to translate multiple frames of input images into the spherical space for spherical spherical spherical expansion of fisheye images, reduce distortion, determine the radius of the input image in each frame, and use coarse grain size and fine grain size step size to search for the optimal radius to achieve efficient stitching.

Benefits of technology

Reduce image distortion and obtain more accurate panoramic images, improve the driving safety and efficiency of vehicle autonomous driving and assisted driving, and enhance users' observation ability of the vehicle's surrounding environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114240769B_ABST
    Figure CN114240769B_ABST
Patent Text Reader

Abstract

The present application provides an image processing method and apparatus for stitching images based on a spherical surface, thereby reducing image distortion and obtaining a panoramic image with better effects. The method includes: first, acquiring multiple frames of input images to be stitched; then determining the position of each frame of the input images in a first coordinate system, where the first coordinate system includes a coordinate system in a spherical space, which is equivalent to establishing a unified world coordinate system, and determining the position of each frame of the input images in the world coordinate system; subsequently, stitching the multiple frames of input images according to their positions in the first coordinate system to obtain a panoramic image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular, to an image processing method and apparatus. Background Art

[0002] Generally, in some scenarios where images are required, such as in autonomous driving, assisted driving, or monitoring scenarios, multiple cameras are usually used to capture scenes in different directions. For example, in autonomous driving and assisted driving related services, in order to perceive the full-direction field of view of the vehicle, multiple cameras are often installed at different positions of the vehicle and multiple captured images are stitched and displayed. In order to minimize the number of cameras used as much as possible and ensure that adjacent cameras have sufficient overlapping areas, fisheye cameras with a larger field of view are often used in autonomous driving and assisted driving services. However, fisheye cameras have severe distortion, so the stitched panoramic image may have distortion. Therefore, how to obtain a more accurate panoramic image has become an urgent problem to be solved. Summary of the Invention

[0003] This application provides an image processing method and apparatus for stitching images based on a spherical surface, thereby reducing the distortion of the images and obtaining a panoramic image with better effects.

[0004] In view of this, in a first aspect, this application provides an image processing method, including: first, obtaining multiple frames of input images to be stitched; then determining the position of each frame of input image in a first coordinate system, where the first coordinate system includes a coordinate system in a spherical space, which is equivalent to establishing a unified world coordinate system, and determining the position of each frame of input image in the world coordinate system; subsequently, stitching the multiple frames of input images according to the position of each frame of input image in the first coordinate system to obtain a panoramic image.

[0005] Therefore, in the embodiments of this application, it is equivalent to translating multiple frames of input images into the same spherical space, and stitching each frame of input image as a spherical texture map, thereby reducing the distortion of the input images and obtaining a more accurate panoramic image. For example, when the input image is a fisheye image, unfolding the fisheye image through a spherical surface can reduce the distortion of the fisheye image, so that the distortion of the stitched panoramic image is also less, obtaining an accurate and clear panoramic image, enabling subsequent corresponding operations to be based on a more accurate panoramic image.

[0006] In a possible embodiment, the foregoing stitching the multiple frames of input images according to the position of each frame of input image in the first coordinate system to obtain a panoramic image may include: mapping the pixel values of the pixel points in each frame of input image to the corresponding spherical surface in the first coordinate system according to the corresponding stitching radius in each frame of input image to obtain a panoramic image, where the stitching radius is determined according to the position of each frame of input image in the first coordinate system.

[0007] Therefore, in the embodiments of the present application, since the first coordinate system is a spherical coordinate system, it can be a sphere or an ellipsoid, etc. Therefore, the input image can be unfolded as a texture map on the spherical surface in the first coordinate system. The shooting distances or angles of different input images may be different. Therefore, the stitching radius corresponding to each frame of the input image can be determined, so that the visual effects of various objects in the finally stitched panoramic image are better and the included information is more accurate.

[0008] In a possible implementation manner, before mapping the pixel values of the pixel points in each frame of the input image to the corresponding spherical surface in the first coordinate system according to the corresponding stitching radius in each frame of the input image to obtain a panoramic image, the above method may further include: determining the stitching radius corresponding to each frame of the input image according to the position of each frame of the input image in the first coordinate system at at least two step sizes. Optionally, the at least two step sizes are different, and the at least two step sizes can be understood as the search granularity used when searching for the stitching radius.

[0009] Therefore, in the embodiments of the present application, when determining the stitching radius of each frame of the input image, the radius search can be performed at different step sizes, so that the optimal stitching radius can be searched more efficiently.

[0010] In a possible implementation manner, the foregoing determining the stitching radius corresponding to each frame of the input image according to the position of each frame of the input image in the first coordinate system at at least two step sizes may include: taking the first image as an example, the first image is any one of multiple frames of the input image. According to the position of the first image in the first coordinate system, multiple coarse-grained radii are obtained at the first step size, that is, some stitching radii are searched coarsely; calculating the first reprojection error corresponding to the multiple coarse-grained radii, where the first reprojection error is the error between the observed position of the first image in the first coordinate system and the predicted position after projecting to the first coordinate system according to the coarse-grained radius; then screening out the first radius from the multiple radii according to the reprojection errors corresponding to the multiple radii; obtaining multiple fine-grained radii at the second step size, where the second step size is smaller than the first step size; calculating the second reprojection error corresponding to the multiple fine-grained radii, where the second reprojection error is the error between the observed position of the first image in the first coordinate system and the position after projecting to the first coordinate system according to the fine-grained radius; and obtaining the radius of the pixel points in each frame of the input image in the first coordinate system according to the second reprojection errors corresponding to the multiple fine-grained radii.

[0011] Therefore, in the embodiments of the present application, first, the optimal coarse-grained radius is searched coarsely at a larger step size, and then the optimal fine-grained radius is searched at an updated granularity within a certain range near the optimal coarse-grained radius, so as to efficiently determine the stitching radius matching each frame of the input image.

[0012] In a possible implementation, determining the position of each input image in the multi-frame input images in the first coordinate system includes: determining the position of the second image in the corresponding second coordinate system, where the second coordinate system is the coordinate system corresponding to the camera that captures the second image, and the second image is any one of the multi-frame input images; and determining the position of the second image in the first coordinate system according to the relative position relationship between the second coordinate system and the first coordinate system.

[0013] In the embodiments of the present application, when determining the position of each input image in the first coordinate system, the coordinate system corresponding to the camera that captures the input image can be determined, and the positions of each pixel point in each input image in the camera coordinate system can be determined, and according to the relative position relationship between the camera coordinate system and the first coordinate system, the position of the input image in the first coordinate system can be accurately determined.

[0014] In a possible implementation, the multi-frame input images are images captured by a fish-eye camera. Generally, the field of view angle of a fish-eye camera is relatively large. If only expanded according to a matrix, the output image is prone to distortion. Therefore, through the method provided by the present application, the fish-eye image is expanded on a sphere and stitched, which can reduce the distortion of the fish-eye image and obtain a more complete panoramic image.

[0015] In a possible implementation, the foregoing multi-frame input images can be images captured by a camera installed in a vehicle, and the stitched panoramic image can be used to plan a driving path for the vehicle when the vehicle is driving autonomously. Therefore, in the embodiments of the present application, a more accurate panoramic image with less distortion can be used to plan a driving path for the vehicle, and a more accurate and effective path can be planned for the vehicle, improving the driving efficiency and driving safety of the vehicle.

[0016] In a possible implementation, the panoramic image can also be displayed on the display screen of the vehicle, so that when the user is driving the vehicle, the surrounding environment of the vehicle can be observed omnidirectionally, reducing the blind area of the vehicle and improving the driving safety of the vehicle.

[0017] In a second aspect, the present application provides an image processing device, including:

[0018] An acquisition module, configured to acquire multi-frame input images;

[0019] A positioning module, configured to determine the position of each input image in the multi-frame input images in a first coordinate system, where the first coordinate system includes a coordinate system in a spherical space;

[0020] A stitching module, configured to stitch the multi-frame input images according to the position of each input image in the first coordinate system to obtain a panoramic image.

[0021] In a possible implementation, the stitching module is specifically configured to: map the pixel values of the pixel points in each input image to the corresponding spherical surface in the first coordinate system according to the corresponding stitching radius in each input image, so as to obtain a panoramic image, where the stitching radius is determined according to the position of each input image in the first coordinate system.

[0022] In a possible implementation, the stitching module is specifically configured to determine the stitching radius corresponding to each input image according to the position of each input image in the first coordinate system in at least two step sizes. Among them, the at least two step sizes can be understood as the search granularity used when searching for the stitching radius.

[0023] In a possible implementation, the stitching module is specifically configured to: obtain a plurality of coarse-grained radii according to the position of the first image in the first coordinate system in a first step size, where the first image is any one of the multiple input images; calculate the first reprojection error corresponding to the plurality of coarse-grained radii, where the first reprojection error is the error between the observed position of the first image in the first coordinate system and the position after projecting to the first coordinate system according to the coarse-grained radius; screen out the first radius from the plurality of radii according to the reprojection errors corresponding to the plurality of radii; obtain a plurality of fine-grained radii according to a second step size, where the second step size is smaller than the first step size; calculate the second reprojection error corresponding to the plurality of fine-grained radii, where the reprojection error is the error between the observed position of the first image in the first coordinate system and the position after projecting to the first coordinate system according to the fine-grained radius; and obtain the radius of the pixel points in each input image in the first coordinate system according to the reprojection errors corresponding to the plurality of fine-grained radii.

[0024] In a possible implementation, the positioning module is specifically configured to: determine the position of the second image in the corresponding second coordinate system, where the second coordinate system is the coordinate system corresponding to the camera that captures the second image, and the second image is any one of the multiple input images; and determine the position of the second image in the first coordinate system according to the relative position relationship between the second coordinate system and the first coordinate system.

[0025] In a possible implementation, the foregoing multiple input images are images collected by a fisheye camera.

[0026] In a possible implementation, the foregoing multiple input images are captured by a camera installed in a vehicle, and the panoramic image is used to plan a driving path for the vehicle when the vehicle is driving autonomously.

[0027] In a third aspect, an embodiment of the present application provides an image processing device, including: a processor and a memory. Among them, the processor and the memory are interconnected by a line, and the processor calls the program code in the memory to execute the functions related to processing in the image processing method shown in any item of the first aspect above. Optionally, the image processing device may be a chip.

[0028] In a fourth aspect, an embodiment of the present application provides a digital processing chip or a chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is used to execute the functions related to processing in any one of the optional implementation manners in the first aspect described above.

[0029] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute the method in the first aspect or any one of the optional implementation manners in the first aspect.

[0030] In a sixth aspect, an embodiment of the present application provides a computer program product containing computer programs / instructions, which when executed by a processor, cause the processor to execute the method in the first aspect or any one of the optional implementation manners in the first aspect. Description of the Drawings

[0031] Figure 1 It is a schematic diagram of a vehicle provided by the present application;

[0032] Figure 2 It is a schematic diagram of the architecture of an image processing system provided by the present application;

[0033] Figure 3 It is a schematic diagram of an application scenario of image processing provided by the present application;

[0034] Figure 4 It is another schematic diagram of an application scenario of image processing provided by the present application;

[0035] Figure 5 It is another schematic diagram of an application scenario of image processing provided by the present application;

[0036] Figure 6 It is another schematic diagram of an application scenario of image processing provided by the present application;

[0037] Figure 7 It is another schematic diagram of an application scenario of image processing provided by the present application;

[0038] Figure 8 It is another schematic diagram of an application scenario of image processing provided by the present application;

[0039] Figure 9 It is a schematic diagram of the process of image processing provided by the present application;

[0040] Figure 10 It is a schematic diagram of the process of another image processing method provided by the present application;

[0041] Figure 11Flow intent of another image processing method provided by this application;

[0042] Figure 12 Application scenario intent of another image processing provided by this application;

[0043] Figure 13 Application scenario intent of another image processing provided by this application;

[0044] Figure 14 Application scenario intent of another image processing provided by this application;

[0045] Figure 15 Flow schematic diagram of another image processing method provided by this application;

[0046] Figure 16 Application scenario schematic diagram of another image processing provided by this application;

[0047] Figure 17 Structural schematic diagram of an image processing device provided by this application;

[0048] Figure 18 Structural schematic diagram of another image processing device provided by this application. Detailed implementation

[0049] Next, the technical solutions in the embodiments of this application will be described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0050] The image processing method provided by this application can be applied to scenarios involving images, such as photographing, autonomous driving, scene monitoring, autonomous driving, drone shooting, etc. The image processing method provided by this application can be executed by an image processing device, and this image processing device can be an electronic device with a photographing function or connected to a photographing device.

[0051] For example, the method provided by this application can be executed by a vehicle or a processing device connected to the vehicle, and the structure of this vehicle can be as Figure 1 shown Figure 1A schematic structural diagram of a vehicle provided by an embodiment of the present application. The vehicle 100 can be configured in an autonomous driving mode. For example, the vehicle 100 can control itself while in the autonomous driving mode, and can determine the current state of the vehicle and its surrounding environment through manual operation, determine whether there are obstacles in the surrounding environment, and control the vehicle 100 based on the information of the obstacles. When the vehicle 100 is in the autonomous driving mode, the vehicle 100 can also be set to operate without interacting with people.

[0052] The vehicle 100 can include various subsystems, such as a propulsion system 102, a sensor system 104, a control system 106, one or more peripheral devices 108, as well as a power source 110, a computer system 112, and a user interface 116. Optionally, the vehicle 100 can include more or fewer subsystems, and each subsystem can include multiple components. In addition, each subsystem and component of the vehicle 100 can be interconnected by wire or wirelessly.

[0053] The propulsion system 102 can include components that provide powered movement for the vehicle 100. In one embodiment, the propulsion system 102 can include an engine 118, an energy source 119, a transmission 120, and wheels / tires 121.

[0054] Among them, the engine 118 can be an internal combustion engine, an electric motor, an air compression engine, or other types of engine combinations. For example, a hybrid engine composed of a gasoline engine and an electric motor, or a hybrid engine composed of an internal combustion engine and an air compression engine. The engine 118 converts the energy source 119 into mechanical energy. Examples of the energy source 119 include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other power sources. The energy source 119 can also provide energy for other systems of the vehicle 100. The transmission 120 can transmit the mechanical power from the engine 118 to the wheels 121. The transmission 120 can include a gearbox, a differential, and a drive shaft. In one embodiment, the transmission 120 can also include other devices, such as a clutch. Among them, the drive shaft can include one or more shafts that can be coupled to one or more wheels 121.

[0055] The sensor system 104 may include several sensors that sense information about the environment surrounding the vehicle 100. For example, the sensor system 104 may include a positioning system 122 (the positioning system may be a global positioning GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU) 124, a radar 126, a lidar 128, and a camera 130. The sensor system 104 may also include sensors that monitor the internal systems of the vehicle 100 (e.g., in-vehicle air quality monitor, fuel gauge, oil temperature gauge, etc.). Sensing data from one or more of these sensors can be used to detect objects and their corresponding characteristics (position, shape, orientation, speed, etc.). Such detection and identification are key functions for the safe operation of the autonomous vehicle 100. The sensors mentioned in the following embodiments of this application may be the radar 126, the lidar 128, or the camera 130, etc.

[0056] Among them, the positioning system 122 can be used to estimate the geographical location of the vehicle 100. The IMU 124 is used to sense the changes in the position and orientation of the vehicle 100 based on inertial acceleration. In one embodiment, the IMU 124 may be a combination of an accelerometer and a gyroscope. The radar 126 can use radio signals to sense objects within the surrounding environment of the vehicle 100, specifically, it can be a millimeter-wave radar or a lidar. In some embodiments, in addition to sensing objects, the radar 126 can also be used to sense the speed and / or forward direction of the objects. The lidar 128 can use lasers to sense objects in the environment where the vehicle 100 is located. In some embodiments, the lidar 128 may include one or more laser sources, a laser scanner, and one or more detectors, as well as other system components. The camera 130 can be used to capture multiple images of the surrounding environment of the vehicle 100. The camera 130 can be a static camera or a video camera.

[0057] The control system 106 controls the operation of the vehicle 100 and its components. The control system 106 may include various components, including a steering system 132, an accelerator 134, a braking unit 136, a computer vision system 140, a line control system 142, and an obstacle avoidance system 144.

[0058] Among them, the steering system 132 is operable to adjust the forward direction of the vehicle 100. For example, in one embodiment, it can be a steering wheel system. The throttle 134 is used to control the operating speed of the engine 118 and thus control the speed of the vehicle 100. The braking unit 136 is used to control the deceleration of the vehicle 100. The braking unit 136 can use friction to slow down the wheels 121. In other embodiments, the braking unit 136 can convert the kinetic energy of the wheels 121 into electric current. The braking unit 136 can also take other forms to slow down the rotational speed of the wheels 121 so as to control the speed of the vehicle 100. The computer vision system 140 can be operated to process and analyze the images captured by the camera 130 in order to identify objects and / or features in the surrounding environment of the vehicle 100. The objects and / or features may include traffic signals, road boundaries, and obstacles. The computer vision system 140 can use object recognition algorithms, structure from motion (SFM) algorithms, video tracking, and other computer vision technologies. In some embodiments, the computer vision system 140 can be used to map the environment, track objects, estimate the speed of objects, and so on. The route control system 142 is used to determine the driving route and driving speed of the vehicle 100. In some embodiments, the route control system 142 may include a lateral planning module 1421 and a longitudinal planning module 1422. The lateral planning module 1421 and the longitudinal planning module 1422 are respectively used to determine the driving route and driving speed of the vehicle 100 by combining data from the obstacle avoidance system 144, the GPS 122, and one or more predetermined maps. The obstacle avoidance system 144 is used to identify, evaluate, and avoid or otherwise cross obstacles in the environment of the vehicle 100. The aforementioned obstacles may specifically be actual obstacles and virtual moving objects that may collide with the vehicle 100. In one example, the control system 106 can additionally or alternatively include components other than those shown and described. Or some of the above-mentioned shown components can also be reduced.

[0059] Vehicle 100 interacts with external sensors, other vehicles, other computer systems, or users through peripheral device 108. Peripheral device 108 may include a wireless data transmission system 146, an in-vehicle computer 148, a microphone 150, and / or a speaker 152. In some embodiments, peripheral device 108 provides a means for the user of vehicle 100 to interact with user interface 116. For example, in-vehicle computer 148 may provide information to the user of vehicle 100. User interface 116 may also operate in-vehicle computer 148 to receive user input. In-vehicle computer 148 may be operated through a touch screen. In other cases, peripheral device 108 may provide a means for vehicle 100 to communicate with other devices located within the vehicle. For example, microphone 150 may receive audio from the user of vehicle 100 (e.g., voice commands or other audio inputs). Similarly, speaker 152 may output audio to the user of vehicle 100. Wireless data transmission system 146 may wirelessly communicate with one or more devices directly or via a communication network. For example, wireless data transmission system 146 may use 3G cellular communication, such as CDMA, EVDO, GSM / GPRS, or 4G cellular communication, such as LTE. Or 5G cellular communication. Wireless data transmission system 146 may utilize wireless local area network (WLAN) communication. In some embodiments, wireless data transmission system 146 may directly communicate with devices using an infrared link, Bluetooth, or ZigBee. Other wireless protocols, such as various vehicle data transmission systems, for example, wireless data transmission system 146 may include one or more dedicated short range communications (DSRC) devices, which may include public and / or private data communication between vehicles and / or roadside stations.

[0060] Power source 110 may supply power to various components of vehicle 100. In one embodiment, power source 110 may be a rechargeable lithium-ion or lead-acid battery. One or more battery packs of such a battery may be configured to supply power to various components of vehicle 100. In some embodiments, power source 110 and energy source 119 may be implemented together, as in some all-electric vehicles.

[0061] Some or all of the functions of vehicle 100 are controlled by computer system 112. Computer system 112 may include at least one processor 113 that executes instructions 115 stored in a non-transitory computer-readable medium such as memory 114. Computer system 112 may also be multiple computing devices that control individual components or subsystems of vehicle 100 in a distributed manner. Processor 113 may be any conventional processor, such as a commercially available central processing unit (CPU). Optionally, processor 113 may be a special-purpose device such as an application-specific integrated circuit (ASIC) or other hardware-based processor. Although Figure 1 The functional diagram shows the processor, memory, and other components of computer system 112 in the same block, but those of ordinary skill in the art should understand that the processor or memory may actually include multiple processors or memories that are not stored within the same physical housing. For example, memory 114 may be a hard disk drive or other storage medium located within a housing different from computer system 112. Thus, references to processor 113 or memory 114 will be understood to include references to a collection of processors or memories that may or may not operate in parallel. Instead of using a single processor to perform the steps described herein, some components such as the steering component and the deceleration component may each have their own processor that only performs calculations related to component-specific functions.

[0062] In various aspects described herein, processor 113 may be located remote from vehicle 100 and communicate wirelessly with vehicle 100. In other aspects, some of the processes described herein are executed on a processor 113 disposed within vehicle 100 while others are executed by a remote processor 113, including taking the necessary steps to perform a single maneuver.

[0063] In some embodiments, the memory 114 may include instructions 115 (e.g., program logic) that may be executed by the processor 113 to perform various functions of the vehicle 100, including those described above. The memory 114 may also include additional instructions, including instructions to send data to, receive data from, interact with, and / or control one or more of the propulsion system 102, the sensor system 104, the control system 106, and the peripheral devices 108. In addition to the instructions 115, the memory 114 may also store data, such as road maps, route information, the position, orientation, speed of the vehicle, and other such vehicle data, as well as other information. Such information may be used by the vehicle 100 and the computer system 112 during operation of the vehicle 100 in autonomous, semi-autonomous, and / or manual modes. The user interface 116 is configured to provide information to or receive information from a user of the vehicle 100. Optionally, the user interface 116 may include one or more input / output devices within the set of peripheral devices 108, such as a wireless data transmission system 146, an on-board computer 148, a microphone 150, or a speaker 152, etc.

[0064] The computer system 112 may control the functions of the vehicle 100 based on inputs received from various subsystems (e.g., the propulsion system 102, the sensor system 104, and the control system 106) and from the user interface 116. For example, the computer system 112 may communicate with other systems or components within the vehicle 100 via a CAN bus. As an example, the computer system 112 may utilize inputs from the control system 106 to control the steering system 132 to avoid obstacles detected by the sensor system 104 and the obstacle avoidance system 144. In some embodiments, the computer system 112 may be operable to provide control over many aspects of the vehicle 100 and its subsystems.

[0065] Optionally, one or more of the above components may be installed or associated separately from the vehicle 100. For example, the memory 114 may exist partially or completely separately from the vehicle 100. The above components may be communicatively coupled together in a wired and / or wireless manner.

[0066] Optionally, the above components are only an example. In actual applications, the components in each of the above modules may be added or deleted according to actual needs. Figure 1It should not be construed as a limitation on the embodiments of the present application. The data transmission method provided by the present application can be executed by a computer system 112, a radar 126, a laser rangefinder 130, or a peripheral device, such as an in-vehicle computer 148 or other in-vehicle terminals. For example, the data transmission method provided by the present application can be executed by the in-vehicle computer 148. The in-vehicle computer 148 can plan a driving path and a corresponding speed curve for the vehicle, generate a control instruction according to the driving path, and send the control instruction to the computer system 112. The computer system 112 controls the steering system 132, the throttle 134, the braking unit 136, the computer vision system 140, the line control system 142, or the obstacle avoidance system 144 in the vehicle control system 106 of the vehicle, so as to realize the automatic driving of the vehicle.

[0067] The above-mentioned vehicle 100 can be a sedan, a truck, a motorcycle, a bus, a ship, an airplane, a helicopter, a lawn mower, a recreational vehicle, a playground vehicle, a construction equipment, a tram, a golf cart, a train, and a trolley, etc. The embodiments of the present application do not make special limitations.

[0068] In some scenarios involving images, such as shooting, assisted driving, autonomous driving, or automatic parking, etc., it may be necessary to use multiple cameras or cameras to take pictures, and then splice the images taken by the multiple cameras or cameras to obtain an image including more information. For example, in autonomous driving and assisted driving related services, in order to perceive the full-direction field of view of the vehicle, multiple cameras are often installed at different positions of the vehicle and multiple captured images are spliced and displayed. In order to minimize the number of cameras used as much as possible and ensure that there is enough overlapping area between adjacent cameras, fish-eye cameras with a larger field of view are often used in autonomous driving and assisted driving services. Usually, fish-eye cameras have severe distortion, and traditional algorithms need to correct them before splicing. Since the viewing angle of the fish-eye camera is very large, even exceeding 180 degrees, directly using a plane-based fish-eye distortion correction method will cause serious stretching and loss of the viewing angle. In addition, since the fish-eye cameras are located at different positions of the vehicle and their optical centers do not coincide, the homography matrix cannot be directly used for registration.

[0069] Some common splicing methods can complete the splicing of multiple-channel fish-eye images in the form of a bird's-eye view. For example, first correct the distortion of the multiple-channel fish-eye images collected, and then use the direct linear transformation algorithm or the inverse projection transformation algorithm based on the viewing angle parameters to convert the fish-eye images from the current viewing angle to the top-down viewing angle. However, the viewing angle of the image will be lost during the conversion process. Finally, the converted top-down images are used with the algorithm idea of traditional image splicing to obtain a full-direction bird's-eye view after splicing. For example, when there is a situation where the surrounding scene is higher than the vehicle body, the distortion of the image after the top-down transformation is large and the observation distance is short, resulting in loss of the viewing angle and loss of image content, and it is impossible to observe the surrounding environment comprehensively.

[0070] Alternatively, each fisheye image can be cylindrically unfolded according to the longitude and latitude coordinates, and then feature points can be extracted and feature matching can be performed on the cylindrically unfolded image. The pose matrix of adjacent cameras can be calculated based on the three-dimensional coordinates of the feature matching pairs, and the panoramic image can be stitched. However, if cylindrical projection is used, the change of the pitch angle of the fisheye camera cannot be efficiently displayed, and stretching-induced distortion will occur at the upper and lower ends of the cylinder, and there is distortion in the cylindrically unfolded image of the fisheye. The camera pose calculated using the feature points extracted from it is not accurate enough.

[0071] Therefore, the present application provides an image processing method that can map multiple images onto a spherical surface, reduce image distortion, and obtain a clearer and more accurate panoramic image.

[0072] First, the image processing method provided by the present application can be applied to various scenarios, such as vehicle surround view systems, autonomous driving, automatic parking, or assisted driving scenarios.

[0073] Exemplarily, the framework applied in the present application can be as Figure 2 shown.

[0074] First, multiple fisheye cameras can be set on the vehicle to collect images of the surrounding environment of the vehicle. The images of the surrounding environment can be collected through the multi-channel fisheye cameras to obtain multiple frames of fisheye images.

[0075] Then, the stitched panoramic image can be obtained through the image processing method provided by the present application. Specifically, it can be stitched through the image processing method provided by the present application. For example, after obtaining multiple frames of input images, the position of each frame of input image in the first coordinate system is determined, and each frame of input image is stitched as a spherical surface of the spherical space according to the position of each frame of input image in the first coordinate system to obtain a panoramic image.

[0076] Refer to the following Figures 9 - 16 introduction, which will not be elaborated here.

[0077] After obtaining the panoramic image, it can be applied to various scenarios. For example, it can be displayed through a user visualization interface, enabling the user to know the surrounding environment of the vehicle through the visualization interface; or obstacle detection can be performed, and the vehicle can be controlled based on the detection result, such as planning the autonomous driving path of the vehicle or controlling the vehicle to brake to avoid obstacles, thereby improving the driving safety of the vehicle.

[0078] For ease of understanding, the application scenarios of the method provided by the present application will be introduced exemplarily below.

[0079] Scenario 1: Vehicle panoramic imaging

[0080] Taking the acquisition of images by a vehicle-mounted four-channel fisheye camera as an example, images of the scene around the vehicle are acquired by the four-channel fisheye camera on the vehicle, and then the multi-channel fisheye images can be stitched by the image processing method provided in this application to obtain a panoramic image including the environment around the vehicle. For example, Figure 3 As shown, a display interface is set on the central control part of the vehicle, and the panoramic image can be displayed on this display interface. For example, when the user is driving the vehicle, the environment around the vehicle can be observed through the in-vehicle display interface. For example, when changing lanes, the environment around the vehicle can be fully observed, thereby reducing the visual blind area and improving the driving safety of the vehicle. Or, in the parking scenario, the user can observe the environment around the vehicle through the in-vehicle display interface, so as to avoid pedestrians or obstacles and complete parking more safely.

[0081] Of course, in addition to displaying the panoramic image on the display interface of the central control part of the vehicle, the panoramic image can also be displayed on components with display interfaces such as the rearview mirror, instrument panel of the vehicle, or through a head-up display (HUD), etc., so that the user can intuitively observe the environment around the vehicle. Details are not described one by one here, and the components for displaying the panoramic image can be specifically determined according to the actual application scenario.

[0082] Scenario 2: Obstacle detection for assisted driving

[0083] Taking the acquisition of images by a vehicle-mounted four-channel fisheye camera as an example, if obstacle detection is performed on the images acquired by the four-channel fisheye camera respectively, duplicate removal needs to be performed on the images acquired by the four-channel fisheye camera, and the amount of calculation is relatively large. Therefore, the images acquired by the four-channel fisheye camera can be stitched to obtain a panoramic image, and detection can be performed in the panoramic image. For example, Figure 4 As shown, the detected targets are highlighted by marking frames, thereby reducing the amount of calculation, simplifying the detection operation, and reducing the detection delay.

[0084] Scenario 3: Mobile phone shooting

[0085] Among them, a wide-angle camera can be set in the user's mobile phone. When the user enables the camera function to take pictures, the mobile phone can be moved to take pictures from multiple angles. The method provided in this application can be used to stitch the pictures taken from multiple angles to obtain a panoramic image and display it on the mobile phone display screen.

[0086] For example, as Figure 5 shown, the user can use the wide-angle camera of the mobile phone to take pictures, move the mobile phone during the shooting process to achieve multi-angle shooting of the current scene, and then stitch multiple images by the method provided in this application. As Figure 6 shown, a part of the viewport area in the stitched image can be displayed on the mobile phone, and the user can switch the viewport area by moving the mobile phone.

[0087] Scenario 4: Augmented Reality (AR)

[0088] When the user starts the AR application, the AR device processes the real-time image captured by the camera, or can also obtain the real-time image from the real-time image buffer and perform AR processing on the obtained real-time image. The specific processing process can be to splice the images captured by multiple cameras through the method provided in this application to obtain a spliced panoramic image, and then a three-dimensional object required by the user can be added to the panoramic image. The user can use an AR device, such as Figure 7 the head-mounted device shown, to view the panoramic image and the added three-dimensional object.

[0089] Scenario 5: Drone shooting

[0090] Among them, multiple cameras can be set in the drone. The itinerary of the drone can be planned to complete automatic flight, or the drone can be controlled by the user. Images of different orientations of the drone can be captured by multiple cameras and spliced through the method provided in this application to obtain a panoramic image, so that the drone can perceive obstacles 360 degrees during flight and thus avoid the obstacles. Or, images of multiple orientations can be captured by multiple cameras set in the drone, and the depth information of 360 degrees near the drone can be obtained through the spliced panoramic image, thereby constructing a three-dimensional map of the flight area of the drone. For example, as Figure 8 shown, the surrounding environment information can be captured by a drone flying in the air, and then the multi-frame images can be spliced through the method provided in this application to obtain the final panoramic image, so that the user can remotely obtain the surrounding environment information from the panoramic image.

[0091] In addition, it can also be applied to more scenarios, such as camera shooting scenarios, target detection scenarios, etc. that require splicing of multi-frame images. This application will not elaborate on them one by one.

[0092] Refer to Figure 9 for the schematic flowchart of an image processing method provided in this application, which is described as follows.

[0093] 901. Obtain multiple frames of input images.

[0094] Among them, the multiple frames of input images can be images captured by one or more imaging devices. The imaging device can include various terminals with imaging functions, such as cameras, mobile phones, monitoring devices, dash cams, or intelligent robots, etc. The camera can specifically include a fish-eye camera, a depth camera, a telephoto camera, or other devices with an image sensor, etc.

[0095] When there are multiple imaging devices, the multiple imaging devices can capture images of different scenes or different orientations in the same scene; when there is one imaging device, the one device can capture images of different orientations by adjusting its own pose, etc., and can be specifically adjusted according to the actual application scenario.

[0096] For example, one or more cameras can be respectively arranged at various positions of a vehicle, and then the environmental information around the vehicle can be collected through the cameras, and multiple input images can be obtained.

[0097] 902. Determine the position of each frame of input image in the first coordinate system.

[0098] Among them, after obtaining multiple frames of input images, the positions of each pixel point of the multiple frames of input images can be mapped to the same coordinate system, that is, the first coordinate system, so as to realize translating the multiple frames of input images to the same space, so that subsequent image stitching can be more accurate.

[0099] Specifically, when determining the position of each frame of input image in the first coordinate system, taking any one frame of input image as an example, which is called the second image for easy distinction, determine the position of the second image in the second coordinate system, where the second coordinate system is the coordinate system corresponding to the camera that captures the second image, determine the relative position relationship between the second coordinate system and the first coordinate system, and determine the position of the second image in the first coordinate system according to the relative position relationship.

[0100] It can be understood that the positions of each pixel point in the image can be determined first in the camera coordinate system, and the camera coordinate system is translated into the established world coordinate system (that is, the first coordinate system) to obtain the relative position relationship between the camera coordinate system and the world coordinate system, and then the positions of each pixel point in each frame of input image in the camera coordinate system can be mapped to the world coordinate system according to the relative position relationship to obtain the position of each frame of input image in the world coordinate system.

[0101] Among them, when mapping the positions of each pixel point of multiple frames of input images to the same coordinate system, one pixel point can be used as a unit to obtain a higher-definition image, or multiple pixel points can be used as a unit for mapping to improve the mapping efficiency, and can be specifically adjusted according to the actual application scenario, and this application does not limit this.

[0102] 903. Stitch multiple frames of input images according to the positions of each frame of input image in the first coordinate system to obtain a panoramic image.

[0103] Among them, after obtaining the positions of each frame of input image in the first coordinate system, each frame of input image can be stitched as a texture map on the spherical surface of a spherical space to obtain a panoramic image.

[0104] The subsequent panoramic image can be used for tasks such as display, detection, or segmentation. For example, in an autonomous driving scenario, the panoramic image can be used to plan the driving path of a vehicle, thereby planning a more accurate and efficient driving path, avoiding obstacles. Or, the panoramic image can be displayed on an in-vehicle display screen, so that when the user is driving, they can obtain the surrounding environment of the vehicle from the displayed panoramic image, reducing the driving blind spot and improving driving safety. Another example is that in a shooting scenario, multiple cameras can be used for shooting, so that a more complete image can be captured, improving the user experience.

[0105] Therefore, in the embodiments of the present application, it is equivalent to translating multiple input images into the same spherical space and splicing each input image as a spherical texture map, so that the image can be fully unfolded, thereby reducing the distortion of the input image and obtaining a more accurate panoramic image. For example, when the input image is a fisheye image, the fisheye image can be unfolded through a sphere, reducing the distortion of the fisheye image, so that the distortion of the spliced panoramic image is also less, obtaining an accurate and clear panoramic image, enabling subsequent corresponding operations to be based on a more accurate panoramic image.

[0106] Specifically, the specific process of splicing multiple input images may include: determining the splicing radius corresponding to each input image according to the position of each input image in the first coordinate system, where the splicing radius is the radius in the first coordinate system when each input image is spliced with an adjacent input image; then, according to the corresponding splicing radius in each input image, mapping the pixel values of the pixel points in each input image to the corresponding sphere in the first coordinate system to obtain a panoramic image. Therefore, in the embodiments of the present application, after calculating the splicing radius, splicing can be performed according to the splicing radius to obtain a complete panoramic image.

[0107] Optionally, when obtaining the splicing radius corresponding to each input image, the splicing radius of each input image can be determined from a certain range according to at least two step sizes.

[0108] Specifically, taking any one of the multi-frame input images (referred to as the first image for easy distinction) as an example, multiple coarse-grained radii can be obtained according to the position of the first image in the first coordinate system at a first step size; calculate the first reprojection error corresponding to the multiple coarse-grained radii, where the first reprojection error is the error between the observed position of the first image in the first coordinate system and the position after projecting to the first coordinate system according to the coarse-grained radius; screen out the first radius from the multiple radii according to the reprojection errors corresponding to the multiple radii; obtain multiple fine-grained radii at a second step size, where the second step size is smaller than the first step size; calculate the second reprojection error corresponding to the multiple fine-grained radii, where the second reprojection error is the error between the observed position of the first image in the first coordinate system and the position after projecting to the first coordinate system according to the fine-grained radius; obtain the radius of the pixel points in each frame of the input image in the first coordinate system according to the second reprojection errors corresponding to the multiple fine-grained radii.

[0109] Therefore, in the embodiments of the present application, the stitching radius matching the input image can be searched according to the coarse-grained step size and the fine-grained step size, so that a suitable stitching radius can be searched efficiently, and thus the input image can be stitched more effectively, avoiding distortion during image stitching, and the resulting panoramic image is more accurate.

[0110] The above describes the process of the image processing method provided by the present application. For easy understanding, the image processing method provided by the present application will be introduced in more detail below in combination with specific application scenarios.

[0111] Exemplarily, taking the fisheye image collected by the fisheye camera installed on the vehicle as the input image for exemplary illustration. Refer to Figure 10 Exemplarily, the image processing method provided by the present application may include multiple steps, such as Figure 6 shown in: image preprocessing, panoramic unfolding of the fisheye image, panoramic image stitching, and seam processing, etc., which will be introduced separately below.

[0112] Image preprocessing: That is, preprocess the fisheye image collected by the fisheye camera. For example, by adding a mask, process the part of the ego vehicle included in the fisheye image and retain the environmental information collected in the fisheye image. Among them, the image preprocessing step is an optional step. For example, if the fisheye image does not include the part of the ego vehicle, the image preprocessing may not be required.

[0113] Panoramic unfolding of the fisheye image: Usually, the fisheye image is a two-dimensional image of a certain size, and each frame of the fisheye image can be unfolded into a panoramic image respectively.

[0114] Panoramic image stitching: That is, stitch the panoramic images after unfolding multiple frames of fisheye images to obtain the final panoramic image.

[0115] Seam processing: Optimize the seams of the stitched panoramic image to reduce the parallax at the seams and improve the user's viewing experience. Among them, seam processing is an optional step. For example, if the same stitching radius is used for each panoramic image processing, the seam processing can be omitted, and an accurate panoramic image can still be obtained.

[0116] More specifically, the process of the image processing method provided in this application can be referred to Figure 11 . First, the fisheye image can be unfolded to obtain the two-dimensional panoramic image coordinates. Then, the two-dimensional panoramic image coordinates are mapped to the unit sphere coordinates, and the unit sphere coordinates are flattened into a unified world coordinate system. Combining the internal parameters or external parameters of the camera, the coordinates of the fisheye image in the unified world coordinate system are obtained. Then, the pixel values of each pixel point in the fisheye image are mapped to the world coordinate system to obtain the final panoramic image.

[0117] Taking the example of setting four fisheye cameras in a vehicle, the above steps will be introduced in more detail by way of example.

[0118] Step 1. Image preprocessing

[0119] Usually, the fisheye cameras are fixed to the vehicle body and their positions relative to the vehicle body will not change. Therefore, a mask corresponding to each fisheye camera can be preset respectively. As Figure 12 shown, filter the vehicle part and the image edge part, so that the obtained fisheye image does not include the vehicle part and can reduce the invalid distorted part. It can be understood that the fisheye image and the mask can be fused. When fusing, the weights of the vehicle part or the edge part included in the fisheye image are set to be lower or 0, so as to reduce the vehicle part or the edge part included in the fisheye image.

[0120] It should be noted that image preprocessing is an optional step. The fisheye images mentioned in the following steps can be either the fisheye images obtained after preprocessing or the images without preprocessing, which can be adjusted according to the actual application scenario, and this application does not make any limitations in this regard.

[0121] Step 2. Unfolding the fisheye image

[0122] The four preprocessed fisheye images can be unfolded respectively with an initial radius to obtain four panoramic images, and the content of the fisheye image only occupies a part of the panoramic image.

[0123] Generally, in order to reduce the loss of viewing angle after the distortion correction of the fisheye image, the fisheye image can be unfolded in the form of a panoramic image. The specific steps include:

[0124] Map the coordinates (x, y) of each pixel point in the two-dimensional fisheye image to the spherical longitude and latitude coordinates, which is expressed as:

[0125] longitude = xπ

[0126] latitude = yπ / 2

[0127] Subsequently, the longitude and latitude coordinates are translated into the unit sphere to obtain (P x , P y , P z ), which is expressed as:

[0128] P x = cos(latitude)cos(longitude)

[0129] P y = cos(latitude)sin(longitude)

[0130] P z = sin(latitude)

[0131] For different dimensions, different radii can be set to obtain new three-dimensional coordinates (P x , P y , P z ), such as expressed as:

[0132] P x = r * P x

[0133] P y = r * P y

[0134] P z = r * P z

[0135] Subsequently, the obtained three-dimensional coordinates (P x , P y , P z ) can be projected into the coordinates of the two-dimensional fisheye image, or it can be understood as determining the coordinates of each point in the fisheye image in the unit sphere. During the imaging process, the fisheye image distortion coefficients are introduced, expressed as:

[0136]

[0137] θ d = (k0θ + k1θ 3 + k2θ 5 + k3θ 7 + k4θ 9 + …)

[0138] φ = atan2(P z , P x )

[0139] x = fθ × cos(φ)

[0140] y = fθ × sin(φ)

[0141] Wherein, θ represents the incident angle, and θ d represents the included angle after distortion, and φ represents the included angle with the two-dimensional plane coordinate axis.

[0142] Exemplarily, as Figure 13 shown, after preprocessing the fish-eye image, the obtained image after preprocessing is unfolded in the panorama to obtain the position of each frame of fish-eye image in the panorama.

[0143] Step Three, Panoramic Image Stitching

[0144] Exemplarily, as Figure 14 shown, four fish-eye cameras can be respectively set at different positions of the vehicle, such as Figure 10 the fish-eye camera 1, fish-eye camera 2, fish-eye camera 3, and fish-eye camera 4 shown therein. Different fish-eye cameras can collect scenes within different fields of view relative to the vehicle to obtain fish-eye images respectively. In order to stitch multiple frames of fish-eye images into a single panoramic image, a unified world coordinate system needs to be established. Therefore, the unit sphere corresponding to each camera can be translated to the center of a unified preset sphere, such as the center of sphere 0, and multiple fish-eye images are stitched through the radius of the finally unfolded panoramic image. By transforming the unit sphere coordinates to the world coordinates and camera coordinates, the position of the image of each fish-eye camera in the final panoramic image is obtained.

[0145] Specifically, when determining the stitching radius, in order to improve the stitching effect, the present application provides a method for adaptively calculating the stitching radius of adjacent fish-eye images from coarse to fine. Exemplarily, taking any two frames of images I1 and I2 to be stitched as an example, the process of calculating the stitching radius can be as Figure 15 shown. I1

[0146] First, within a certain range, search for different positions of the images to be stitched in the world coordinate system at different radii in a coarse-grained manner, calculate the reprojection error, and select the smallest R as the coarse-grained radius.

[0147] Subsequently, search for different positions of the images to be stitched in the world coordinate system at different radii in a fine-grained manner, calculate the reprojection error, and select the smallest R as the coarse-grained radius.

[0148] Wherein, the difference between the coarse-grained search and the fine-grained search lies in searching with different step sizes. The step size of the coarse-grained search is greater than that of the fine-grained search. Therefore, in the embodiment of the present application, the optimal stitching radius is searched through different granularities, so that a better panoramic image can be stitched.

[0149] Specifically, the process of calculating the reprojection error may include:

[0150] First, according to the center O(x, y) of the unit sphere and the radius R of the panoramic image in the unified coordinate system, the unit sphere is translated to the center 0, and the three-dimensional coordinates at the large sphere are calculated:

[0151] P x = R × P x + C x

[0152] P y = R × P y + C y

[0153] P z = R × P z + C z

[0154] Subsequently, according to the pre-calibrated external parameters of the fisheye camera, including the rotation matrix r and the translation matrix T of the camera, the three-dimensional coordinates where the four cameras are located are converted to the camera coordinate system for imaging:

[0155] _x = r 11 P x + r 12 P y + r 13 P z + T1

[0156] _y = r 21 P x + r 22 P y + r 23 P z + T2

[0157] _z = r 31 P x + r 32 P y + r 33 P z + T3

[0158] Subsequently, referring to the aforementioned method of unfolding the fisheye image, the coordinates corresponding to each pixel point in the fisheye image are projected in the panoramic image, and the stitching result is obtained by using remapping.

[0159] Then, the reprojection error is calculated. For example, assuming that the first image stitched to the panoramic image is I1 and the second image stitched to the panoramic image is I2, then the error is E = (I1 - I2) / M, where M represents the number of pixels in the overlapping area.

[0160] When the reprojection error or the change value of the reprojection error is greater than a certain value, the aforementioned step two can be iteratively executed until the reprojection error is less than a preset value or the reprojection error is less than a preset value, etc., so as to obtain the final panoramic image.

[0161] Exemplarily, as Figure 16 shown, after unfolding and stitching four frames of fisheye images into panoramic images respectively, a final and more accurate panoramic image is obtained.

[0162] Therefore, the method provided by this application can adaptively determine the stitching radius, stitch multiple frames of fisheye images based on the stitching radius, effectively solve the problem of perspective distortion caused by plane image distortion correction, and even if the field of view of the fisheye camera exceeds 180 degrees, it can be stitched based on the matching stitching radius, and can adapt to more application scenarios. Moreover, this application proposes a grid search method from coarse-grained to fine-grained to determine the optimal stitching radius for adjacent cameras, effectively alleviating the parallax problem of non-concentric fisheye stitching and effectively eliminating the ghosting or misalignment problem of non-concentric fisheye stitching. In addition, the application scenario of obtaining panoramic images through multi-channel camera image stitching is wide, such as it can be applied in application scenarios such as vehicle-mounted panoramic imaging, autonomous driving, and assisted driving, with strong generalization ability.

[0163] The above has introduced the process of the image processing method provided by this application in detail. Next, in combination with the above method process, the device for executing this method process will be introduced.

[0164] First, refer to Figure 17 , the structural schematic diagram of an image processing device provided by this application. The image processing device may include:

[0165] An acquisition module 1701, configured to acquire multiple frames of input images;

[0166] A positioning module 1702, configured to determine the position of each frame of input image in the multiple frames of input images in a first coordinate system, where the first coordinate system includes a coordinate system in a spherical space;

[0167] A stitching module 1703, configured to stitch the multiple frames of input images according to the position of each frame of input image in the first coordinate system to obtain a panoramic image.

[0168] In a possible implementation manner, the stitching module 1703 is specifically configured to: map the pixel values of the pixel points in each frame of input image to the corresponding spherical surface in the first coordinate system according to the corresponding stitching radius in each frame of input image, so as to obtain a panoramic image, where the stitching radius is determined according to the position of each frame of input image in the first coordinate system.

[0169] In a possible implementation, the stitching module 1703 is specifically configured to determine the stitching radius corresponding to each input image according to the position of each input image in the first coordinate system at at least two step sizes. Herein, the at least two step sizes can be understood as the search granularity used when searching for the stitching radius.

[0170] In a possible implementation, the stitching module 1703 is specifically configured to: obtain a plurality of coarse-grained radii according to the position of the first image in the first coordinate system at the first step size, where the first image is any one of the multiple input images; calculate the first reprojection error corresponding to the plurality of coarse-grained radii, and the reprojection error is the error between the observed position of the first image in the first coordinate system and the position after being projected onto the first coordinate system according to the coarse-grained radius; screen out the first radius from the plurality of radii according to the reprojection errors corresponding to the plurality of radii; obtain a plurality of fine-grained radii at the second step size, where the second step size is smaller than the first step size; calculate the second reprojection error corresponding to the plurality of fine-grained radii, and the reprojection error is the error between the observed position of the first image in the first coordinate system and the position after being projected onto the first coordinate system according to the fine-grained radius; and screen out the radius of the pixel points in each input image in the first coordinate system according to the reprojection errors corresponding to the plurality of fine-grained radii.

[0171] In a possible implementation, the positioning module 1702 is specifically configured to: determine the position of the second image in the corresponding second coordinate system, where the second coordinate system is the coordinate system corresponding to the camera that captures the second image, and the second image is any one of the multiple input images; and determine the position of the second image in the first coordinate system according to the relative position relationship between the second coordinate system and the first coordinate system.

[0172] In a possible implementation, the foregoing multiple input images are images captured by a fisheye camera.

[0173] In a possible implementation, the foregoing multiple input images are captured by a camera installed in a vehicle, and the panoramic image is used to plan a driving path for the vehicle during autonomous driving.

[0174] Please refer to Figure 18 , the structural schematic diagram of another image processing device provided by this application is as follows.

[0175] The image processing device may include a processor 1801 and a memory 1802. The processor 1801 and the memory 1802 are interconnected by a line. Among them, program instructions and data are stored in the memory 1802.

[0176] The memory 1802 stores the program instructions and data corresponding to the steps in the foregoing Figures 9 - 16 .

[0177] The processor 1801 is configured to execute the method steps performed by the image processing apparatus shown in any of the foregoing Figures 9 - 16 embodiments.

[0178] Optionally, the image processing apparatus may further include a transceiver 1803 configured to receive or transmit data.

[0179] An embodiment of the present application further provides a computer-readable storage medium storing a program for generating a vehicle driving speed. When the program runs on a computer, it causes the computer to execute the steps in the method described in the foregoing Figures 9 - 16 embodiments.

[0180] Optionally, the Figure 18 image processing apparatus shown above is a chip.

[0181] An embodiment of the present application further provides an image processing apparatus, which may also be referred to as a digital processing chip or a chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is configured to execute the method steps performed by the image processing apparatus shown in any of the foregoing Figures 9 - 16 embodiments.

[0182] An embodiment of the present application further provides a digital processing chip. The digital processing chip integrates circuits for implementing the foregoing processor 1801 or the functions of the processor 1801 and one or more interfaces. When the digital processing chip integrates a memory, the digital processing chip can complete the method steps of any one or more of the foregoing embodiments. When the digital processing chip does not integrate a memory, it can be connected to an external memory through a communication interface. The digital processing chip implements the actions performed by the image processing apparatus in the foregoing embodiments according to the program code stored in the external memory.

[0183] An embodiment of the present application further provides a computer program product which, when running on a computer, causes the computer to execute the steps performed by the image processing apparatus in the method described in the foregoing Figures 9 - 16 embodiments.

[0184] The image processing apparatus provided in the embodiments of the present application may be a chip, which includes a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, or a circuit. The processing unit can execute computer-executable instructions stored in a storage unit to cause the chip in the server to execute the foregoing Figures 9 - 16The neural network training method described in the illustrated embodiment. Optionally, the storage unit is a storage unit within the chip, such as a register, cache, etc. The storage unit can also be a storage unit outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0185] Specifically, the aforementioned processing unit or processor can be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0186] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.

[0187] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for this application, software programs are a better implementation method in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of this application.

[0188] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0189] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0190] The terms "first", "second", "third", "fourth", etc. (if any) in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0191] Finally, it should be noted that the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, and all of them should be covered within the protection scope of this application.

Claims

1. An image processing method, characterized in that, Comprising: Obtaining multiple frames of input images; Determining the position of each frame of input image in the multiple frames of input images in a first coordinate system, the first coordinate system including a coordinate system in a spherical space; Stitching the multiple frames of input images according to the positions of each frame of input image in the first coordinate system to obtain a panoramic image; Wherein, the stitching the multiple frames of input images according to the positions of each frame of input image in the first coordinate system to obtain a panoramic image includes: Determining a stitching radius corresponding to each frame of input image according to the position of each frame of input image in the first coordinate system at at least two step lengths, where the at least two step lengths are search granularities used for searching the stitching radius; Mapping the pixel values of the pixel points in each frame of input image to the corresponding spherical surface in the first coordinate system according to the corresponding stitching radius in each frame of input image to obtain the panoramic image.

2. The method according to claim 1, characterized in that, The determining a stitching radius corresponding to each frame of input image according to the position of each frame of input image in the first coordinate system at at least two step lengths includes: Searching for a plurality of coarse-grained radii according to the position of a first image in the first coordinate system at a first step length, the first image being any one of the multiple frames of input images; Calculating a first reprojection error corresponding to the plurality of coarse-grained radii, the first reprojection error being the error between the observed position of the first image in the first coordinate system and the position after projection to the first coordinate system according to the coarse-grained radius; Selecting a first radius from the plurality of radii according to the reprojection errors corresponding to the plurality of radii; Obtaining a plurality of fine-grained radii at a second step length, the second step length being less than the first step length; Calculating a second reprojection error corresponding to the plurality of fine-grained radii, the second reprojection error being the error between the observed position of the first image in the first coordinate system and the position after projection to the first coordinate system according to the fine-grained radius; Obtaining the radius of the pixel points in each frame of input image in the first coordinate system according to the second reprojection error corresponding to the plurality of fine-grained radii.

3. The method according to claim 1 or 2, characterized in that, The determining the position of each frame of input image in the multiple frames of input images in the first coordinate system includes: Determining the position of a second image in a corresponding second coordinate system, the second coordinate system being the coordinate system corresponding to the camera that captures the second image, the second image being any one of the multiple frames of input images; Determining the position of the second image in the first coordinate system according to the relative position relationship between the second coordinate system and the first coordinate system.

4. The method according to claim 1 or 2, characterized in that The first coordinate system includes a coordinate system in an ellipsoidal spherical space.

5. The method according to claim 1 or 2, characterized in that, The multiple frames of input images are images captured by a fish-eye camera.

6. The method according to claim 1 or 2, wherein The multiple frames of input images are captured by a camera provided in a vehicle, and the panoramic image is used for autonomous driving of the vehicle.

7. An image processing apparatus, characterized in that, Comprising: An obtaining module, configured to obtain multiple frames of input images; A positioning module, configured to determine the position of each input image in the multi-frame input images in a first coordinate system, where the first coordinate system includes a coordinate system in a spherical space; A stitching module, configured to stitch the multi-frame input images according to the positions of each input image in the first coordinate system to obtain a panoramic image; The stitching module is specifically configured to: Determine the stitching radius corresponding to each input image according to the position of each input image in the first coordinate system at at least two step sizes, where the at least two step sizes are search granularities used for searching the stitching radius; Map the pixel values of the pixel points in each input image to the corresponding spherical surface in the first coordinate system according to the corresponding stitching radius in each input image to obtain the panoramic image.

8. The device according to claim 7, wherein The stitching module is specifically configured to: Search for a plurality of coarse-grained radii according to the position of the first image in the first coordinate system at a first step size, where the first image is any one of the multi-frame input images; Calculate a first reprojection error corresponding to the plurality of coarse-grained radii, where the first reprojection error is the error between the observed position of the first image in the first coordinate system and the position after being projected to the first coordinate system according to the coarse-grained radius; Screen out a first radius from the plurality of radii according to the reprojection errors corresponding to the plurality of radii; Obtain a plurality of fine-grained radii at a second step size, where the second step size is smaller than the first step size; Calculate a second reprojection error corresponding to the plurality of fine-grained radii, where the second reprojection error is the error between the observed position of the first image in the first coordinate system and the position after being projected to the first coordinate system according to the fine-grained radius; Obtain the radius of the pixel points in each input image in the first coordinate system according to the second reprojection error corresponding to the plurality of fine-grained radii.

9. The device according to claim 7 or 8, characterized in that The positioning module is specifically configured to: Determine the position of a second image in a corresponding second coordinate system, where the second coordinate system is the coordinate system corresponding to the camera that captures the second image, and the second image is any one of the multi-frame input images; Determine the position of the second image in the first coordinate system according to the relative position relationship between the second coordinate system and the first coordinate system.

10. The device according to claim 7 or 8, characterized in that, The first coordinate system includes a coordinate system in an ellipsoidal spherical space.

11. The device according to claim 7 or 8, characterized in that, The multi-frame input images are images collected by a fish-eye camera.

12. The device according to claim 7 or 8, wherein The multi-frame input images are captured by a camera installed in a vehicle, and the panoramic image is used for autonomous driving of the vehicle.

13. An image processing apparatus, characterized in that, Comprising one or more processors, the one or more processors are coupled to a memory, and the memory stores a program. When the program instructions stored in the memory are executed by the one or more processors, the steps of the method according to any one of claims 1 to 6 are implemented.

14. A computer-readable storage medium, characterized in that, Comprising a program, when it is executed by a processing unit, it executes the method according to any one of claims 1 to 6.

15. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

16. A chip, characterized in that, The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is configured to execute the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Panoramic image splicing rendering method and device

    CN110728619A

  • Panoramic shooting method, electronic equipment and storage medium

    CN113454980A

  • Method of Adaptive Image Stitching and Image Processing Device

    US20200104977A1