Image processing method and corresponding device
By using raster images to match the camera images in traffic scenes, determining the camera position and realizing image stitching, the complex problems of multi-camera image management and stitching are solved, and the stitching efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202311557960.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-05-23
AI Technical Summary
In traffic scenes, storyboard images captured by multiple cameras are difficult to manage and splice at the same time, resulting in complex data management and the inability to accurately judge the correlation between storyboards.
Image stitching is achieved by acquiring images captured by multiple cameras and matching features with pre-generated raster images. The first pose of at least two cameras is determined.
It improves the efficiency and accuracy of image stitching, reduces the use of storage space, and realizes fast matching and optimization of camera poses.
Smart Images

Figure CN120034637A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and in particular to an image processing method and a corresponding device. Background Art
[0002] In traffic scenes, multiple cameras are usually deployed at intersections / highways to capture images at different locations on the intersections / highways, thereby providing data support for event prediction and signal optimization at key intersections. In the traditional mode, multiple cameras have multiple split-lens images, and it is difficult for the data management center to take into account all split-lens data at the same time; at the same time, the camera perspectives are scattered, making it difficult to accurately judge the relationship between the split-lens and determine the specific location of the incident. The key technology to solve these problems is to stitch images taken by different cameras.
[0003] There are two main methods for image stitching at present. One is the manual method, that is, the key points of the image to be matched are selected and matched by manual point selection. The process is time-consuming and laborious. Every time the camera rotates or moves, manual intervention is required. The other is the automated method, that is, based on high-precision maps or three-dimensional maps, the position of the camera is determined by matching the image to be matched with the high-precision map or three-dimensional map, and then stitching and fusion are performed. However, high-precision maps and three-dimensional maps have large data volumes, low stitching efficiency, and poor matching effects in complex scenes. Summary of the invention
[0004] The present application provides an image processing method for improving the efficiency and accuracy of image stitching. The present application also provides a corresponding device, system, computer-readable storage medium, computer program product, etc.
[0005] A first aspect of the present application provides an image processing method, comprising: acquiring multiple images taken by multiple cameras; determining the first position of at least two cameras among the multiple cameras based on the multiple images and the multiple raster images, wherein the multiple images and the multiple raster images are associated with the same target scene, wherein the first position of the first camera is related to the position of a virtual camera corresponding to the first raster image, the first raster image is a raster image among the multiple raster images corresponding to the image taken by the first camera, and the first camera is any one of the multiple cameras; and stitching the multiple images according to the first position of the at least two cameras.
[0006] The image processing method in the present application can be applied to traffic scenes, such as intersections / highways. The process can be executed by a cloud device or a road-side device.
[0007] In the present application, the target scene may be an intersection or a section of road. A target scene is usually configured with multiple cameras (also called cameras), wherein each camera captures images of the target scene in different directions and positions.
[0008] In the present application, the raster image may be pre-generated, and the raster image may be obtained by rendering the three-dimensional model of the target scene according to the position and posture of the virtual camera. Each raster image will correspond to the position and posture of a virtual camera, and the image features under the position and posture of the virtual camera. The raster image, the position and posture of the virtual camera, and the image features can be understood as a set of data with a corresponding relationship. In the present application, when generating a raster image, the lane lines or other features with identification functions in the three-dimensional model can be rendered in particular, and other features can be ignored or simply rendered. In this way, a raster image with a smaller amount of data can be obtained, and such a raster image takes up less storage space.
[0009] In this application, the pose of a camera refers to the position and attitude of the camera, which may include the position (x, y, z) in a three-dimensional coordinate system, as well as the azimuth, pitch, and roll angles. Usually, the pose of a camera is described by the rotation matrix (R) and translation vector (t) of the camera, rather than by the above position coordinates and angle coordinates. The pose of a virtual camera refers to the pose assumed when generating a raster image.
[0010] In the present application, by matching the characteristic features of the image captured by the camera with the raster image, the first position of the camera can be obtained through the position of the virtual camera corresponding to the raster image. The raster image corresponding to the image captured by the camera can be the raster image with the highest similarity to the image captured by the camera among multiple raster images.
[0011] In the first aspect, since the data volume of the raster image is small, the storage space occupied can be reduced. In addition, the raster image mainly includes lane line features, building features, etc., and other features are relatively few, so it can achieve rapid matching with the image taken by the camera, thereby improving the efficiency and accuracy of image stitching.
[0012] In a possible implementation, the above steps: determining the first pose of at least two cameras among the multiple cameras based on the multiple images and the multiple raster images, including: matching the multiple images with the corresponding raster images for marker features to estimate the first pose of each camera among the multiple cameras; correspondingly, splicing the multiple images according to the first poses of the at least two cameras, including: splicing the multiple images according to the estimated first pose of each camera.
[0013] In this possible implementation, the marker feature may include at least one of a lane feature, a building feature, a vehicle feature, a pedestrian feature, and a guardrail feature. By matching the features of the marker features, the speed of image matching can be improved. In addition, by stitching multiple images according to the first position of each camera, the accuracy of image stitching can be improved.
[0014] In a possible implementation, the above steps: matching the multiple images with the corresponding raster images for marker features respectively to estimate the first pose of each camera in the multiple cameras, include: matching the first lane line feature associated with the first raster image with the second lane line feature in the image of the first camera; adjusting the pose of the virtual camera corresponding to the first raster image according to the optimization target to obtain the first pose of the first camera, the optimization target is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization target is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
[0015] In this possible implementation, when determining the first pose of the first camera, the residual between the image of the first camera and the lane line features in the first grid image can be used as the basis for back propagation optimization based on the pose of the virtual camera corresponding to the first grid image, so that the lane line features between the image of the first camera and the first grid image are as similar as possible, or the distance between the lane line features in space is as close as possible, and the output pose is the first pose of the first camera. In the present application, the pose of the virtual camera corresponding to the first grid image is optimized according to the optimization target to obtain the pose of the first camera, which can improve the accuracy of the first pose.
[0016] In a possible implementation, the above step: stitching multiple images according to the estimated first pose of each camera, includes: optimizing the estimated first pose of each camera to obtain the second pose of each camera; stitching multiple images according to the second pose of each camera.
[0017] In this possible implementation, the first pose may be optimized to obtain a second pose with higher accuracy, thereby improving the accuracy of image stitching.
[0018] In one possible implementation, the above steps: optimizing the estimated first pose of each camera to obtain the second pose of each camera, including: generating a rendered image of the target scene at the corresponding first pose according to the first pose of each camera; matching the lane line features in the image of each camera with the lane line features in the corresponding rendered image to obtain the second pose of each camera.
[0019] In this possible implementation, the rendered image generated according to the first pose of each camera is closer to the image captured by the camera than the raster image. Therefore, the second pose obtained using the rendered image is more accurate, which can further improve the accuracy of image stitching.
[0020] In one possible implementation, the above step of: matching the lane line features in the image of each camera with the lane line features in the corresponding rendered image includes: using the feature points of the lane lines in the image of each camera as guide points, determining the feature points of the lane lines corresponding to the guide points in the corresponding rendered image, and matching them.
[0021] In this possible implementation, it is only necessary to determine the feature points of the lane line in the camera image, and then use these feature points as guide points to match the corresponding feature points in the corresponding rendered image, and then solve them through the (perspective-n-point, PnP) algorithm to obtain the second pose. In this application, global matching is not required, only local matching of the feature points of the lane line is required, which can reduce the amount of calculation and improve matching efficiency.
[0022] In a possible implementation, stitching a plurality of images according to the second posture of each camera includes: determining a stitching relationship between the plurality of images according to the second posture of each camera; and stitching the plurality of images according to the stitching relationship.
[0023] In this possible implementation, the stitching relationship can be a mapping relationship between the camera identifier and the stitching area identifier. One camera can correspond to a stitching area on the display area. By placing the camera image in the corresponding stitching area, image stitching can be achieved, thereby improving the efficiency of image stitching.
[0024] In one possible implementation, the mapping relationship between the identification of the above-mentioned camera and the identification of the stitching area can be: through the second posture of each camera and the distortion parameter of each camera, determine the camera to which the 3D point corresponding to each pixel on the display area in the space of the target scene is projected, and then, according to the stitching area to which each pixel point belongs, determine the mapping relationship between the camera and the stitching area.
[0025] In a possible implementation, the above-mentioned process of image stitching using a mapping relationship may be: when the 3D point corresponding to the target pixel in the space of the target scene is projected to one camera, the pixel information corresponding to the target pixel in the target image of one camera is rendered to the target pixel; when the 3D point corresponding to the target pixel in the space of the target scene is projected to at least two cameras, the pixel information corresponding to the target pixel in the target image of the target camera is rendered to the target pixel, and the target camera is the camera with the smallest distance to the 3D point corresponding to the target pixel in the space of the target scene among the at least two cameras. In this way, duplication can be removed and the image stitching quality can be improved.
[0026] In a possible implementation manner, the method further includes: adjusting the first camera according to the first pose of the first camera.
[0027] In this possible implementation, after the first pose of the first camera is determined, the first camera may be controlled to move or rotate so that the first camera reaches the position and posture indicated by the first pose.
[0028] In one possible implementation, the image from the first camera is an image taken by the first camera at a first moment, and the method further includes: acquiring an image taken by the first camera at a second moment, the second moment being later than the first moment; performing lane line feature matching on the image taken by the first camera at the second moment and the second grid image to estimate the position and posture of the first camera at the second moment, the position and posture of the first camera at the second moment being different from the first position of the first camera to indicate that the first camera has rotated or moved.
[0029] In this possible implementation, after the first camera is rotated or moved, no human intervention is required, and the data after the movement can still be used to accurately stitch images.
[0030] A second aspect of the present application provides an image processing method, comprising: estimating a first pose of the first camera through an image from the first camera; generating a rendered image of a target scene at the first pose, where the target scene is a scene associated with the first camera; and matching the image from the first camera with the rendered image to optimize the first pose to obtain a second pose of the first camera.
[0031] In the second aspect, the accuracy of the posture can be improved by optimizing the first posture to obtain the second posture.
[0032] In one possible implementation, the image from the first camera is included in multiple images, each of the multiple images comes from a different camera associated with the target scene, the first camera is any one of the multiple cameras associated with the target scene, and the method further includes: stitching the multiple images according to the second posture of each camera.
[0033] In a possible implementation, the above step: estimating the first pose of the first camera through the image from the first camera, includes: performing lane line feature matching on the image from the first camera and the corresponding first grid image to estimate the first pose of the first camera; wherein the first pose of the first camera is related to the pose of the virtual camera corresponding to the first grid image, and the first grid image is a grid image among multiple grid images corresponding to the image taken by the first camera.
[0034] In a possible implementation, the marker feature includes at least one of a lane feature, a building feature, a vehicle feature, a pedestrian feature, and a guardrail feature.
[0035] In one possible implementation, the above steps: matching the image from the first camera with the corresponding first grid image for marker features to estimate the first pose of the first camera, include: matching the first lane line feature associated with the first grid image with the second lane line feature in the image of the first camera for lane line features; adjusting the pose of the virtual camera corresponding to the first grid image according to an optimization target to obtain the first pose of the first camera, wherein the optimization target is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization target is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
[0036] In one possible implementation, the above step of: matching the image from the first camera with the rendered image includes: using the feature points of the lane lines in the image of the first camera as guide points, determining the feature points of the lane lines corresponding to the guide points in the corresponding rendered image, and matching them.
[0037] In a possible implementation, stitching a plurality of images according to the second posture of each camera includes: determining a stitching relationship between the plurality of images according to the second posture of each camera; and stitching the plurality of images according to the stitching relationship.
[0038] In a possible implementation manner, the method further includes: adjusting the first camera according to the second posture of the first camera.
[0039] In one possible implementation, the image from the first camera is an image taken by the first camera at a first moment, and the method further includes: acquiring an image taken by the first camera at a second moment, the second moment being later than the first moment; performing lane line feature matching on the image taken by the first camera at the second moment and the second grid image to estimate the posture of the first camera at the second moment, the posture of the first camera at the second moment being different from the first posture of the first camera to indicate that the first camera has rotated or moved.
[0040] The third aspect of the present application provides an image processing device, which has the function of implementing the image processing method of the first aspect or any possible implementation of the first aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions, such as: an acquisition unit, a first processing unit, and a second processing unit.
[0041] In a fourth aspect, the present application provides an image processing device, which has the function of implementing the image processing method of the second aspect or any possible implementation of the second aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions, such as: a first processing unit, a second processing unit, and a third processing unit.
[0042] In a fifth aspect, the present application provides an electronic device, comprising a communication interface, a processor and a memory, wherein the communication interface and the processor are coupled to the memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the electronic device executes a method as in the first aspect or any possible implementation method of the first aspect.
[0043] In the present application, the processor may include at least one of a central processing unit (CPU) and a graphics processing unit (GPU); wherein both the CPU and the GPU may execute the image processing process described in the above-mentioned first aspect or any possible implementation of the first aspect, or the CPU and the GPU cooperate to execute the image processing process described in the above-mentioned first aspect or any possible implementation of the first aspect.
[0044] In a sixth aspect of the present application, there is provided an electronic device, comprising a communication interface, a processor and a memory, wherein the communication interface and the processor are coupled to the memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the electronic device executes a method as in the second aspect or any possible implementation method of the second aspect.
[0045] In the present application, the processor may include at least one of a central processing unit (CPU) and a graphics processing unit (GPU); wherein both the CPU and the GPU may execute the image processing process described in the above-mentioned second aspect or any possible implementation of the second aspect, or the CPU and the GPU cooperate to execute the image processing process described in the above-mentioned second aspect or any possible implementation of the second aspect.
[0046] In the seventh aspect of the present application, a chip system is provided, which includes one or more interface circuits and one or more processors; the interface circuit and the processor are interconnected by lines; the interface circuit is used to receive signals from the memory of the electronic device and send signals to the processor, and the signals include computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the method as described in the first aspect or any possible implementation of the first aspect, and the processor is at least one of a CPU and a GPU.
[0047] In an eighth aspect of the present application, a chip system is provided, which includes one or more interface circuits and one or more processors; the interface circuit and the processor are interconnected by lines; the interface circuit is used to receive signals from the memory of the electronic device and send signals to the processor, and the signals include computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the method as described in the second aspect or any possible implementation of the second aspect, and the processor is at least one of a CPU and a GPU.
[0048] A ninth aspect of the present application provides a computer-readable storage medium, which stores instructions. When the instructions are executed on an electronic device, the electronic device executes a method as described in the first aspect or any possible implementation of the first aspect.
[0049] The tenth aspect of the present application provides a computer-readable storage medium, which stores instructions. When the instructions are executed on an electronic device, the electronic device executes a method as described in the second aspect or any possible implementation of the second aspect.
[0050] In an eleventh aspect of the present application, a computer program product is provided, which includes a computer program code. When the computer program code runs on a computer, the computer executes the method of the first aspect or any possible implementation of the first aspect.
[0051] The twelfth aspect of the present application provides a computer program product, which includes a computer program code. When the computer program code runs on a computer, the computer executes the method of the second aspect or any possible implementation of the second aspect.
[0052] The thirteenth aspect of the present application provides an image processing system, including: an electronic device, the electronic device is used to execute the method of the above-mentioned first aspect or any possible implementation method of the first aspect, and the electronic device is a cloud-side device or a road-side device.
[0053] A fourteenth aspect of the present application provides an image processing system, including: an electronic device, the electronic device is used to execute the method of the above-mentioned second aspect or any possible implementation method of the second aspect, and the electronic device is a cloud-side device or a road-side device.
[0054] The relevant features and effects of the second aspect of the present application, as well as any possible implementation of the second aspect to the fourteenth aspect, can be understood by referring to the corresponding introduction in the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1A This is a schematic diagram of a scenario architecture provided by an embodiment of the present application;
[0056] Figure 1B This is another schematic diagram of a scenario architecture provided by an embodiment of the present application;
[0057] Figure 1C This is another schematic diagram of a scenario architecture provided by an embodiment of the present application;
[0058] Figure 1D is a structural schematic diagram of an electronic device provided in an embodiment of the present application;
[0059] Figure 2 It is a structural diagram of a cloud device / road-end device provided in an embodiment of the present application;
[0060] Figure 3 It is a schematic diagram of a process for generating a raster image provided by an embodiment of the present application;
[0061] Figure 4 This is a schematic diagram of a scenario example provided in an embodiment of the present application;
[0062] Figure 5 is a schematic diagram of an embodiment of the image processing method provided by an embodiment of the present application;
[0063] Figure 6 is an exemplary schematic diagram of an image processing method provided in an embodiment of the present application;
[0064] Figure 7 is another exemplary schematic diagram of the image processing method provided in an embodiment of the present application;
[0065] Figure 8 is another exemplary schematic diagram of the image processing method provided in an embodiment of the present application;
[0066] Fig. 9 is another exemplary schematic diagram of the image processing method provided in an embodiment of the present application;
[0067] Fig.10is another exemplary schematic diagram of the image processing method provided in an embodiment of the present application;
[0068] Fig.11 is a structural schematic diagram of an image processing device provided in an embodiment of the present application;
[0069] Fig.12 It is another structural schematic diagram of the image processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0070] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of the present application, rather than all embodiments. It is known to those skilled in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0071] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0072] The embodiment of the present application provides an image processing method for improving the efficiency and accuracy of image stitching. The present application also provides corresponding devices, systems, computer-readable storage media, and computer program products, etc. The following are detailed descriptions.
[0073] The solution provided by the embodiment of the present application can be applied to smart transportation, such as: at an intersection or on a road, a camera (also called a camera) can be used to shoot the traffic conditions at the intersection or on the road, and then the camera transmits the captured image to the cloud, and the cloud-side device on the cloud stitches the images captured by multiple cameras at an intersection, and then sends the stitched image to the display of the road management center for display. Alternatively, the camera transmits the captured image to the roadside device, and the roadside device stitches the images captured by multiple cameras at an intersection, and then sends the stitched image to the display of the road management center for display.
[0074] For cloud-based processing, see Figure 1A To understand, such as Figure 1AAs shown, taking an intersection as an example, the intersection has four cameras, each camera can capture the traffic conditions within a certain range of the intersection, and then each camera transmits the image to the cloud, such as: the images transmitted to the cloud by the four cameras are image 1, image 2, image 3 and image 4 respectively, and the cloud-side device on the cloud will stitch the four images together according to their postures to obtain the stitched image of the intersection, and then send the stitched image of the intersection to the display of the road surface management center for display.
[0075] For situations handled by the roadside, please refer to Figure 1B To understand, such as Figure 1B As shown, taking an intersection as an example, the intersection has four cameras, each camera can capture the traffic conditions within a certain range of the intersection, and then each camera transmits the image to the road-side device, such as: the images transmitted by the four cameras to the road-side device are image 1, image 2, image 3 and image 4 respectively. The road-side device will stitch the four images together according to their postures to obtain the stitched image of the intersection, and then send the stitched image of the intersection to the display of the road surface management center for display.
[0076] The road-end device may be a terminal device arranged at the intersection, and the terminal device may be a camera with an image processing function, or may be a device specifically responsible for image processing.
[0077] Above Figure 1A The cloud-side device in the mid-cloud can be a working node or a scheduling node in the cloud system. After the scheduling node receives the image from the camera, the scheduling node can perform the corresponding image processing process. The scheduling node can also schedule the images of the four cameras to a working node in the cloud system, and the working node will perform the corresponding image processing process.
[0078] The function of the scheduling node can be implemented by software or hardware.
[0079] As an example of a software functional unit, the scheduling node may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the above computing instance may be one or more. For example, the scheduling node may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple data centers with close geographical locations. Among them, usually a region may include multiple AZs.
[0080] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in one region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.
[0081] As an example of a hardware functional unit, the scheduling node may include at least one computing device, such as a server, etc. Alternatively, the scheduling node may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0082] The multiple computing devices included in the scheduling node can be distributed in the same area or in different areas. The multiple computing devices included in the scheduling node can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the scheduling node can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0083] A working node can be a physical machine, or a computing instance such as a virtual machine (VM) or a container. A working node can include one or more central processing units (CPU) and graphics processing units (GPU), etc. A working node can also be a CPU or a GPU.
[0084] Above Figure 1A The cloud end can be located in the cloud system. The architecture of the cloud system can be found in Figure 1C Understand. Figure 1C As shown in Figure 1, the cloud system includes a cloud platform and basic resources. The cloud platform includes a cloud platform manager. The scheduling nodes introduced above can be Figure 1C The basic resources may include multiple servers, each of which may be a working node, or each server may include multiple working nodes.
[0085] exist Figure 1C The working node in the system may be a computing device card or a virtual machine (VM). The computing device card may be at least one of a central processing unit (CPU), a graphics processing unit (GPU) and a neural network processing unit (NPU).
[0086] The cloud platform manager will maintain or regularly collect information about each work node in the basic resources, such as the resource usage on each work node (resource usage rate or resource idle rate), etc. This information can be used as auxiliary decision-making information when scheduling images.
[0087] The cloud platform manager can receive images from the camera, and then the cloud platform manager can perform corresponding image processing. The cloud platform manager can also Figure 1AThe images of the four cameras are dispatched to a working node, which performs the corresponding image processing process.
[0088] Above Figures 1A to 1C In the scenario of the image processing system introduced, whether the image processing is performed in the cloud, on the road side, or through the combination of end and cloud, the image processing system introduced above can include an electronic device, which can be a cloud-side device or a road-side device. When the electronic device is a cloud-side device, after obtaining the image taken by the camera from the road side, the subsequent image processing process can be completed. When the electronic device is a road-side device, when the road-side device is a certain camera, after the road-side device obtains the image from other cameras, it can perform subsequent image processing on the images of other cameras and the images taken by itself. If the road-side device is a dedicated image processing device, after the road-side device obtains the image taken by the camera, the subsequent image processing process can be completed.
[0089] The structure of the electronic device can participate Figure 1D It is understood that in one possible embodiment, Figure 1D As shown, the electronic device 1000 may include: a central processing unit 1001, a graphics processor 1002, a display device 1003 (may not be included) and a memory 1004. Optionally, the electronic device 1000 may also include at least one communication bus ( Figure 1D ), which is not shown in the figure, is used to realize the connection and communication between various components.
[0090] It should be understood that the various components in the electronic device 1000 may also be coupled through other connectors, and the other connectors may include various interfaces, transmission lines or buses, etc. The various components in the electronic device 1000 may also be connected in a radial manner centered on the central processor 1001. In various embodiments of the present application, coupling refers to being electrically connected or connected to each other, including being directly connected or indirectly connected through other devices.
[0091] There are many ways to connect the CPU 1001 and the GPU 1002, and they are not limited to Figure 1D The central processing unit 1001 and the graphics processing unit 1002 in the electronic device 1000 may be located on the same chip or may be separate chips.
[0092] The functions of the central processing unit 1001 , the graphics processing unit 1002 , the display device 1003 and the memory 1004 are briefly introduced below.
[0093] Central processing unit 1001: used to run operating system 1005 and application 1006. Application 1006 can be a graphics application, such as a video player, etc. Operating system 1005 provides a system graphics library interface, and application 1006 generates an instruction stream for rendering graphics or image frames and the required related rendering data through the system graphics library interface and the driver provided by operating system 1005, such as graphics library user mode driver and / or graphics library kernel mode driver. Among them, the system graphics library includes but is not limited to: open graphics library for embedded system (OpenGL ES), Khronos platform graphics interface or Vulkan (a cross-platform drawing application interface) and other system graphics libraries. The instruction stream contains a series of instructions, which are usually call instructions for the system graphics library interface.
[0094] Optionally, the central processing unit 1001 may include at least one of the following types of processors: an application processor, one or more microprocessors, a digital signal processor (DSP), a microcontroller unit (MCU), or an artificial intelligence processor, etc.
[0095] The CPU 1001 may further include necessary hardware accelerators, such as application specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or integrated circuits for implementing logic operations. The processor 1001 may be coupled to one or more data buses for transmitting data and instructions between various components of the electronic device 1000.
[0096] Graphics processor 1002: used to receive the graphics instruction stream sent by processor 1001, generate a rendering target through a rendering pipeline, and display the rendering target to display device 1003 through the layer synthesis display module of the operating system. The rendering pipeline can also be called a rendering pipeline, a pixel pipeline or a pixel pipeline, which is a parallel processing unit inside the graphics processor 1002 for processing graphics signals. The graphics processor 1002 may include multiple rendering pipelines, and multiple rendering pipelines can process graphics signals in parallel independently of each other. For example, the rendering pipeline can perform a series of operations in the process of rendering graphics or image frames, and typical operations may include: vertex processing, primitive processing, rasterization, fragment processing, etc.
[0097] Alternatively, the graphics processor 1002 may include a general-purpose graphics processor that executes software, such as a GPU or other types of dedicated graphics processing units.
[0098] Display device 1003: used to display various images generated by electronic device 1000, which may be a graphical user interface (GUI) of an operating system or image data (including still images and video data) processed by graphics processor 1002.
[0099] Optionally, the display device 1003 may include any suitable type of display screen, such as a liquid crystal display (LCD), a plasma display, or an organic light-emitting diode (OLED) display.
[0100] The memory 1004 is a transmission channel between the CPU 1001 and the GPU 1002 and may be a double data rate synchronous dynamic random access memory (DDR SDRAM) or other types of cache.
[0101] The structure of the cloud device or roadside device mentioned above can also be found in Figure 2 To understand. Figure 2 A possible logical structure diagram of a cloud device / roadside device provided in an embodiment of the present application. Figure 2As shown, the cloud device / roadside device 20 provided in the embodiment of the present application includes: a processor 201, a communication interface 202, a memory 203 and a bus 204. The processor 201, the communication interface 202 and the memory 203 are interconnected via the bus 204. In the embodiment of the present application, the processor 201 is used to control and manage the actions of the cloud device / roadside device 20. For example, the processor 201 is used to stitch images from multiple cameras to obtain a stitched image. The communication interface 202 is used to support the cloud device 20 to communicate. For example, the communication interface 202 can receive images transmitted from the camera and send stitched images. The memory 203 is used to store program codes and data of the cloud device / roadside device 20.
[0102] Among them, the processor 201 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the contents disclosed in this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 204 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 2 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0103] The following is a description of the image processing method provided in the embodiment of the present application. The content of the method involving cloud execution can be executed by the cloud, or by a component of the cloud (such as a processor, chip, or chip system, etc.). The content involving the execution of the roadside device can be executed by the roadside device, or by a component of the roadside device (such as a processor, chip, or chip system, etc.).
[0104] The image processing method provided in the embodiment of the present application uses a raster image, which can be generated in the cloud or on an independent server. Figure 3 To understand, the process includes:
[0105] 301. Collect data of the target scene and construct a three-dimensional model of the target scene.
[0106] The target scene data may include an image of the target scene captured by a panoramic camera.
[0107] 302. Render the three-dimensional model under different positions of the virtual camera to obtain multiple raster images.
[0108] The raster image includes marker features, and the marker features include at least one of lane line features, building features, vehicle features, pedestrian features, and guardrail features.
[0109] 303. Extract the features of the markers in each raster image.
[0110] At least one of lane line features, building features, vehicle features, pedestrian features, and guardrail features of various shapes can be extracted from the raster image through a deep learning algorithm.
[0111] 304. Construct the correspondence between each raster image, the corresponding marker feature and the position and posture of the virtual camera.
[0112] This correspondence can be found in Figure 4 To understand. Figure 4 In the example, the feature of the marker is the lane line feature, such as Figure 4 As shown, there are three virtual cameras, each of which renders three raster images at different positions. By extracting the lane line features in the raster images, the symbols of the lane line features can be extracted.
[0113] In this application, the raster image, the corresponding lane line features and the virtual camera's position and posture can also be referred to as a feature map. Taking a 500-meter intersection as an example, the original point cloud or high-precision map of this 500-meter intersection will occupy several GB of storage space. The feature map of this application only needs to occupy about 10MB of storage space, which greatly reduces the storage space occupied and enables the image processing process to be performed on the roadside.
[0114] It should be noted that, in the present application, the above steps 301 to 304 only need to be executed once at the beginning. When the camera rotates or the position changes, there is no need to re-execute, and the raster image generated for the first time can still be used.
[0115] The correspondence between the raster image, the corresponding lane line features and the virtual camera's position and posture generated by the embodiment of the present application, or simply the feature map, can be stored in the cloud or in a roadside device. If the feature map is generated in the cloud, it can also be sent to the roadside device for use.
[0116] The following describes the image processing method provided in the embodiment of the present application in conjunction with the accompanying drawings. The process can be executed by a cloud device or a road-side device.
[0117] like Figure 5 As shown, an embodiment of the image processing method provided in the embodiment of the present application includes:
[0118] 501. Obtain multiple images taken by multiple cameras associated with the target scene.
[0119] The target scene may be an intersection or a section of road. A target scene is usually configured with multiple cameras (also called cameras), wherein each camera acquires images of the target scene in different directions and positions by shooting.
[0120] Each of the multiple images can be understood as an image taken by different cameras at the same time point.
[0121] 502. Determine the first position of at least two cameras among the multiple cameras according to the multiple images and the multiple raster images.
[0122] Multiple images and multiple raster images are associated with the same target scene, wherein the first position of the first camera is related to the position of the virtual camera corresponding to the first raster image, the first raster image is a raster image among the multiple raster images corresponding to the image taken by the first camera, and the first camera is any one of the multiple cameras.
[0123] Optionally, the process of determining the first pose in step 502 may be: performing marker feature matching on the multiple images and the corresponding grid images respectively to estimate the first pose of each camera in the multiple cameras.
[0124] The pose of a camera refers to the position and attitude of the camera, which can include the position (x, y, z) in a three-dimensional coordinate system, as well as the azimuth, pitch, and roll angles. Usually, the pose of a camera is described by the rotation matrix (R) and translation vector (t) of the camera, rather than by the above position coordinates and angle coordinates. The pose of a virtual camera refers to the pose assumed when generating a raster image.
[0125] In step 502, because the raster image and the image taken by the camera are from the same target scene, the image taken by the camera can be matched with a raster image with the highest similarity from multiple raster images, and then the first pose of the camera is determined by the raster image.
[0126] Taking any one of the multiple cameras as an example, call this camera the first camera, and match the image taken by the first camera with multiple raster images, it can be determined that the first raster image has the highest similarity with the image taken by the first camera, then the first position of the first camera can be determined by the position of the virtual camera corresponding to the first raster image.
[0127] The process of determining the first raster image from multiple raster images can be: taking the marker feature as a lane line feature as an example, the lane line feature in the image of the first camera is matched with the lane line feature in each raster image for similarity, thereby determining that the first similarity between the first lane line feature in the first raster image and the second lane line feature in the image of the first camera is higher than the second similarity; wherein the second similarity is the similarity between the second lane line feature and the third lane line feature, and the third lane line feature is the lane line feature in the raster image other than the first raster image among the multiple raster images.
[0128] The process of determining the first pose of the first camera through the pose of the virtual camera corresponding to the first raster image can be: adjusting the pose of the virtual camera corresponding to the first raster image according to the optimization target to obtain the first pose of the first camera, and the optimization target is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization target is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
[0129] The process may be: through the residual between the lane line features in the image of the first camera and the first grid image, based on the pose of the virtual camera corresponding to the first grid image, back propagation optimization is performed to make the lane line features between the image of the first camera and the first grid image as similar as possible, or to make the distance between the lane line features in space as close as possible. At this time, the output pose of the first camera is the first pose of the first camera. This process can be referred to Figure 6 To understand, such as Figure 6 As shown in FIG. 1 , the first grid image 601 and the image 602 taken by the first camera are processed by a convolutional neural network (CNN) to extract the lane line features in the image, and then the position (R r , t r ) is substituted into the residual formula below as the initial value, and the first pose (R, t) of the first camera is obtained by back propagation optimization (Backward).
[0130]
[0131] Among them, residual represents the residual, R r represents the rotation matrix of the virtual camera corresponding to the first raster image 601, t r represents the translation vector of the virtual camera corresponding to the first raster image 601, F q represents the lane line features of the image of the first camera, F r represents the lane line feature corresponding to the first raster image, s i Indicates the scaling factor.
[0132] 503. Stitching the multiple images according to the estimated first poses of at least two cameras.
[0133] The process can be to stitch multiple images using the first pose of each camera, or to stitch multiple images using the first pose of some cameras. For example, stitch four images taken by four cameras using the first pose of three cameras.
[0134] In the embodiment of the present application, the image taken by the camera is matched with the raster image for lane line features, and the first position of the camera can be obtained through the position of the virtual camera corresponding to the raster image, thereby realizing the splicing of multiple images taken by multiple cameras. Because the amount of data of the raster image is very small, the storage space occupied can be reduced. In addition, the raster image mainly includes lane line features and other features are relatively few, so it can achieve fast matching with the image taken by the camera, thereby improving the efficiency and accuracy of image splicing.
[0135] Optionally, in the embodiment of the present application, the first pose can also be optimized to obtain a second pose with higher accuracy, and then multiple images can be spliced according to the second pose of each camera.
[0136] The process of optimizing the first pose to obtain the second pose may include: generating a rendered image of the target scene at the corresponding first pose according to the first pose of each camera; matching the lane line features in the image of each camera with the lane line features in the corresponding rendered image. The matching process may be to use the feature points of the lane lines in the image of each camera as guide points, determine the feature points of the lane lines corresponding to the guide points in the corresponding rendered image, and obtain the second pose by solving the (perspective-n-point, PnP) algorithm to obtain the second pose of each camera.
[0137] Among them, the PnP algorithm is a method for solving the correspondence between 3D and 2D points. When multiple 3D spatial points and their positions in an image are known, the position and posture of the camera that took the image can be estimated.
[0138] The above process of optimizing the first pose to obtain the second pose can also be referred to Figure 7 To understand, such as Figure 7As shown, taking the first camera as an example, according to the first pose of the first camera, a rendered image 701 of the target scene in the corresponding first pose is generated, and then the feature points of the lane lines in the rendered image 701 can be extracted through a deep neural network or a convolutional neural network, and these feature points are used as guide points 702, and then the feature points 704 of the lane lines corresponding to each guide point 702 are searched from the image 703 taken by the first camera, and then the multiple feature points 704 are solved by the PnP algorithm to determine a more accurate second pose of the first camera. At the same time, the algorithms involved in the above-mentioned image processing process can also be pruned, so that the efficiency of image processing can be improved.
[0139] In the embodiment of the present application, in the process of optimizing the first pose to obtain the second pose, it is only necessary to determine the feature points of the lane line in the camera image, and then use these feature points as guide points to match them with the corresponding feature points in the corresponding rendered image. By solving the PnP algorithm, the second pose can be obtained. No global matching is required, only local matching of the feature points of the lane line is required, which can reduce the amount of calculation and improve matching efficiency.
[0140] After obtaining the second pose of the camera through the above optimization, the second pose can be used to perform image stitching according to the stitching relationship. The stitching relationship can be a mapping relationship between the camera identifier and the stitching area identifier. One camera can correspond to a stitching area on the display area. By placing the camera image in the corresponding stitching area, image stitching can be achieved. The process can be referred to Figure 8 To understand, such as Figure 8 As shown: the display area includes four stitching areas, namely stitching area 1, stitching area 2, stitching area 3 and stitching area 4. Among them, stitching area 1 corresponds to camera 3, stitching area 2 corresponds to camera 4, stitching area 3 corresponds to camera 1, and stitching area 4 corresponds to camera 2. In this way, the image from camera 1 can be mapped to stitching area 3, the image from camera 2 can be mapped to stitching area 4, the image from camera 3 can be mapped to stitching area 1, and the image from camera 4 can be mapped to stitching area 2, thereby completing the stitching of the images of the four cameras.
[0141] The mapping relationship between the camera identifier and the stitching area identifier can be: through the second posture of each camera and the distortion parameter of each camera, determine the camera to which the 3D point corresponding to each pixel on the display area is projected in the space of the target scene, and then determine the mapping relationship between the camera and the stitching area according to the stitching area to which each pixel belongs. Figure 8 The correspondence between the camera and the stitching area shown is determined by the second pose.
[0142] Above Figure 8The stitching principle described is a simple illustration. In fact, when stitching, images from different cameras may overlap. In order to improve the quality of image stitching, the process of image stitching can be: when the 3D point corresponding to the target pixel in the space of the target scene is projected to one camera, the pixel information corresponding to the target pixel in the target image of one camera is rendered to the target pixel; when the 3D point corresponding to the target pixel in the space of the target scene is projected to at least two cameras, the pixel information corresponding to the target pixel in the target image of the target camera is rendered to the target pixel, and the target camera is the camera with the smallest distance to the 3D point corresponding to the target pixel in the space of the target scene among the at least two cameras. In this way, duplicates can be removed and the quality of image stitching can be improved.
[0143] For more information about the situation when the 3D point corresponding to the target pixel in the target scene space is projected onto at least two cameras, please refer to Fig. 9 To understand, such as Fig. 9 As shown, the pixel point 901 in the stitching area 1 is projected onto the camera 3 and the camera 4 at the 3D point 902 in the space. The distance between the 3D point 902 and the camera 3 is farther than the distance between the 3D point 902 and the camera 4. Then, the pixel information corresponding to the pixel point 901 in the image taken by the camera 4 is rendered for the pixel point 901.
[0144] For a schematic diagram of the image stitching scene, please refer to Fig.10 To understand, such as Fig.10 As shown, four images taken at different angles by four cameras, after the determination of the first pose, optimization of the second pose, and the image stitching process described above, can obtain a stitched image 100 of the intersection. The stitched image 100 can be displayed on a display in a road surface management center, which is conducive to quickly checking the traffic conditions at the intersection.
[0145] The above introduction is about the stitching of images at a certain time point. If the stitching images of continuous time are continuously displayed on the display, the traffic situation of the intersection can be viewed in the form of video.
[0146] It should be noted that, in the embodiment of the present application, after the first posture or the second posture is determined, the posture of the camera can be adjusted according to the first posture or the second posture.
[0147] In addition, in the embodiment of the present application, it is not necessary to know the position and posture of the camera in advance, and accurate stitching of images from different cameras can be achieved according to the above process. Even if the camera rotates or moves after taking the image at the first moment, the above cloud device or roadside device can still perform the image processing process introduced above on the image taken by the camera at the second moment, and complete the stitching of the image at the second moment.
[0148] The above describes the process of stitching images from different cameras based on the first pose of the camera or the second pose of the camera. In fact, the image processing method provided in the embodiment of the present application can also be: estimating the first pose of the first camera through the image from the first camera; generating a rendered image of the target scene under the first pose, where the target scene is the scene associated with the first camera; matching the lane line features in the image from the first camera with the lane line features in the rendered image to optimize the first pose to obtain the second pose of the first camera. In this way, by optimizing the first pose to obtain the second pose, the accuracy of the pose can be improved.
[0149] That is to say, the solution provided by this application does not limit the number of cameras in the target scene, and pose optimization can also be performed when there is only one camera. For the pose optimization process, please refer to the previous Figure 7 The contents introduced in this part have been understood and will not be repeated here.
[0150] In addition, the front Figures 3 to 10 The processes of determining the first pose, optimizing the first pose to obtain the second pose, and image stitching described above are all applicable to the present application.
[0151] The above introduces the image processing method. The following introduces the image processing device provided in the embodiment of the present application with reference to the accompanying drawings.
[0152] like Fig.11 As shown, the image processing device 110 provided in the embodiment of the present application includes:
[0153] The acquisition unit 1101 is used to acquire multiple images taken by multiple cameras.
[0154] The first processing unit 1102 is used to determine the first position of at least two cameras among the multiple cameras based on the multiple images and the multiple raster images, wherein the multiple images and the multiple raster images are associated with the same target scene, wherein the first position of the first camera is related to the position of the virtual camera corresponding to the first raster image, the first raster image is a raster image among the multiple raster images corresponding to the image taken by the first camera, and the first camera is any one of the multiple cameras.
[0155] The second processing unit 1103 is used to stitch multiple images according to the first poses of at least two cameras.
[0156] Optionally, the first processing unit 1102 is specifically configured to perform marker feature matching on the multiple images and the corresponding raster images respectively, so as to estimate the first position of each camera in the multiple cameras.
[0157] Optionally, the second processing unit 1103 is specifically configured to stitch the multiple images together according to the estimated first pose of each camera.
[0158] Optionally, the marker features may include at least one of lane features, building features, vehicle features, pedestrian features, and guardrail features.
[0159] Optionally, the first processing unit 1101 is used to perform lane line feature matching on a first lane line feature associated with the first raster image and a second lane line feature in the image of the first camera; and adjust the pose of the virtual camera corresponding to the first raster image according to an optimization target to obtain the first pose of the first camera, wherein the optimization target is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization target is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
[0160] Optionally, the second processing unit 1102 is used to optimize the estimated first pose of each camera to obtain a second pose of each camera; and stitch the multiple images according to the second pose of each camera.
[0161] Optionally, the second processing unit 1102 is used to generate a rendered image of the target scene in the corresponding first pose according to the first pose of each camera; and match the lane line features in the image of each camera with the lane line features in the corresponding rendered image to obtain the second pose of each camera.
[0162] Optionally, the second processing unit 1102 is configured to use the feature points of the lane lines in the image of each camera as the guide points, determine the feature points of the lane lines corresponding to the guide points in the corresponding rendered image, and perform matching.
[0163] Optionally, the second processing unit 1102 is used to determine a stitching relationship between the multiple images according to the second posture of each camera; and stitch the multiple images according to the stitching relationship.
[0164] Optionally, the first processing unit 1101 is further configured to adjust the first camera according to the first pose of the first camera.
[0165] Optionally, the image from the first camera is an image taken by the first camera at a first moment, and the first processing unit 1101 is further used to obtain an image taken by the first camera at a second moment, where the second moment is later than the first moment; the image taken by the first camera at the second moment is matched with the second grid image for lane line features to estimate the posture of the first camera at the second moment, and the posture of the first camera at the second moment is different from the first posture of the first camera to indicate that the first camera has rotated or moved.
[0166] The functions of each unit in the image processing device 110 introduced in the present application can be understood by referring to the corresponding contents in the method embodiment introduced above, and will not be repeated here.
[0167] like Fig.12 As shown, another image processing device 120 provided in an embodiment of the present application includes:
[0168] The first processing unit 1201 is used to estimate the first pose of the first camera through the image from the first camera.
[0169] The second processing unit 1202 is used to generate a rendered image of a target scene in a first posture, where the target scene is a scene associated with the first camera.
[0170] The third processing unit 1203 is used to match the image rendering image from the first camera to optimize the first pose to obtain the second pose of the first camera.
[0171] Optionally, the image from the first camera is included in multiple images, each of the multiple images comes from a different camera associated with the target scene, the first camera is any one of the multiple cameras associated with the target scene, and the third processing unit is further used to stitch the multiple images according to the second posture of each camera.
[0172] Optionally, the first processing unit 1201 is used to match lane line features of the image from the first camera with the corresponding first raster image to estimate the first pose of the first camera; wherein the first pose of the first camera is determined based on the pose of the virtual camera corresponding to the first raster image, and the first raster image is the raster image with the highest similarity to the image of the first camera among multiple raster images associated with the target scene.
[0173] Optionally, the marker features include at least one of lane line features, building features, vehicle features, pedestrian features, and guardrail features.
[0174] Optionally, the first processing unit 1201 is used to perform lane line feature matching on a first lane line feature associated with the first raster image and a second lane line feature in the image of the first camera; and adjust the position and posture of the virtual camera corresponding to the first raster image according to an optimization target to obtain the first position and posture of the first camera, wherein the optimization target is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization target is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
[0175] Optionally, the third processing unit 1203 is configured to use the feature points of the lane lines in the image of the first camera as the guide points, determine the feature points of the lane lines corresponding to the guide points in the corresponding rendered image, and perform matching.
[0176] Optionally, the third processing unit 1203 is used to determine a stitching relationship between the multiple images according to the second posture of each camera; and stitch the multiple images according to the stitching relationship.
[0177] Optionally, the third processing unit 1203 is further configured to adjust the first camera according to the second posture of the first camera.
[0178] Optionally, the image from the first camera is an image taken by the first camera at a first moment, and the first processing unit 1201 is further used to obtain an image taken by the first camera at a second moment, where the second moment is later than the first moment; the image taken by the first camera at the second moment is matched with the second grid image for lane line features to estimate the posture of the first camera at the second moment, and the posture of the first camera at the second moment is different from the first posture of the first camera to indicate that the first camera has rotated or moved.
[0179] The functions of each unit in the image processing device 120 introduced in the present application can be understood by referring to the corresponding contents in the method embodiment introduced above, and will not be repeated here.
[0180] The present application also provides a computer-readable storage medium in which a program is stored. When the program is run on a computer, the computer executes the above-mentioned Figure 3-Figure 10 The illustrated embodiments describe the method.
[0181] The embodiment of the present application also provides an image processing device, which can also be called a digital processing chip or chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface. The program instructions are executed by the processing unit. The processing unit is used to execute the aforementioned Figure 3-Figure 10 The method shown in any one of the embodiments.
[0182] The embodiment of the present application also provides a digital processing chip. The digital processing chip integrates a circuit and one or more interfaces for implementing the above-mentioned processor 201 or the functions of the processor 201. When the digital processing chip integrates a memory, the digital processing chip can complete the method steps of any one or more of the above-mentioned embodiments. When the digital processing chip does not integrate a memory, it can be connected to an external memory through a communication interface. The digital processing chip implements the method in the above-mentioned embodiment according to the program code stored in the external memory.
[0183] In an embodiment of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores instructions. When the instructions are executed on an electronic device, the electronic device executes the above-mentioned Figure 3-Figure 10 The steps in the method described in any embodiment.
[0184] The present application also provides a computer program product which, when executed on a computer, enables the computer to execute the above-mentioned Figure 3-Figure 10 The steps in the method described in any embodiment.
[0185] The image processing device provided in the embodiment of the present application may be a chip, and the chip includes: a processing unit and a communication unit, the processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit. The processing unit may execute the computer execution instructions stored in the storage unit, so that the chip in the computer device executes the above Figure 3-Figure 10 The method of image processing described in the illustrated embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0186] Specifically, the aforementioned processing unit or processor may include a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0187] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0188] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0189] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0190] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)), etc.
Claims
1. A method for image processing, It is characterized in that include: Get multiple images taken by multiple cameras; Determining first poses of at least two cameras among the multiple cameras according to the multiple images and the multiple raster images, wherein the multiple images and the multiple raster images are associated with the same target scene, wherein the first pose of the first camera is related to the pose of a virtual camera corresponding to the first raster image, the first raster image is a raster image among the multiple raster images corresponding to the image captured by the first camera, and the first camera is any one of the multiple cameras; The multiple images are stitched according to the first poses of the at least two cameras.
2. The method according to claim 1, It is characterized in that The step of determining the first poses of at least two cameras among the plurality of cameras according to the plurality of images and the plurality of grid images comprises: Matching the multiple images with the corresponding raster images respectively to estimate the first pose of each camera in the multiple cameras; Correspondingly, stitching the multiple images according to the first poses of the at least two cameras includes: The multiple images are stitched together according to the estimated first pose of each camera.
3. The method according to claim 2, It is characterized in that The marker features include at least one of lane features, building features, vehicle features, pedestrian features, and guardrail features.
4. The method according to claim 3, It is characterized in that The method of matching the multiple images with the corresponding grid images for the features of the markers to estimate the first position of each camera in the multiple cameras includes: Performing lane line feature matching on a first lane line feature associated with the first raster image and a second lane line feature in the image of the first camera; The pose of the virtual camera corresponding to the first raster image is adjusted according to an optimization goal to obtain a first pose of the first camera, wherein the optimization goal is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization goal is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
5. The method according to any one of claims 2 to 4, It is characterized in that The step of stitching the multiple images according to the estimated first pose of each camera comprises: Optimizing the estimated first pose of each camera to obtain a second pose of each camera; The multiple images are stitched together according to the second posture of each camera.
6. The method according to claim 5, It is characterized in that The step of optimizing the estimated first pose of each camera to obtain a second pose of each camera includes: Generating a rendered image of the target scene at the corresponding first pose according to the first pose of each camera; The lane line features in the image of each camera are matched with the lane line features in the corresponding rendered image to obtain the second pose of each camera.
7. The method according to claim 6, It is characterized in that The matching of the lane line features in the image of each camera with the lane line features in the corresponding rendered image includes: The feature points of the lane lines in the images of each camera are used as guide points, and the feature points of the lane lines corresponding to the guide points are determined in the corresponding rendered images and matched.
8. The method according to any one of claims 5 to 7, It is characterized in that The step of stitching the multiple images according to the respective second poses of each camera comprises: Determining a stitching relationship between the multiple images according to the respective second postures of each camera; The multiple images are stitched together according to the stitching relationship.
9. The method according to any one of claims 1 to 8, It is characterized in that The method further comprises: The first camera is adjusted according to the first pose of the first camera.
10. The method according to any one of claims 2 to 8, It is characterized in that The image from the first camera is an image taken by the first camera at a first moment, and the method further includes: Acquire an image captured by the first camera at a second moment, where the second moment is later than the first moment; Lane line features are matched between the image captured by the first camera at the second moment and the second grid image to estimate the position and posture of the first camera at the second moment. The position and posture of the first camera at the second moment is different from the first position and posture of the first camera to indicate that the first camera is rotated or moved.
11. A method of image processing, It is characterized in that include: estimating a first pose of the first camera using an image from the first camera; generating a rendered image of a target scene at the first pose, the target scene being a scene associated with the first camera; The image from the first camera is matched with the rendered image to optimize the first pose to obtain a second pose of the first camera.
12. The method according to claim 11, It is characterized in that The image from the first camera is included in a plurality of images, each of the plurality of images is from a different camera associated with the target scene, and the first camera is any one of the plurality of cameras associated with the target scene; The method further comprises: The multiple images are stitched together according to the second posture of each camera.
13. The method according to claim 11 or 12, It is characterized in that The estimating the first pose of the first camera by using an image from the first camera comprises: The image from the first camera is matched with the corresponding first raster image for marker features to estimate the first pose of the first camera; wherein the first pose of the first camera is related to the pose of the virtual camera corresponding to the first raster image, and the first raster image is a raster image among multiple raster images corresponding to the image taken by the first camera.
14. The method according to claim 13, It is characterized in that The marker features include at least one of lane features, building features, vehicle features, pedestrian features, and guardrail features.
15. The method according to claim 14, It is characterized in that The step of performing marker feature matching on the image from the first camera and the corresponding first grid image to estimate the first pose of the first camera includes: Performing lane line feature matching on a first lane line feature associated with the first raster image and a second lane line feature in the image of the first camera; The pose of the virtual camera corresponding to the first raster image is adjusted according to an optimization goal to obtain a first pose of the first camera, wherein the optimization goal is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization goal is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
16. The method according to any one of claims 11 to 15, It is characterized in that The matching the image from the first camera with the rendered image comprises: The feature points of the lane lines in the image of the first camera are used as guide points, and the feature points of the lane lines corresponding to the guide points are determined in the corresponding rendered image and matched.
17. The method according to claim 12, It is characterized in that The step of stitching the multiple images according to the respective second poses of each camera comprises: Determining a stitching relationship between the multiple images according to the respective second postures of each camera; The multiple images are stitched together according to the stitching relationship.
18. The method according to any one of claims 11 to 17, It is characterized in that The method further comprises: The first camera is adjusted according to the second position and posture of the first camera.
19. The method according to any one of claims 11 to 17, It is characterized in that The image from the first camera is an image taken by the first camera at a first moment, and the method further includes: Acquire an image captured by the first camera at a second moment, where the second moment is later than the first moment; Lane line features are matched between the image captured by the first camera at the second moment and the second grid image to estimate the position and posture of the first camera at the second moment. The position and posture of the first camera at the second moment is different from the first position and posture of the first camera to indicate that the first camera is rotated or moved.
20. An image processing device, It is characterized in that include: An acquisition unit, used for acquiring a plurality of images taken by a plurality of cameras; A first processing unit is configured to determine first poses of at least two cameras among the multiple cameras according to the multiple images and the multiple raster images, wherein the multiple images and the multiple raster images are associated with the same target scene, wherein the first pose of the first camera is related to the pose of a virtual camera corresponding to the first raster image, the first raster image is a raster image among the multiple raster images corresponding to the image captured by the first camera, and the first camera is any one of the multiple cameras; The second processing unit is used to stitch the multiple images according to the first poses of the at least two cameras.
21. An image processing device, It is characterized in that include: A first processing unit, configured to estimate a first pose of the first camera using an image from the first camera; a second processing unit, configured to generate a rendered image of a target scene in the first pose, wherein the target scene is a scene associated with the first camera; A third processing unit is used to match the image from the first camera with the rendered image to optimize the first posture to obtain a second posture of the first camera.
22. An electronic device, It is characterized in that include: A communication interface, a processor and a memory, wherein the communication interface and the processor are coupled to the memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the electronic device executes the method as described in any one of claims 1 to 10, or executes the method as described in any one of claims 11 to 19.
23. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed on an electronic device, the electronic device executes the method as claimed in any one of claims 1 to 10, or executes the method as claimed in any one of claims 11 to 19.
24. A computer program product, It is characterized in that The computer program product comprises a computer program code, and when the computer program code is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 10, or execute the method according to any one of claims 11 to 19.
25. A chip system, It is characterized in that The chip system includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected through lines; the interface circuits are used to receive signals from the memory of the electronic device and send signals to the processor, and the signals include computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the method as described in any one of claims 1 to 10, or executes the method as described in any one of claims 11 to 19.
26. An image processing system, It is characterized in that include: An electronic device, wherein the electronic device is used to execute the method as described in any one of claims 1 to 10, or execute the method as described in any one of claims 11 to 19, and the electronic device is a cloud-side device or a road-side device.