Image processing method and corresponding apparatus
By using raster images to match the characteristic features of the camera image in traffic scenes, determining the camera position and realizing image stitching, the problems of low efficiency and poor accuracy of traditional methods are solved, and the efficiency and accuracy of image stitching are improved.
Patent Information
- Application Number
- PCT/CN2024/129676
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-21
- Filing Date
- 2024-11-04
- Publication Date
- 2025-05-30
AI Technical Summary
In traffic scenarios, traditional image stitching methods are inefficient and have poor accuracy, especially in complex scenarios, it is difficult to accurately judge the relationship between story shots, and it is impossible to determine the specific location of the incident.
By acquiring images captured by multiple cameras and matching characteristic features with pre-generated raster images, the camera position is determined, and image stitching is realized. Raster images are obtained by rendering the three-dimensional model of the target scene. They have small data volume, little storage space, and can quickly match.
It improves the efficiency and accuracy of image stitching, reduces the use of storage space, and realizes accurate image matching and stitching in complex scenes.
Smart Images

Figure CN2024129676_30052025_PF_FP_ABST
Abstract
Description
Image processing method and corresponding device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 21, 2023, with application number 202311557960.3 and application name “A method for image processing and corresponding device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of image processing technology, and in particular to an image processing method and corresponding device. Background Art
[0003] In traffic scenarios, multiple cameras are typically deployed at intersections or highways, capturing images at various locations along the road. This provides data support for incident prediction and signal optimization at key intersections. Traditionally, multiple cameras generate multiple shots, making it difficult for data management centers to simultaneously maintain data from all shots. Furthermore, the scattered camera perspectives make it difficult to accurately determine the connections between shots and pinpoint the exact location of an incident. The key technology to address these issues is stitching images captured by different cameras.
[0004] Currently, two main approaches are used for image stitching. One is manual, requiring manual selection and matching of key points in the image to be matched. This process is time-consuming and labor-intensive, requiring manual intervention every time the camera rotates or moves. The other is automated, based on high-precision or 3D maps. By matching the image to be matched with the high-precision or 3D map, the camera's position is determined and stitched together. However, high-precision and 3D maps require large amounts of data, resulting in low stitching efficiency and poor matching results in complex scenes.
[0005] Summary of the Invention
[0006] The present application provides an image processing method for improving the efficiency and accuracy of image stitching. The present application also provides corresponding devices, systems, computer-readable storage media, and computer program products.
[0007] A first aspect of the present application provides an image processing method, comprising: acquiring multiple images captured by multiple cameras; determining the first poses of at least two cameras among the multiple cameras based on the multiple images and multiple raster images, wherein the multiple images and the multiple raster images are associated with the same target scene, wherein the first pose of the first camera is related to the pose of a virtual camera corresponding to the first raster image, the first raster image is a raster image among the multiple raster images corresponding to the image captured by the first camera, and the first camera is any one of the multiple cameras; and stitching the multiple images according to the first poses of the at least two cameras.
[0008] The image processing method in this application can be applied to traffic scenarios, such as intersections / highways. The process can be executed by a cloud device or a road-side device.
[0009] In this application, the target scene can be an intersection or a section of road. A target scene is usually configured with multiple cameras (also called cameras), where each camera captures images of the target scene in different directions and positions.
[0010] In the present application, the raster image may be pre-generated, and the raster image may be obtained by rendering the three-dimensional model of the target scene according to the posture of the virtual camera. Each raster image corresponds to the posture of a virtual camera and the image features under the posture of the virtual camera. The raster image, the posture of the virtual camera, and the image features can be understood as a set of data with a corresponding relationship. In the present application, when generating the raster image, the lane lines or other features with a marking function in the three-dimensional model can be rendered in a focused manner, and other features can be ignored or simply rendered. In this way, a raster image with a smaller amount of data can be obtained, and such a raster image takes up less storage space.
[0011] In this application, the camera pose refers to the camera's position and attitude, which can include the position (x, y, z) in a three-dimensional coordinate system, as well as the azimuth, pitch, and roll angles. Typically, the camera pose is described by the camera's rotation matrix (R) and translation vector (t), rather than the position and angle coordinates described above. The virtual camera pose refers to the pose assumed when generating the raster image.
[0012] In this application, by matching the characteristics of the image captured by the camera with the raster image, the first pose of the camera can be obtained through the pose of the virtual camera corresponding to the raster image. The raster image corresponding to the image captured by the camera can be the raster image with the highest similarity to the image captured by the camera among multiple raster images.
[0013] In this first aspect, because raster images have a small data size, storage space usage can be reduced. Furthermore, raster images primarily contain features such as lane markings and buildings, with relatively few other features. This allows for rapid matching with camera-captured images, thereby improving the efficiency and accuracy of image stitching.
[0014] In one possible implementation, the above steps: determining the first pose of at least two cameras among the multiple cameras based on the multiple images and the multiple raster images, including: performing marker feature matching on the multiple images and the corresponding raster images respectively to estimate the first pose of each camera among the multiple cameras; correspondingly, stitching the multiple images according to the first poses of the at least two cameras, including: stitching the multiple images according to the estimated first pose of each camera.
[0015] In this possible implementation, the marker features may include at least one of lane features, building features, vehicle features, pedestrian features, and guardrail features. Feature matching based on the marker features can improve image matching speed. Furthermore, stitching multiple images based on each camera's first pose can improve image stitching accuracy.
[0016] In one possible implementation, the above steps: performing marker feature matching on multiple images with corresponding raster images respectively to estimate the first pose of each camera in the multiple cameras, including: performing lane line feature matching on a first lane line feature associated with the first raster image with a second lane line feature in the image of the first camera; adjusting the pose of the virtual camera corresponding to the first raster image according to an optimization target to obtain the first pose of the first camera, where the optimization target is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization target is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
[0017] In this possible implementation, when determining the first pose of the first camera, a backpropagation optimization can be performed based on the residual between the lane line features in the image of the first camera and the first grid image, using the pose of the virtual camera corresponding to the first grid image as the basis. This ensures that the lane line features between the image of the first camera and the first grid image are as similar as possible, or that the distance between the lane line features in space is as close as possible. The output pose is then the first pose of the first camera. In this application, optimizing the pose of the virtual camera corresponding to the first grid image according to the optimization objective to obtain the pose of the first camera can improve the accuracy of the first pose.
[0018] In one possible implementation, the above step of stitching multiple images based on the estimated first pose of each camera includes: optimizing the estimated first pose of each camera to obtain the second pose of each camera; and stitching multiple images based on the second pose of each camera.
[0019] In this possible implementation, the first pose can also be optimized to obtain a second pose with higher accuracy, thereby improving the accuracy of image stitching.
[0020] In one possible implementation, the above steps: optimizing the estimated first pose of each camera to obtain the second pose of each camera, including: generating a rendered image of the target scene at the corresponding first pose based on the first pose of each camera; matching the lane line features in the image of each camera with the lane line features in the corresponding rendered image to obtain the second pose of each camera.
[0021] In this possible implementation, the rendered image generated from each camera's first pose is closer to the image captured by the camera than the raster image. Therefore, the second pose obtained using the rendered image is more accurate, further improving the accuracy of image stitching.
[0022] In one possible implementation, the above step of matching the lane line features in each camera's image with the lane line features in the corresponding rendered image includes using the lane line feature points in each camera's image as guide points, determining the lane line feature points corresponding to the guide points in the corresponding rendered image, and matching them.
[0023] In this possible implementation, we only need to determine the lane feature points in the camera image, then use these feature points as guide points to match them with the corresponding feature points in the rendered image. Using a perspective-n-point (PnP) algorithm, we can obtain the second pose. This application does not require global matching, but only requires local matching of lane feature points, which can reduce the computational effort and improve matching efficiency.
[0024] In a possible implementation, stitching the multiple images according to the second posture of each camera includes: determining a stitching relationship between the multiple images according to the second posture of each camera; and stitching the multiple images according to the stitching relationship.
[0025] In this possible implementation, the stitching relationship can be a mapping relationship between the camera identifier and the stitching area identifier. One camera can correspond to a stitching area on the display area. By placing the camera image in the corresponding stitching area, image stitching can be achieved, thereby improving the efficiency of image stitching.
[0026] In one possible implementation, the mapping relationship between the identification of the above-mentioned camera and the identification of the stitching area can be: through the second pose of each camera and the distortion parameter of each camera, determine the camera to which the 3D point corresponding to each pixel on the display area in the space of the target scene is projected, and then, according to the stitching area to which each pixel point belongs, determine the mapping relationship between the camera and the stitching area.
[0027] In one possible implementation, the image stitching process using the mapping relationship can be as follows: when a target pixel in the target scene's spatially corresponding 3D point is projected onto one camera, the pixel information corresponding to the target pixel in the target image of one camera is rendered onto the target pixel; and when the target pixel in the target scene's spatially corresponding 3D point is projected onto at least two cameras, the pixel information corresponding to the target pixel in the target image of the target camera is rendered onto the target pixel, where the target camera is the camera with the shortest distance to the 3D point corresponding to the target pixel in the target scene's spatially corresponding 3D point of the at least two cameras. This can eliminate duplication and improve image stitching quality.
[0028] In a possible implementation, the method further includes: adjusting the first camera according to the first pose of the first camera.
[0029] In this possible implementation, after the first pose of the first camera is determined, the first camera may be controlled to move or rotate so that the first camera reaches the position and posture indicated by the first pose.
[0030] In one possible implementation, the image from the first camera is an image taken by the first camera at a first moment, and the method further includes: obtaining an image taken by the first camera at a second moment, where the second moment is later than the first moment; performing lane line feature matching on the image taken by the first camera at the second moment with the second grid image to estimate the position of the first camera at the second moment, where the position of the first camera at the second moment is different from the first position of the first camera, indicating that the first camera has rotated or moved.
[0031] In this possible implementation, after the first camera is rotated or moved, no human intervention is required, and accurate image stitching can still be performed using the data after the movement.
[0032] A second aspect of the present application provides an image processing method, comprising: estimating a first pose of the first camera through an image from the first camera; generating a rendered image of a target scene at the first pose, where the target scene is a scene associated with the first camera; and matching the image from the first camera with the rendered image to optimize the first pose to obtain a second pose of the first camera.
[0033] In the second aspect, the accuracy of the posture can be improved by optimizing the first posture to obtain the second posture.
[0034] In one possible implementation, the image from the first camera is included in multiple images, each of the multiple images comes from a different camera associated with the target scene, and the first camera is any one of the multiple cameras associated with the target scene. The method also includes: stitching the multiple images according to the second pose of each camera.
[0035] In one possible implementation, the step of estimating the first pose of the first camera using the image from the first camera includes performing lane line feature matching on the image from the first camera and the corresponding first raster image to estimate the first pose of the first camera; wherein the first pose of the first camera is related to the pose of the virtual camera corresponding to the first raster image, and the first raster image is a raster image among the multiple raster images that corresponds to the image captured by the first camera.
[0036] In one possible implementation, the marker feature includes at least one of a lane feature, a building feature, a vehicle feature, a pedestrian feature, and a guardrail feature.
[0037] In one possible implementation, the above steps of: performing marker feature matching on the image from the first camera and the corresponding first raster image to estimate the first pose of the first camera include: performing lane line feature matching on the first lane line feature associated with the first raster image and the second lane line feature in the image of the first camera; adjusting the pose of the virtual camera corresponding to the first raster image according to the optimization target to obtain the first pose of the first camera, wherein the optimization target is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization target is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
[0038] In one possible implementation, the above step of matching the image from the first camera with the rendered image includes using the feature points of the lane lines in the image of the first camera as guide points, determining the feature points of the lane lines corresponding to the guide points in the corresponding rendered image, and matching them.
[0039] In a possible implementation, stitching the multiple images according to the second posture of each camera includes: determining a stitching relationship between the multiple images according to the second posture of each camera; and stitching the multiple images according to the stitching relationship.
[0040] In a possible implementation, the method further includes: adjusting the first camera according to the second posture of the first camera.
[0041] In one possible implementation, the image from the first camera is an image taken by the first camera at a first moment, and the method further includes: obtaining an image taken by the first camera at a second moment, where the second moment is later than the first moment; performing lane line feature matching on the image taken by the first camera at the second moment with the second grid image to estimate the position of the first camera at the second moment, where the position of the first camera at the second moment is different from the first position of the first camera, indicating that the first camera has rotated or moved.
[0042] In a third aspect, the present application provides an image processing device having the function of implementing the image processing method of the first aspect or any possible implementation of the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above-mentioned functions, such as: an acquisition unit, a first processing unit, and a second processing unit.
[0043] In a fourth aspect, the present application provides an image processing device having the function of implementing the image processing method of the second aspect or any possible implementation of the second aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above-mentioned functions, such as: a first processing unit, a second processing unit, and a third processing unit.
[0044] In a fifth aspect, the present application provides an electronic device comprising a communication interface, a processor and a memory, wherein the communication interface and the processor are coupled to the memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the electronic device executes a method as described in the first aspect or any possible implementation of the first aspect.
[0045] In the present application, the processor may include at least one of a central processing unit (CPU) and a graphics processing unit (GPU); wherein both the CPU and the GPU may execute the image processing process described in the above-mentioned first aspect or any possible implementation of the first aspect, or the CPU and the GPU cooperate to execute the image processing process described in the above-mentioned first aspect or any possible implementation of the first aspect.
[0046] In a sixth aspect of the present application, an electronic device is provided, comprising a communication interface, a processor and a memory, wherein the communication interface and the processor are coupled to the memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the electronic device executes a method as described in the second aspect or any possible implementation of the second aspect.
[0047] In the present application, the processor may include at least one of a central processing unit (CPU) and a graphics processing unit (GPU); wherein both the CPU and the GPU may execute the image processing process described in the above-mentioned second aspect or any possible implementation of the second aspect, or the CPU and the GPU cooperate to execute the image processing process described in the above-mentioned second aspect or any possible implementation of the second aspect.
[0048] In a seventh aspect, the present application provides a chip system comprising one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected by lines; the interface circuits are used to receive signals from a memory of an electronic device and send signals to the processor, the signals comprising computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes a method as described in the first aspect or any possible implementation of the first aspect, and the processor is at least one of a CPU and a GPU.
[0049] In an eighth aspect of the present application, a chip system is provided, which includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected by lines; the interface circuits are used to receive signals from the memory of the electronic device and send signals to the processor, and the signals include computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the method as described in the second aspect or any possible implementation of the second aspect, and the processor is at least one of a CPU and a GPU.
[0050] In a ninth aspect, the present application provides a computer-readable storage medium, which stores instructions. When the instructions are executed on an electronic device, the electronic device executes a method as described in the first aspect or any possible implementation of the first aspect.
[0051] In a tenth aspect, the present application provides a computer-readable storage medium, which stores instructions. When the instructions are executed on an electronic device, the electronic device executes a method as described in the second aspect or any possible implementation of the second aspect.
[0052] In an eleventh aspect, the present application provides a computer program product, which includes a computer program code. When the computer program code runs on a computer, the computer executes the method of the first aspect or any possible implementation of the first aspect.
[0053] The twelfth aspect of the present application provides a computer program product, which includes computer program code. When the computer program code runs on a computer, it enables the computer to execute the method of the second aspect or any possible implementation of the second aspect.
[0054] The thirteenth aspect of the present application provides an image processing system, including: an electronic device, the electronic device is used to execute the method of the above-mentioned first aspect or any possible implementation method of the first aspect, and the electronic device is a cloud-side device or a road-side device.
[0055] In a fourteenth aspect, the present application provides an image processing system, comprising: an electronic device, the electronic device being used to execute the method of the above-mentioned second aspect or any possible implementation method of the second aspect, the electronic device being a cloud-side device or a road-side device.
[0056] The relevant features and effects of the second aspect of this application, as well as any possible implementation of the second aspect to the fourteenth aspect, can be understood by referring to the corresponding introduction in the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] FIG1A is a schematic diagram of a scenario architecture provided by an embodiment of the present application;
[0058] FIG1B is a schematic diagram of another scenario architecture provided by an embodiment of the present application;
[0059] FIG1C is a schematic diagram of another scenario architecture provided by an embodiment of the present application;
[0060] FIG1D is a schematic structural diagram of an electronic device provided in an embodiment of the present application;
[0061] FIG2 is a schematic diagram of the structure of a cloud device / road-end device provided in an embodiment of the present application;
[0062] FIG3 is a schematic diagram of a process for generating a raster image according to an embodiment of the present application;
[0063] FIG4 is a schematic diagram of an example scenario provided by an embodiment of the present application;
[0064] FIG5 is a schematic diagram of an embodiment of an image processing method provided by an embodiment of the present application;
[0065] FIG6 is a schematic diagram of an example of an image processing method provided by an embodiment of the present application;
[0066] FIG7 is another exemplary schematic diagram of the image processing method provided by an embodiment of the present application;
[0067] FIG8 is another exemplary schematic diagram of an image processing method provided by an embodiment of the present application;
[0068] FIG9 is another exemplary schematic diagram of an image processing method provided by an embodiment of the present application;
[0069] FIG10 is another exemplary schematic diagram of an image processing method provided by an embodiment of the present application;
[0070] FIG11 is a schematic structural diagram of an image processing apparatus provided in an embodiment of the present application;
[0071] FIG12 is another schematic diagram of the structure of the image processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0072] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present application, rather than all the embodiments. Those skilled in the art will appreciate that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0073] The terms "first," "second," and the like in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.
[0074] The present application provides an image processing method for improving the efficiency and accuracy of image stitching. The present application also provides corresponding devices, systems, computer-readable storage media, and computer program products. These are described in detail below.
[0075] The solution provided by the embodiments of the present application can be applied to smart transportation. For example, at an intersection or on a highway, a camera (also called a camera) can be used to capture traffic conditions at the intersection or highway. The camera then transmits the captured image to the cloud. A cloud-side device in the cloud stitches the images captured by multiple cameras at the intersection, and then sends the stitched image to a display at a road management center for display. Alternatively, the camera transmits the captured image to a roadside device, which stitches the images captured by multiple cameras at the intersection, and then sends the stitched image to a display at a road management center for display.
[0076] The situation processed by the cloud can be understood by referring to Figure 1A. As shown in Figure 1A, taking an intersection as an example, the intersection has four cameras, each of which can capture the traffic conditions within a certain range of the intersection, and then each camera transmits the image to the cloud. For example, the images transmitted to the cloud by the four cameras are image 1, image 2, image 3 and image 4 respectively. The cloud-side device on the cloud will splice the four images according to their postures to obtain the spliced image of the intersection, and then send the spliced image of the intersection to the display of the road management center for display.
[0077] The situation processed by the road side can be understood by referring to Figure 1B. As shown in Figure 1B, taking an intersection as an example, the intersection has four cameras, each of which can capture the traffic conditions within a certain range of the intersection, and then each camera transmits the image to the road side device. For example, the images transmitted to the road side device by the four cameras are image 1, image 2, image 3 and image 4 respectively. The road side device will splice the four images according to their postures to obtain the spliced image of the intersection, and then send the spliced image of the intersection to the display of the road surface management center for display.
[0078] The road-end device may be a terminal device configured at the intersection, and the terminal device may be a camera with an image processing function, or may be a device specifically responsible for image processing.
[0079] The cloud-side device in the cloud in Figure 1A above can be a working node or a scheduling node in the cloud system. After the scheduling node receives the image from the camera, the scheduling node can perform the corresponding image processing process. The scheduling node can also schedule the images of the four cameras to a working node in the cloud system, and the working node can perform the corresponding image processing process.
[0080] The function of the scheduling node can be implemented by software or hardware.
[0081] As an example of a software functional unit, a scheduling node may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, a scheduling node may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.
[0082] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Inter-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0083] As an example of a hardware functional unit, a scheduling node may include at least one computing device, such as a server. Alternatively, the scheduling node may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0084] The multiple computing devices included in a scheduling node can be distributed in the same zone or in different zones. The multiple computing devices included in a scheduling node can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in a scheduling node can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0085] A worker node can be a physical machine, a virtual machine (VM), or a container. It can include one or more central processing units (CPUs) and graphics processing units (GPUs). A worker node can also be a CPU or a GPU.
[0086] The cloud in Figure 1A above may be located in a cloud system. The architecture of the cloud system can be understood with reference to Figure 1C . As shown in Figure 1C , the cloud system includes a cloud platform and basic resources. The cloud platform includes a cloud platform manager. The scheduling node described above may be the cloud platform manager in Figure 1C . Basic resources may include multiple servers, each of which may be a worker node, or each server may include multiple worker nodes.
[0087] The working node in Figure 1C can be a computing device card or a virtual machine (VM). The computing device card can be at least one of a central processing unit (CPU), a graphics processing unit (GPU), and a neural network processing unit (NPU).
[0088] The cloud platform manager maintains or regularly collects information about each worker node in the underlying resources, such as resource usage (resource utilization rate or resource idle rate) on each worker node. This information can be used as auxiliary decision-making information when scheduling images.
[0089] The cloud platform manager can receive images from the camera, and then the cloud platform manager can perform the corresponding image processing process. The cloud platform manager can also dispatch the images of the four cameras in Figure 1A to a working node, and the working node can perform the corresponding image processing process.
[0090] In the scenarios of the image processing systems described in Figures 1A to 1C above, whether image processing is performed on the cloud, on the roadside, or through a combination of end and cloud, the image processing systems described above can include an electronic device, which can be a cloud-side device or a roadside device. When the electronic device is a cloud-side device, after obtaining the image taken by the camera from the roadside, the subsequent image processing process can be completed. When the electronic device is a roadside device, and the roadside device is a certain camera, after the roadside device obtains the image from other cameras, it can perform subsequent image processing on the images of other cameras and the images taken by itself. If the roadside device is a dedicated image processing device, then after the roadside device obtains the image taken by the camera, the subsequent image processing process can be completed.
[0091] The structure of the electronic device can be understood by referring to FIG1D . In one possible embodiment, as shown in FIG1D , the electronic device 1000 may include: a central processing unit 1001, a graphics processing unit 1002, a display device 1003 (optional), and a memory 1004. Optionally, the electronic device 1000 may further include at least one communication bus (not shown in FIG1D ) for enabling connection and communication between various components.
[0092] It should be understood that the various components in the electronic device 1000 may also be coupled via other connectors, which may include various interfaces, transmission lines, or buses. The various components in the electronic device 1000 may also be connected in a radial manner centered around the central processing unit 1001. In various embodiments of the present application, coupling refers to mutual electrical connection or communication, including direct connection or indirect connection through other devices.
[0093] There are many ways to connect the CPU 1001 and the GPU 1002, and they are not limited to the one shown in Figure 1D. The CPU 1001 and the GPU 1002 in the electronic device 1000 can be located on the same chip or on separate chips.
[0094] The functions of the central processing unit 1001 , the graphics processing unit 1002 , the display device 1003 and the memory 1004 are briefly introduced below.
[0095] Central processing unit 1001: used to run operating system 1005 and application 1006. Application 1006 can be a graphics application, such as a video player. Operating system 1005 provides a system graphics library interface. Application 1006 generates an instruction stream for rendering graphics or image frames and the required related rendering data through this system graphics library interface and the driver provided by operating system 1005, such as the graphics library user-mode driver and / or the graphics library kernel-mode driver. The system graphics library includes, but is not limited to, system graphics libraries such as OpenGL ES (Open Graphics Library for Embedded Systems), the Khronos platform graphics interface, or Vulkan (a cross-platform graphics application interface). The instruction stream contains a series of instructions, which are generally calls to the system graphics library interface.
[0096] Optionally, the central processing unit 1001 may include at least one of the following types of processors: an application processor, one or more microprocessors, a digital signal processor (DSP), a microcontroller unit (MCU), or an artificial intelligence processor.
[0097] The CPU 1001 may further include necessary hardware accelerators, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or an integrated circuit for implementing logic operations. The CPU 1001 may be coupled to one or more data buses for transmitting data and instructions between the various components of the electronic device 1000.
[0098] Graphics processor 1002: Receives the graphics instruction stream sent by processor 1001, generates a rendering target through a rendering pipeline, and displays the rendering target on display device 1003 through the layer synthesis display module of the operating system. A rendering pipeline, which may also be referred to as a rendering pipeline, pixel pipeline, or pixel pipeline, is a parallel processing unit within graphics processor 1002 for processing graphics signals. Graphics processor 1002 may include multiple rendering pipelines, which can process graphics signals independently and in parallel. For example, a rendering pipeline can perform a series of operations in the process of rendering graphics or image frames. Typical operations may include vertex processing, primitive processing, rasterization, fragment processing, and so on.
[0099] Alternatively, the graphics processor 1002 may include a general-purpose graphics processor that executes software, such as a GPU or other type of dedicated graphics processing unit.
[0100] Display device 1003 : used to display various images generated by electronic device 1000 , which may be a graphical user interface (GUI) of an operating system or image data (including still images and video data) processed by graphics processor 1002 .
[0101] Optionally, the display device 1003 may include any suitable type of display screen, such as a liquid crystal display (LCD), a plasma display, or an organic light-emitting diode (OLED) display.
[0102] The memory 1004 is a transmission channel between the CPU 1001 and the GPU 1002 and can be a double data rate synchronous dynamic random access memory (DDR SDRAM) or other types of cache.
[0103] The structure of the above-mentioned cloud device or roadside device can also be understood by referring to Figure 2. Figure 2 is a possible logical structure diagram of the cloud device / roadside device provided in an embodiment of the present application. As shown in Figure 2, the cloud device / roadside device 20 provided in an embodiment of the present application includes: a processor 201, a communication interface 202, a memory 203 and a bus 204. The processor 201, the communication interface 202 and the memory 203 are interconnected via the bus 204. In an embodiment of the present application, the processor 201 is used to control and manage the actions of the cloud device / roadside device 20. For example, the processor 201 is used to splice images from multiple cameras to obtain a spliced image. The communication interface 202 is used to support the cloud device 20 to communicate. For example, the communication interface 202 can receive images transmitted from the camera and send spliced images. The memory 203 is used to store program code and data of the cloud device / roadside device 20.
[0104] The processor 201 may be a central processing unit (CPU), a general-purpose processor (GPOR), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device (PLD), a transistor logic device (TLD), a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. A processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like. The bus 204 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, for example. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG. 2 shows only one thick line, but this does not imply that there is only one bus or only one type of bus.
[0105] The following describes the image processing method provided in an embodiment of the present application. The content of the method involving cloud execution can be executed by the cloud or by components of the cloud (such as a processor, chip, or chip system). The content involving roadside device execution can be executed by the roadside device or by components of the roadside device (such as a processor, chip, or chip system).
[0106] The image processing method provided in the embodiment of the present application uses a raster image, which can be generated in the cloud or on a separate server. The process of generating a raster image can be understood by referring to Figure 3. The process includes:
[0107] 301. Collect data of the target scene and build a three-dimensional model of the target scene.
[0108] The target scene data may include an image of the target scene captured by a panoramic camera.
[0109] 302. Render the three-dimensional model at different positions of the virtual camera to obtain multiple raster images.
[0110] The raster image includes marker features, and the marker features include at least one of lane line features, building features, vehicle features, pedestrian features, and guardrail features.
[0111] 303. Extract the features of the markers in each raster image.
[0112] At least one of lane line features, building features, vehicle features, pedestrian features, and guardrail features of various shapes can be extracted from the raster image through a deep learning algorithm.
[0113] 304. Construct the correspondence between each raster image, the corresponding marker feature and the position of the virtual camera.
[0114] This correspondence can be understood by referring to Figure 4. In Figure 4, taking the lane line feature as an example, as shown in Figure 4, there are three virtual cameras, each of which renders three raster images at different positions. By extracting the lane line features in the raster images, the symbols of the lane line features can be extracted.
[0115] In this application, the raster image, the corresponding lane features, and the virtual camera's position are also referred to as a feature map. For example, for a 500-meter intersection, the original point cloud or high-precision map of this 500-meter intersection would occupy several GB of storage space. However, the feature map in this application only requires approximately 10 MB of storage space, significantly reducing storage space usage and allowing image processing to be performed roadside.
[0116] It should be noted that, in the present application, the above steps 301 to 304 only need to be executed once at the beginning. When the camera rotates or the position changes, there is no need to re-execute, and the raster image generated for the first time can still be used.
[0117] The raster image generated by the embodiments of the present application, the corresponding relationship between lane features and the virtual camera's position, or simply the feature map, can be stored in the cloud or on a roadside device. If the feature map is generated in the cloud, it can also be sent to the roadside device for use.
[0118] The following describes the image processing method provided in an embodiment of the present application with reference to the accompanying drawings. This process can be executed by a cloud device or a road-side device.
[0119] As shown in FIG5 , an embodiment of the image processing method provided in the embodiment of the present application includes:
[0120] 501. Obtain multiple images captured by multiple cameras associated with the target scene.
[0121] The target scene can be an intersection or a section of road. A target scene is usually configured with multiple cameras (also called cameras), wherein each camera captures images of the target scene in different directions and positions.
[0122] Each of the multiple images can be understood as an image taken by a different camera at the same time.
[0123] 502. Determine the first pose of at least two cameras among the multiple cameras based on the multiple images and the multiple raster images.
[0124] Multiple images and multiple raster images are associated with the same target scene, wherein the first pose of the first camera is related to the pose of the virtual camera corresponding to the first raster image, the first raster image is a raster image among the multiple raster images corresponding to the image taken by the first camera, and the first camera is any one of the multiple cameras.
[0125] Optionally, the process of determining the first pose in step 502 may be: performing marker feature matching on the multiple images and the corresponding raster images respectively to estimate the first pose of each camera in the multiple cameras.
[0126] The camera pose refers to the camera's position and orientation, which can include the position (x, y, z) in a three-dimensional coordinate system, as well as the azimuth, pitch, and roll angles. Typically, the camera pose is described by its rotation matrix (R) and translation vector (t), rather than by the position and angle coordinates described above. The virtual camera pose refers to the pose assumed when generating raster images.
[0127] In step 502, because the raster image and the image taken by the camera are from the same target scene, the image taken by the camera can be matched to a raster image with the highest similarity from multiple raster images, and then the first pose of the camera is determined by the raster image.
[0128] Taking any one of the multiple cameras as an example, call this camera the first camera, and match the image taken by the first camera with multiple raster images, it can be determined that the first raster image has the highest similarity with the image taken by the first camera. Then, the first position of the first camera can be determined by the position of the virtual camera corresponding to the first raster image.
[0129] The process of determining the first raster image from multiple raster images can be: taking the marker feature as an example, the lane line feature in the image of the first camera is matched with the lane line feature in each raster image for similarity, thereby determining that the first similarity between the first lane line feature in the first raster image and the second lane line feature in the image of the first camera is higher than the second similarity; wherein the second similarity is the similarity between the second lane line feature and the third lane line feature, and the third lane line feature is the lane line feature in the raster image other than the first raster image in the multiple raster images.
[0130] The process of determining the first pose of the first camera through the pose of the virtual camera corresponding to the first raster image can be: adjusting the pose of the virtual camera corresponding to the first raster image according to the optimization target to obtain the first pose of the first camera, the optimization target is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization target is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
[0131] The process can be: through the residual between the lane line features in the image of the first camera and the first grid image, based on the position of the virtual camera corresponding to the first grid image, back propagation optimization is performed to make the lane line features between the image of the first camera and the first grid image as similar as possible, or to make the distance between the lane line features in space as close as possible. At this time, the output position of the first camera is the first position of the first camera. This process can be understood with reference to Figure 6. As shown in Figure 6, the first grid image 601 and the image 602 taken by the first camera are processed by a convolutional neural network (CNN) respectively to extract the lane line features in the image, and then the position (R r , t r ) is substituted into the residual formula below as the initial value, and the first pose (R, t) of the first camera is obtained by backpropagation optimization (Backward).
[0132] Among them, residual represents the residual, R r represents the rotation matrix of the virtual camera corresponding to the first raster image 601, t rrepresents the translation vector of the virtual camera corresponding to the first raster image 601, F q represents the lane line features of the first camera image, F r Indicates the lane line feature corresponding to the first raster image, s i Indicates the scaling factor.
[0133] 503. Stitching multiple images based on the estimated first poses of at least two cameras.
[0134] This process can be done by stitching multiple images using the first pose of each camera, or by stitching multiple images using the first pose of some cameras. For example, stitching four images taken by four cameras using the first poses of three cameras.
[0135] In the embodiments of the present application, by matching lane line features between the camera image and the raster image, the first camera pose can be obtained from the pose of the virtual camera corresponding to the raster image, thereby achieving the stitching of multiple images taken by multiple cameras. Because the data volume of raster images is very small, the storage space occupied can be reduced. In addition, the raster image mainly includes lane line features and has fewer other features, so it can achieve rapid matching with the camera image, thereby improving the efficiency and accuracy of image stitching.
[0136] Optionally, in an embodiment of the present application, the first pose can be optimized to obtain a second pose with higher accuracy, and then multiple images can be stitched together based on the second pose of each camera.
[0137] The process of optimizing the first pose to obtain the second pose may include: generating a rendered image of the target scene in the corresponding first pose based on the first pose of each camera; matching the lane line features in the image of each camera with the lane line features in the corresponding rendered image. The matching process may be to use the feature points of the lane lines in the image of each camera as guide points, determine the feature points of the lane lines corresponding to the guide points in the corresponding rendered image, and solve them through a (perspective-n-point, PnP) algorithm to obtain the second pose, so as to obtain the second pose of each camera.
[0138] Among them, the PnP algorithm is a method for solving the correspondence between 3D and 2D points. When multiple 3D spatial points and their positions in an image are known, the position of the camera that took the image can be estimated.
[0139] The above process of optimizing the first pose to obtain the second pose can also be understood by referring to Figure 7. As shown in Figure 7, taking the first camera as an example, according to the first pose of the first camera, a rendered image 701 of the target scene is generated at the corresponding first pose, and then the feature points of the lane lines in the rendered image 701 can be extracted through a deep neural network or a convolutional neural network. These feature points are used as guide points 702, and then the feature points 704 of the lane lines corresponding to each guide point 702 are searched from the image 703 taken by the first camera. Then, the multiple feature points 704 are solved by the PnP algorithm to determine a more accurate second pose of the first camera. At the same time, the algorithms involved in the above image processing process can be pruned, so that the efficiency of image processing can be improved.
[0140] In the embodiment of the present application, in the process of optimizing the first pose to obtain the second pose, it is only necessary to determine the feature points of the lane line in the camera image, and then use these feature points as guide points to match them with the corresponding feature points in the corresponding rendered image. By solving the PnP algorithm, the second pose can be obtained. There is no need for global matching, only local matching of the feature points of the lane line is required, which can reduce the amount of calculation and improve matching efficiency.
[0141] After obtaining the second pose of the camera through the above optimization, the second pose can be used to perform image stitching according to the stitching relationship. Among them, the stitching relationship can be a mapping relationship between the camera identifier and the stitching area identifier. One camera can correspond to a stitching area on the display area. By placing the camera image into the corresponding stitching area, image stitching can be achieved. This process can be understood with reference to Figure 8. As shown in Figure 8, the display area includes four stitching areas, namely stitching area 1, stitching area 2, stitching area 3, and stitching area 4. Among them, stitching area 1 corresponds to camera 3, stitching area 2 corresponds to camera 4, stitching area 3 corresponds to camera 1, and stitching area 4 corresponds to camera 2. In this way, the image from camera 1 can be mapped to stitching area 3, the image from camera 2 can be mapped to stitching area 4, the image from camera 3 can be mapped to stitching area 1, and the image from camera 4 can be mapped to stitching area 2, thereby completing the stitching of the images from the four cameras.
[0142] The mapping relationship between the camera identifiers and the stitching area identifiers can be achieved by determining the camera to which the corresponding 3D point in the target scene space of each pixel on the display area is projected using the second pose of each camera and the distortion parameters of each camera. Furthermore, the mapping relationship between the camera and the stitching area is determined based on the stitching area to which each pixel belongs. In other words, the correspondence between the cameras and the stitching area shown in FIG8 is determined using the second pose.
[0143] The stitching principle described in Figure 8 is a simple illustration. In practice, images from different cameras may overlap during stitching. To improve image stitching quality, the following process can be used: when a target pixel in the target scene's spatial 3D point is projected onto one camera, the pixel information corresponding to the target pixel in the target image of one camera is rendered onto the target pixel; when the target pixel in the target scene's spatial 3D point is projected onto at least two cameras, the pixel information corresponding to the target pixel in the target image of the target camera is rendered onto the target pixel. The target camera is the one with the smallest distance to the target pixel's spatial 3D point in the target scene among the at least two cameras. This eliminates duplication and improves image stitching quality.
[0144] Regarding the situation when the 3D point corresponding to the target pixel point in the space of the target scene is projected onto at least two cameras, please refer to Figure 9 for understanding. As shown in Figure 9, the pixel point 901 in the stitching area 1 is projected onto camera 3 and camera 4 at the 3D point 902 in space. The distance between the 3D point 902 and camera 3 is farther than the distance to camera 4. Then, the pixel information corresponding to the pixel point 901 in the image taken by camera 4 is rendered for the pixel point 901.
[0145] The scene diagram of image stitching can be understood by referring to FIG10. As shown in FIG10, four images taken by four cameras at different angles, after the determination of the first pose, optimization of the second pose, and the image stitching process described above, can obtain a stitched image 100 of the intersection. The stitched image 100 can be displayed on the display of the road management center, which is conducive to quickly checking the traffic conditions of the intersection.
[0146] The above description is about the splicing of images at a certain point in time. If the spliced images at a certain time are continuously displayed on the display, the traffic situation at the intersection can be viewed in the form of a video.
[0147] It should be noted that, in the embodiment of the present application, after the first posture or the second posture is determined, the posture of the camera can be adjusted according to the first posture or the second posture.
[0148] In addition, in an embodiment of the present application, accurate stitching of images from different cameras can be achieved according to the above process without knowing the camera's position in advance. Even if the camera rotates or moves after taking the image at the first moment, the above cloud device or roadside device can still perform the image processing process described above on the image taken by the camera at the second moment to complete the stitching of the image at the second moment.
[0149] The above describes the process of stitching images from different cameras based on the camera's first pose or the camera's second pose. In fact, the image processing method provided in the embodiment of the present application can also include: estimating the first pose of the first camera using the image from the first camera; generating a rendered image of the target scene at the first pose, where the target scene is the scene associated with the first camera; and matching the lane line features in the image from the first camera with the lane line features in the rendered image to optimize the first pose to obtain the second pose of the first camera. In this way, by optimizing the first pose to obtain the second pose, the pose accuracy can be improved.
[0150] In other words, the solution provided in this application does not limit the number of cameras in the target scene, and posture optimization can also be performed with only one camera. The posture optimization process can be understood by referring to the content introduced in Figure 7 above, and will not be repeated here.
[0151] In addition, the processes of determining the first pose, optimizing the first pose to obtain the second pose, and image stitching described in the previous Figures 3 to 10 can all be applied to this application.
[0152] The above describes the image processing method. The following describes the image processing device provided in the embodiment of the present application with reference to the accompanying drawings.
[0153] As shown in FIG11 , the image processing apparatus 110 provided in an embodiment of the present application includes:
[0154] The acquisition unit 1101 is configured to acquire multiple images captured by multiple cameras.
[0155] A first processing unit 1102 is configured to determine a first pose of at least two cameras among the multiple cameras based on the multiple images and the multiple raster images, where the multiple images and the multiple raster images are associated with the same target scene, wherein the first pose of the first camera is related to the pose of a virtual camera corresponding to the first raster image, the first raster image is a raster image among the multiple raster images that corresponds to an image captured by the first camera, and the first camera is any one of the multiple cameras.
[0156] The second processing unit 1103 is configured to stitch multiple images together according to the first poses of at least two cameras.
[0157] Optionally, the first processing unit 1102 is specifically configured to perform marker feature matching on the multiple images and the corresponding raster images respectively, so as to estimate the first pose of each camera in the multiple cameras.
[0158] Optionally, the second processing unit 1103 is specifically configured to stitch the multiple images together according to the estimated first pose of each camera.
[0159] Optionally, the marker feature may include at least one of a lane feature, a building feature, a vehicle feature, a pedestrian feature, and a guardrail feature.
[0160] Optionally, the first processing unit 1101 is used to perform lane line feature matching on the first lane line feature associated with the first raster image and the second lane line feature in the image of the first camera; adjust the posture of the virtual camera corresponding to the first raster image according to the optimization target to obtain the first posture of the first camera, and the optimization target is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization target is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
[0161] Optionally, the second processing unit 1102 is configured to optimize the estimated first pose of each camera to obtain a second pose of each camera; and stitch the multiple images according to the second pose of each camera.
[0162] Optionally, the second processing unit 1102 is used to generate a rendered image of the target scene in the corresponding first pose according to the first pose of each camera; and match the lane line features in the image of each camera with the lane line features in the corresponding rendered image to obtain the second pose of each camera.
[0163] Optionally, the second processing unit 1102 is configured to use the feature points of the lane lines in the image of each camera as guidance points, determine the feature points of the lane lines corresponding to the guidance points in the corresponding rendered image, and perform matching.
[0164] Optionally, the second processing unit 1102 is configured to determine a stitching relationship between the multiple images according to the second posture of each camera; and stitch the multiple images according to the stitching relationship.
[0165] Optionally, the first processing unit 1101 is further configured to adjust the first camera according to the first pose of the first camera.
[0166] Optionally, the image from the first camera is an image taken by the first camera at a first moment, and the first processing unit 1101 is further used to obtain an image taken by the first camera at a second moment, where the second moment is later than the first moment; and perform lane line feature matching on the image taken by the first camera at the second moment with the second grid image to estimate the posture of the first camera at the second moment, and the posture of the first camera at the second moment is different from the first posture of the first camera, indicating that the first camera has rotated or moved.
[0167] The functions of the various units in the image processing device 110 introduced in this application can be understood by referring to the corresponding contents in the method embodiment introduced above, and will not be repeated here.
[0168] As shown in FIG12 , another image processing apparatus 120 provided in an embodiment of the present application includes:
[0169] The first processing unit 1201 is configured to estimate the first pose of the first camera using an image from the first camera.
[0170] The second processing unit 1202 is configured to generate a rendered image of a target scene at a first pose, where the target scene is the scene associated with the first camera.
[0171] The third processing unit 1203 is configured to match the image rendered from the first camera to optimize the first pose to obtain a second pose of the first camera.
[0172] Optionally, the image from the first camera is included in multiple images, each of the multiple images comes from a different camera associated with the target scene, the first camera is any one of the multiple cameras associated with the target scene, and the third processing unit is further used to stitch the multiple images according to the second pose of each camera.
[0173] Optionally, the first processing unit 1201 is used to match lane line features between the image from the first camera and the corresponding first raster image to estimate the first pose of the first camera; wherein the first pose of the first camera is determined based on the pose of the virtual camera corresponding to the first raster image, and the first raster image is the raster image with the highest similarity to the image of the first camera among multiple raster images associated with the target scene.
[0174] Optionally, the marker feature includes at least one of a lane feature, a building feature, a vehicle feature, a pedestrian feature, and a guardrail feature.
[0175] Optionally, the first processing unit 1201 is used to perform lane line feature matching on a first lane line feature associated with the first raster image and a second lane line feature in the image of the first camera; and adjust the position of the virtual camera corresponding to the first raster image according to the optimization target to obtain the first position of the first camera, wherein the optimization target is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization target is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
[0176] Optionally, the third processing unit 1203 is configured to use the feature points of the lane lines in the image of the first camera as the guiding points, determine the feature points of the lane lines corresponding to the guiding points in the corresponding rendered image, and perform matching.
[0177] Optionally, the third processing unit 1203 is configured to determine a stitching relationship between the multiple images according to the second posture of each camera; and stitch the multiple images according to the stitching relationship.
[0178] Optionally, the third processing unit 1203 is further configured to adjust the first camera according to the second posture of the first camera.
[0179] Optionally, the image from the first camera is an image taken by the first camera at a first moment, and the first processing unit 1201 is further used to obtain an image taken by the first camera at a second moment, where the second moment is later than the first moment; and perform lane line feature matching on the image taken by the first camera at the second moment with the second grid image to estimate the posture of the first camera at the second moment, and the posture of the first camera at the second moment is different from the first posture of the first camera, indicating that the first camera has rotated or moved.
[0180] The functions of the various units in the image processing device 120 introduced in this application can be understood by referring to the corresponding contents in the method embodiment introduced above, and will not be repeated here.
[0181] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores a program, which, when executed on a computer, enables the computer to execute the method described in the embodiments shown in the aforementioned Figures 3 to 10.
[0182] An embodiment of the present application also provides an image processing device, which can also be called a digital processing chip or chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is used to execute the method shown in any of the embodiments in Figures 3 to 10 above.
[0183] The present application also provides a digital processing chip. The digital processing chip integrates circuitry and one or more interfaces for implementing the aforementioned processor 201 or the functions of processor 201. When the digital processing chip integrates memory, it can perform the method steps of any one or more of the aforementioned embodiments. When the digital processing chip does not integrate memory, it can be connected to an external memory via a communication interface. The digital processing chip implements the methods of the aforementioned embodiments based on program code stored in the external memory.
[0184] An embodiment of the present application also provides a computer-readable storage medium, which stores instructions. When the instructions are executed on an electronic device, the electronic device executes the steps in the method described in any of the embodiments in Figures 3 to 10 above.
[0185] An embodiment of the present application also provides a computer program product, which, when executed on a computer, enables the computer to execute the steps of the method described in any one of the embodiments in FIG. 3 to FIG. 10 .
[0186] The image processing device provided in the embodiments of the present application may be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, or a circuit. The processing unit may execute computer-executable instructions stored in the storage unit, so that the chip in the computer device executes the image processing method described in the embodiments shown in Figures 3 to 10 above. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0187] Specifically, the aforementioned processing unit or processor may include a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0188] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0189] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general-purpose hardware, and of course can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., including a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0190] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0191] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a server, or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
Claims
1. A method for image processing, characterized in that: include: Get multiple images taken by multiple cameras; Determining first poses of at least two cameras among the multiple cameras according to the multiple images and the multiple raster images, wherein the multiple images and the multiple raster images are associated with the same target scene, wherein the first pose of the first camera is related to the pose of a virtual camera corresponding to the first raster image, the first raster image is a raster image among the multiple raster images corresponding to the image taken by the first camera, and the first camera is any one of the multiple cameras; The multiple images are stitched according to the first poses of the at least two cameras.
2. The method according to claim 1, characterized in that The step of determining the first poses of at least two cameras among the plurality of cameras according to the plurality of images and the plurality of grid images comprises: Matching the multiple images with the corresponding raster images respectively to estimate the first pose of each camera in the multiple cameras; Correspondingly, stitching the multiple images according to the first poses of the at least two cameras includes: The multiple images are stitched together according to the estimated first pose of each camera.
3. The method according to claim 2, characterized in that The marker features include at least one of lane features, building features, vehicle features, pedestrian features, and guardrail features.
4. The method according to claim 3, characterized in that The method of matching the multiple images with the corresponding grid images for the features of the markers to estimate the first position of each camera in the multiple cameras includes: Performing lane line feature matching on a first lane line feature associated with the first raster image and a second lane line feature in the image of the first camera; The pose of the virtual camera corresponding to the first raster image is adjusted according to an optimization goal to obtain a first pose of the first camera, wherein the optimization goal is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization goal is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
5. The method according to any one of claims 2 to 4, characterized in that: The step of stitching the multiple images according to the estimated first pose of each camera comprises: Optimizing the estimated first pose of each camera to obtain a second pose of each camera; The multiple images are stitched together according to the second posture of each camera.
6. The method according to claim 5, characterized in that The step of optimizing the estimated first pose of each camera to obtain a second pose of each camera includes: Generating a rendered image of the target scene at the corresponding first pose according to the first pose of each camera; The lane line features in the image of each camera are matched with the lane line features in the corresponding rendered image to obtain the second pose of each camera.
7. The method according to claim 6, characterized in that The matching of the lane line features in the image of each camera with the lane line features in the corresponding rendered image includes: The feature points of the lane lines in the images of each camera are used as guide points, and the feature points of the lane lines corresponding to the guide points are determined in the corresponding rendered images and matched.
8. The method according to any one of claims 5 to 7, characterized in that: The step of stitching the multiple images according to the respective second poses of each camera comprises: Determining a stitching relationship between the multiple images according to the respective second posture of each camera; The multiple images are stitched together according to the stitching relationship.
9. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: The first camera is adjusted according to the first pose of the first camera.
10. The method according to any one of claims 2 to 8, characterized in that: The image from the first camera is an image taken by the first camera at a first moment, and the method further includes: Acquire an image captured by the first camera at a second moment, where the second moment is later than the first moment; The image captured by the first camera at the second moment is matched with the second grid image to estimate the lane line feature of the first camera. At the position and posture of the first camera at the second moment, the position and posture of the first camera at the second moment is different from the first position and posture of the first camera, so as to indicate that the first camera is rotated or moved.
11. A method for image processing, characterized in that: include: estimating a first pose of the first camera using an image from the first camera; generating a rendered image of a target scene at the first pose, the target scene being a scene associated with the first camera; The image from the first camera is matched with the rendered image to optimize the first pose to obtain a second pose of the first camera.
12. The method according to claim 11, characterized in that The image from the first camera is included in a plurality of images, each of the plurality of images is from a different camera associated with the target scene, and the first camera is any one of the plurality of cameras associated with the target scene; The method further comprises: The multiple images are stitched together according to the second posture of each camera.
13. The method according to claim 11 or 12, characterized in that: The step of estimating the first pose of the first camera by using an image from the first camera comprises: The image from the first camera is matched with the corresponding first raster image for marker features to estimate the first pose of the first camera; wherein the first pose of the first camera is related to the pose of the virtual camera corresponding to the first raster image, and the first raster image is a raster image among multiple raster images corresponding to the image taken by the first camera.
14. The method according to claim 13, characterized in that The marker features include at least one of lane features, building features, vehicle features, pedestrian features, and guardrail features.
15. The method according to claim 14, characterized in that The step of performing marker feature matching on the image from the first camera and the corresponding first grid image to estimate the first pose of the first camera includes: Performing lane line feature matching on a first lane line feature associated with the first raster image and a second lane line feature in the image of the first camera; The pose of the virtual camera corresponding to the first raster image is adjusted according to an optimization goal to obtain a first pose of the first camera, wherein the optimization goal is that the similarity between the first lane line feature and the second lane line feature is greater than a first threshold, or the optimization goal is that the distance between the first lane line feature and the second lane line feature is less than a second threshold.
16. The method according to any one of claims 11 to 15, characterized in that: The matching the image from the first camera with the rendered image comprises: The feature points of the lane lines in the image of the first camera are used as guide points, and the feature points of the lane lines corresponding to the guide points are determined in the corresponding rendered image and matched.
17. The method according to claim 12, characterized in that The step of stitching the multiple images according to the respective second poses of each camera comprises: Determining a stitching relationship between the multiple images according to the respective second posture of each camera; The multiple images are stitched together according to the stitching relationship.
18. The method according to any one of claims 11 to 17, characterized in that: The method further comprises: The first camera is adjusted according to the second position and posture of the first camera.
19. The method according to any one of claims 11 to 17, characterized in that: The image from the first camera is an image taken by the first camera at a first moment, and the method further includes: Acquire an image captured by the first camera at a second moment, where the second moment is later than the first moment; Lane line features are matched between the image captured by the first camera at the second moment and the second grid image to estimate the position and posture of the first camera at the second moment. The position and posture of the first camera at the second moment is different from the first position and posture of the first camera to indicate that the first camera is rotated or moved.
20. An image processing device, characterized in that: include: An acquisition unit, used for acquiring a plurality of images taken by a plurality of cameras; a first processing unit, configured to determine first poses of at least two cameras among the plurality of cameras according to the plurality of images and the plurality of raster images, wherein the plurality of images and the plurality of raster images are associated with the same target scene, wherein the first pose of the first camera is associated with the first raster image; The first grid image is related to the posture of a virtual camera corresponding to the grid image, the first grid image is a grid image among the multiple grid images corresponding to the image taken by the first camera, and the first camera is any one of the multiple cameras; The second processing unit is used to stitch the multiple images according to the first poses of the at least two cameras.
21. An image processing device, characterized in that: include: A first processing unit, configured to estimate a first pose of the first camera using an image from the first camera; a second processing unit, configured to generate a rendered image of a target scene in the first pose, wherein the target scene is a scene associated with the first camera; A third processing unit is used to match the image from the first camera with the rendered image to optimize the first posture to obtain a second posture of the first camera.
22. An electronic device, characterized in that: include: A communication interface, a processor and a memory, wherein the communication interface and the processor are coupled to the memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the electronic device executes the method as described in any one of claims 1 to 10, or executes the method as described in any one of claims 11 to 19.
23. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on an electronic device, the electronic device executes the method as claimed in any one of claims 1 to 10, or executes the method as claimed in any one of claims 11 to 19.
24. A computer program product, characterized in that The computer program product comprises a computer program code, and when the computer program code is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 10, or execute the method according to any one of claims 11 to 19.
25. A chip system, characterized in that: The chip system includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected through lines; the interface circuits are used to receive signals from the memory of the electronic device and send signals to the processor, and the signals include computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the method as described in any one of claims 1 to 10, or executes the method as described in any one of claims 11 to 19.
26. An image processing system, characterized in that: include: An electronic device, wherein the electronic device is used to execute the method as described in any one of claims 1 to 10, or execute the method as described in any one of claims 11 to 19, and the electronic device is a cloud-side device or a road-side device.
Citation Information
Patent Citations
Panoramic video stitching method and device for keeping pedestrian integrity
CN114022562A
Multi-sensor calibration method and device based on feature map
CN114202577A
Image splicing method and device and related product
CN115375594A
Image splicing method and device, electronic equipment and storage medium
CN116402687A
Auto water blocking apparatus
KR1020240120258A