Method and system for associating two or more images
By using a projection matrix that approximates polynomial equations, the problem of pixel mapping between fisheye images is solved, and automated object associations in multi-objective multi-camera tracking is achieved, reducing computational complexity and dependence on calibration parameters.
Patent Information
- Application Number
- CN202380070855.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-23
- Filing Date
- 2023-09-14
- Publication Date
- 2025-05-13
AI Technical Summary
Prior art When processing fisheye images, it is difficult to achieve pixel-to-pixel mapping between two or more images, resulting in distortion and computational complexity in multi-objective multi-camera tracking.
The object association between two or more images is achieved by using a projection matrix, especially an approximation of the projection matrix through a polynomial equation. The method identifies corresponding points between images and approximates the projection matrix based on these points, thereby achieving cross-camera tracking of objects.
This method realizes image mapping and object association without human intervention, avoids dependence on calibration parameters, reduces computational complexity, and does not require significant image overlap areas.
Smart Images

Figure CN119998847A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates generally to computer vision and, more particularly, to a method and system for associating objects between two or more images captured by separate image sensors. Background Art
[0002] In many computer vision applications, it is necessary to track multiple objects across multiple cameras, a process known as multi-target multi-camera (MTMC) tracking. MTMC tracking is used in several applications, including driving, where multiple traffic participants (such as vehicles and pedestrians) as well as infrastructure are tracked across multiple cameras. In many MTMC tracking methods, wide-angle or ultra-wide field of view cameras (also known as fisheye cameras) are preferred because they have a wider field of view compared to rectilinear cameras, and therefore images captured by fisheye cameras (also known as fisheye images) contain more information than rectilinear images. The basic requirement for MTMC tracking is to represent the scene captured by multiple cameras in a common coordinate system. Such a representation requires a pixel-to-pixel mapping between the two images, which is particularly difficult for fisheye images due to the distortion present in such fisheye images.
[0003] Current methods for solving the problems associated with pixel-to-pixel mapping of fisheye images use image correction or image registration, both of which have several disadvantages. Image correction involves aligning the image planes of two fisheye cameras using camera calibration parameters, which is undesirable because the camera calibration parameters are difficult to calculate, require human intervention to reliably detect calibration points, and must be recalculated for each new setting and periodically recalculated due to potential shifts in camera positions or lens aberrations. In addition, efficient image correction requires a large overlap area. On the other hand, image registration involves identifying corresponding points (also called key points) between the two image planes after distortion correction, and performing pixel-to-pixel mapping between the images using the transformation calculated between the identified corresponding points. Affine transformations are usually used to approximate the geometric transformation between two image planes, but because affine transformations approximate rigid body transformations, affine transformations cannot be applied to non-rigid transformations between the image planes of a rectilinear camera and the image plane of a fisheye camera or between the image planes of two fisheye cameras. In addition, efficient image registration requires a large overlap area, and key point matching between two images is also challenging because the camera views of the same object are different. Summary of the invention
[0004] Embodiments of the present invention improve the mapping of images by associating between two or more images using a projection matrix for transforming between two image planes or images. In some embodiments, the projection matrix can use polynomial equations and coefficients to approximate the non-rigid transformation between them. The projection matrix can be used for pixel-to-pixel mapping between images to facilitate object association between images, thereby enabling tracking of multiple objects across two or more cameras, which can be used for subsequent applications such as driving, computer vision applications, and security or surveillance applications.
[0005] In order to solve the above technical problems, the present invention provides a vehicle, the vehicle including a system for associating objects on two or more images, wherein the vehicle system includes at least a first image sensor and a second image sensor, one or more processors, and a memory storing executable instructions for the one or more processors to execute, the executable instructions including instructions for executing a computer-implemented method, the computer-implemented method including: receiving a first image from the first image sensor and a second image from the second image sensor; identifying corresponding points between the first image and the second image; approximating a projection matrix between the first image and the second image based on the identified corresponding points; and At least one identified object in the first image is associated with a corresponding identified object in the second image based at least on the approximated projection matrix.
[0006] The vehicle system of the present invention performs a computer-implemented method that is fully automated and does not require human intervention. Associating objects on two or more images by identifying corresponding points between the images and approximating the projection matrix can be triggered automatically, thereby avoiding the need for manual triggering or recalculation. The computer-implemented method has several advantages over previous solutions. Unlike image correction, the computer-implemented method of the present invention does not rely on calibration parameters, which are labor-intensive and prone to errors when performed automatically. The computer-implemented method of the present invention is computationally more efficient because it does not require distortion correction. In addition, the computer-implemented method of the present invention is also advantageous because it does not require significant overlap between the field of view of the image sensor and / or the images captured by the image sensor.
[0007] A preferred vehicle of the present invention is a vehicle as described above, wherein the one or more processors and the memory storing executable instructions for execution by the one or more processors include instructions for performing a computer-implemented method comprising: receiving a first image from the first image sensor and a second image from the second image sensor; identifying corresponding points between the first image and the second image; approximating a projection matrix between the first image and the second image based on the identified corresponding points; associating at least one identified object in the first image with a corresponding identified object in the second image based at least on the approximated projection matrix; and The positions of the identified objects are plotted on a common image plane, wherein the common image plane preferably covers a 360° surround view of the objects around the vehicle.
[0008] An advantage of the above aspects of the invention is that one common image plane including the position of the identified object may be more representative of the position of the object than an image taken by a single fisheye camera with a large field of view which distorts the object and its position.
[0009] The above-described advantageous aspects of the inventive vehicle also apply to all aspects of the inventive computer-implemented method described below. All below-described advantageous aspects of the inventive computer-implemented method also apply to all aspects of the above-described vehicle of the invention.
[0010] The present invention also relates to a computer-implemented method for associating objects on two or more images, the method comprising: receiving a first image from a first image sensor and a second image from a second image sensor; identifying corresponding points between the first image and the second image; approximating a projection matrix between the first image and the second image based on the identified corresponding points; and At least one identified object in the first image is associated with a corresponding identified object in the second image based at least on the approximated projection matrix.
[0011] The computer-implemented method of the present invention is fully automated and does not require human intervention. Associating objects on two or more images by approximating the projection matrix can be automatically triggered, thereby avoiding the need for manual triggering or recalculation. In addition, the computer-implemented method of the present invention has several advantages over previous solutions. Unlike image correction, the computer-implemented method of the present invention does not rely on calibration parameters, which are labor-intensive and prone to errors when performed automatically. The computer-implemented method of the present invention is more computationally efficient because it does not require distortion correction. In addition, the computer-implemented method of the present invention is also advantageous because it does not require a significant overlap between the field of view of the image sensor and / or the images captured by the image sensor.
[0012] A preferred method of the present invention is a computer-implemented method as described above, wherein the first image sensor and / or the second image sensor is a fisheye camera, and / or wherein the first image sensor and the second image sensor are positioned so that there is an overlap between at least one first image of the scene captured by the first image sensor and at least one second image of the scene captured by the second image sensor.
[0013] An advantage of the above aspects of the invention is that the fisheye camera has a wider field of view. The wider the field of view, the greater the overlap area between the images generated by the cameras. This will enable fewer cameras to be used to cover a larger area. For example, fewer cameras can be used to obtain a 360° view around the vehicle, which may be useful for driving applications.
[0014] A preferred method of the present invention is a computer implemented method as described above or as preferably described above, wherein: The corresponding points are unique points, preferably points corresponding to vertices or intersections; and / or These corresponding points are key points.
[0015] An advantage of the above aspects of the invention is that unique points (such as points corresponding to vertices or intersections) and key points (unique points in an image) can be identified in an image regardless of orientation or distortion. Using points corresponding to vertices or intersections is advantageous because these points are well defined and easily detected, thereby ensuring that the same point is accurately detected and selected in both the first image and the second image.
[0016] A preferred method of the present invention is a computer-implemented method as described above or as preferably described above, wherein the keypoints are selected using a neural network, and preferably using a neural network comprising a convolutional neural network backbone and a keypoint extractor, wherein the convolutional neural network backbone preferably comprises a constrained deformable convolution module.
[0017] The advantages of the above aspects of the invention are that the use of a neural network to select keypoints allows non-linear and complex relationships to be learned and modeled and subsequently applied to new data sets or inputs. In addition, the neural network has the ability to self-learn and produce outputs that are not limited to the inputs provided. A convolutional neural network (CNN) backbone is preferred because CNN has high accuracy in image recognition and can automatically filter images and detect features without any human supervision, and a CNN with constrained deformable convolution is preferred because it takes into account distortion in fisheye images.
[0018] A preferred method of the present invention is a computer-implemented method as described above or as preferably described above, wherein the step of approximating the projection matrix comprises approximating at least one projection matrix comprising at least one polynomial equation, wherein preferably the degree of the at least one polynomial equation is n and the number of corresponding points identified is between n+1 and n+7, wherein n is an integer equal to or greater than 2, and / or wherein preferably the degree of the at least one polynomial equation is 2 and the number of corresponding points identified is between 3 and 9.
[0019] The above aspects of the present invention are advantageous because, unlike affine transformations in image registration, which are limited to rigid body transformations, the use of polynomial equations allows for approximation of non-rigid transformations and / or projections and takes into account nonlinear geometry or distortions present in fisheye images. Using a polynomial equation of degree n and having n+1 to n+7 corresponding points is advantageous because the approximated polynomial equation may be accurate and sufficient for its purpose without overfitting. Using a polynomial equation of degree 2 and having 3 to 9 corresponding points is also advantageous because in many cases a polynomial equation of degree 2 may be accurate enough for its purpose without overfitting.
[0020] A preferred method of the invention is a computer implemented method as described above or as preferably described above, wherein the projection matrix comprises two polynomial equations: a first polynomial equation for coordinates on the x-axis and a second polynomial equation for coordinates on the y-axis.
[0021] The advantage of the above aspect of the invention is that the accuracy of the method is improved due to the calculation of different polynomial equations for the x-axis and the y-axis. The different polynomial equations for the x-axis and the y-axis take into account the distortion of the image (especially in the fisheye image) along the x-axis and the y-axis.
[0022] A preferred method of the present invention is a computer-implemented method as described above or as preferably described above, wherein associating at least one identified object in the first image with a corresponding identified object in the second image comprises: identifying at least one object in the first image and at least one object in the second image; generating at least one first bounding box in the first image and generating at least one second bounding box in the second image, each of the at least one first bounding box and the at least one second bounding box representing a spatial location of the identified object; and Each first bounding box is associated with a corresponding second bounding box based at least on a relationship between a projection of the first bounding box and the corresponding second bounding box, wherein the projection is based on the approximated projection matrix.
[0023] An advantage of the above aspects of the present invention is that the approximated projection matrix allows the locations of multiple objects or bounding boxes to be accurately projected from a first image to a second image for comparison and correlation as a group or on a large scale.
[0024] A preferred method of the present invention is a computer-implemented method as described above or as preferably described above, wherein associating each first bounding box with a corresponding second bounding box comprises: projecting the at least one first bounding box onto the second image based on the approximated projection matrix to generate at least one projected first bounding box on the second image, wherein each projected first bounding box corresponds to one of the at least one first bounding boxes; determining a relationship between each projected first bounding box and each second bounding box; and Each first bounding box is associated with the corresponding projected first bounding box based on the determined relationship between the corresponding second bounding box.
[0025] An advantage of the above aspects of the present invention is that the association or matching between multiple objects or bounding boxes in a first image and multiple objects or bounding boxes in a second image is optimized as a group or on a large scale based on the locations of the multiple objects or bounding boxes.
[0026] A preferred method of the present invention is a computer-implemented method as described above or as preferably described above, wherein the relationship between each projected first bounding box and each second bounding box comprises: The overlap between each projected first bounding box and each second bounding box; and / or The distance between each projected first bounding box and each second bounding box, wherein the distance is preferably the distance between the center of each projected first bounding box and the center of each second bounding box.
[0027] An advantage of the above aspects of the present invention is that the positions of multiple objects or bounding boxes may be projected and matched simultaneously across two or more image sensors.
[0028] A preferred method of the present invention is a computer-implemented method as described above or as preferably described above, wherein associating each first bounding box with a corresponding second bounding box further comprises: identifying at least one feature of each identified object in the first image and the second image, wherein the at least one feature is preferably an appearance feature vector; and At least one feature of each identified object in the first image is compared to at least one feature of each identified object in the second image.
[0029] An advantage of the above aspect of the invention is that in addition to the position of the objects or bounding boxes, the visual features or appearance of the objects are compared to ensure improved accuracy of the computer-implemented method because the visual similarity of the objects is also considered during association. This is particularly advantageous in situations where there is a low overlap area between the first image and the second image.
[0030] A preferred method of the present invention is a computer-implemented method as described above or as preferably described above, wherein the method further includes drawing the position of the identified object on a common image plane, wherein the common image plane preferably covers a 360° surround view of the object surrounding the first image sensor, the second image sensor and the additional image sensors.
[0031] An advantage of the above aspects of the present invention is that mapping the position of objects on a common image plane provides an accurate representation of the spatial position of the objects within a scene. A common image plane including the positions of the identified objects may be more representative of the positions of the objects than an image captured by a single fisheye camera with a large field of view that distorts the objects and their positions.
[0032] A particularly preferred method of the invention is a computer implemented method as described above or as preferably described above, wherein: The first image sensor and the second image sensor are fisheye cameras and are positioned such that there is an overlap between at least one first image of a scene captured by the first image sensor and at least one second image of the scene captured by the second image sensor; The corresponding points are selected using a neural network comprising a convolutional neural network backbone and a keypoint extractor, wherein the neural network is trained on a scene classification dataset comprising images classified into scene categories, wherein each scene category comprises a sequence of images captured from a single camera capturing overlapping areas of the scene and a sequence of images captured in the same time period by other cameras having overlapping fields of view with the single camera, and wherein the convolutional neural network backbone comprises a constrained deformable convolution module; The approximated projection matrix includes a first polynomial equation for coordinates on the x-axis and a second polynomial equation for coordinates on the y-axis, wherein the first polynomial equation and the second polynomial equation are each of degree 2 and the number of identified corresponding points is 3; Associating at least one identified object in the first image with a corresponding identified object in the second image comprises: identifying at least one object in the first image and at least one object in the second image; generating at least one first bounding box in the first image and generating at least one second bounding box in the second image, each of the at least one first bounding box and the at least one second bounding box representing a spatial location of the identified object; projecting the at least one first bounding box onto the second image based on the approximated projection matrix to generate at least one projected first bounding box on the second image, wherein each projected first bounding box corresponds to one of the at least one first bounding boxes; determining an extent of overlap between each projected first bounding box and each second bounding box; and Each first bounding box is associated with a corresponding projected first bounding box based on a determined overlap range between the corresponding second bounding box.
[0033] The above-described advantageous aspects of the vehicle or computer-implemented method of the present invention also apply to all aspects of the training dataset of the present invention described below. All the below-described advantageous aspects of the training dataset of the present invention also apply to all aspects of the above-described vehicle or computer-implemented method of the present invention.
[0034] The invention further relates to a training data set for a neural network, in particular a neural network according to the invention, wherein the training data set comprises images classified into scene categories, wherein each scene category comprises: a sequence of images captured from a first image sensor capturing overlapping areas of a scene; Optionally, a sequence of images captured during the same time period by other image sensors having an overlapping field of view with the first image sensor; and Optionally, the image is changed from a sequence of images captured from the first image sensor and / or a sequence of images captured in the same time period by other image sensors having an overlapping field of view with the first image sensor.
[0035] The training dataset of the present invention is advantageous because it includes a large number of images for each scene category for training a neural network for keypoint recognition. The training dataset includes a wide variety of images for each scene category from limited images and limited image sensors.
[0036] The above-described advantageous aspects of the vehicle, computer-implemented method or training data set of the present invention also apply to all aspects of the system of the present invention described below. All the below-described advantageous aspects of the system of the present invention also apply to all aspects of the above-described vehicle, computer-implemented method or training data set of the present invention.
[0037] The present invention also relates to a system suitable for use in a vehicle, the system comprising at least a first image sensor and a second image sensor, one or more processors and a memory storing executable instructions for execution by the one or more processors, the executable instructions including instructions for performing a computer-implemented method according to the present invention.
[0038] The above-described advantageous aspects of the vehicle, computer-implemented method, training data set or system of the present invention also apply to all aspects of the computer program, machine-readable medium or data signal described below of the present invention. All the below-described advantageous aspects of the computer program, machine-readable medium or data signal of the present invention also apply to all aspects of the above-described vehicle, computer-implemented method, training data set or system of the present invention.
[0039] The present invention also relates to a computer program, a machine-readable medium or a data carrier signal, which includes instructions that, when executed on one or more processors, cause the one or more processors to perform a computer-implemented method according to the present invention. The machine-readable medium may include any medium and / or mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). The machine-readable medium may be any medium, such as, for example, a read-only memory (ROM); a random access memory (RAM); a universal serial bus (USB) stick; a compact disc (CD); a digital video disc (DVD); a data storage device; a hard disk; an electrical, acoustic, optical or other form of propagated signal (e.g., a digital signal, a data carrier signal, a carrier wave) or any other medium on which program elements as described above may be transmitted and / or stored.
[0040] As used in this summary, in the following description, in the following claims, and in the accompanying drawings, the term "scene" refers to a unique physical environment that can be captured by one or more image sensors. A scene can include one or more objects that can be visually captured by one or more image sensors, whether such objects are stationary or moving.
[0041] As used in this disclosure, in the following description, in the following claims, and in the accompanying drawings, the term "fisheye camera" refers to an image sensor, camera, and / or video camera equipped with a fisheye lens or wide-angle lens with a field of view of not less than 60 degrees, and the term "fisheye image" refers to an image captured or generated by a fisheye camera. Fisheye images can also be characterized as spherical or hemispherical images.
[0042] As used in this summary, in the following description, in the following claims, and in the accompanying drawings, the term "bounding box" refers to the bounding area of an object, and may include a bounding box, a bounding circle, a bounding ellipse, or any other suitable shape representing an object. A bounding box associated with an object may have a rectangular shape, a square shape, a polygonal shape, a patch shape, or any other suitable shape.
[0043] As used in this disclosure, in the following description, in the following claims and in the accompanying drawings, the term "vehicle" refers to any mobile agent capable of moving, including cars, trucks, buses, farm machinery, forklifts, robots, whether or not such mobile agent is capable of carrying or transporting goods, animals or people.
[0044] As used in this summary, in the following description, in the following claims, and in the accompanying drawings, the term "keypoint" refers to an area in an image that is particularly unique and identifies a unique feature. Keypoints are used to identify key areas of an object that are used as a basis for later matching and identifying that object in another image.
[0045] As used in this disclosure, in the following description, in the following claims, and in the accompanying drawings, the term "keypoint descriptor" or "local descriptor" refers to an image patch surrounding a keypoint that is a high-dimensional point in a feature space. A keypoint descriptor or local descriptor includes edge and / or color information that is invariant to small affine transformations and preserves spatial relationships. A keypoint descriptor or local descriptor may also include shape, texture, and / or semantic information.
[0046] As used in this disclosure, in the following description, in the following claims, and in the accompanying drawings, the term “feature” refers to a variable, attribute, property, or characteristic in a data set, and the term “feature map,” “activation map,” or “convolutional features” may refer to a set of features output by a layer of a neural network after a filter (also called a kernel or feature detector), the set of features comprising a vector of weights and biases applied to an input data set. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] These and other features, aspects and advantages will become better understood with reference to the following description, appended claims, and accompanying drawings, in which:
[0048] Figure 1 is a schematic illustration of a system for associating objects on two or more images according to an embodiment of the present disclosure;
[0049] Figure 2 is a schematic illustration of a method for associating objects on two images according to an embodiment of the present disclosure;
[0050] Figure 3 A first image and a second image corresponding to an example of an embodiment according to the present disclosure are shown;
[0051] Figure 4 illustrates example corresponding points identified and matched between an example first image and a second image according to an embodiment of the present disclosure;
[0052] Figure 5 is a schematic illustration of an exemplary method of associating each of one or more identified objects in a first image with a corresponding object in a second image according to an embodiment of the present disclosure;
[0053] Figure 6 is a schematic illustration of a method of associating each first bounding box with a corresponding second bounding box according to an embodiment of the present disclosure;
[0054] Figure 7 shows an example first and second image after projection according to an embodiment of the present disclosure;
[0055] Figure 8 is a schematic illustration of a top view of a vehicle with an image sensor installed according to an embodiment of the present disclosure; and
[0056] Fig. 9 is a schematic illustration of the architecture of a trained neural network for identifying corresponding points according to an embodiment of the present disclosure.
[0057] In the drawings, the same parts are denoted by the same reference numerals.
[0058] Those skilled in the art will appreciate that any block diagram herein represents a conceptual view of an illustrative system that embodies the principles of the present subject matter. Similarly, it will be understood that any flow charts, flow diagrams, state transition diagrams, pseudocodes, etc. represent various processes that can be substantially represented in a computer-readable medium and executed by a computer or processor, whether or not such a computer or processor is explicitly shown. DETAILED DESCRIPTION
[0059] In the above summary of the invention, in this specification, in the following claims and in the drawings, reference is made to specific features (including method steps) of the invention. It should be understood that the disclosure of the invention in this specification includes all possible combinations of such specific features. For example, where a specific feature is disclosed in the context of a specific aspect or embodiment of the invention or a specific claim, the feature may also be used in combination with and / or in the context of other specific aspects and embodiments of the invention, and generally, as far as possible.
[0060] In this document, the word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any embodiment or implementation of the present subject matter described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0061] Although the present disclosure is susceptible to various modifications and alternative forms, specific embodiments thereof have been shown by way of example in the accompanying drawings and will be described in detail below. However, it should be understood that this is not intended to limit the present disclosure to the disclosed form, but rather, the present disclosure will cover all modifications, equivalents and alternatives falling within the scope of the present disclosure.
[0062] Figure 1 1 is a schematic illustration of a system for associating objects on two or more images according to an embodiment of the present disclosure. The system 100 for associating objects on two or more images may include a first image sensor 104, a second image sensor 108, one or more processors 112, and one or more displays 116. Although only the first image sensor 104 and the second image sensor 108 are shown, in some embodiments, the system 100 may include more than two image sensors.
[0063] According to some embodiments, the first image sensor 104 and the second image sensor 108 may be visible light sensors that capture information related to the color of objects in the scene. In some embodiments, the first image sensor 104 and / or the second image sensor 108 may be a camera or a video camera. In some embodiments, the first image sensor 104 may be operable to provide at least one first image, and the second image sensor 108 may be operable to provide at least one second image. In some embodiments, the first image sensor 104 and / or the second image sensor 108 may be an image sensor, camera, and / or video camera equipped with a standard lens. Preferably, the first image sensor 104 and / or the second image sensor 108 may be a fisheye camera, which is an image sensor, camera, and / or video camera equipped with a fisheye lens or a wide-angle lens with a field of view of not less than 60 degrees.
[0064] According to some embodiments, the first image sensor 104 may be positioned to capture a scene from a first direction, and the second image sensor 108 may be positioned to capture the same scene from a second direction, the first image sensor 104 and the second image sensor 108 being positioned such that there is an overlap between at least one first image of the scene captured by the first image sensor 104 and at least one second image of the scene captured by the second image sensor 108. In general, the larger the overlapping area, the more accurate the result, and the first image sensor 104 and the second image sensor 108 may be adjusted and customized based on the user's desired accuracy and desired scene coverage. For example, in the case where the overlap exceeds 50% of the area of interest, the distortion of the image plane may be mathematically modeled with higher accuracy and lower scene coverage in the method of the present disclosure. For example, in the case where the overlap is less than 50% of the area of interest, the distortion may still be modeled with lower accuracy and higher scene coverage in the method of the present disclosure. In some embodiments, the scene may be a scene around a vehicle, a scene along a corridor, a scene around a building, a scene within a room, or any other scene that may be useful for the identification and / or tracking of objects, people, or participants. The first image sensor 104 and the second image sensor 108 can be installed anywhere, at any position, and at any height, depending on the scene they are used to capture. In some embodiments, the first image sensor 104 and the second image sensor 108 can be installed on a vehicle and positioned to capture the scene around the vehicle. In some embodiments, the first image sensor 104 and the second image sensor 108 can be installed on the outside of a building and positioned to capture the scene around the building. In some embodiments, the first image sensor 104 and the second image sensor 108 can be installed along a corridor and positioned to capture the scene of the corridor.
[0065] According to some embodiments, one or more processors 112 may be coupled to the first image sensor 104 and the second image sensor 108 to receive at least one image captured by the first image sensor 104 and at least one second image captured by the second image sensor 108. The one or more processors 112 may be operable to identify and associate one or more objects found in both the at least one first image and the at least one second image. The association method will be described in detail later.
[0066] According to some embodiments, the one or more processors 112 may be coupled to one or more displays 116. In some embodiments, the one or more displays 116 may display at least one first image captured by the first image sensor 104 and / or at least one second image captured by the second image sensor 108 received by the one or more processors 112. In some embodiments, the one or more displays 116 may display markings, shadows, or any other indicators generated by the one or more processors 112. Examples of such markings or indications may include tracker identity tags, bounding boxes, projected bounding boxes, key points, and points.
[0067] Figure 2 200 is a schematic illustration of a method for associating objects on two images according to an embodiment of the present disclosure. Although the present disclosure discusses a method for associating two images, the method can be extended to associating objects on more than two images as long as there is an overlapping area between the images. The method 200 for associating objects on two images can be implemented by any architecture and / or computing system. For example, various architectures using, for example, multiple integrated circuit (IC) chips and / or packages, and / or various computing devices and / or consumer electronics (CE) devices (such as multi-function devices, tablet computers, smart phones, etc.) can implement the techniques and / or arrangements described herein.
[0068] According to some embodiments, the method 200 for associating objects on two images may begin at operation 204, where a first image is received from a first image sensor 104 and a second image is received from a second image sensor 108. Preferably, the first image and the second image are images taken across different views of the same scene. Preferably, there is an overlap between the first image and the second image, where an overlapping or common area is captured in both the first image and the second image. The overlapping area may be any proportion of the first image and / or the second image. Preferably, the overlapping area exceeds 50% of the first image and / or the second image, although the method has acceptable accuracy even when the overlapping area is less than 50%. In some embodiments, the first image and the second image may be taken sequentially by the first image sensor 104 and the second image sensor 108. Preferably, the first image and the second image are taken simultaneously by the first image sensor 104 and the second image sensor 108. Preferably, the first image and the second image have the same timestamp.
[0069] Figure 3An example of a first image 304 and an example of a second image 308 according to an embodiment of the present disclosure are shown. As shown, the first image 304 and the second image 308 include image content on a fisheye image plane. In other embodiments, the first image 304 and the second image 308 may include image content on a rectilinear image plane. Figure 3 As shown, the first image 304 and the second image 308 cover different views of the same scene, wherein an overlapping area 312 is captured in both the first image 304 and the second image 308 .
[0070] Return to Figure 2 , method 200 may include operation 208, in which corresponding points between the first image and the second image are identified. Corresponding points are image points that exist or are found in both the first image captured by the first image sensor 104 and the second image captured by the second image sensor 108. In some embodiments, each identified corresponding point may include image point data, which may include any suitable data or data structure indicating the identified image point, such as the position of each identified point and / or a point vector. The point vector may include any data structure, such as a value vector indicating characteristics of a particular point (e.g., a measure of various parameter characteristics of a point). For example, a point position may be a pixel position.
[0071] According to some embodiments, operation 208 may include identifying image points in the first image and the second image, and matching the image points in the first image and the second image to identify corresponding image points. In some embodiments, the identified image points may be generally geometrically invariant (i.e., invariant to image translation, rotation, and scaling) and photometrically invariant (i.e., invariant to changes in brightness, contrast, and color), which may increase the ease of identifying and matching corresponding points in the first image and the second image. In some embodiments, the identified image points may be points corresponding to vertices or intersections, which may increase the ease of identifying and matching corresponding points in the first image and the second image. In some embodiments, the identified image points may be unique points, also referred to as key points, or pixels having highly unique or distinguishable visual features.
[0072] According to some embodiments, the recognition of image points may be performed manually or using any suitable technique(s) to detect or identify suitable image-based features, which are detected based on features extracted using image information such as pixel values. Examples of methods that may be used to recognize image points include, but are not limited to, methods disclosed in David G. Lowe's "Distinctive Image Features from Scale-Invariant Keypoints," Bay et al.'s "SURF: Speeded Up Robust Features," BRISK (Binary Robust Invariant Scalable Keypoints) disclosed in Leutenegger et al.'s "BRISK: Binary Robust Invariant Scalable Keypoints," and ORB (Oriented FAST and Rotated BRIEF) disclosed in Rublee et al.'s "ORB: an efficient alternative to SIFT or SURF."
[0073] According to some embodiments, image points defined in the first image may be compared with image points identified in the second image to identify corresponding or matching points found in both the first image and the second image. The comparison and matching of points may be performed using any known image analysis method for cross-matching purposes, such as similarity-based matching or template matching. The comparison and cross-matching may be repeated between each possible pair of points until all identified points have been processed. An example of an algorithm that may be employed is the random sample consensus (RANSAC) algorithm disclosed by Fischler, AM and Bolles, RC in "Random Sample Consensus: A Paradigm for Model Fitting with Applications to Image Analysis and Automated Cartography".
[0074] According to some embodiments, operation 208 of identifying corresponding points between the first image and the second image may be performed using a neural network that is trained to identify and match key points using features within the images. Fig. 9The architecture and training examples of neural networks for identifying corresponding points are discussed in further detail.
[0075] Figure 4 1 shows example corresponding points identified and matched between an example first image 304 and a second image 308 according to an embodiment of the present disclosure. As shown, a plurality of image points 404 are identified in the first image 304, and a plurality of image points 408 are identified in the second image 308. Figure 4 , each image point 404 or 408 is indicated by a dot representing a pixel location in the first image 304 or the second image 308. Each image point 404 identified in the first image 304 can be matched with a corresponding image point 408 in the second image 308. As shown, the image points identified in the first image 304 and the second image 308 are unique points, such as the corners of a sign, the corners of a building, the corners of a road marking, or the corners of a truck.
[0076] Return to Figure 2 , method 200 may include operation 212, wherein a projection matrix is approximated between the first image and the second image based on the corresponding points identified in operation 208. The projection matrix represents a geometric relationship and / or a transformation relationship between points of the first image and corresponding points of the second image.
[0077] According to some embodiments, approximating the projection matrix preferably comprises approximating at least one polynomial equation. A polynomial equation is an algebraic equation having the general form, P(x) = a n x n + a n-1 x n-1 + … + a2x 2 + a1x + a0 (1) Among them, a0,…,a n are coefficients of the polynomial equation, x is an indeterminate value that may be substituted for any value, and the exponent n on the indeterminate value x may be any integer representing the power of the indeterminate value x.
[0078] According to some embodiments, the first image and / or the second image may be a fisheye image. In such an embodiment, the geometric relationship between the first image and the second image is nonlinear, and the degree of the polynomial equation may be n, and the number of corresponding points identified is between n+1 and n+7, where n is an integer equal to or greater than 2. In some embodiments, the degree of the polynomial equation is 2, and the number of corresponding points identified may be between 3 and 9 to reduce the required computing power while maintaining sufficient accuracy without overfitting. In some embodiments, the degree of the polynomial may be independent of the number of corresponding points.
[0079] According to some embodiments, the number of corresponding points used to approximate the projection matrix in operation 212 may be less than the number of corresponding points identified in operation 208. In some embodiments, the corresponding points used to approximate the projection matrix in operation 212 may be selected in a distributed manner to cover a maximum area in the overlap between the first image and the second image.
[0080] According to some embodiments, the projection matrix may include two polynomial equations: a first polynomial equation for coordinates on the x-axis and a second polynomial equation for coordinates on the y-axis. The location or position of each pixel or point on the image may be represented as a 2-dimensional (2D) coordinate, which may be represented as (x, y), where x represents the x-coordinate and y represents the y-coordinate. In some embodiments, the x-coordinate may be transformed based on the first polynomial equation, and the y-coordinate may be transformed based on the second polynomial equation.
[0081] For example, the projection matrix consisting of two polynomial equations can be expressed as: Where P(x) represents the first polynomial equation of the coordinates on the x-axis, P(y) represents the second polynomial equation of the coordinates on the y-axis, a0,…,a n are the coefficients of the first polynomial equation, b0,…,b n are coefficients of the second polynomial equation, x is an uncertain value of the first polynomial equation that can replace the x-coordinate, y is an uncertain value of the second polynomial equation that can replace the y-coordinate, and the exponent n of the uncertain value x or y can be any integer representing the degree of the uncertain value x or y.
[0082] According to some embodiments, the coefficients a0, ..., a1 of the at least one polynomial equation may be determined by any known curve fitting method to identify the best fitting polynomial equation for a series of data points. n An example of a known method is the 2D polynomial transform function of the skimage library in python.
[0083] According to some embodiments, method 200 may include operation 216, wherein each of the at least one identified object in the first image is associated with a corresponding identified object in the second image based on the projection matrix approximated in operation 212. Any known object matching method may be used to associate each of the one or more identified objects in the first image with the corresponding identified object in the second image. The association may be based on location using the approximated projection matrix, and additionally based on visual appearance. Figure 5 Examples of methods of associating each of one or more identified objects in a first image with a corresponding identified object in a second image are discussed.
[0084] According to some embodiments, method 200 may optionally include operation 220, wherein the position of the object identified in the first image and the second image is drawn on a common image plane. In some embodiments, a common image plane may cover a 360° surround view of the object around the first image sensor, the second image sensor, and the other image sensors. This drawing of the position of the identified object in the common image plane allows accurate drawing of the position of the object in the scene. The common image plane may be an image plane of the first image, an image plane of the second image, or a separate image plane. Preferably, drawing is performed by projecting the four corner points of each object and / or bounding box using an approximate polynomial equation obtained for the x-axis and y-axis. In some embodiments, any number of image sensors may be installed, as long as adjacent image sensor pairs have overlapping fields of view, so that objects in adjacent image sensors may be associated with each other, and objects may be tracked across multiple cameras and may be tracked in a time series.
[0085] Figure 5 is a schematic illustration of an exemplary method of associating each of at least one identified object in a first image with a corresponding object in a second image according to an embodiment of the present disclosure. In some embodiments, a method 500 of associating each of at least one identified object in a first image with a corresponding object in a second image may be employed in operation 216 of method 200. The method 500 of associating each of at least one identified object in a first image with a corresponding object in a second image may begin with operation 504, wherein at least one object is identified in the first image and at least one object is identified in the second image. The at least one object may be identified using any known object detection or image instance segmentation method.
[0086] According to some embodiments, method 500 may include operation 508, wherein at least one first bounding box is generated in the first image and at least one second bounding box is generated in the second image. The bounding box represents the spatial location of the identified object. The object can be identified, and the bounding box can be defined using any known object detection method (such as a convolutional neural network). CNN is a multi-layer feedforward neural network formed by stacking many hidden layers on top of each other in sequence. The sequential design can allow CNN to learn hierarchical features. The hidden layer is typically a convolutional layer followed by an activation layer, some of which are followed by a pooling layer. CNN can be configured to recognize patterns in data. The convolutional layer can include convolutional kernels that are used to find patterns across input data. The convolutional kernel can return a large positive value for a portion of the input data that matches the kernel pattern, or can return a smaller value for another portion of the input data that does not match the kernel pattern. CNN is preferred because CNN can be able to extract informative features from training data without manually processing the training data. In cases involving large unstructured data, such as image classification, speech recognition, and natural language processing, CNN can produce accurate results. Moreover, CNNs are computationally efficient because CNNs are able to assemble patterns of increasing complexity using relatively small kernels in each hidden layer. CNNs are also advantageous because CNNs have high accuracy in image recognition and can automatically filter images and detect features without any human supervision. An example of a convolutional neural network is YOLOv4, where an example of YOLOv4 can be found in "YOLOv4: Optimal Speed and Accuracy of Object Detection" by Bochkovskiy et al., where an example of the architecture of YOLOv4 can be found at least in Section 3, and an example of training of YOLOv4 can be found at least in Section 4.1. In some embodiments, where the first image and / or the second image is a fisheye image, the convolutional filter of the convolutional network can be a restricted deformable convolution (or RDC) module for effectively modeling geometric transformations present in the fisheye image, where the deformation is adjusted based on the input features in a local, dense, and adaptive manner. The shape of the RDC is learned to adapt to changes in the features, where the shape of the kernel is adapted to unknown complex transformations in the input. In particular, the RDC module learns a sampling matrix with position offsets, where these offsets are learned from the aforementioned feature maps via additional convolutional layers.Detailed information about the RDC module can be found in “Restricted Deformable Convolution based Road Scene Semantic Segmentation Using Surround View Cameras” by Deng et al.
[0087] According to some embodiments, a first tracker identifier (ID) may be assigned to each bounding box detected within an image using any known tracker, such as the MOTDT tracker disclosed in "Real-time Multiple People Tracking with Deeply Learned Candidate Selection and Person Re-identification" by Chen et al. The first tracker identifier may be a local track ID, which is a track ID associated with an object identified in an image captured by a separate image sensor.
[0088] According to some embodiments, method 500 may include operation 512, wherein each first bounding box in the first image is associated with a corresponding second bounding box in the second image. In other words, each first bounding box in the first image is associated with a second bounding box in the second image, and the first bounding box and its corresponding second bounding box represent the spatial position of the same object found in the first image captured by the first image sensor and the second image captured by the second image sensor. In some embodiments, a second trajectory identifier (ID) may be assigned to each first bounding box in the first image, and the same second trajectory identifier (ID) may be assigned to its associated corresponding second bounding box in the second image. The second trajectory identifier may be a global trajectory ID, which is a trajectory ID associated with an object identified within a scene. When the position of the identified object is subsequently drawn on a common image plane, the global trajectory ID may be used to mark such an object.
[0089] According to some embodiments, associating each first bounding box with the corresponding second bounding box may be based at least on a relationship between a projection of the first bounding box and the corresponding second bounding box, wherein the projection of the first bounding box is based on an approximate projection matrix, which is related to Figure 6Further detailed description. According to some embodiments, associating each first bounding box with a corresponding second bounding box may further include determining visual similarities or dissimilarity between objects contained within the bounding boxes by identifying at least one feature of each identified object in the first image and at least one feature of each identified object in the second image, and comparing at least one feature of each identified object in the first image with at least one feature of each identified object in the second image. This may be advantageous for improving the accuracy of the results in situations where there is a low overlap area between the first image and the second image.
[0090] Figure 6 6 is a schematic illustration of a method of associating each first bounding box with a corresponding second bounding box according to an embodiment of the present disclosure. In some embodiments, a method 600 of associating each first bounding box with a corresponding second bounding box may be employed in operation 512 of method 500. The method 600 of associating each first bounding box with a corresponding second bounding box may begin with operation 604, wherein at least one first bounding box from a first image is projected onto a second image to form at least one projected bounding box, wherein each projected first bounding box corresponds to one of the at least one first bounding boxes. The projection is based on the projection matrix approximated in operation 212 of method 200, wherein the position or coordinates of each first bounding box may be transformed using the approximated projection matrix and then projected onto the second image.
[0091] Figure 7 An example first image 304 and a second image 308 are shown after being projected according to operation 604 of method 600 in accordance with an embodiment of the present disclosure. Figure 7 As shown, the first image 304 includes four first bounding boxes 704a to 704d, and the second image 308 includes four second bounding boxes 708a to 708d, each of which represents the spatial position of a person present in the overlapping area 312 of the first image 304 and the second image 308. A track ID is assigned to each of the first bounding boxes 704a to 704d and the second bounding boxes 708a to 708d. Figure 7 As shown in the first image 304 in FIG. 1 , the first bounding box 704a is assigned a track ID 5, the first bounding box 704b is assigned a track ID 6, the first bounding box 704c is assigned a track ID 7, and the first bounding box 704d is assigned a track ID 8. Figure 7 As shown in the second image 308 in FIG. 1 , the second bounding box 708 a is assigned track ID 1, the second bounding box 708 b is assigned track ID 2, the second bounding box 708 c is assigned track ID 3, and the second bounding box 708 d is assigned track ID 4. Figure 7As shown, each of the first bounding boxes 704a to 704d is projected onto the second image 308 as projected first bounding boxes 704a' to 704d'. Each of the projected first bounding boxes 704a' to 704d' corresponds to the first bounding box 704. The projected first bounding box 704a' corresponds to the first bounding box 704a, the projected first bounding box 704b' corresponds to the first bounding box 704b, the projected first bounding box 704c' corresponds to the first bounding box 704c, and the projected first bounding box 704d' corresponds to the first bounding box 704d. Figure 7 As shown, although the first bounding boxes 704a' to 704d' of each projection correspond to the first bounding boxes 704a to 704d, the sizes and relative positions of the bounding boxes may or may not be the same depending on the projection matrix. Figure 7 In the example shown, since both the first image 304 and the second image 308 are fisheye images, the geometric relationship between the pixels of the first image 304 and the second image 308 is non-linear, and therefore each projected first bounding box 704a' to 704d' may be different in size from its corresponding first bounding box 704a to 704d.
[0092] Back to Figure 6 , method 600 may include operation 608, wherein a relationship between each projected first bounding box and each second bounding box is calculated. Each determined relationship is a cost for each projected first bounding box-second bounding box pair. Figure 7 In the example shown in , separate relationships or costs between the projected bounding box 704a' and each of the second bounding boxes 708a-708d, and separate relationships or costs between the projected first bounding boxes 704b', 704c', and 704d' and each of the second bounding boxes 708a-708d can be determined.
[0093] In some embodiments, the relationship between each projected first bounding box and each second bounding box may include an overlap range between each projected first bounding box and each second bounding box. In some embodiments, the overlap range between each projected first bounding box and each second bounding box may be determined based on an intersection-over-union (IOU) evaluation metric, which is calculated by dividing the overlap area between the projected first bounding box and the second bounding box by the union area (i.e., the area covered by both the projected first bounding box and the second bounding box). The higher the IOU between the projected first bounding box and the second bounding box, the higher the probability that the first bounding box corresponding to the projected first bounding box and the second bounding box covers the same object within the scene, and the lower the cost.
[0094] According to some embodiments, the relationship between each projected first bounding box and each second bounding box may include a distance between each projected first bounding box and each second bounding box. Preferably, the distance is between the center of each projected first bounding box and the center of each second bounding box. In some embodiments, the Euclidean distance between each projected first bounding box and each second bounding box may be calculated. The shorter the Euclidean distance between the projected first bounding box and the second bounding box, the higher the probability that the first bounding box corresponding to the projected first bounding box and the second bounding box covers the same object within the scene, and the lower the cost.
[0095] According to some embodiments, method 600 may include operation 612, wherein each first bounding box is associated with a corresponding second bounding box based on the determined relationship between the corresponding projected first bounding box and the corresponding second bounding box. In some embodiments, the determined relationship or cost may be arranged in a cost matrix, which may be used for association in subsequent operations. For example, the cost matrix may be a 2-dimensional matrix, one dimension of which is the first bounding box of the projection, and the second dimension is the second bounding box. Each projected first bounding box-second bounding box pair includes a relationship or cost included in the cost matrix. The best match between the projected first bounding box and the second bounding box may be determined by identifying the first bounding box-second bounding box pair of the lowest cost in the matrix. In some embodiments, the association is determined using the Hungarian algorithm (also known as the Kuhn Munkres algorithm), which is a combined optimization algorithm that can optimize the global cost of the first bounding box and the second bounding box across all projections to minimize the global cost. The first bounding box-second bounding box combination of the projection that minimizes the global cost in the cost matrix may be determined and used as an association. Any other method may also be used to associate each projected first bounding box with each second bounding box.
[0096] According to an embodiment of the present disclosure, the operation 612 of associating each first bounding box with the corresponding second bounding box may further include identifying at least one feature of each identified object in the first image and the second image, and comparing at least one feature of each identified object in the first image with at least one feature of each identified object in the second image. The identification and comparison of the features of the identified objects allows the accuracy of the results to be improved because the visual similarity of the objects within the first bounding box and its associated second bounding box is also considered. In some embodiments, any known algorithm or machine learning method can be used to perform the identification and comparison of features separately. For example, any known feature recognition algorithm or machine learning method (such as a pre-trained convolutional neural network (CNN)) can be used to perform the identification or extraction of at least one feature of the identified object in the first image and the second image, and the cost can be calculated based on the similarity of the features. In some embodiments where the first image and / or the second image is a fisheye image, the convolution filter of the CNN may include the above-mentioned RDC module. In some embodiments, any known generative model-based re-recognition can be used to perform the identification and comparison of features. In some embodiments, a Siamese network can be used to perform the identification and comparison of features, which is a class of neural networks containing one or more identical networks. Pairs of identified objects are fed into the one or more identical networks. The one or more identical networks compute features of an input object and use differences or dot products of these features to compute similarities of the features.
[0097] Figure 8 is a schematic illustration of an example of an arrangement of image sensors mounted on a vehicle according to an embodiment of the present disclosure. According to some embodiments, system 100 may be adapted for use in a vehicle, the system comprising at least a first image sensor and a second image sensor, one or more processors, and a memory storing executable instructions for execution by the one or more processors, the executable instructions including instructions for performing method 200. In some embodiments, first image sensor 804 and second image sensor 808 may be mounted on vehicle 800 such that first image sensor 804 and second image sensor 808 capture images of a scene around vehicle 800. For example, as Figure 8 As shown in , the first image sensor 804 can be mounted in front of the vehicle 800 so that the first image sensor 804 captures an image of the scene in front of the vehicle 800 within a first field of view (FOV) 812. For example, Figure 8, the second image sensor 808 may be mounted on the right side of the vehicle 800 such that the second image sensor 808 captures images of the scene on the right side of the vehicle 800 within a second field of view (FOV) 816. The first image sensor 804 and the second image sensor 808 are positioned such that there is an overlap region 820 between the first FOV 812 of the first image sensor 804 and the second FOV 816 of the second image sensor 808. The overlap region 820 enables an object in a first image received from the first image sensor 804 to be associated with a corresponding object in a second image received from the second image sensor 808.
[0098] According to some embodiments, the system 100 and method 200 may be adapted to associate objects on more than two images by including more image sensors, as long as the field of view of the additional image sensors overlaps with at least one other image sensor. In some embodiments, the method 200 may be adapted such that step 220 includes mapping the locations of the identified objects on a common image plane, wherein the common image plane preferably covers a 360° surround view of the objects around the vehicle. For example, the system 100 mounted on the vehicle may further include a third image sensor 824 and a fourth image sensor 828. For example, as Figure 8 , the third image sensor 824 can be mounted on the left side of the vehicle 800 so that the third image sensor 824 captures an image of a scene on the left side of the vehicle 800 within a third field of view (FOV) 832. The third image sensor 824 can be positioned so that there is an overlap area 836 between the first FOV 812 of the first image sensor 804 and the third FOV 832 of the third image sensor 824. For example, Figure 8, the fourth image sensor 840 may be mounted behind the vehicle 800 such that the fourth image sensor 840 captures images of a scene behind the vehicle 800 within a fourth field of view (FOV) 844. The fourth image sensor 840 may be positioned such that there is an overlap region 848 between the third FOV 832 of the third image sensor 824 and the fourth FOV 844 of the fourth image sensor 840, and an overlap region 852 between the second FOV 816 of the second image sensor 808 and the fourth FOV 844 of the fourth image sensor 840. In this example, the method 200 may be applied between images received from the first image sensor 804 and images received from the second image sensor 808, between images received from the first image sensor 804 and images received from the third image sensor 816, between images received from the third image sensor 816 and images received from the fourth image sensor 840, and between images received from the second image sensor 808 and images received from the fourth image sensor. In this example, one common plane mapping the locations of the identified objects would cover a 360° surround view of the objects surrounding the vehicle 800 .
[0099] Fig. 9 9 is a schematic diagram of the architecture of a trained neural network for identifying corresponding points according to an embodiment of the present disclosure. In some embodiments, the trained neural network 900 may be employed in operation 208 of method 200. It is emphasized that Fig. 9 The architecture shown in is an example, and other suitable architectures may be adopted depending on the user's application and requirements. In some embodiments, a trained neural network 900 may be used to identify key points in an input image 904. A neural network may include input nodes (i.e., layers), hidden nodes, and output nodes.
[0100] According to some embodiments, the neural network 900 may be trained on a scene classification dataset. In some embodiments, the scene classification dataset may include images of different scenes captured from different cameras mounted on the vehicle. In some embodiments, the scene classification dataset may include images classified into scene categories. For example, the scene classification dataset may include images classified into at least 500 categories, wherein each category includes at least 70 images. In some embodiments, the number of categories may be based on possible environments in which the system 100 may be employed. In some embodiments, a single scene category representing the content of the same scene may include a sequence of images captured from an image sensor that captures overlapping areas of the scene, and may further include a sequence of images captured in the same time period by other image sensors having overlapping fields of view with the first-mentioned image sensor. In some embodiments, a single scene category may further include images that have been changed from the aforementioned images using image scaling and random cropping to increase the number of images in the category, as long as the overlapping area is retained. In some embodiments, when training the neural network 900, pairs of images from the same category may be designated as matching pairs, while pairs of images from different categories may be designated as non-matching pairs.
[0101] According to some embodiments, the trained neural network 900 may include a convolutional neural network (CNN) backbone 908 for extracting features, and a key point extractor 916 for identifying key points from the features extracted by the CNN backbone 908. In some embodiments, the trained neural network 900 may be trained with a batch size of 16 in 25 epochs, optimized using stochastic gradient descent (SGD) with a momentum of 0.9. In some embodiments, a decaying learning rate may be used until the rate reaches zero and the learning rate initialization may be between 0.0003 and 0.01. In some embodiments, training of the key point extractor 916 and the CNN backbone 908 may be performed independently to learn stronger local descriptors. In some embodiments, the CNN backbone 908 may be trained and refined first, and the key point extractor 916 may be trained subsequently so that only the weights of the key point extractor 916 are learned without interfering with the trained and refined CNN backbone 908.
[0102] Examples of CNN backbones 908 include ResNet-50 and ResNet-101, which can be found in “Deep Residual Learning for Image Recognition”, where the architectures of ResNet-50 and ResNet-101 can be found in at least Section 3, Table 1, and Figure 5, and the training of ResNet-50 and ResNet-101 can be found in at least Section 3.4. Examples of the architectures of ResNet-50 and ResNet-101 are reproduced in Table 1 below. In some embodiments where the first image and / or the second image is a fisheye image, the convolutional filters of the CNN backbone 908 may include the above-mentioned RDC module. Table 1: Architecture of ResNet-50 and ResNet-101
[0103] According to some embodiments, the CNN backbone 908 may include one or more shallow layers 924 and one or more deep layers 930. For example, in the case where the CNN backbone 908 is ResNet-50 or ResNet-101, the layers conv1, conv2_x, conv3_x, and conv4_x may be shallow layers 924, and the layer conv5_x may be a deep layer 930. In some embodiments, the output from the shallow layer 924 may be input into the keypoint extractor 916 to identify a feature subset of densely extracted features extracted by the shallow layer 924 of the CNN backbone 908 as a keypoint. For example, in the case where the CNN backbone 908 is ResNet-50 or ResNet-101, the output from the layer conv4_x may be input into the keypoint extractor 916 to identify the keypoint. In some embodiments, the output from the shallow layer 924 may include 1024 channels and may be represented as 14×14×1024, which means that each element in the 14×14 matrix has 1024 feature vectors. In some embodiments, the 196 elements in the 14×14 matrix may correspond to 196 overlapping patches (or local descriptors) from a 512×512 input image, such that each patch may have a feature vector or local descriptor of size 1024. The output from the shallow layer 924 is preferably used to extract key points because the feature map generated by the shallow layer 924 includes information about local features in the input image, compared to the output from the deep layer 930, which includes information about global features describing the entire image. The deeper the layer, the more abstract and high-level the semantic attributes. The shallow layer 924 generates shallow feature maps that represent low-level local features that include descriptors and geometric information about specific image regions such as objects, while the deep layer 930 generates deep feature maps that represent global features that summarize the image content, but do not contain information about the spatial arrangement of visual elements.
[0104] According to some embodiments, the CNN backbone 908 can be pre-trained on a labeled image classification and localization dataset (such as the ImageNet dataset available at https: / / image-net.org / download.php.). Dense features can be extracted from the input image 904 by applying a fully convolutional network (FCN) obtained from ResNet-50, using the output of the conv4_x convolutional block of ResNet-50. In order to handle scale changes, an image pyramid can be constructed, and FCN can be explicitly applied on each layer. The obtained feature map can be considered as a dense grid of local descriptors, and the features can be located based on their receptive fields, which can be calculated based on the configuration of the convolutional and pooling layers of the FCN. In some embodiments, the pixel position of the feature can be determined by obtaining the center of the receptive field of the feature. In some embodiments, in order to enhance the discriminability of the local descriptor, the ResNet-50 model can be fine-tuned for scene classification by training the ResNet-50 model with the above-mentioned scene classification dataset and using the cross entropy loss so that the local descriptor learns a better representation of the scene without using object-level and slice-level labels.
[0105] According to some embodiments, the key point extractor 916 may receive as input the output of the shallow layer 924 from the CNN backbone 908. In some embodiments, the key point extractor 916 may include an attention model 936 and a dimension reducer 942, the attention model selecting a feature subset from the densely extracted features from the shallow layer 924 of the CNN backbone 908, and the dimension reducer is used to reduce the dimensionality of the features. In some embodiments, the dimension reducer 942 may include PCA (principal component analysis) or an autoencoder. In some embodiments, the dimension reducer 942 may generate a local descriptor of the key points of the feature subset selected by the attention model 936.
[0106] According to some embodiments, the keypoint extractor 916 may include PCA and L2 normalization as the dimension reducer 942. In particular, the dimension of the feature map may be L2 normalized, and its dimension may be reduced from 1024 to 40 by PCA followed by additional L2 normalization.
[0107] According to some embodiments, the keypoint extractor 916 may include an autoencoder as a dimensionality reducer 942. In some embodiments, the autoencoder may reduce the dimensionality of the feature map from 1024 to 128. In some embodiments, the autoencoder may be a convolutional autoencoder module that learns a suitable low-dimensional representation of the feature map. In some embodiments, a cross entropy classification loss L with a weight λ=10 may be used. r To train the autoencoder.
[0108] According to some embodiments, the attention model 936 can be used to predict which features of the densely extracted features are discriminative for objects in the image by learning the weights of the local descriptors. In some embodiments, the attention model 936 can determine which features are discriminative based on the key point detection score. In some embodiments, the attention model 936 can remove redundant local descriptors based on the relevance of the local descriptors so that the remaining local descriptors are key point descriptors 950, wherein such key point descriptors 950 correspond to key points identified in the image. In some embodiments, the attention model 936 may include two convolutional layers, which have no stride, use convolutional filters of size 1×1, with ReLU as the activation function of the first convolutional layer, and softplus as the activation function of the second convolutional layer. In some embodiments, the attention model can be trained by using a weighted sum of local descriptors and then by using a standard softmax cross entropy loss with weight β=1. In some embodiments, attention can be focused on the key point descriptors 950 to determine the exact location of the key points identified in the image.
[0109] The steps shown are explained to explain the exemplary embodiments shown, and it should be expected that ongoing technical development will change the way to perform specific functions. The examples presented herein are for illustrative purposes, not for limitation. Further, for ease of description, the boundaries of functional building blocks are arbitrarily defined herein. As long as the specified functions and their relationships are properly performed, alternative boundaries can be defined. Based on the teachings contained herein, alternatives (including equivalents, extensions, variations, deviations, etc. of those described herein) will be obvious to (multiple) technical personnel in the relevant fields. Such alternatives fall within the scope and spirit of the disclosed embodiments. The term "comprises, comprising, includes" or any other variation thereof is intended to cover non-exclusive inclusions, so that the arrangement, device or method including a series of parts or steps includes not only those parts or steps, but also other parts or steps that are not explicitly listed or inherent to such an arrangement or device or method. In other words, in the absence of more constraints, one or more elements in a system or device preceded by "include" do not exclude the presence of other elements or additional elements in the system or method. It must also be noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0110] Finally, the language used in the specification is selected primarily for readability and instructional purposes, and is not necessarily selected to define or limit the subject matter of the present invention. Therefore, the scope of the present invention is not limited by this detailed description, but is limited by any claims issued based on the application hereof. Therefore, the embodiments of the present invention are intended to illustrate, but not limit, the scope of the present invention set forth in the appended claims.
Claims
1. A vehicle comprising a vehicle system for associating objects on two or more images, wherein: The vehicle system includes at least a first image sensor (104) and a second image sensor (108), one or more processors, and a memory storing executable instructions for execution by the one or more processors, the executable instructions including instructions for performing a computer-implemented method, the computer-implemented method comprising: receiving a first image from the first image sensor (104) and a second image from the second image sensor (108); identifying corresponding points between the first image and the second image; approximating a projection matrix between the first image and the second image based on the identified corresponding points; and At least one identified object in the first image is associated with a corresponding identified object in the second image based at least on the approximated projection matrix.
2. The vehicle according to claim 1, wherein: The one or more processors and the memory storing executable instructions for execution by the one or more processors, the executable instructions comprising instructions for performing a computer-implemented method comprising: receiving a first image from the first image sensor (104) and a second image from the second image sensor (108); identifying corresponding points between the first image and the second image; approximating a projection matrix between the first image and the second image based on the identified corresponding points; associating at least one identified object in the first image with a corresponding identified object in the second image based at least on the approximated projection matrix; and The positions of the identified objects are plotted on a common image plane, wherein the common image plane preferably covers a 360° surround view of the objects around the vehicle.
3. A computer-implemented method for associating objects on two or more images, the method comprising: receiving a first image from a first image sensor (104) and a second image from a second image sensor (108); identifying corresponding points between the first image and the second image; approximating a projection matrix between the first image and the second image based on the identified corresponding points; as well as At least one identified object in the first image is associated with a corresponding identified object in the second image based at least on the approximated projection matrix.
4. The computer-implemented method of claim 3, wherein: The first image sensor (104) and / or the second image sensor (108) are fisheye cameras, and / or wherein the first image sensor and the second image sensor are positioned such that there is an overlap between at least one first image of a scene captured by the first image sensor and at least one second image of the scene captured by the second image sensor.
5. A computer-implemented method according to claim 3 or 4, wherein: The corresponding points are unique points, preferably points corresponding to vertices or intersections; and / or These corresponding points are key points.
6. A computer-implemented method according to any one of claims 3 to 5, wherein: The keypoints are selected using a neural network, and preferably a neural network comprising a convolutional neural network backbone and a keypoint extractor, wherein the convolutional neural network backbone preferably comprises a constrained deformable convolution module.
7. A computer-implemented method according to any one of claims 3 to 6, wherein: The step of approximating the projection matrix comprises approximating at least one projection matrix comprising at least one polynomial equation, wherein, preferably, the degree of the at least one polynomial equation is n and the number of identified corresponding points is between n+1 and n+7, wherein n is an integer equal to or greater than 2, and / or wherein, preferably, the degree of the at least one polynomial equation is 2 and the number of identified corresponding points is between 3 and 9.
8. A computer-implemented method according to any one of claims 3 to 7, wherein: The projection matrix includes two polynomial equations: a first polynomial equation for coordinates on the x-axis and a second polynomial equation for coordinates on the y-axis.
9. The computer-implemented method according to any one of claims 3 to 8, wherein: Associating at least one object in the first image with a corresponding object in the second image includes: identifying at least one object in the first image and at least one object in the second image; generating at least one first bounding box in the first image and generating at least one second bounding box in the second image, each of the at least one first bounding box and the at least one second bounding box representing a spatial location of the identified object; and Each first bounding box is associated with a corresponding second bounding box based at least on a relationship between a projection of the first bounding box and the corresponding second bounding box, wherein the projection is based on the approximated projection matrix.
10. The computer-implemented method of claim 9, wherein: Associating each first bounding box with a corresponding second bounding box comprises: projecting the at least one first bounding box onto the second image based on the approximated projection matrix to generate at least one projected first bounding box on the second image, wherein each projected first bounding box corresponds to one of the at least one first bounding boxes; determining a relationship between each projected first bounding box and each second bounding box; and Each first bounding box is associated with the corresponding projected first bounding box based on the determined relationship between the corresponding second bounding box.
11. The computer-implemented method of claim 10, wherein: The relationship between each projected first bounding box and each second bounding box includes: The overlap between each projected first bounding box and each second bounding box; and / or The distance between each projected first bounding box and each second bounding box, wherein the distance is preferably the distance between the center of each projected first bounding box and the center of each second bounding box.
12. The computer-implemented method according to any one of claims 9 to 11, wherein: Associating each first bounding box with a corresponding second bounding box further comprises: identifying at least one feature of each identified object in the first image and the second image, wherein the at least one feature is preferably an appearance feature vector; and At least one feature of each identified object in the first image is compared to at least one feature of each identified object in the second image.
13. The computer-implemented method of any one of claims 3 to 12, further comprising mapping the positions of the identified objects on a common image plane, wherein This common image plane preferably covers a 360° surround view of objects around the first image sensor, the second image sensor and the further image sensors.
14. A computer-implemented method according to any one of claims 3 to 13, wherein: The first image sensor and the second image sensor are fisheye cameras and are positioned such that there is an overlap between at least one first image of a scene captured by the first image sensor and at least one second image of the scene captured by the second image sensor; The corresponding points are selected using a neural network comprising a convolutional neural network backbone and a keypoint extractor, wherein the neural network is trained on a scene classification dataset comprising images classified into scene categories, wherein each scene category comprises a sequence of images captured from a single camera capturing overlapping areas of the scene and a sequence of images captured in the same time period by other cameras having overlapping fields of view with the single camera, and wherein the convolutional neural network backbone comprises a constrained deformable convolution module; The approximated projection matrix includes a first polynomial equation for coordinates on the x-axis and a second polynomial equation for coordinates on the y-axis, wherein the first polynomial equation and the second polynomial equation are each of degree 2 and the number of identified corresponding points is 3; Associating at least one identified object in the first image with a corresponding identified object in the second image comprises: identifying at least one object in the first image and at least one object in the second image; generating at least one first bounding box in the first image and generating at least one second bounding box in the second image, each of the at least one first bounding box and the at least one second bounding box representing a spatial location of the identified object; projecting the at least one first bounding box onto the second image based on the approximated projection matrix to generate at least one projected first bounding box on the second image, wherein each projected first bounding box corresponds to one of the at least one first bounding boxes; determining an extent of overlap between each projected first bounding box and each second bounding box; and Each first bounding box is associated with a corresponding projected first bounding box based on a determined overlap range between the corresponding second bounding box.
15. A training data set for a neural network, in particular a neural network according to claim 6 or 14, wherein: The training dataset includes images classified into scene categories, where each scene category includes: a sequence of images captured from a first image sensor capturing overlapping areas of a scene; Optionally, a sequence of images captured during the same time period by other image sensors having an overlapping field of view with the first image sensor; and Optionally, the image is changed from a sequence of images captured from the first image sensor and / or a sequence of images captured in the same time period by other image sensors having an overlapping field of view with the first image sensor.
16. A system suitable for use in a vehicle, the system comprising at least a first image sensor and a second image sensor, one or more processors, and a memory storing executable instructions for execution by the one or more processors, the executable instructions including instructions for performing a computer-implemented method according to any one of the preceding claims 3 to 14.
17. A computer program, a machine-readable storage medium or a data signal comprising instructions which, when executed on one or more processors, cause the one or more processors to perform the steps of the computer-implemented method according to any one of the preceding claims 3 to 14.