Object detection device and object detection method
Through cameras and deep neural network learning models, high-reliability image detection of objects in front of the vehicle is achieved, solving the problem of single function of traditional ranging equipment and supporting forward collision warning of advanced driver assistance systems.
Patent Information
- Application Number
- CN202111531600.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2041-12-14
AI Technical Summary
Traditional vehicle ranging equipment only provides simple distance sensing functions and cannot provide the type and motion status of the target object. It is prone to misjudgment and is costly.
Using a camera and a deep neural network learning model, the system can identify and calculate the position, size, and type of the target object through image detection, and then output the actual distance of the target object in combination with the deep neural network learning model.
It provides high-reliability forward object detection, suitable for vehicle forward object detection, and supports forward collision warning in advanced driver assistance systems, with the advantages of immediacy and low computational complexity.
Smart Images

Figure CN114266885B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sensing technology, and more particularly to an object detection device and an object detection method. Background Art
[0002] With the rapid growth of traffic volume on roads, the incidence of road traffic accidents has been increasing year by year, with rear-end collisions being a particularly significant increase. Consequently, most traditional vehicles are equipped with ranging devices, such as radar, to detect surrounding obstacles and provide forward ranging capabilities. However, these devices only provide simple distance sensing and fail to provide more comprehensive information, such as the target object's type and motion state. Traditional ranging devices also suffer from the disadvantages of being prone to misjudgment and high setup costs. Summary of the Invention
[0003] The present invention provides an object detection device and an object detection method, which can provide a highly reliable front object detection function through image detection.
[0004] The object detection device of the present invention includes a camera, a storage unit, and a processor. The camera continuously obtains a plurality of original sensing images. The storage unit stores a plurality of modules. The processor is coupled to the storage unit and executes the plurality of modules to perform the following operations: the processor defines the entire image area of each of the plurality of first sensing images in the plurality of original sensing images as a first range of interest; the processor defines the partial image area of each of the plurality of second sensing images in the plurality of original sensing images as a second range of interest, and crops a plurality of third sensing images according to the second range of interest of each of the plurality of second sensing images; the processor inputs the plurality of first sensing images and the plurality of third sensing images into a deep neural network learning model, so that the deep neural network learning model outputs image information of the target object images in the plurality of first sensing images and the plurality of third sensing images respectively; the processor obtains the actual object distance of the target object in the target object image based on the image information of the target object image.
[0005] The object detection method of the present invention includes the following steps: acquiring a plurality of original sensing images through a camera; defining, by a processor, the entire image area of each of a plurality of first sensing images among the plurality of sensing images as a first range of interest; defining, by the processor, a partial image area of each of a plurality of second sensing images among the plurality of original sensing images as a second range of interest, and cropping a plurality of third sensing images based on the second range of interest of each of the plurality of second sensing images; inputting, by the processor, the plurality of first sensing images and the plurality of third sensing images into a deep neural network learning model so that the deep neural network learning model outputs image information of target object images in the plurality of first sensing images and the plurality of third sensing images; and obtaining, based on the image information of the target object images, an actual object distance of the target object in the target object images.
[0006] Based on the above, the object detection device and object detection method of the present invention can perform image processing and image analysis on the sensing image provided by the camera to obtain the position information and image size of the target object image.
[0007] In order to make the above features and advantages of the present invention more clearly understood, embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 FIG. 4 is a circuit diagram of an object detection device according to an embodiment of the present invention.
[0009] Figure 2 is a flow chart of an object detection method according to an embodiment of the present invention.
[0010] Figure 3 FIG. 4 is a schematic diagram of a sensing image according to an embodiment of the present invention.
[0011] Figure 4 FIG. 4 is a schematic diagram of cropping a sensing image according to an embodiment of the present invention.
[0012] Figure 5 FIG. 4 is a schematic diagram of analyzing an object image in a sensing image according to an embodiment of the present invention.
[0013] Figure 6 FIG. 4 is a flow chart of calculating the horizon height coordinates in a sensing image according to an embodiment of the present invention.
[0014] Figure 7 FIG. 4 is a flowchart of calculating the actual object distance according to an embodiment of the present invention.
[0015] Among them, the brief description of the symbols in the accompanying drawings is as follows:
[0016] 100: Object detection device; 110: Processor; 120: Storage unit; 121: Deep neural network learning model; 130: Camera; 300_1-300_N, 301_1-301_N, 302_1-302_M, 303, 303_1-303_P, 304_1-304_P, 505: Sensed image; 506: Target object image; Wc, Wf, Wo: Width; Hc, Hf, Ho, Yh: Height; I1, I2: Range of interest; S210-S250, S610-S630, S710-S720: Steps. DETAILED DESCRIPTION
[0017] In order to make the content of the present invention more clearly understood, the following embodiments are given as examples of how the present invention can be truly implemented. In addition, wherever possible, elements / components / steps with the same reference numerals in the drawings and embodiments represent the same or similar components.
[0018] Figure 1 FIG is a circuit diagram of an object detection device according to an embodiment of the present invention. Figure 1 The object detection device 100 includes a processor 110, a storage unit 120, and a camera 130. The storage unit 120 can store a deep neural network learning model 121 and multiple modules. The processor 110 is coupled to the storage unit 120 and the camera 130. In this embodiment, the object detection device 100 is suitable for being set in front of a vehicle (e.g., at the front of the vehicle) to provide an object detection function in front of the vehicle (e.g., front vehicle detection), but the present invention is not limited to this. In this embodiment, the camera 130 can continuously obtain multiple raw sensing images. The processor 110 can receive the multiple raw sensing images and execute the deep neural network learning model 121 and other modules to perform image processing and image analysis on the multiple sensing images. The object detection device 100 can identify object images in the sensing images and can obtain the position information, image size, object type, and actual object distance of the object images.
[0019] In this embodiment, the processor 110 may be, for example, a central processing unit (CPU), a microprocessor (MCU), or a field programmable gate array (FPGA), or other similar processing circuit or control circuit, and the present invention is not limited thereto. In this embodiment, the storage unit 120 may be, for example, a memory, and is used to store the deep neural network learning model 121, other related modules, image data, and related software programs or algorithms for access and execution by the processor 110. The camera 130 may be a CMOS image sensor (CIS) or a charge coupled device (CCD) camera.
[0020] Figure 2 is a flow chart of an object detection method according to an embodiment of the present invention. Figure 3 FIG is a schematic diagram of a sensing image according to an embodiment of the present invention. Figures 1 to 3 The object detection device 100 may execute steps S210 to S250 to implement object detection. In step S210, the object detection device 100 may continuously acquire a plurality of raw sensing images 300_1 to 300_N via the camera 130, where N is a positive integer. In step S220, the object detection device 100 may, via the processor 110, scale the plurality of sensing images 300_1 to 300_N according to a scaling ratio r to generate a plurality of scaled sensing images 301_1 to 301_N. In this embodiment, the image size of the raw sensing images 300_1 to 300_N may be, for example, 1920×1080 pixels, and the image size of the scaled sensing images 301_1 to 301_N may be, for example, 1024×576 pixels. However, the image size of the raw sensing images and the scaling ratio r of the present invention are not limited thereto. In one embodiment, the scaling ratio r may be, for example, 0.5. Furthermore, in another embodiment, the object detection device 100 may not scale the original sensing images 300_1 - 300_N (ie, the scaling ratio r may be set to 1).
[0021] In step S230, the object detection device 100, through the processor 110, defines the entire image area of each of the plurality of first sensor images 302_1-302_M among the plurality of scaled original sensor images 301_1-301_N as a first range of interest I1, where M is a positive integer. In step S240, the object detection device 100, through the processor 110, defines a partial image area of each of the plurality of second sensor images 303_1-303_P among the plurality of scaled original sensor images 301_1-301_N as a second range of interest I2. The object detection device 100 also uses the processor 110 to crop the plurality of third sensor images 304_1-304_P based on the second range of interest I2 of each of the plurality of second sensor images 303_1-303_P, where P is a positive integer. The second range of interest I2 can be, for example, a predetermined range in the center of the sensor image, allowing the object detection device 100 to focus on the target object directly in front of the camera 130.
[0022] Matching reference Figure 4 , Figure 4 is a schematic diagram of cropping a sensor image according to an embodiment of the present invention. For example, the second sensor image 303 (hereinafter collectively referred to as 303_1 to 303_P) may have an image size (in pixels) of width Wf × height Hf, and the second region of interest I2 may have an image size (in pixels) of width Wc × height Hc. Therefore, the distance between the lower edge of the second region of interest I2 and the lower image boundary of the second sensor image 303, and the distance between the upper edge of the second region of interest I2 and the upper image boundary of the second sensor image 303, are both (Hf - Hc) / 2. The distance between the left edge of the second region of interest I2 and the left image boundary of the second sensor image 303, and the distance between the right edge of the second region of interest I2 and the right image boundary of the second sensor image 303, are both (Wf - Wc) / 2. Therefore, the processor 110 can crop the second sensor image 303 according to the aforementioned image size parameters and distance parameters to generate a corresponding third sensor image. However, the position and range of the second region of interest I2 of the present invention are not limited to the aforementioned example. In one embodiment, the second region of interest I2 may be cropped from other regions of the complete image according to different object detection requirements.
[0023] In this embodiment, the first sensor images 302_1-302_M may be, for example, images from odd-numbered frames among the scaled plurality of sensor images 301_1-301_N, and the second sensor images 303_1-303_P may be, for example, images from even-numbered frames among the scaled plurality of sensor images 301_1-301_N. In other words, the odd-numbered frames retain a full image area to minimize loss of critical information when the distance to a target object in front is relatively close (e.g., a large truck), thereby maximizing the ability to capture the complete outline of the object image. However, in one embodiment, based on different object detection requirements, the odd-numbered and even-numbered frames may be sensor images cropped from the scaled plurality of original sensor images 301_1-301_N according to two different ranges of interest. In another embodiment, the processor 110 may further separate the scaled original sensor images 301_1 to 301_N into more groups for different cropping (e.g., three groups corresponding to the 1st, 4th, 7th, ... frames, the 2nd, 5th, 8th, ... frames, and the 3rd, 6th, 9th, ... frames) according to more ranges of interest (e.g., three groups corresponding to the 1st, 4th, 7th, ... frames, and the 3rd, 6th, 9th, ... frames, respectively), rather than being limited to the aforementioned classification method of odd-numbered frames and even-numbered frames.
[0024] Next, before the plurality of first sensor images 302_1-302_M and the plurality of third sensor images 304_1-304_P are input into the deep neural network learning model 121, the processor 110 may first resize the plurality of first sensor images 302_1-302_M and the plurality of third sensor images 304_1-304_P to the same image size before inputting them into the deep neural network learning model 121. In one embodiment, the plurality of first sensor images 302_1-302_M and the plurality of third sensor images 304_1-304_P may be uniformly scaled down to a pixel area size of 512×288 pixels, for example, but the present invention is not limited thereto.
[0025] In step S250, the object detection device 100 may input the plurality of first sensor images 302_1-302_M and the plurality of third sensor images 304_1-304_P into the deep neural network learning model 121 via the processor 110, so that the deep neural network learning model 121 outputs a plurality of position information and a plurality of image sizes of the target object image in each of the plurality of first sensor images 302_1-302_M and the plurality of third sensor images 304_1-304_P. The target object in each of the first sensor images 302_1-302_M and each of the third sensor images 304_1-304_P may be one or more. In this embodiment, the deep neural network learning model 121 may be pre-trained to enable it to recognize target object images in images and output image information of the target object image in each sensor image, such as position information and image size. Notably, the position information may be the coordinates of a vertex of the target object image in each sensor image, and the image size may be the width and height of the target object image in each sensor image.
[0026] For example, with reference Figure 5 , Figure 5 FIG. 4 is a schematic diagram of analyzing an object image in a sensing image according to an embodiment of the present invention. Figure 5 Taking a single target object image in a sensed image as an example, in other embodiments, the sensed image may also contain multiple target object images. Processor 110 can identify target object image 506 in sensed image 505. Assuming the region of interest (ROI) is the entire image area (Wf × Hf) as the output of deep neural network learning model 121, the vertex coordinates of target object image 506 in sensed image 505 are (Xo, Yo) = (x × Wf, y × Hf), where the coordinate origin (0, 0) is the upper left corner of sensed image 505. The width of target object image 506 in sensed image 505 is Wo = w × Wf, and the height is Ho = h × Hf. Taking the output of deep neural network learning model 121 as an example, where the region of interest is a Wc×Hc area cropped from the center of the full image area, the vertex coordinates of target object image 506 in sensed image 505 are (Xo, Yo) = (x×Wc+(Wf-Wc) / 2, y×Hc+(Hf-Hc) / 2). The width of target object image 506 in sensed image 505 is Wo = w×Wc, and the height is Ho = h×Hc.
[0027] Specifically, the deep neural network learning model 121 outputs image information (x, y, w, h) for each target object image identified in the plurality of first sensor images 302_1-302_M and the plurality of third sensor images 304_1-304_P. (x, y) corresponds to the normalized position information of the target object image in the sensor image, and (w, h) corresponds to the normalized image size of the target object image in the sensor image. Therefore, the processor 110 can obtain the position information and image size of each target object image in the sensor image based on the above formulas. In other words, the processor 110 can calculate the position information and image size of each target object image in the first sensor images 302_1-302_M and the third sensor images 304_1-304_P in the scaled original sensor images 301_1-301_N, respectively, based on the above formulas.
[0028] Furthermore, in one embodiment, the image information output by the deep neural network learning model 121 may also include the object type of each target object image in the plurality of first sensor images 302_1-302_M and the plurality of third sensor images 304_1-304_P, such as a small car or a large truck. Furthermore, in another embodiment, the processor 110 may further execute an image tracking module to track each target object image in the plurality of first sensor images 302_1-302_M and the plurality of third sensor images 304_1-304_P, thereby ensuring stable detection of the target object image and enhancing the reliability of the detection result. The image tracking module may be stored in the storage unit 120 and may, for example, utilize the Lucas-Kanade optical flow algorithm, but the present invention is not limited thereto. In another embodiment, the processor 110 may further execute an image smoothing module to detect the position and size of each target object image in the first and third sensor images, thereby smoothing the detected image positions and sizes of the target object images in the plurality of sensor images. Based on this, a stable position, height, and width of the target object image, as well as object type information, may be obtained. The image smoothing module may be stored in the storage unit 120 and may utilize, for example, a Kalman filtering algorithm, but the present invention is not limited thereto.
[0029] Figure 6 FIG1 is a flow chart of calculating the horizon height coordinates in a sensing image according to an embodiment of the present invention. Figure 1 、 Figure 5 as well as Figure 6Following step S250, the object detection device 100 may execute steps S610-S630 on each scaled raw sensor image 301_1-301_N to calculate the horizon height coordinates in each scaled raw sensor image 301_1-301_N. In step S610, the object detection device 100 may, through the processor 110, obtain the actual physical width of each target object based on the object type of each target object image in the sensor image. For example, if the object type of the target object image is a small family car, the processor 110 may obtain the actual physical width of the target object (empirical value Wp) as 260 centimeters (cm). If the object type of the target object image is a mini family car, the processor 110 may obtain the actual physical width of the target object (empirical value Wp) as 180 centimeters. Alternatively, if the target object image is of a mini family car, the processor 110 may determine that the actual physical width of the target object (empirical value Wp) is 300 cm. If the target object image is of a compact family car, the processor 110 may determine that the actual physical width of the target object (empirical value Wp) is 350 cm.
[0030] In step S620, the object detection device 100 may calculate, via the processor 110, the horizon height coordinate Yh (in pixels) corresponding to each target object image in the sensed image based on the installation height (Hc) (in centimeters) of the camera 130, the bottom height coordinate (Yo) of each target object image, the image width (Wo) of each target object image, and the actual physical width (Wp) of each target object. In this embodiment, the processor 110 may execute the following formula (1) to obtain the horizon height coordinate Yh corresponding to each target object image.
[0031] Yh=Yo-Hc×Wo / Wp…………Formula (1)
[0032] In step S630, the object detection device 100 may smooth the plurality of horizon height coordinates through the processor 110. For example, the smoothing process may be performed on the horizon height coordinates corresponding to each target object image in the sensed image. The smoothing process may further be performed on the horizon height coordinates in the plurality of sensed images (e.g., sensed images of preceding and following frames) to eliminate errors in the calculated horizon positions corresponding to each target object. In this embodiment, the smoothing process may be performed, for example, by using information such as the calculated horizon positions corresponding to the plurality of target object images or the horizon positions of the scaled original sensed images of preceding and following frames to perform an arithmetic average operation or a weighted average operation on the horizon height coordinates obtained by the above formula (1) to obtain the current frame horizon height coordinate Yh_f.
[0033] Figure 7FIG. 1 is a flowchart of calculating the actual object distance according to an embodiment of the present invention. Figure 1 、 Figure 5 as well as Figure 7 Following step S630, the object detection device 100 may execute the following steps S710-S720 to calculate the actual object distance of each target object in the sensed image. In step S710, the object detection device 100 may calculate the zoomed focal length information of the camera 130 based on the focal length F of the camera 130 and the aforementioned zoom ratio r via the processor 110. In this embodiment, the processor 110 may, for example, execute the following formula (2) to obtain the zoomed focal length information (F') of the camera 130 (in pixels).
[0034] F ′ =F×r…………Formula (2)
[0035] In step S720, the object detection device 100 may calculate, via the processor 110, the actual object distance (d) (in centimeters) of each target object based on the current frame horizon height coordinate Yh_f, the zoomed focal length information (F') of the camera 130, the installation height (Hc) of the camera 130, and the bottom height coordinate (Yo) of each target object image. In this embodiment, the processor 110 may, for example, perform the following equation (3) to obtain the actual object distance (d) of each target object.
[0036] d=F'×Hc / (Yo-Yh_f)…………Formula (3)
[0037] However, in one embodiment, the object detection device 100 may also not be Figure 6 and Figure 7 The actual object distance (d) is obtained by the process of using the camera 130 installed on the vehicle. Take the case where the camera 130 is installed on a vehicle as an example. If the camera 130 is uniformly installed at a fixed position on a vehicle of the same design (for example, the object detection device 100 is uniformly installed by the vehicle manufacturer), the parameters such as the focal length and installation position of the camera 130 are fixed. Therefore, the processor 110 can also directly measure / correct the image width (Wo) of the target object image or / and the correspondence between the bottom edge height coordinate (Yo) of the target object image and the actual object distance (d) during the vehicle production process, wherein, for example, a lookup table is established based on the correspondence. In this way, the processor 110 can search the lookup table according to at least one of the bottom edge height coordinate (Yo) of the target object image and the image width (Wo) of the target object image to directly obtain the actual object distance (d) of the target object.
[0038] In summary, the object detection device and object detection method of the present invention can be effectively applied to the detection of vehicles in front of the vehicle by means of real-time image detection, thereby providing highly reliable front object detection and distance detection functions. In addition, the object detection device and object detection method of the present invention can also be used in an Advanced Driving Assistant System (ADAS), such as a Forward Collision Warning System (FCW), to provide assisted driving and collision warning functions. The present invention also has the advantages of low computational complexity and independence from calibration, is suitable for the computing power of the vehicle-mounted system, and can meet the real-time requirements of object detection.
[0039] The above description is only a preferred embodiment of the present invention, but it is not intended to limit the scope of the present invention. Anyone familiar with this technology can make further improvements and changes on this basis without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be based on the scope defined by the claims of this application.
Claims
1. An object detection device, characterized in that: include: A camera obtains a plurality of raw sensory images; A storage unit for storing multiple modules; as well as A processor is coupled to the storage unit and executes the plurality of modules to perform the following operations: The processor scales the plurality of original sensing images according to a scaling ratio to generate a plurality of scaled sensing images; The processor defines an entire image area of each of a plurality of first sensing images among the plurality of sensing images as a first range of interest, wherein the plurality of first sensing images are sensing images of odd-numbered frames among the plurality of sensing images; The processor defines a partial image area of each of a plurality of second sensing images among the plurality of sensing images as a second range of interest, and crops each of the plurality of second sensing images according to the second range of interest of each of the plurality of second sensing images to generate a plurality of third sensing images corresponding to the plurality of second sensing images, wherein the plurality of second sensing images are sensing images of even-numbered frames among the plurality of sensing images; The processor inputs the plurality of first sensing images and the plurality of third sensing images into a deep neural network learning model, so that the deep neural network learning model outputs image information of the target object image in the plurality of first sensing images and the plurality of third sensing images, wherein the image information includes an image size and an object type of the target object image in the plurality of first sensing images and the plurality of third sensing images, and the image size includes a width and a height of the target object image in the plurality of first sensing images and the plurality of third sensing images; and The zoomed focal length information of the camera is calculated based on the focal length of the camera and the zoom ratio. The actual physical width of the target object is obtained based on the object type. The horizon height coordinate corresponding to the target object image is calculated based on the installation height of the camera, the height coordinate of the target object image, the image width of the target object image, and the actual physical width. The horizon height coordinates corresponding to multiple target object images are smoothed to obtain the horizon height coordinate of the current frame. The actual object distance of the target object in the target object image is calculated based on the horizon height coordinate of the current frame, the zoomed focal length information, the installation height, and the height coordinate.
2. The object detection device according to claim 1, wherein: The processor first adjusts the plurality of first sensing images and the plurality of third sensing images to the same image size and then inputs the images into the deep neural network learning model.
3. The object detection device according to claim 1, wherein: The second range of interest is a portion of the image area in the middle of the plurality of second sensing images.
4. The object detection device according to claim 1, wherein: The processor executes an image tracking module to track the target object image in the plurality of first sensing images and the plurality of third sensing images respectively.
5. The object detection device according to claim 1, wherein: The image information further includes position information of the target object image in the plurality of first sensing images and the plurality of third sensing images.
6. An object detection method, characterized in that: include: Acquire multiple original sensor images through a camera; scaling the plurality of original sensing images according to a scaling ratio to generate a plurality of scaled sensing images; defining an entire image area of each of a plurality of first sensing images among the plurality of sensing images as a first range of interest, wherein the plurality of first sensing images are sensing images of odd-numbered frames among the plurality of sensing images; defining a partial image area of each of a plurality of second sensing images among the plurality of sensing images as a second range of interest, and cropping each of the plurality of second sensing images according to the second range of interest of each of the plurality of second sensing images to generate a plurality of third sensing images corresponding to the plurality of second sensing images, wherein the plurality of second sensing images are sensing images of even-numbered frames among the plurality of sensing images; Inputting the plurality of first sensing images and the plurality of third sensing images into a deep neural network learning model, so that the deep neural network learning model outputs image information of the target object image in the plurality of first sensing images and the plurality of third sensing images, wherein the image information includes an image size and an object type of the target object image in the plurality of first sensing images and the plurality of third sensing images, and the image size includes a width and a height of the target object image in the plurality of first sensing images and the plurality of third sensing images; Calculating zoomed focal length information of the camera according to the focal length of the camera and the zoom ratio; Get the actual physical width of the target object according to the object type; Calculating a horizon height coordinate corresponding to the target object image according to the installation height of the camera, the height coordinate of the target object image, the image width of the target object image, and the actual physical width; Smoothing the horizon height coordinates corresponding to the plurality of target object images to obtain the horizon height coordinates of the current frame; and The actual object distance of the target object in the target object image is calculated according to the current frame horizon height coordinate, the zoomed focal length information, the installation height, and the height coordinate.
7. The object detection method according to claim 6, wherein: The step of inputting the plurality of first sensing images and the plurality of third sensing images into the deep neural network learning model comprises: The plurality of first sensing images and the plurality of third sensing images are first adjusted to the same image size and then input into the deep neural network learning model.
8. The object detection method according to claim 6, wherein: The second range of interest is a portion of the image area in the middle of the plurality of second sensing images.
9. The object detection method according to claim 6, wherein: The step of inputting the plurality of first sensing images and the plurality of third sensing images into the deep neural network learning model comprises: An image tracking module is executed to track the target object image in the plurality of first sensing images and the plurality of third sensing images respectively.
10. The object detection method according to claim 6, wherein: The image information further includes position information of the target object image in the plurality of first sensing images and the plurality of third sensing images.
Citation Information
Patent Citations
Enhanced object detection for autonomous vehicles based on field view
US20200175326A1
Objection detection using images and message information
US20220405952A1
Continuous training of an object detection and classification model for varying environmental conditions
US20230139682A1
Navigation systems and methods for determining object dimensions
WO2021136967A2