Cross-scale panoramic perception system and cross-scale target detection method for panoramic images

CN115331074BActive Publication Date: 2026-09-22YANGTZE DELTA REGION INST OF TSINGHUA UNIV ZHEJIANG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210907265.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2026-09-22
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

[0003]双鱼眼镜头和多目摄像机都具有结构简单、部署方便等特点,但是其均存在价格昂贵、边缘成像质量差、分辨率低等缺点,导致对目标检测精度较差;多普通摄像机全景布控要求在不同地点进行相机布控,在大场景下布控点过多导致成本过高、存在死角、部署结构比较离散

Benefits of technology

[0033]本发明实施例提出了一种面向大场景全景目标检测的非结构化亿像素级别跨尺度全景感知系统和目标检测方法;非结构化亿像素级别感知系统结构简单、部署方便,解决了跨尺度的高分辨全景成像的问题,满足了大场景全景高精度目标检测的数据采集和训练条件,最后通过本发明实施例设计的跨尺度目标检测方法能够实现亿像素级的全景目标检测;本发明实施例提供的目标检测方法通过局部场景感知图像进行目标的大尺度检测,然后通过感知设备标定关系映射到全局场景感知图像,最终实现跨尺度高精度的目标检测;本发明相比现有的全景目标检测解决方案的像素和精度更高。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115331074B_ABST
    Figure CN115331074B_ABST
Patent Text Reader

Abstract

The application discloses a cross-scale panoramic perception system and a cross-scale target detection method of panoramic images. The cross-scale panoramic perception system comprises multiple groups of fixing devices arranged in a regular polygon, each group of fixing devices is fixed with a camera array fixing support, and each camera array fixing support is fixed with a group of camera arrays; the multiple groups of camera arrays correspond to multiple peripheral regions covering the circumference of the regular polygon respectively; each group of camera arrays comprises two local scene perception devices and one global scene perception device; and the global scene perception regions of two adjacent groups of camera arrays have a perception overlap region. The perception system provided by the application has simple structure and is convenient to deploy, solves the problem of cross-scale high-resolution panoramic imaging, meets the data collection and training conditions of large-scene panoramic high-precision target detection, and finally, the cross-scale target detection method designed by the application can realize high-precision panoramic target detection of the order of 100 million pixels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and in particular to a cross-scale panoramic perception system and a cross-scale target detection method for panoramic images. Background Technology

[0002] Based on current technology research, there are three main ways to achieve panoramic target detection in large scenes at the visual sensing level: one is to obtain 360-degree field of view imaging through dual fisheye lenses; the second is to obtain panoramic imaging through panoramic control of multiple ordinary cameras; and the third is to achieve panoramic imaging through multi-view cameras.

[0003] Both fisheye lenses and multi-view cameras are characterized by simple structure and convenient deployment, but they also have disadvantages such as high price, poor edge imaging quality, and low resolution, resulting in poor target detection accuracy. Panoramic deployment of multiple ordinary cameras requires camera deployment at different locations. In large scenes, too many deployment points lead to excessive costs, blind spots, and a relatively discrete deployment structure.

[0004] Therefore, the inventors recognized that there is an urgent need for a large-scene panoramic hardware solution and algorithm that is simple in structure, easy to deploy, and has high-precision target detection performance to meet the needs of large-scene panoramic target detection. Summary of the Invention

[0005] Based on this, and to address the aforementioned technical problems, a cross-scale panoramic perception system and a cross-scale target detection method for panoramic images are provided.

[0006] In a first aspect, a cross-scale panoramic perception system includes multiple sets of fixed devices arranged in a regular polygon, each set of fixed devices having a camera array fixing bracket fixed on it, and each camera array fixing bracket having a set of camera arrays fixed on it; the multiple sets of camera arrays respectively cover multiple peripheral areas of the regular polygon circumference.

[0007] Each camera array includes three sensing and imaging devices, namely two local scene sensing devices and one global scene sensing device, with the global scene sensing device fixed between the two local scene sensing devices.

[0008] The global scene perception areas of two adjacent camera arrays have overlapping perception areas; for each camera array, the field of view of the global scene perception device covers the field of view of the two adjacent local perception devices, and the vertical field of view of the global perception device is greater than or equal to twice that of the local scene perception device.

[0009] Optionally, the regular polygon is a regular octagon, and there are a total of eight sets of fixing devices.

[0010] Further optionally, the field-of-view optical axes of all sensing and imaging devices are coplanar with the regular octagon, and the backward extensions of the field-of-view optical axes of all global scene sensing devices pass through the center of the regular octagon.

[0011] Further optionally, the horizontal viewing angle of each global scene perception device is greater than or equal to 60°, and the resolution of the perceived image of each global scene perception device and each local scene perception device is greater than or equal to 14 million pixels.

[0012] Secondly, a cross-scale target detection method for panoramic images at the 100-megapixel level includes:

[0013] Step 1: Construct a target detection training data acquisition device, which includes a set of camera arrays in the cross-scale panoramic perception system provided in the first aspect and a camera array fixing bracket for fixing the camera arrays; calibrate the position of the perception device using the constructed target detection training data acquisition device, and train a cross-scale target detection model using the constructed target detection training data acquisition device.

[0014] Step two: Using the cross-scale panoramic perception system provided in the first aspect, images are acquired to obtain 8 global perception images and 16 local perception images, and the transformation matrix of the adjacent images of the global perception images is calculated; the 8 global perception images are then stitched together using the transformation matrix of the adjacent images of the global perception images to obtain a 360° panoramic stitched image I. G ;

[0015] Step 3: For the 16 locally sensed images acquired in Step 2, target detection is performed using the cross-scale target detection model trained in Step 1 to obtain the target's coordinate position in each locally sensed image. Based on the target's coordinate position in each locally sensed image, the target's coordinate position in the corresponding global sensed image is obtained. The target positions in the 8 global sensed images are then transformed using the transformation matrix of the adjacent images to obtain a 360° panoramic stitched image I. G The exact location of all targets.

[0016] Optionally, step one, which involves calibrating the location of the constructed target detection training data acquisition device and training a cross-scale target detection model using the constructed target detection training data acquisition device, includes:

[0017] The position calibration of two local scene perception devices and a global scene perception device in the target detection training data acquisition device is performed using the feature point matching method, and the mapping matrices M1 and M2 of the two local scene perception devices in the camera array relative to the global scene perception device are obtained respectively.

[0018] Image data was collected at a specific location using a target detection training data acquisition device to obtain a local perception image dataset and a global perception image dataset. The targets collected in the local perception image dataset and the global perception image dataset were labeled with local bounding boxes. The labeled local perception image dataset was divided into a training set, a test set and a validation set according to a preset ratio. The labeled global perception image dataset was not used for training, but was used as a GrounTruth to constrain the training of the cross-scale target detection model.

[0019] Utilizing existing object detection networks and their loss function L det The training set is used for training, and the loss function during training is defined as follows:

[0020] L = L det (local)+λL aet (global)

[0021] Where L det (local) represents the loss function used by existing object detection networks to train on locally perceptual image datasets, L det (global) represents the loss function using a labeled globally-aware dataset as a GrounTruth constraint, and λ represents L det (global) weighting coefficient; L det The formula for calculating (global) is:

[0022] L det (global) = L det (PredM i )

[0023] Where Pred is the target location predicted based on the local perception image; PredM i This indicates that Pred passes through the mapping matrix M. j Transform to a globally perceived image;

[0024] According to the pre-set training strategy, the model is trained iteratively until the loss function L converges, thus obtaining a cross-scale object detection model.

[0025] Further, optionally, step two specifically includes:

[0026] The cross-scale panoramic perception system provided in the first aspect is constructed, and image data from 8 camera arrays are synchronously acquired using a software timestamp-based synchronization triggering method. The 8 global images and 16 local images perceived by the 8 camera arrays at time t are denoted as follows: and

[0027] For a global sensing image sequence (G1, G2, ..., G8) at time t, feature extraction and matching algorithms are used to perform feature detection and matching on adjacent images in the sequence, respectively, to obtain the transformation matrix (TR) of the adjacent images of the global sensing image. 21 TR 32 TR 87 );

[0028] Using the transformation matrix (TR) of the adjacent images 21 TR 32 TR 87 Eight globally perceived images are stitched together to obtain a 360° panoramic stitched image I. G .

[0029] Further optionally, step three, obtaining the target's coordinates in the corresponding global sensing image based on the target's coordinates in each local sensing image, specifically involves mapping the target's coordinates in each local sensing image back to the corresponding global sensing image using mapping matrices M1 and M2, thus obtaining the target's coordinates in the corresponding global sensing image; the transformation of the target's position in the eight global sensing images using the transformation matrix of the adjacent images specifically involves using the transformation matrix (TR) of the adjacent images. 21 TR 32 TR 87 This transforms the target position in eight globally perceived images.

[0030] Further optionally, the feature point matching method is the SIFT algorithm; the specific locations include pedestrian plazas and intersections.

[0031] Further optionally, the preset ratio is 8:1:1; the target detection network is a Yolov5 neural network.

[0032] The present invention has at least the following beneficial effects:

[0033] This invention proposes an unstructured, megapixel-level, cross-scale panoramic perception system and target detection method for large-scene panoramic target detection. The unstructured, megapixel-level perception system is simple in structure and easy to deploy, solving the problem of high-resolution panoramic imaging across scales and meeting the data acquisition and training conditions for high-precision target detection in large-scene panoramic environments. Finally, the cross-scale target detection method designed in this invention can achieve megapixel-level panoramic target detection. The target detection method provided in this invention performs large-scale target detection using local scene perception images, and then maps them to global scene perception images through the calibration relationship of the perception devices, ultimately achieving high-precision target detection across scales. Compared with existing panoramic target detection solutions, this invention has higher pixel count and accuracy. Attached Figure Description

[0034] Figure 1 A schematic diagram of an unstructured, multi-scale panoramic perception system provided in one embodiment of the present invention;

[0035] Figure 2 This is a schematic diagram of the structure of a cross-scale target detection training data acquisition device provided in one embodiment of the present invention;

[0036] Figure 3 This is a flowchart illustrating a cross-scale target detection method for panoramic images at the 100-megapixel level, provided as an embodiment of the present invention.

[0037] Figure 4 This is a block diagram of the module architecture of a cross-scale target detection device for panoramic images with a resolution of hundreds of millions of pixels, provided as an embodiment of the present invention.

[0038] Explanation of reference numerals in the attached figures:

[0039] 1. Fixing device;

[0040] 2. Camera array; 21. First local scene sensing device; 22. Second local scene sensing device; 23. Global scene sensing device;

[0041] 3. Camera array mounting bracket. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0043] Example 1

[0044] In this embodiment, an unstructured cross-scale panoramic perception system is provided. The system adopts a polygon array design to realize panoramic perception imaging, including multiple sets of fixing devices arranged in regular polygons. Each set of fixing devices is fixed with a camera array fixing bracket, and each camera array fixing bracket is fixed with a set of camera arrays. The multiple sets of camera arrays respectively cover multiple peripheral areas of the regular polygons.

[0045] Each camera array includes three sensing and imaging devices, namely two local scene sensing devices and one global scene sensing device, with the global scene sensing device fixed between the two local scene sensing devices.

[0046] The global scene perception areas of two adjacent camera arrays have overlapping perception areas; for each camera array, the field of view of the global scene perception device covers the field of view of the two adjacent local perception devices, and the vertical field of view of the global perception device is greater than or equal to twice that of the local scene perception device.

[0047] Among them, regular polygons can be, but are not limited to, octagons, decagons, etc.

[0048] The unstructured megapixel-level perception system provided in this invention has a simple hardware structure and is easy to deploy. It solves the problem of high-resolution panoramic imaging across scales, meets the data acquisition and training conditions for large-scene panoramic high-precision target detection, and provides a foundation for large-scene panoramic high-precision target detection.

[0049] Example 2

[0050] In this embodiment, an unstructured, multi-scale panoramic perception system is provided, such as... Figure 1 As shown in (a), it includes eight sets of fixing devices 1 arranged in a regular octagon; that is, around the fixing devices 1, eight sets of camera arrays 2 are arranged around the polygonal array, and the outer periphery of the regular octagon is covered by the eight sets of camera arrays 2 respectively. Figure 1 In (a), 2 represents a camera array, and each camera array 2 contains three sensing and imaging devices. Specifically, as shown... Figure 1 As shown in (a) and (b), each set of fixing devices 1 is fixed with a camera array fixing bracket 3, and each set of camera arrays 2 is fixed with a set of camera arrays 2. Each set of camera arrays 2 includes three sensing and imaging devices. The three sensing and imaging devices specifically include two local scene sensing devices (first local scene sensing device 21 and second local scene sensing device 22) and one global scene sensing device 23. The global scene sensing device 23 is fixed between the first local scene sensing device 21 and the second local scene sensing device 22.

[0051] For each camera array 2, the field of view of the global scene sensing device 23 must be able to cover the field of view of two local scene sensing devices. This system includes 8 global sensing devices and 16 local sensing devices, and the resolution of the sensing images of each global scene sensing device 23 and each local scene sensing device is no less than 14 million pixels.

[0052] The characteristic of this system is that the global scene perception areas of the global scene perception devices 23 of two adjacent sets of camera arrays 2 need to have a certain degree of overlap, so as to... Figure 1Taking the 8-camera array 2 as an example, the horizontal field of view of a single global sensing device theoretically needs to reach more than 45°. In order to avoid blind spots and facilitate subsequent algorithm development, a certain size of the sensing overlap area needs to be guaranteed. Therefore, the horizontal field of view of each global sensing device should be no less than 60°. For a set of camera array 2, the vertical field of view of the global sensing device should be no less than twice that of the local sensing device. The local sensing devices of the system can be adjusted at any angle to achieve unstructured deployment.

[0053] In addition, such as Figure 2 As shown, a large-scene, cross-scale target detection training data acquisition device is provided to meet the training needs of panoramic perception target detection models. Similar to the unstructured, cross-scale panoramic perception system described above, it includes a set of camera array mounting brackets 3 and a set of camera arrays 2, specifically including two local scene perception devices (a first local scene perception device 21 and a second local scene perception device 22) and a global scene perception device 23. The configuration requirements of this training data acquisition device are similar to those for... Figure 1 The configuration requirements for the local scene sensing devices (first local scene sensing device 21 and second local scene sensing device 22) and the global scene sensing device 23 in the unstructured cross-scale panoramic perception system are the same. The field-of-view optical axes of the first local scene sensing device 21 and the second local scene sensing device 22 can be flexibly adjusted, as long as the global scene sensing device 23 can cover the field of view of the first local scene sensing device 21 and the second local scene sensing device 22.

[0054] exist Figure 2 In the diagram, A represents the assumed target of the perceived scene, B represents the schematic imaging result of the global scene perception target from the global scene perception device 23, C and D represent the schematic imaging results of the local scene perception targets from the first local scene perception device 21 and the second local scene perception device 22, respectively, and E represents the regions corresponding to C and D in B. Figure 2 It is evident that the target scale captured by the global scene perception device 23 and the local scene perception device are significantly different. The target size is smaller under global scene perception, resulting in poor target detection accuracy in the global scene perception image. Large-scale target detection can be performed using the local scene perception image, and the target can be mapped to the global scene perception image through the calibration relationship of the perception device, thereby achieving high-precision target detection across scales.

[0055] The unstructured megapixel-level perception system provided in this invention has a simple hardware structure and is easy to deploy. It solves the problem of high-resolution panoramic imaging across scales, meets the data acquisition and training conditions for large-scene panoramic high-precision target detection, and provides a foundation for large-scene panoramic high-precision target detection.

[0056] Example 3

[0057] In this embodiment, as Figure 3 As shown, a cross-scale target detection method for panoramic images with resolutions of hundreds of millions of pixels is provided, including the following steps:

[0058] Step S301, refer to the above embodiment two. Figure 1 and Figure 2 The hardware description, building the above Figure 2 The target detection training data acquisition device shown includes a set of camera arrays and a camera array fixing bracket for fixing the camera arrays in the cross-scale panoramic perception system provided in the above embodiment; the position of the sensing device is calibrated using the constructed target detection training data acquisition device, and a cross-scale target detection model is trained using the constructed target detection training data acquisition device.

[0059] Step S301 involves training a large-scale cross-object detection model. This includes calibrating the location of the constructed object detection training data acquisition device and training the cross-scale object detection model using this device.

[0060] (1) The position of the two local scene perception devices and the global scene perception device in the target detection training data acquisition device is calibrated using the feature point matching method, and the positions of the two local scene perception devices in the camera array are obtained. Figure 2 21 and 22 in the figure are respectively relative to the global scene perception device ( Figure 2 Mapping matrices M1 and M2 in (23); wherein the feature point matching method can be, but is not limited to, the SIFT algorithm;

[0061] (2) Use the constructed target detection training data acquisition hardware to carry out diverse large-scene data acquisition, that is, to collect image data in specific locations, including pedestrian squares, intersections, etc., to obtain local perception image datasets and global perception image datasets, and to label the targets collected in the local perception image datasets and global perception image datasets with local bounding boxes; divide the labeled local perception image datasets into training set, test set and validation set according to a preset ratio (8:1:1); the labeled global perception image datasets do not participate in training, but are used as GrounTruth to constrain the training of cross-scale target detection models;

[0062] (3) Utilize existing object detection networks such as the YOLOv5 neural network and its loss function L det The training set is used for training, and the loss function during training is defined as follows:

[0063] L = L det (local)+λLdet (global)

[0064] Where L de (local) represents the loss function used by existing object detection networks such as the Yolov5 neural network to train on a locally perceptual image dataset. det (global) represents the loss function using the labeled global-aware dataset as the GrounTruth constraint, which is the loss function based on the globally-aware labeled data constraint. λ represents L det (global) weighting coefficient; L det The formula for calculating (global) is as follows:

[0065] L det (global) = L det (PredM i )

[0066] Where Pred is the target location predicted based on the local perception image, which is mapped by the matrix M. j Transform to a globally perceived image, then calculate the loss between the transformed image and the globally perceived annotation result; PredM i This indicates that Pred passes through the mapping matrix M. i Transform to a globally perceived image; mapping matrix M i That is, mapping matrices M1 and M2;

[0067] (4) According to the pre-set training strategy (including learning rate, training batch, training epoch, optimizer, etc.), iteratively train the model until the loss function L converges to obtain a large-scale cross-scale target detection model.

[0068] Step S302: Using the cross-scale panoramic perception system provided in Embodiment 2 above, images are acquired to obtain 8 global perception images and 16 local perception images, and the transformation matrix of the adjacent images of the global perception images is calculated; the 8 global perception images are stitched together using the transformation matrix of the adjacent images of the global perception images to obtain a 360° panoramic stitched image I. G .

[0069] Step S302 involves performing a 100-megapixel-level panoramic image fusion calculation. Step S302 specifically includes:

[0070] (1) Construct the above embodiment two. Figure 1 The provided cross-scale panoramic perception system uses a software timestamp-based synchronization triggering method to synchronously acquire image data from eight camera arrays. Assuming that at a certain time t, the eight global images and sixteen local images sensed by the eight camera arrays are denoted as follows: and

[0071] (2) For the global sensing image sequence (G1, G2, ..., G8) at time t, feature extraction and matching algorithms are used to perform feature detection and matching on adjacent images in the sequence to obtain the transformation matrix (TR) of the adjacent images of the global sensing image. 21 TR 32 TR 87 This matrix can be used to stitch adjacent images together;

[0072] (3) Using the transformation matrix (TR) of the adjacent images 21 TR 32 TR 87 Eight globally perceived images are stitched together to obtain a 360° panoramic stitched image I. G .

[0073] Step S303: For the 16 local sensing images acquired in step S302, target detection is performed using the cross-scale target detection model trained in step S301 to obtain the target's coordinate position in each local sensing image. Based on the target's coordinate position in each local sensing image, the target's coordinate position in the corresponding global sensing image is obtained. The target positions in the 8 global sensing images are then transformed using the transformation matrix of the adjacent images to obtain a 360° panoramic stitched image I. G It pinpoints the exact location of all targets, enabling cross-scale target detection with a 100-megapixel panoramic view.

[0074] Step S303, which involves obtaining the target's coordinates in the corresponding global sensing image based on its coordinates in each local sensing image, specifically involves mapping the target's coordinates in each local sensing image back to the corresponding global sensing image using mapping matrices M1 and M2, thus obtaining the target's coordinates in the corresponding global sensing image. The step of transforming the target's position in the eight global sensing images using the transformation matrix of the adjacent images specifically involves using the transformation matrix (TR) of the adjacent images. 21 TR 32 TR 87 This transforms the target position in eight globally perceived images.

[0075] It should be understood that, although Figure 3 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 3At least some of the steps in the process may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.

[0076] Conventional target detection typically involves locating and recognizing targets within a local scene perceived by a monocular camera. High-precision panoramic target detection based on large scenes requires global perception of large scenes and high-precision detection of targets at multiple scales, which places higher demands on both panoramic imaging hardware and target detection algorithms.

[0077] Therefore, this invention proposes an unstructured, megapixel-level, cross-scale panoramic perception system hardware and target detection method for large-scene panoramic target detection. The unstructured, megapixel-level perception system hardware has a simple structure and is easy to deploy, solving the problem of high-resolution panoramic imaging across scales and meeting the data acquisition and training conditions for high-precision target detection in large-scene panoramic environments. Finally, the cross-scale target detection method designed in this invention can achieve megapixel-level panoramic target detection. The target detection method provided in this invention performs large-scale target detection using local scene perception images and maps them to global scene perception images through the calibration relationship of the perception devices, ultimately achieving high-precision cross-scale target detection, with higher pixel count and accuracy compared to existing panoramic target detection devices.

[0078] In one embodiment, such as Figure 4 As shown, a cross-scale target detection device for panoramic images with resolutions of hundreds of millions of pixels is provided, comprising the following modules:

[0079] The cross-scale target detection model training module 401 is used to build a target detection training data acquisition device. The training data acquisition device includes a set of camera arrays in the cross-scale panoramic perception system provided in Embodiment 2 above and a camera array fixing bracket for fixing the camera arrays. The module calibrates the position of the perception device in the built target detection training data acquisition device and trains the cross-scale target detection model using the built target detection training data acquisition device.

[0080] The panoramic stitching image generation module 402 is used to acquire images using the cross-scale panoramic perception system provided in Embodiment 2, obtaining 8 global perception images and 16 local perception images, and calculating the transformation matrix of the adjacent images of the global perception images; using the transformation matrix of the adjacent images of the global perception images, the 8 global perception images are stitched together to obtain a 360° panoramic stitched image I. G ;

[0081] The target location detection module 403 detects targets in each of the 16 local sensing images acquired by the panoramic stitching image generation module using a cross-scale target detection model trained by the cross-scale target detection model training module, thereby obtaining the target's coordinate position in each local sensing image. Based on the target's coordinate position in each local sensing image, the module obtains the target's coordinate position in the corresponding global sensing image. Finally, it transforms the target positions in the eight global sensing images using the transformation matrix of the adjacent images to obtain a 360° panoramic stitched image I. G The exact location of all targets.

[0082] Specific limitations regarding a cross-scale target detection device for megapixel-level panoramic images can be found in the above description of a cross-scale target detection method for megapixel-level panoramic images, and will not be repeated here. Each module in the aforementioned cross-scale target detection device for megapixel-level panoramic images can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0083] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program relating to all or part of the processes in the methods of the above embodiments.

[0084] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon relating to all or part of the processes in the methods of the above embodiments.

[0085] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0086] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0087] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A cross-scale target detection method for panoramic images with hundreds of millions of pixels, characterized in that, It is based on a cross-scale panoramic perception system, which includes multiple sets of fixed devices arranged in a regular polygon. Each set of fixed devices has a camera array fixing bracket, and each camera array fixing bracket has a set of camera arrays. The multiple sets of camera arrays respectively cover multiple peripheral areas of the regular polygon. Each camera array includes three sensing and imaging devices, namely two local scene sensing devices and one global scene sensing device, with the global scene sensing device fixed between the two local scene sensing devices. The global scene perception areas of two adjacent sets of camera arrays have overlapping perception areas; for each set of camera arrays, the field of view of the global scene perception device covers the field of view of two adjacent local perception devices, and the vertical field of view of the global perception device is greater than or equal to twice that of the local scene perception device. The cross-scale target detection method for megapixel-level panoramic images includes: Step 1: Construct a target detection training data acquisition device, which includes a set of camera arrays in the cross-scale panoramic perception system and a camera array fixing bracket for fixing the camera arrays; calibrate the position of the perception device using the constructed target detection training data acquisition device, and train a cross-scale target detection model using the constructed target detection training data acquisition device. Step two: Images are acquired using the aforementioned cross-scale panoramic perception system, resulting in 8 global perception images and 16 local perception images. The transformation matrix of the adjacent images of the global perception images is then calculated. The 8 global perception images are then stitched together using the transformation matrix of the adjacent images to obtain a 360° panoramic stitched image. ; Step 3: For the 16 locally sensed images acquired in Step 2, target detection is performed using the cross-scale target detection model trained in Step 1 to obtain the target's coordinate position in each locally sensed image. Based on the target's coordinate position in each locally sensed image, the target's coordinate position in the corresponding global sensed image is obtained. The target positions in the 8 global sensed images are then transformed using the transformation matrix of the adjacent images to obtain a 360° panoramic stitched image. The exact locations of all targets.

2. The method for cross-scale target detection of panoramic images at the 100-megapixel level according to claim 1, characterized in that, Step one, which involves calibrating the location of the constructed target detection training data acquisition device and training a cross-scale target detection model using the constructed target detection training data acquisition device, includes: The location of two local scene perception devices and a global scene perception device in the target detection training data acquisition device is calibrated using the feature point matching method, resulting in mapping matrices of the two local scene perception devices relative to the global scene perception device in the camera array. and ; Image data was collected at a specific location using a target detection training data acquisition device to obtain a local perception image dataset and a global perception image dataset. The targets collected in the local perception image dataset and the global perception image dataset were labeled with local bounding boxes. The labeled local perception image dataset was divided into a training set, a test set and a validation set according to a preset ratio. The labeled global perception image dataset was not used for training, but was used as a GrounTruth to constrain the training of the cross-scale target detection model. Utilizing existing object detection networks and their loss functions The training set is used for training, and the loss function during training is defined as follows: in This represents the loss function used by existing object detection networks to train on locally perceptual image datasets. This indicates that the labeled global-aware dataset is used as the loss function for the GrounTruth constraint. express Weighting coefficients; The calculation formula is: in The target location is predicted based on the locally perceived image; express After mapping matrix Transform to a globally perceived image; Based on a pre-defined training strategy, the model is trained iteratively until the loss function is reached. The convergence yields a cross-scale target detection model.

3. The cross-scale target detection method for megapixel-level panoramic images according to claim 2, characterized in that, Step two specifically includes: The image data of eight camera arrays will be synchronously acquired using a software timestamp-based synchronization triggering method. The 8 global images and 16 local images sensed by the 8 camera arrays at time are denoted as follows: and ; for Global perception image sequence at time step Feature extraction and matching algorithms are used to detect and match features in adjacent images of the sequence, respectively, to obtain the transformation matrix of adjacent images of the globally perceived image. ; Using the transformation matrix of the adjacent images Eight global perception images are stitched together to obtain a 360° panoramic stitched image. .

4. The cross-scale target detection method for megapixel-level panoramic images according to claim 3, characterized in that, Step three describes obtaining the target's coordinates in the corresponding global perception image based on its coordinates in each local perception image. Specifically, this is done using a mapping matrix. The coordinates of the target in each local sensing image are mapped back to the corresponding global sensing image to obtain the target's coordinates in the corresponding global sensing image; the target positions in the eight global sensing images are transformed using the transformation matrix of the adjacent images. It transforms the target position in 8 global perception images.

5. The cross-scale target detection method for megapixel-level panoramic images according to claim 2, characterized in that, The feature point matching method is the SIFT algorithm; the specific locations include pedestrian plazas and intersections.

6. The cross-scale target detection method for megapixel-level panoramic images according to claim 2, characterized in that, The preset ratio is 8:1:1; the target detection network is a Yolov5 neural network.

7. The method for cross-scale target detection of panoramic images at the 100-megapixel level according to claim 1, characterized in that, The regular polygon is a regular octagon, and there are a total of eight sets of fixing devices.

8. The cross-scale target detection method for megapixel-level panoramic images according to claim 7, characterized in that, The optical axes of the field of view of all sensing and imaging devices are coplanar with the regular octagon, and the backward extensions of the optical axes of the field of view of all global scene sensing devices pass through the center of the regular octagon.

9. The cross-scale target detection method for megapixel-level panoramic images according to claim 8, characterized in that, Each global scene perception device has a horizontal viewing angle greater than or equal to 60°, and the perceived image resolution of each global scene perception device and each local scene perception device is greater than or equal to 14 million pixels.

Citation Information

Patent Citations

  • Panoramic imaging system and method

    CN108399600A

  • Stitching method and apparatus for panoramic stereo video system

    CN108886611A