Target identification method, device and equipment based on panoramic spliced image

By chunking recognition and batch processing of high-resolution panoramic stitching images, the speed and accuracy of small target recognition in panoramic stitching images are solved, and efficient target recognition on resource-constrained devices are achieved.

CN120259732AActive Publication Date: 2025-07-04MINGFEI WEIYE TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510288331.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-04
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing object detection methods are difficult to balance the detection accuracy and speed of small objects in high-resolution panoramic stitching images, especially on devices with limited computing power, which is difficult to achieve real-time recognition.

Method used

The image to be identified is blocked according to the preset area grid, and a single block area is identified using the block recognition model. At the same time, the batch recognition model is used to batch identify the extracted target area, and a suitable batch recognition model is dynamically selected based on the number of target areas to be identified.

Benefits of technology

It significantly shortens the recognition time, improves detection speed and accuracy, and meets the real-time/near-real-time image target recognition requirements on embedded devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259732A_ABST
    Figure CN120259732A_ABST
Patent Text Reader

Abstract

The invention provides a target recognition method, device and equipment based on a panoramic spliced image, and relates to the technical field of computer vision and image processing, and the method comprises the steps: carrying out the blocking of a to-be-recognized image according to a preset region grid, and obtaining a to-be-recognized image region; identifying the to-be-identified image region by using the block identification model, obtaining all identification results, performing external expansion on each identification result to obtain all current target regions, and determining all to-be-identified target regions according to all identified regions and all current target regions; and determining a target batch recognition model according to the number of all the to-be-recognized target areas, recognizing all the to-be-recognized target areas by using the target batch recognition model to obtain all target recognition objects, and displaying all the target recognition objects in the to-be-recognized image. According to the method, the block recognition model is used for recognizing the single block region, and meanwhile, the batch recognition model is used for performing batch recognition on the extracted target region, so that the calculation amount of each time of recognition is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and image processing, and in particular to a method, device and equipment for object recognition based on panoramic stitching images. Background Art

[0002] In the field of image recognition, small target recognition (such as small object detection in drone monitoring, small lesion recognition in medical images, etc.) has always been a difficult problem. Small targets usually occupy a small area in the image and are often integrated with the background, resulting in low accuracy of general target detection methods when recognizing small targets. Deep learning models have significantly improved the performance of small target detection through methods such as multi-scale feature fusion and feature pyramid networks, but the balance between detection accuracy and detection speed for small targets in high-resolution panoramic images is still a challenge that needs to be solved.

[0003] Single Shot MultiBox Detector (SSD) is an efficient target detection method. It can detect targets of different sizes better by fusing feature maps of different scales. Because it predicts at different resolutions, it requires more detailed anchor box design and feature adjustment. Since YOLO v5, the real-time target detection algorithm YOLO can be configured to add a P2 detection layer to enhance the detection capability of small targets. SAHI, a slice-assisted hyper-inference library for small target detection, is mainly used to improve the detection accuracy when the target size in the data set is small. SAHI can be applied to various existing target detection networks, especially for scenes such as drone aerial images. Existing target detection methods, such as SSD and YOLO series algorithms, are competitive in detection accuracy and speed, but they still have shortcomings when processing high-resolution panoramic stitched images. For devices with limited computing power, it is still a challenge to perform real-time inference on multiple high-resolution images at the same time to meet the needs of small target recognition.

[0004] Therefore, a more efficient and accurate method is needed for identifying small targets in high-resolution panoramic stitching images. However, there is currently no technical solution that can solve the above technical problems, and there is no target recognition method, device or equipment based on panoramic stitching images. Summary of the invention

[0005] The present invention provides a target recognition method, device and equipment based on panoramic stitching images, aiming to solve the problem of small target recognition in high-resolution panoramic stitching images.

[0006] In a first aspect, the present invention provides a method for object recognition based on panoramic stitching images, comprising:

[0007] For the image to be recognized obtained at any moment, divide the image to be recognized according to a preset regional grid to obtain the image region to be recognized corresponding to the moment in the image to be recognized;

[0008] Use the block recognition model to recognize the image region to be recognized, obtain all the recognition results output by the block recognition model, expand each recognition result to obtain all the current target regions corresponding to the moment, and determine all the target regions to be recognized according to all the recognized regions corresponding to all the preset moments before the moment and all the current target regions;

[0009] Determine the target batch recognition model according to the number of all the target regions to be recognized, use the target batch recognition model to recognize all the target regions to be recognized, obtain all the target recognition objects output by the target batch recognition model, and display all the target recognition objects in the image to be recognized;

[0010] Wherein, each moment corresponds to processing one block in the preset regional grid, and the sum of the number of all the preset moments and the number of moments of the moment is the same as the number of blocks in the preset regional grid; all the recognized regions are obtained by using the block recognition model to recognize the image to be recognized corresponding to each preset moment, obtaining all the historical recognition results corresponding to each preset moment, and expanding all the historical recognition results corresponding to each preset moment to obtain all the historical target regions corresponding to all the preset moments.

[0011] According to the target recognition method based on the panoramic stitching image provided by the present invention, before dividing the image to be recognized according to the preset regional grid, the method further includes:

[0012] Use a panoramic pod to obtain camera images of cameras in different directions;

[0013] Stitch all the camera images to obtain a panoramic stitching image, and perform redundant removal on the overlapping parts of the panoramic stitching image to obtain the image to be recognized.

[0014] According to the target recognition method based on the panoramic stitching image provided by the present invention, the preset regional grid has 15 blocks. Among them, the cameras in different directions include a middle camera, a left camera, a right camera, an upper camera, and a lower camera, and the camera images obtained by each camera are divided into 3 blocks;

[0015] The step of dividing the image to be recognized according to the preset regional grid to obtain the image region to be recognized corresponding to the moment in the image to be recognized includes:

[0016] Partition each camera image in the image to be recognized according to a preset regional grid, and obtain the partition region corresponding to the partition at the moment.

[0017] Determine the partition region as the image region to be recognized corresponding to the moment in the image to be recognized.

[0018] According to the object recognition method based on panoramic stitching images provided by the present invention, before using the block recognition model to recognize the image region to be recognized and obtain all recognition results output by the block recognition model, the method further includes:

[0019] Train a first initial recognition model using the original sample image corresponding to the original resolution and the sample recognition result corresponding to each original sample image to obtain the block recognition model.

[0020] Both the image to be recognized and the image region to be recognized are of the original resolution.

[0021] According to the object recognition method based on panoramic stitching images provided by the present invention, before using the target batch recognition model to recognize all the image regions to be recognized, the method further includes:

[0022] Train a second initial recognition model using the sample recognition image corresponding to the preset resolution and the sample recognition object corresponding to each sample recognition image to obtain the target batch recognition model.

[0023] The preset resolution is less than the original resolution.

[0024] According to the object recognition method based on panoramic stitching images provided by the present invention, the expanding each recognition result to obtain all current target regions corresponding to the moment includes:

[0025] Determine the center point of the recognition result, and coincide the center point of a matrix region with a preset size with the center point of the recognition result.

[0026] Determine the matrix region with the preset size as the current target region corresponding to the recognition result.

[0027] Traverse all recognition results to obtain all current target regions corresponding to the moment.

[0028] According to the object recognition method based on panoramic stitching images provided by the present invention, the determining the target batch recognition model according to the number of all the image regions to be recognized includes:

[0029] In the case where the number of all the image regions to be recognized is less than or equal to 8, determine the target batch recognition model as an 8-batch recognition model.

[0030] When the number of all target regions to be recognized is greater than 8 and less than or equal to 16, determine the target batch recognition model as a 16-batch recognition model;

[0031] When the number of all target regions to be recognized is greater than 16 and less than or equal to 32, determine the target batch recognition model as a 32-batch recognition model;

[0032] When the number of all target regions to be recognized is greater than 32 and less than or equal to 64, determine the target batch recognition model as a 64-batch recognition model;

[0033] When the number of all target regions to be recognized is greater than 64, determine the target batch recognition model as a 128-batch recognition model.

[0034] According to the target recognition method based on panoramic stitching images provided by the present invention, where the target recognition object is a vehicle, before all target recognition objects are displayed in the image to be recognized, the method further includes:

[0035] For any target region to be recognized, if there is more than one recognized target recognition object, filter all target recognition objects in the target region to be recognized and retain one target recognition object.

[0036] In a second aspect, a target recognition device based on panoramic stitching images is provided, including:

[0037] A block division unit, which is configured to divide an image to be recognized obtained at any moment into blocks according to a preset regional grid to obtain the image region to be recognized corresponding to the moment in the image to be recognized;

[0038] A recognition unit, which is configured to use a block recognition model to recognize the image region to be recognized, obtain all recognition results output by the block recognition model, expand each recognition result to obtain all current target regions corresponding to the moment, and determine all target regions to be recognized according to all recognized regions corresponding to all preset moments before the moment and all current target regions;

[0039] A determination unit, which is configured to determine a target batch recognition model according to the number of all target regions to be recognized, use the target batch recognition model to recognize all target regions to be recognized, obtain all target recognition objects output by the target batch recognition model, and display all target recognition objects in the image to be recognized;

[0040] Among them, each moment corresponds to a block in the preset area grid, and the sum of the number of all preset moments and the number of moments of the moment is the same as the number of blocks in the preset area grid; all the identified areas are obtained by using the block recognition model to recognize the image to be recognized corresponding to each preset moment, obtaining all historical recognition results corresponding to each preset moment, and expanding all historical recognition results corresponding to each preset moment to obtain all historical target areas corresponding to all preset moments.

[0041] In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the target recognition method based on the panoramic stitching image is implemented.

[0042] In the present invention, the image to be recognized is divided into blocks according to a preset area grid, and the block recognition model is used to recognize a single divided area. At the same time, the batch recognition model is used to batch recognize the extracted target areas, which greatly reduces the amount of calculation for each recognition. This method effectively shortens the recognition time, improves the detection speed, dynamically selects a suitable batch recognition model according to the number of target areas to be recognized, further optimizes the recognition process, and improves the overall recognition efficiency; the image to be recognized and the area of the image to be recognized both maintain the original resolution, ensuring the accuracy of the recognition result. By optimizing the recognition process and model selection, the recognition time is significantly shortened. While ensuring a certain detection accuracy, it has a high detection speed and can meet the real-time / near-real-time image target recognition requirements on embedded devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1 is one of the flow diagrams of the target recognition method based on the panoramic stitching image provided by the present invention;

[0045] Figure 2 is the second flow diagram of the target recognition method based on the panoramic stitching image provided by the present invention;

[0046] Figure 3 is a schematic diagram of the image area directly below the panoramic stitching image provided by the present invention;

[0047] Figure 4It is a schematic diagram of the side image area based on the panoramic stitched image provided by the present invention;

[0048] Figure 5 It is a schematic diagram of dividing the directly downward image area provided by the present invention;

[0049] Figure 6 It is a schematic diagram of dividing the side image area provided by the present invention;

[0050] Figure 7 It is a schematic structural diagram of the target recognition device based on the panoramic stitched image provided by the present invention;

[0051] Figure 8 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0052] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0053] The present invention provides an acceleration method for small target recognition of panoramic stitched images at the original resolution, which can be widely applied to industries such as remote sensing image analysis, ecological environment monitoring, industrial production lines and item detection, military and national defense. For example, for a certain type of panoramic agile pod, its panoramic image is stitched by 5 images with a size of 2880*1860 pixels each taken by 5 cameras with 5 million pixels. After aligning the video frames of the cameras, 5 images need to be processed each time for recognition. If the 5 images are downsampled by 3 times in both length and width to 960*620 pixels for recognition, the recognition accuracy of small targets will be reduced; if no downsampling is performed and recognition is carried out at the original resolution, the inference calculation amount of its target recognition is about 65 times that of recognizing an image with a size of 640*640 pixels, and it is difficult to meet the application requirements of low latency.

[0054] Based on this technical problem, the present invention provides a target recognition method based on panoramic stitched images, Figure 1 It is one of the flow schematic diagrams of the target recognition method based on panoramic stitched images provided by the present invention. The target recognition method based on panoramic stitched images includes:

[0055] Step 101: For the image to be recognized obtained at any moment, divide the image to be recognized according to a preset regional grid to obtain the image area to be recognized corresponding to the moment in the image to be recognized;

[0056] Step 102: Use the block recognition model to recognize the to-be-recognized image region, obtain all recognition results output by the block recognition model, expand each recognition result, obtain all current target regions corresponding to the moment, and determine all to-be-recognized target regions according to all recognized regions corresponding to all preset moments before the moment and all the current target regions;

[0057] Step 103: Determine the target batch recognition model according to the number of all to-be-recognized target regions, use the target batch recognition model to recognize all to-be-recognized target regions, obtain all target recognition objects output by the target batch recognition model, and display all target recognition objects in the to-be-recognized image;

[0058] Wherein, each moment corresponds to a block in the preset regional grid, and the sum of the number of all preset moments and the number of moments is the same as the number of blocks in the preset regional grid; all the recognized regions are obtained by using the block recognition model to recognize the to-be-recognized image corresponding to each preset moment, obtaining all historical recognition results corresponding to each preset moment, and expanding all historical recognition results corresponding to each preset moment to obtain all historical target regions corresponding to all preset moments.

[0059] In step 101, before dividing the to-be-recognized image according to the preset regional grid, the method further includes:

[0060] Use a panoramic pod to obtain camera images of cameras in different orientations;

[0061] Stitch all the camera images to obtain a panoramic stitched image, and remove the redundancy of the overlapping parts of the panoramic stitched image to obtain the to-be-recognized image.

[0062] Optionally, the panoramic pod is deployed in a suitable position to ensure that a panoramic view of the required area can be captured. Multiple cameras in the panoramic pod (which may include middle, left, right, upper, and lower cameras) are triggered simultaneously or sequentially to capture images in different orientations. Each camera captures the image in its perspective and transmits these images to the processing system. The processing system first registers all the collected camera images, that is, determines their relative positions and poses for correct stitching, and uses a preset image stitching algorithm to seamlessly stitch all the registered images into a panoramic stitched image. In the panoramic stitched image, detect the redundant regions generated by the overlap of different camera images, and selectively remove these redundant parts according to the detection results of the overlapping regions to obtain a clear and non-overlapping to-be-recognized image.

[0063] Figure 3It is a schematic diagram of the image area directly below the panoramic stitched image provided by the present invention. Figure 4 It is a schematic diagram of the side image area based on the panoramic stitched image provided by the present invention. After installing this pod on the drone, the central camera points vertically downward, and the 4 side cameras tilt 45 degrees forward, backward, left, and right respectively. The resolution of all 5 cameras is 2880*1860 pixels, and the stitched panoramic image is as follows Figure 3 As shown, the area with yellow dots corresponds to the image captured by the central camera. For the stitched panoramic image, only the central square area is the effectively used area, and the rest of the area overlaps with the images captured by the side cameras. Therefore, these areas can be cropped, and only the central 1728*1728 pixel area is retained.

[0064] As Figure 4 shown, the area with yellow dots corresponds to the images captured by the side cameras. The distance seen in the upper half is very far. Through experimental statistics, the size of vehicle targets is generally below 16*16 pixels, and the tilt angle is large and it is very likely to be blocked, making it difficult to identify under the existing conditions. Considering that in actual needs, more attention is paid to the vehicle targets nearby, so the upper half with about one-third of the height can be discarded, and only the lower 2688*1120 pixel area is retained.

[0065] Optionally, the preset area grid has 15 blocks. Among them, the cameras in different orientations include the middle camera, the left camera, the right camera, the upper camera, and the lower camera, and the camera image obtained by each camera is divided into 3 blocks;

[0066] Said dividing the image to be recognized according to the preset area grid to obtain the image area to be recognized corresponding to the moment in the image to be recognized includes:

[0067] Dividing each camera image in the image to be recognized according to the preset area grid to obtain the block area corresponding to the block at the moment;

[0068] Determining the block area as the image area to be recognized corresponding to the moment in the image to be recognized.

[0069] Optionally, the present invention first defines a preset regional grid containing 15 blocks. This grid is designed to cover the entire image to be recognized, ensuring that for each frame of the image, it is processed according to such a preset regional grid, and each part can be effectively processed, thereby obtaining the divided block regions corresponding to the frame image at each moment. Considering that the panoramic image is composed of the images of five cameras in different orientations, namely the middle camera, the left camera, the right camera, the upper camera, and the lower camera, each camera image is evenly divided into 3 blocks. In this way, the entire panoramic image is divided into 15 blocks, which matches the preset regional grid. At any moment, according to the preset regional grid, the image obtained at that moment is selected for block processing, thereby obtaining the divided block region image corresponding to the divided block at that moment. Those skilled in the art understand that since the present invention involves dynamic image processing, a to-be-recognized image will be obtained at each moment, and then the preset regional grid is processed for the to-be-recognized image corresponding to that moment. Also, since each moment corresponds to processing one block in the preset regional grid, the sum of the number of all preset moments and the number of moments of that moment is the same as the number of blocks in the preset regional grid. Thus, the corresponding block that needs to be recognized and processed can be determined from the to-be-recognized image at that moment, and this block is "the to-be-recognized image region corresponding to the moment in the to-be-recognized image".

[0070] The present invention divides the panoramic image into 15 blocks and uses a block recognition model to process one of the divided blocks. This method significantly reduces the computational complexity of single processing and improves the overall processing efficiency. Each time it only processes one divided block region, enabling the algorithm to focus more on the local details of the image, thereby improving the accuracy of target recognition. Since only a small area is processed each time, this method effectively utilizes computational resources and avoids unnecessary waste, and is particularly suitable for resource-constrained environments.

[0071] In step 102, the present invention uses a pre-trained block recognition model to perform object recognition on the current image region to be recognized, collects all the recognition results output by the model, and expands each recognition result, which can be a certain range of expansion based on the boundary of the recognized object to obtain the current target region. Then, in combination with all the recognized regions at all previous preset times, these recognized regions are obtained by the same process, that is, using the block recognition model to recognize and expand the images to be recognized corresponding to a preset number of times before this moment. Those skilled in the art understand that, for example, if there are 15 blocks, 15 times can be set at this time, where each time corresponds to processing the image to be recognized corresponding to the specified block in the preset regional grid. At this time, cyclic recognition can be performed. For example, after all the recognition results at the current time are recognized, all the historical recognition results at the previous 14 times before the current time are determined, and then the all historical recognition results corresponding to the previous 14 times are expanded. After obtaining all the historical target regions corresponding to the previous 14 times, they are determined as all the recognized regions. Finally, all the recognized regions plus all the current target regions corresponding to the current time are jointly used as all the target regions to be recognized.

[0072] Optionally, before using the block recognition model to recognize the image region to be recognized and obtain all the recognition results output by the block recognition model, the method further includes:

[0073] Training a first initial recognition model using the original sample images corresponding to the original resolution and the sample recognition results corresponding to each original sample image to obtain the block recognition model;

[0074] Both the image to be recognized and the image region to be recognized are of the original resolution.

[0075] Optionally, the present invention collects sample images of the original resolution, which will be used to train the block recognition model. For each original sample image, the corresponding sample recognition results are labeled or generated. These results can be the object categories, position information, etc. in the image, specifically depending on the requirements of the recognition task. A suitable first initial recognition model is selected. This model can be any deep learning model suitable for the image recognition task, such as a convolutional neural network. The first initial recognition model is trained using the sample images of the original resolution and the corresponding sample recognition results. During the training process, the model will learn to extract useful features from the images of the original resolution and perform recognition based on these features. Through multiple iterations and optimizations, the trained block recognition model is obtained. The trained block recognition model is used to recognize the image region to be recognized, and the model will extract the features of this region and output the recognition results according to these features.

[0076] Optionally, before identifying all the target regions to be identified using the target batch identification model, the method further includes:

[0077] Training a second initial identification model using a sample identification image corresponding to a preset resolution and a sample identification object corresponding to each sample identification image to obtain the target batch identification model;

[0078] The preset resolution is less than the original resolution.

[0079] Optionally, collect sample identification images with a preset resolution, the resolution of these images is lower than the original resolution, usually to reduce computational complexity and improve processing speed. For each sample identification image, determine the corresponding sample identification object, such as a target vehicle. Select a suitable second initial identification model as the training starting point, which can be a convolutional neural network in deep learning or other models suitable for image recognition. Use the sample identification images with the preset resolution and the corresponding sample identification objects to train the second initial identification model. When there is a batch of target regions to be identified, use the trained target batch identification model for identification.

[0080] Optionally, expanding each recognition result to obtain all current target regions corresponding to the moment includes:

[0081] Determine the center point of the recognition result, and coincide the center point of a matrix region with a preset size with the center point of the recognition result;

[0082] Determine the matrix region with the preset size as the current target region corresponding to the recognition result;

[0083] Traverse all recognition results to obtain all current target regions corresponding to the moment.

[0084] Figure 5 is a schematic diagram of dividing the image region directly below provided by the present invention, Figure 6 is a schematic diagram of dividing the side image region provided by the present invention. In Figure 5 and Figure 6 the redundant region is the region on the image captured by the camera that does not need to be identified, such as Figure 5 and Figure 6 shown by the gray regions; the divided regions are the remaining regions after removing the redundant regions, which are the regions to be identified and are divided, such as Figure 5 and Figure 6 shown by the white regions, and there are 3 divided regions in both cases, but their shapes and sizes are different; the target region is to expand a certain range outside the identified target box, such as Figure 5 and Figure 6 shown by the cyan regions in.

[0085] Optionally, for each recognition result, first determine its center point, which can be obtained by calculating the geometric center of the recognition result area, that is, calculating the average value of the boundary coordinates of the area. Define a matrix area with a preset size. The size of this area can be set according to actual needs, and may depend on the size and shape of the target object or specific requirements of the recognition task. Coincide the center point of this preset-size matrix area with the center point of the recognition result to ensure that the matrix area expands with the recognition result as the center. Once the center point of the matrix area coincides with the center point of the recognition result, this matrix area is defined as the current target area corresponding to the recognition result.

[0086] In step 103, the present invention selects a suitable batch recognition model according to the number of target areas to be recognized, uses the selected batch recognition model to recognize all target areas to be recognized, obtains all target recognition objects, and displays these recognition objects in the original image to be recognized. In such an embodiment, the block recognition model is a model for recognizing a single sub-block area at a time, while the batch recognition model is a model for recognizing multiple target areas in batches at a time. The batch recognition model can process multiple target areas at one time, improving the recognition efficiency. Selecting a model according to the number of target areas can achieve more accurate and efficient recognition.

[0087] Optionally, determining the target batch recognition model according to the number of all target areas to be recognized includes:

[0088] In the case where the number of all target areas to be recognized is less than or equal to 8, determine the target batch recognition model as an 8-batch recognition model;

[0089] In the case where the number of all target areas to be recognized is greater than 8 and less than or equal to 16, determine the target batch recognition model as a 16-batch recognition model;

[0090] In the case where the number of all target areas to be recognized is greater than 16 and less than or equal to 32, determine the target batch recognition model as a 32-batch recognition model;

[0091] In the case where the number of all target areas to be recognized is greater than 32 and less than or equal to 64, determine the target batch recognition model as a 64-batch recognition model;

[0092] In the case where the number of all target areas to be recognized is greater than 64, determine the target batch recognition model as a 128-batch recognition model.

[0093] By dynamically selecting an appropriate batch recognition model according to the number of target regions to be recognized, the present invention can make more effective use of computing resources, avoid resource waste or overload, provide multiple batch model options, and can be flexibly adjusted according to the actual situation to meet different recognition requirements for different numbers. Selecting a batch model that matches the number of target regions to be recognized can ensure that the model operates in the best performance state, thereby improving the recognition accuracy and speed.

[0094] Optionally, the target recognition object is a vehicle. Before all target recognition objects are displayed in the image to be recognized, the method further includes:

[0095] For any target region to be recognized, if more than one target recognition object is recognized, filter all the target recognition objects in the target region to be recognized and retain one target recognition object.

[0096] Optionally, for any target region to be recognized, first check the number of target recognition objects recognized in this region. If more than one target recognition object is recognized in a certain target region to be recognized, start a filtering mechanism, such as selecting the recognition object with the highest confidence, selecting the recognition object closest to the center of the region, using more complex decision logic, such as comprehensively considering confidence, position, vehicle speed and other features. According to the result of the filtering mechanism, only retain one target recognition object and remove or ignore other recognized objects in this region.

[0097] The present invention aims to solve the problem of difficult real-time high-accuracy target recognition of multiple high-resolution images on resource-constrained embedded devices. Based on the relatively mature YOLO deep learning algorithm, method design and experimental verification are carried out. First, two types of target recognition models are trained and model deployment and inference optimization are performed. Each time of recognition, multiple high-resolution images are read, and then these images are divided into blocks to obtain block regions. One block region is directly recognized, and for the remaining block regions, multiple target regions are extracted by expanding a certain size outside the position and size of the target box recognized at the previous moment. These target regions are recognized, and finally, the result positions of the target recognition are integrated and back-calculated onto the original image.

[0098] The present invention divides the image to be identified into blocks according to a preset area grid, uses a block recognition model to identify a single block area, and uses a batch recognition model to batch recognize the extracted target areas, thereby greatly reducing the amount of calculation for each identification. This method effectively shortens the recognition time and improves the detection speed. It dynamically selects a suitable batch recognition model according to the number of target areas to be identified, further optimizes the recognition process, and improves the overall recognition efficiency. The image to be identified and the image area to be identified both maintain the original resolution, ensuring the accuracy of the recognition result. By optimizing the recognition process and model selection, the recognition time is significantly shortened. While ensuring a certain detection accuracy, it has a high detection speed and can meet the real-time / near real-time image target recognition requirements on embedded devices.

[0099] Figure 2 FIG. 2 is a flow chart of a target recognition method based on panoramic stitching images provided by the present invention. Figure 2 As shown:

[0100] Step 1: Create training sets for block regions and target regions. The long sides of all images in the block region training set should be as equal or similar as possible. For example, the long sides of the images collected in the experiment are all 1440 pixels. The original image size taken by the camera during the flight acquisition is 2880*1860 pixels. Each image is evenly cropped into 4 images, each of which is 1440*930 pixels. The long sides of the images in the target region training set should be as close as possible to the long sides of the cropped target regions in actual applications, because the smaller size scaling before inference can keep the feature information of the original image as much as possible and reduce the loss of recognition accuracy. For example, in the experiment, 60% of the images were randomly selected from the block region training set, and the sample images of the target region were cropped by expanding the target frame by 16 pixels and adding random position offsets. According to experimental statistics, the length and width of these target region images are mostly distributed between 42*121 pixels, and the added random position offset meets the application requirements of actual scenes.

[0101] Step 2: Train the block recognition model and the batch recognition model. The block recognition model is trained at the original image resolution, and the batch recognition model is trained at a resolution close to the average length and width of the image in the target area training set. For example, in the experiment, the block recognition model is trained at a size of 1440*1440 pixels, and the batch recognition model is trained at a size of 64*64 pixels.

[0102] Step 3: Export and deploy the block recognition model and batch recognition model. According to the actual application requirements, export the block recognition model into one or more single-batch inference models with different sizes, and export the batch recognition model into multiple inference models with the same size but different batches. For example, export the block recognition model into two single-batch models with sizes of 1728*576 and 896*1120, and export the batch recognition model into 5 models with 8 batches, 16 batches, 32 batches, 64 batches, and 128 batches. A total of 7 inference models are exported. Deploy the exported multiple inference models to the target inference platform. For example, deploy the 7 inference models exported in the previous step to the Jetson Orin NX embedded platform, load them all using TensorRT in the program, and prepare the necessary memory and video memory space.

[0103] Step 4: Read N high-resolution images each time for recognition, and read a set of images to be recognized into the memory. For example, read 5 images of 2880*1860 from the shared memory of the camera and record the memory pointer address.

[0104] Step 5: Crop the redundant areas in the images that do not need to be recognized. This step is not necessary and should be selected according to actual needs.

[0105] Step 6: Divide the remaining areas of the images to be recognized into blocks. For example, the area to be recognized in the image captured by the central camera is 1728*1728 pixels in size and is divided into three blocks of 1728*576 pixels. The areas to be recognized in the images captured by the 4 side cameras are 2688*1120 pixels in size, and each image is also divided into three blocks of 896*1120 pixels.

[0106] Step 7: Directly recognize one of the blocks with the block recognition model. For example, a total of 15 image areas are obtained after the previous step. Recognize the first block at time T1, the second block at time T2, the third block at time T3... the 15th block at time T15, the first block at time T16, the second block at time T17, the third block at time T18... and so on, in a cycle. Note that the size of the block area to be recognized should correspond to the input image size of the block recognition model.

[0107] Step 8: Extract multiple target regions from the remaining blocks. Taking the positions and sizes of the recognition results at the previous moment as references, expand them by a certain size and crop multiple target regions from the image at the current moment. If there are no recognized target bounding boxes at time T1, target regions cannot be cropped, and this step and the next step are not executed; at time T2, multiple target regions are cropped with the multiple recognition results in the first block at time T1 as references; at time T3, multiple target regions are cropped with the multiple recognition results in the first and second blocks at time T2 as references; at time T4, multiple target regions are cropped with the multiple recognition results in the first, second, and third blocks at time T3 as references... At time T15, multiple target regions are cropped with the multiple recognition results in the first to fourteenth blocks at time T14 as references; at time T16, multiple target regions are cropped with the multiple recognition results in the second to fifteenth blocks at time T15 as references; at time T17, multiple target regions are cropped with the multiple recognition results in the first and third to fifteenth blocks at time T16 as references, and repeat the execution.

[0108] Step 9: Use the batch recognition model to recognize the target regions. For example, when the number of target regions is less than or equal to 8, use the batch recognition model with 8 batches for recognition; when the number of target regions is less than or equal to 16, use the batch recognition model with 16 batches for recognition; when the number of target regions is less than or equal to 32, use the batch recognition model with 32 batches for recognition... and so on. When the number of target regions is greater than 128, first use the batch recognition model with 128 batches to recognize the first 128, and then recognize the rest.

[0109] Step 10: Combine the recognition results of all block regions, and reverse-calculate the target positions and sizes in the original image for the recognition results in Step 7 and Step 9. Whether the target is recognized by the block recognition model or the batch recognition model, the position of the target in the block region needs to be added to the cropping position of the block region in the original image; when the number of recognized targets on a target region is greater than 1, factors such as the maximum movement speed of the target also need to be considered, and filtered according to the target positions and size dimensions of the recognition results at the previous moment, and only 1 target is retained.

[0110] For a certain model of panoramic agile pod, identification tests are carried out on Jetson Orin NX. When identifying ground vehicle targets, before adopting this method, if no downsampling is performed for identification, the inference time for each identification of 5 images is about 207 milliseconds. After adopting this method, the inference time for identifying 5 images by name is about 35 milliseconds. For every 15 loops, a block area identification can be performed. Therefore, the time interval between two adjacent identifications of the same block area is about 500 milliseconds. During the flight experiment, when all identifications are carried out at the original image resolution, the identification accuracy is about 92.2%. After adopting this method, the identification accuracy is about 87.1%, a decrease of about 5.1%. However, it is much higher than the identification accuracy of small targets by methods such as downsampling, simplifying the network structure, or model quantization. The loss of target identification accuracy caused by adopting this method is mainly reflected in the missed detection rate and false detection rate of small targets during target area identification. Near the central area of an image with 64 * 64 pixels, identifying and finding a target larger than 16 * 16 pixels can generally ensure a relatively high accuracy. Therefore, this method or its variants have a certain degree of generality and can be generalized to other types of target identification or other deep learning tasks.

[0111] Figure 7 FIG. is a schematic structural diagram of an object recognition device based on panoramic mosaic images provided by the present invention. The object recognition device based on panoramic mosaic images includes a block unit 1. The block unit 1 is configured to, for a to-be-recognized image obtained at any moment, divide the to-be-recognized image according to a preset regional grid to obtain a to-be-recognized image area corresponding to the moment in the to-be-recognized image. The working principle of the block unit 1 can refer to the foregoing step 101 and will not be elaborated here.

[0112] The object recognition device based on panoramic mosaic images further includes an identification unit 2. The identification unit 2 is configured to use a block recognition model to identify the to-be-recognized image area, obtain all recognition results output by the block recognition model, expand each recognition result to obtain all current target areas corresponding to the moment, and determine all to-be-recognized target areas according to all recognized areas corresponding to all preset moments before the moment and all current target areas. The working principle of the identification unit 2 can refer to the foregoing step 102 and will not be elaborated here.

[0113] The object recognition device based on panoramic mosaic images further includes a determination unit 3. The determination unit 3 is configured to determine a target batch recognition model according to the number of all to-be-recognized target areas, use the target batch recognition model to identify all to-be-recognized target areas, obtain all target recognition objects output by the target batch recognition model, and display all target recognition objects in the to-be-recognized image. The working principle of the determination unit 3 can refer to the foregoing step 103 and will not be elaborated here.

[0114] Among them, each moment corresponds to a block in the preset area grid, and the sum of the number of all preset moments and the number of moments of the moment is the same as the number of blocks in the preset area grid; all the identified areas are determined by using the block recognition model to recognize the image to be recognized corresponding to each preset moment, obtaining all historical recognition results corresponding to each preset moment, and expanding each historical recognition result corresponding to each preset moment to obtain all historical target areas corresponding to all preset moments.

[0115] The present invention divides the image to be recognized according to a preset area grid, uses a block recognition model to recognize a single divided area, and simultaneously uses a batch recognition model to batch-recognize the extracted target areas, greatly reducing the computational amount of each recognition. This method effectively shortens the recognition time, improves the detection speed, dynamically selects a suitable batch recognition model according to the number of target areas to be recognized, further optimizes the recognition process, and improves the overall recognition efficiency; the image to be recognized and the area of the image to be recognized both maintain the original resolution, ensuring the accuracy of the recognition result. By optimizing the recognition process and model selection, the recognition time is significantly shortened. While ensuring a certain detection accuracy, it has a high detection speed and can meet the requirements of real-time / near-real-time image target recognition on embedded devices.

[0116] Figure 8 It is a schematic structural diagram of the electronic device provided by the present invention. As Figure 8As shown in the figure, the electronic device may include: a processor 110, a communication interface 120, a memory 130, and a communication bus 140. Among them, the processor 110, the communication interface 120, and the memory 130 complete mutual communication through the communication bus 140. The processor 110 may call logic instructions in the memory 130 to execute an object recognition method based on a panoramic stitched image. The method includes: for a to-be-recognized image obtained at any moment, dividing the to-be-recognized image according to a preset regional grid to obtain a to-be-recognized image region corresponding to the moment in the to-be-recognized image; using a block recognition model to recognize the to-be-recognized image region, obtaining all recognition results output by the block recognition model, expanding each recognition result to obtain all current target regions corresponding to the moment, and determining all to-be-recognized target regions according to all recognized regions corresponding to all preset moments before the moment and all the current target regions; determining a target batch recognition model according to the number of all to-be-recognized target regions, using the target batch recognition model to recognize all to-be-recognized target regions, obtaining all target recognition objects output by the target batch recognition model, and displaying all target recognition objects in the to-be-recognized image; where each moment corresponds to one block in the preset regional grid, and the sum of the number of all preset moments and the number of the moment is the same as the number of blocks in the preset regional grid; the all recognized regions are determined by using the block recognition model to recognize the to-be-recognized image corresponding to each preset moment, obtaining all historical recognition results corresponding to each preset moment, and expanding all historical recognition results corresponding to each preset moment to obtain all historical target regions corresponding to all preset moments.

[0117] In addition, the logic instructions in the above-mentioned memory 130 may be implemented in the form of software functional units. When sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of the technical solution may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0118] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute an object recognition method based on a panoramic stitched image provided by each of the above methods. The method includes: for a to-be-recognized image obtained at any moment, dividing the to-be-recognized image according to a preset regional grid to obtain a to-be-recognized image region corresponding to the moment in the to-be-recognized image; using a block recognition model to recognize the to-be-recognized image region, obtaining all recognition results output by the block recognition model, expanding each recognition result to obtain all current target regions corresponding to the moment, determining all to-be-recognized target regions according to all recognized regions corresponding to all preset moments before the moment and all the current target regions; determining a target batch recognition model according to the number of all to-be-recognized target regions, using the target batch recognition model to recognize all the to-be-recognized target regions, obtaining all target recognition objects output by the target batch recognition model, and displaying all the target recognition objects in the to-be-recognized image; wherein, each moment corresponds to a block in the preset regional grid, and the sum of the number of all preset moments and the number of the moment is the same as the number of blocks in the preset regional grid; all the recognized regions are determined by using the block recognition model to recognize the to-be-recognized image corresponding to each preset moment, obtaining all historical recognition results corresponding to each preset moment, and expanding all historical recognition results corresponding to each preset moment to obtain all historical target regions corresponding to all preset moments.

[0119] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the target recognition method based on panoramic stitching images provided by the above-mentioned various methods. The method includes: for the image to be recognized obtained at any moment, dividing the image to be recognized according to a preset regional grid to obtain the image region to be recognized corresponding to the moment in the image to be recognized; using a block recognition model to recognize the image region to be recognized, obtaining all recognition results output by the block recognition model, expanding each recognition result to obtain all current target regions corresponding to the moment, and determining all target regions to be recognized according to all recognized regions corresponding to all preset moments before the moment and all current target regions; determining a target batch recognition model according to the number of all target regions to be recognized, using the target batch recognition model to recognize all target regions to be recognized, obtaining all target recognition objects output by the target batch recognition model, and displaying all target recognition objects in the image to be recognized; wherein, each moment corresponds to a block in the preset regional grid, and the sum of the number of all preset moments and the moment is the same as the number of blocks in the preset regional grid; all recognized regions are determined by using the block recognition model to recognize the image to be recognized corresponding to each preset moment, obtaining all historical recognition results corresponding to each preset moment, and expanding all historical recognition results corresponding to each preset moment to obtain all historical target regions corresponding to all preset moments.

[0120] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A target recognition method based on panoramic stitching images, characterized in that, Including: For a to-be-recognized image obtained at any moment, dividing the to-be-recognized image according to a preset regional grid to obtain a to-be-recognized image region corresponding to the moment in the to-be-recognized image; Using a block recognition model to recognize the to-be-recognized image region, obtaining all recognition results output by the block recognition model, expanding each recognition result, obtaining all current target regions corresponding to the moment, and determining all to-be-recognized target regions according to all recognized regions corresponding to all preset moments before the moment and all the current target regions; Determining a target batch recognition model according to the number of all to-be-recognized target regions, using the target batch recognition model to recognize all to-be-recognized target regions, obtaining all target recognition objects output by the target batch recognition model, and displaying all target recognition objects in the to-be-recognized image; Wherein, each moment corresponds to a block in the preset regional grid, and the sum of the number of all preset moments and the number of the moment is the same as the number of blocks in the preset regional grid; all the recognized regions are determined by using the block recognition model to recognize the to-be-recognized image corresponding to each preset moment, obtaining all historical recognition results corresponding to each preset moment, and expanding all historical recognition results corresponding to each preset moment to obtain all historical target regions corresponding to all preset moments.

2. The object recognition method based on panoramic stitching images according to claim 1, wherein Before dividing the to-be-recognized image according to the preset regional grid, the method further includes: Using a panoramic pod to obtain camera images of cameras in different orientations; Stitching all the camera images to obtain a panoramic stitched image, and removing redundancy of overlapping parts of the panoramic stitched image to obtain the to-be-recognized image.

3. The object recognition method based on panoramic stitching images according to claim 2, wherein, The preset regional grid has 15 blocks. Among them, the cameras in different orientations include a middle camera, a left camera, a right camera, an upper camera, and a lower camera, and the camera image obtained by each camera is divided into 3 blocks; The dividing the to-be-recognized image according to the preset regional grid to obtain the to-be-recognized image region corresponding to the moment in the to-be-recognized image includes: Dividing each camera image in the to-be-recognized image according to the preset regional grid to obtain a divided region corresponding to the divided block at the moment; Determining the divided region as the to-be-recognized image region corresponding to the moment in the to-be-recognized image.

4. The object recognition method based on panoramic stitching images according to claim 1, characterized in that, Before using the block recognition model to recognize the to-be-recognized image region and obtaining all recognition results output by the block recognition model, the method further includes: Training a first initial recognition model by using an original sample image corresponding to an original resolution and a sample recognition result corresponding to each original sample image to obtain the block recognition model; Both the to-be-recognized image and the to-be-recognized image region are of the original resolution.

5. The object recognition method based on panoramic stitching images according to claim 4, wherein Before using the target batch recognition model to recognize all to-be-recognized target regions, the method further includes: Train the second initial recognition model using the sample recognition images corresponding to the preset resolution and the sample recognition objects corresponding to each sample recognition image to obtain the target batch recognition model; The preset resolution is less than the original resolution.

6. The object recognition method based on panoramic mosaic images according to claim 1, wherein The expanding each recognition result to obtain all current target regions corresponding to the moment includes: Determine the center point of the recognition result, and coincide the center point of the matrix region with the preset size with the center point of the recognition result; Determine the matrix region with the preset size as the current target region corresponding to the recognition result; Traverse all recognition results to obtain all current target regions corresponding to the moment.

7. The object recognition method based on panoramic stitching images according to claim 1, wherein The determining the target batch recognition model according to the number of all target regions to be recognized includes: When the number of all target regions to be recognized is less than or equal to 8, determine the target batch recognition model as an 8-batch recognition model; When the number of all target regions to be recognized is greater than 8 and less than or equal to 16, determine the target batch recognition model as a 16-batch recognition model; When the number of all target regions to be recognized is greater than 16 and less than or equal to 32, determine the target batch recognition model as a 32-batch recognition model; When the number of all target regions to be recognized is greater than 32 and less than or equal to 64, determine the target batch recognition model as a 64-batch recognition model; When the number of all target regions to be recognized is greater than 64, determine the target batch recognition model as a 128-batch recognition model.

8. The object recognition method based on panoramic stitching images according to claim 1, wherein When the target recognition object is a vehicle, before displaying all target recognition objects in the image to be recognized, the method further includes: For any target region to be recognized, if there is more than one recognized target recognition object, filter all target recognition objects in the target region to be recognized and retain one target recognition object.

9. An object recognition device based on a panoramic stitched image, characterized in that, Including: A block unit, which is used to block the image to be recognized according to a preset regional grid for any image to be recognized obtained at any moment, to obtain the image region to be recognized corresponding to the moment in the image to be recognized; A recognition unit, which is used to use the block recognition model to recognize the image region to be recognized, obtain all recognition results output by the block recognition model, expand each recognition result to obtain all current target regions corresponding to the moment, and determine all target regions to be recognized according to all recognized regions corresponding to all preset moments before the moment and all current target regions; A determination unit, which is used to determine the target batch recognition model according to the number of all target regions to be recognized, use the target batch recognition model to recognize all target regions to be recognized, obtain all target recognition objects output by the target batch recognition model, and display all target recognition objects in the image to be recognized; Among them, each moment corresponds to a block in the preset area grid, and the sum of the number of all preset moments and the number of moments of the moment is the same as the number of blocks in the preset area grid; the all identified areas are determined by using the block recognition model to recognize the image to be recognized corresponding to each preset moment, obtaining all historical recognition results corresponding to each preset moment, and expanding each of the all historical recognition results corresponding to each preset moment to obtain all historical target areas corresponding to all preset moments.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the target recognition method based on the panoramic stitching image according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Digital recognition method and device, electronic equipment and storage medium

    CN110443159A

  • Target detection method and device based on image partition

    CN115035300A

  • Target identification method and device, domain controller and operation machine

    CN115330988A

  • Virtual reality teaching gesture recognition method based on dual-camera multi-branch network

    CN116466816A

  • Multi-target recognition method and apparatus based on video super-resolution

    WO2024109902A1