Target identification method of fisheye image, electronic device, and storage medium

By detecting and segmenting the fisheye camera image, combining multiple feature extraction and semantic segmentation techniques, the problem of traffic target recognition in the fisheye camera image is solved, and the rapid and accurate recognition effect is achieved, meeting the needs of real-time perception.

WO2025130866A1PCT designated stage expired Publication Date: 2025-06-26NIO TECH ANHUI CO LTD

Patent Information

Application Number
PCT/CN2024/139942
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-12-17
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The identification of traffic targets in the fisheye camera image has interference factors such as scale change, target occlusion, and light changes. The object deformation caused by the fisheye camera increases the difficulty of identification, and the prior art is difficult to meet the real-time needs.

Method used

A target recognition method for fisheye images is proposed. By acquiring the fisheye images, the target detection and segmentation are performed, and the target detection results and segmentation results are verified to obtain the final target recognition results. This method can quickly and accurately identify traffic targets in fisheye camera images using multiple feature extraction, fusion feature images and semantic segmentation.

Benefits of technology

It realizes rapid and accurate identification of traffic targets in the fisheye camera image, meets the need to perceive the driving equipment environment information in real time, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024139942_26062025_PF_FP_ABST
    Figure CN2024139942_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image identification, and specifically provides a target identification method of a fisheye image, an electronic device, and a storage medium, aiming to solve the technical problem of how to quickly and accurately identify a traffic target in a fisheye camera image. To achieve this objective, the target identification method of the fisheye image in the present application comprises: acquiring a fisheye image to be identified; performing target detection on the fisheye image to be identified to obtain a target detection result; segmenting the fisheye image to be identified to obtain a segmentation result; and obtaining a target identification result on the basis of the target detection result and the segmentation result. By means of the embodiments, the target identification result can be jointly obtained on the basis of the target detection result and the segmentation result of the fisheye image, the traffic target in the fisheye camera image can be quickly and accurately identified, and the objective of sensing environment information of a driving device in real time is achieved, improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Fisheye image target recognition method, electronic device and storage medium

[0001] This application claims priority to Chinese patent application No. 202311773751.2, filed on December 21, 2023, entitled “Target recognition method, electronic device and storage medium for fisheye images”. The entire contents of the above Chinese patent application are incorporated into this application by reference. Technical Field

[0002] The present application relates to the field of image recognition technology, and in particular to a method for identifying a target using a fisheye image, an electronic device, and a storage medium. Background Art

[0003] Environmental perception based on fisheye cameras is part of intelligent driving and has important applications in scenarios such as automatic parking and traffic congestion assistance systems. Among them, the recognition of traffic targets such as pedestrians, vehicles, and bicycles is crucial.

[0004] However, traffic object recognition in fisheye camera images is a challenging task. First, traffic objects are subject to interference factors such as scale change, object occlusion, and lighting changes, which require high robustness of the algorithm. Second, the object deformation caused by the fisheye camera further increases the difficulty of fisheye image recognition. Furthermore, such scenarios require high real-time performance of the algorithm, and many methods cannot meet the application requirements.

[0005] Accordingly, this field requires a new technical solution to solve the above problems. Summary of the Invention

[0006] In order to overcome the above-mentioned defects, the present application is proposed to provide a fisheye image target recognition method, electronic device and storage medium that solve or at least partially solve the technical problem of how to quickly and accurately identify traffic targets in fisheye camera images.

[0007] In a first aspect, a method for object recognition in a fisheye image is provided, the method comprising:

[0008] S1, obtaining a fisheye image to be identified;

[0009] S2. performing target detection on the fisheye image to be identified to obtain a target detection result;

[0010] S3, segmenting the fisheye image to be identified to obtain a segmentation result;

[0011] S4. Obtain a target recognition result based on the target detection result and the segmentation result.

[0012] In one technical solution of the above-mentioned method for object recognition in fisheye images, performing object detection on the fisheye image to be recognized to obtain the object detection result includes:

[0013] Performing multiple feature extractions on the fisheye image to be identified to obtain multiple feature images;

[0014] Performing target detection on the multiple feature images to obtain target detection frames and attribute information of the target detection frames;

[0015] The attribute information of the target detection frame includes at least one of the category, coordinates, size and confidence of the target detection frame.

[0016] In one technical solution of the above-mentioned method for object recognition in fisheye images, segmenting the fisheye image to be recognized to obtain a segmentation result includes:

[0017] Processing the multiple feature images to obtain a fused feature image;

[0018] Performing semantic segmentation on the fused feature image to obtain a segmentation mask image;

[0019] The segmentation mask image marks the mask category of each pixel.

[0020] In one technical solution of the above-mentioned method for object recognition in fisheye images, obtaining the object recognition result based on the object detection result and the segmentation result includes:

[0021] The target detection result is verified based on the segmentation result to obtain the target recognition result.

[0022] In one technical solution of the above-mentioned method for object recognition in fisheye images, verifying the object detection result based on the segmentation result to obtain the object recognition result includes:

[0023] Determine whether the mask category of the segmentation mask image is the same as the category of the target detection frame, and whether the intersection-over-union ratio of the mask area of ​​the category to the area of ​​the target detection frame is greater than a preset threshold;

[0024] If yes, retain the target detection frame of the category, and obtain the target recognition result based on the retained target detection frame;

[0025] The target recognition result includes the target detection frame, the category of the detection frame and the confidence level.

[0026] In one technical solution of the above-mentioned method for object recognition in fisheye images, performing multiple feature extractions on the fisheye image to be recognized to obtain multiple feature images includes:

[0027] Performing multiple first feature extraction operations on the fisheye image to be identified to obtain multiple first feature images with different scales and different numbers of channels;

[0028] A second feature extraction operation is performed on the plurality of first feature images with different scales and different numbers of channels to obtain a plurality of second feature images with different scales and the same number of channels.

[0029] In one technical solution of the above-mentioned method for object recognition in fisheye images, processing the multiple feature images to obtain a fused feature image includes:

[0030] performing an upsampling operation on all or part of the second feature images having different scales and the same number of channels based on an interpolation method to obtain a plurality of third feature images having the same scale and the same number of channels;

[0031] The plurality of third feature images having the same scale and the same number of channels are fused to obtain the fused feature image.

[0032] In one technical solution of the above-mentioned method for object recognition in fisheye images, steps S2-S4 are performed based on a trained object recognition model.

[0033] In one technical solution of the above-mentioned object recognition method for fisheye images, the object recognition model includes a plurality of object detection modules, a segmentation module and a verification module;

[0034] The target detection module is configured to perform the target detection on the fisheye image to be identified to obtain the target detection result;

[0035] The segmentation module is configured to perform the segmentation on the fisheye image to be identified to obtain the segmentation result;

[0036] The verification module is configured to verify the target detection result based on the segmentation result to obtain the target recognition result.

[0037] In one technical solution of the above-mentioned object recognition method for fisheye images, the object recognition model further includes a feature extraction module, and the feature extraction module includes a lightweight backbone network and a bidirectional feature pyramid network;

[0038] The lightweight backbone network is configured to perform the multiple first feature extraction operations on the fisheye image to be identified to obtain the multiple first feature images with different scales and different numbers of channels;

[0039] The bidirectional feature pyramid network is configured to perform the second feature extraction operation on the multiple first feature images with different scales and different numbers of channels to obtain the multiple second feature images with different scales and the same number of channels.

[0040] In one technical solution of the above-mentioned object recognition method for fisheye images, the object recognition model is trained based on training samples, and the training samples are fisheye images marked with object detection labels and semantic segmentation labels.

[0041] In a second aspect, an electronic device is provided, comprising a processor and a storage device, wherein the storage device is suitable for storing a plurality of program codes, and the program codes are suitable for being loaded and run by the processor to execute the method for target recognition of fisheye images described in any one of the technical solutions of the above-mentioned method for target recognition of fisheye images.

[0042] In a third aspect, a computer-readable storage medium is provided, which stores a plurality of program codes, wherein the program codes are suitable for being loaded and run by a processor to execute the method for target recognition of fisheye images described in any one of the technical solutions of the method for target recognition of fisheye images.

[0043] The above one or more technical solutions of this application have at least one or more of the following beneficial effects:

[0044] In implementing the technical solution of this application, a fisheye image to be identified is first acquired. Target detection is then performed on the fisheye image to be identified to obtain a target detection result. The fisheye image to be identified is then segmented to obtain a segmentation result. Finally, a target recognition result is obtained based on the target detection and segmentation results. Through the above-described implementation, a target recognition result can be obtained based on both the target detection and segmentation results of the fisheye image, allowing for rapid and accurate identification of traffic targets in the fisheye camera image, achieving real-time perception of driving equipment environmental information and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The disclosure of this application will become more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application. Among them:

[0046] FIG1 is a schematic flow chart of the main steps of a method for object recognition using a fisheye image according to an embodiment of the present application;

[0047] FIG2 is a schematic flow chart of the main steps of performing target detection on a fisheye image to be identified and obtaining a target detection result according to an embodiment of the present application;

[0048] FIG3 is a flow chart showing the main steps of performing multiple feature extractions on a fisheye image to be identified to obtain multiple feature images according to an embodiment of the present application;

[0049] FIG4 is a flow chart showing the main steps of segmenting a fisheye image to be identified and obtaining a segmentation result according to an embodiment of the present application;

[0050] FIG5 is a flow chart showing the main steps of processing multiple feature images to obtain a fused feature image according to an embodiment of the present application;

[0051] FIG6 is a flowchart illustrating the main steps of verifying the target detection result based on the segmentation result to obtain the target recognition result according to an embodiment of the present application;

[0052] FIG7 is a schematic diagram of a target recognition result according to an embodiment of the present application;

[0053] FIG8 is a schematic diagram of the structure of a target recognition model according to an embodiment of the present application;

[0054] FIG9 is a schematic flow chart of the main steps of training a target recognition model according to an embodiment of the present application;

[0055] FIG10 is a schematic diagram of the main flow of a method for object recognition using a fisheye image according to an embodiment of the present application;

[0056] FIG11 is a schematic diagram of the main structure of an electronic device according to an embodiment of the present application.

[0057] List of reference numerals:

[0058] 1101: processor; 1102: storage device. DETAILED DESCRIPTION

[0059] Some embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the scope of protection of the present application.

[0060] In the description of this application, "module" and "processor" may include hardware, software, or a combination of both. A module may include hardware circuitry, various suitable sensors, communication ports, and memory. It may also include software components, such as program code, or a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. A processor has data and / or signal processing capabilities. A processor may be implemented in software, hardware, or a combination of both. Non-transitory computer-readable storage media include any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, etc. The term "A and / or B" refers to all possible combinations of A and B, such as only A, only B, or both A and B. The terms "at least one of A or B" or "at least one of A and B" have similar meanings to "A and / or B" and may include only A, only B, or both A and B. The singular forms "a" and "the" may also include the plural forms.

[0061] Here we first explain some terms involved in this application.

[0062] Fisheye image: An image captured by a fisheye camera with an extremely short focal length and an extremely wide-angle lens. The fisheye image has severe distortion around it.

[0063] PeleeNet: A lightweight backbone network that uses depthwise separable convolution and lightweight feature pyramid networks to reduce model parameters and computational complexity while maintaining high accuracy.

[0064] PAN: A bidirectional feature pyramid network that adds a bottom-up branch based on the feature pyramid network FPN to achieve better feature fusion and target positioning.

[0065] CNN: Convolutional Neural Networks (CNNs) are a type of feedforward neural network with a deep structure that incorporates convolutional computations. They are a representative algorithm for deep learning. Convolutional neural networks possess representation learning capabilities and can perform shift-invariant classification of input information based on their hierarchical structure. Therefore, they are also called shift-invariant artificial neural networks (SIANNs).

[0066] Downsampling: also known as feature extraction, such as extracting a small part from the majority set. In the image field, downsampling actually means reducing the image. The main purpose is to make the image fit the size of the display area and generate a thumbnail of the corresponding image.

[0067] Upsampling: After the image is feature extracted, the image is enlarged and restored to its original size.

[0068] As described in the background technology, environmental perception based on fisheye cameras is an integral part of intelligent driving and has important applications in scenarios such as automatic parking and traffic congestion assistance systems. Among them, the recognition of traffic targets such as pedestrians, vehicles, and bicycles is crucial.

[0069] However, traffic object recognition in fisheye camera images is a challenging task. Existing technical solutions have the following shortcomings:

[0070] 1. Target detection based on neural networks to identify and locate traffic targets. This method works well on images from ordinary cameras, but fisheye camera images are distorted, causing severe deformation of the target, making target detection extremely difficult.

[0071] 2. First, correct the fisheye image for distortion before performing target detection. This method alleviates the impact of image distortion on detection accuracy, but the process is complex and time-consuming, making it difficult to meet the requirements of real-time perception. In addition, the dedistortion process easily leads to the loss of boundary information.

[0072] 3. Pixel-by-pixel recognition through image segmentation. This method offers high accuracy and adapts to varying degrees of image distortion. However, semantic segmentation lacks instance-level information and cannot isolate individual objects for accurate recognition. Instance segmentation also comes with high annotation costs.

[0073] In order to solve the above problems, the present application provides a target recognition method for fisheye images, an electronic device and a storage medium.

[0074] Referring to FIG1 , FIG1 is a flow chart showing the main steps of a method for object recognition using a fisheye image according to an embodiment of the present application. As shown in FIG1 , the method for object recognition using a fisheye image in the embodiment of the present application mainly includes the following steps S101 to S104 .

[0075] Step S101: obtaining a fisheye image to be identified;

[0076] Step S102: performing target detection on the fisheye image to be identified to obtain a target detection result;

[0077] Step S103: Segment the fisheye image to be identified to obtain a segmentation result;

[0078] Step S104: Obtaining a target recognition result based on the target detection result and the segmentation result.

[0079] Based on the method described in steps S101 to S104 above, a target recognition result can be obtained based on the target detection result and segmentation result of the fisheye image, and traffic targets in the fisheye camera image can be quickly and accurately identified, thereby achieving the purpose of real-time perception of driving equipment environment information and improving the user experience.

[0080] The above steps S101 to S104 are further explained below.

[0081] In some implementations of the above step S101 , the fisheye camera image may be preprocessed, including standardization, normalization, and image compression, to obtain a fisheye image to be identified.

[0082] The pre-processing may include one or more of the following steps:

[0083] Image standardization: fisheye images are standardized by subtracting the mean and dividing by the standard deviation according to certain attribute values ​​to make the data conform to the normal distribution for subsequent processing.

[0084] Image normalization: Normalize the fisheye image and scale the pixel values ​​of the image to a specified range, usually [0, 1], [0, 255], etc., to facilitate subsequent processing.

[0085] Image compression: The fisheye image is compressed to reduce the image file size for easy storage and transmission. For example, each image is compressed to a size of 480×480×3, where 480×480 is the image scale and 3 is the number of channels.

[0086] Image cropping: Crop the fisheye image to a standard pixel size for subsequent processing.

[0087] Image enhancement: fisheye images are enhanced, including contrast enhancement, sharpening, denoising and other operations to improve image quality and clarity.

[0088] By preprocessing fisheye camera images, the quality and clarity of the images can be improved, which is convenient for subsequent processing and analysis. At the same time, the size of the image files can be reduced, which is convenient for storage and transmission.

[0089] It should be noted that the above examples of preprocessing fisheye images are only illustrative. In practical applications, those skilled in the art may process fisheye images according to specific scenarios, and this is not limited here.

[0090] The above is a further description of step S101 , and the following further describes step S102 .

[0091] In some implementations of the above step S102, refer to FIG2, which is a flow chart of the main steps of performing target detection on the fisheye image to be identified and obtaining the target detection result according to an embodiment of the present application. As shown in FIG2, it mainly includes the following steps S201 to S202:

[0092] Step S201: performing multiple feature extractions on the fisheye image to be identified to obtain multiple feature images;

[0093] Specifically, referring to FIG3 , FIG3 is a schematic flow chart of the main steps of performing multiple feature extractions on a fisheye image to be identified to obtain multiple feature images according to an embodiment of the present application. As shown in FIG3 , step S201 mainly includes the following steps S2011 to S2012:

[0094] Step S2011: performing multiple first feature extraction operations on the fisheye image to be identified to obtain multiple first feature images with different scales and different numbers of channels;

[0095] Among them, the first feature extraction operation can include extracting key points, descriptors, gradients, rotation invariance, etc. in the fisheye image. Performing multiple first feature extractions can extract richer and multi-dimensional feature information, and obtain multiple first feature images with different scales and different numbers of channels.

[0096] Specifically, the first feature extraction operation may be performed four times on the fisheye image to obtain four first feature images, for example, 120×120×128, 60×60×256, 30×30×512, and 15×15×704, respectively.

[0097] The above four first feature images cover feature images of different scales and different numbers of channels. If more feature extraction operations are performed, the accuracy of the recognition results may not be greatly improved, and the calculation overhead and time are greater; if fewer feature extraction operations are performed, the recognition results may be inaccurate.

[0098] Step S2012: performing a second feature extraction operation on a plurality of first feature images having different scales and different numbers of channels to obtain a plurality of second feature images having different scales and the same number of channels.

[0099] After obtaining multiple first feature images with different scales and different numbers of channels, the second feature extraction operation can be performed on the multiple first feature images respectively to extract higher-level features. These features contain more semantic information and can aggregate information in high-dimensional space, thereby improving the accuracy of target recognition and semantic segmentation.

[0100] Specifically, further feature extraction can be performed on the four first feature images to obtain four second feature images with different scales and the same number of channels, for example, 120×120×64, 60×60×64, 30×30×64, and 15×15×64, respectively.

[0101] Based on the above implementation, features of multiple scales can be extracted and fused to solve the problem of difficulty in multi-scale target detection.

[0102] The above is a further explanation of step S201.

[0103] Step S202: performing target detection on multiple feature images to obtain target detection frames and attribute information of the target detection frames.

[0104] Specifically, target detection can be performed on multiple second feature images based on algorithms such as YOLO, SSD, and Faster R-CNN. For each second feature image, several target detection frames and their attribute information may be detected. The final target detection result includes the target detection frames of all second feature images and the attribute information of each target detection frame.

[0105] Furthermore, the attribute information of the target detection frame includes at least one of the category, coordinates, size and confidence of the target detection frame.

[0106] Among them, the category of the target detection box corresponds to the category of the detected traffic target, such as background, pedestrian, vehicle, etc.

[0107] The confidence score refers to the model's evaluation score of the accuracy of the detection box. It is calculated by the prediction probability module within the model and indicates the probability that the detection box contains a certain type of target object. It is usually a number between 0 and 1. When the confidence score is higher than a certain threshold (such as 0.7), the model will consider the detection box to be valid and use it as part of the target detection result. If the confidence score is lower than the threshold, the detection box will be considered invalid and will be removed from the detection result.

[0108] The coordinates include the pixel coordinates of the upper left corner and the lower right corner of the target detection box, which are usually expressed in the form of a rectangular box, such as (x, y) and (w, h), which represent the pixel coordinates of the upper left corner and the lower right corner of the rectangular box, respectively.

[0109] Size represents the width and height of the object detection box, usually expressed in pixels.

[0110] Furthermore, in this application, the category, confidence and coordinates of the target detection frame can be obtained.

[0111] It should be noted that the above examples of attribute information of the target detection frame are only for illustrative purposes. In actual applications, those skilled in the art can make settings according to specific scenarios, and no limitation is made here.

[0112] The above is a further description of step S102 , and the following further describes step S103 .

[0113] In some implementations of the above step S103, refer to FIG4, which is a flow chart of the main steps of segmenting the fisheye image to be identified and obtaining the segmentation result according to an embodiment of the present application. As shown in FIG4, it mainly includes the following steps S401 to S402:

[0114] Step S401: Process multiple feature images to obtain a fused feature image;

[0115] Specifically, referring to FIG5 , FIG5 is a flow chart showing the main steps of processing multiple feature images to obtain a fused feature image according to an embodiment of the present application. As shown in FIG5 , step S401 mainly includes the following steps S4011 to S4012:

[0116] Step S4011: performing an upsampling operation on all or part of the second feature images with different scales and the same number of channels based on an interpolation method to obtain a plurality of third feature images with the same scale and the same number of channels;

[0117] Specifically, the second feature image may be up-sampled based on interpolation methods such as nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation, among which the bilinear interpolation method has the best performance.

[0118] Furthermore, in some embodiments, all the second feature images can be upsampled to obtain third feature images with the same scale and the same number of channels, or only a few second feature images with smaller scales can be upsampled to unify the scales, for example, the scales of the above four second feature images can be unified to 120×120×64.

[0119] Among them, the upsampling operation is to restore the feature image after feature extraction to the size of the original fisheye image through interpolation.

[0120] Step S4012: performing fusion processing on multiple third feature images with the same scale and the same number of channels to obtain a fused feature image.

[0121] Specifically, the above-mentioned multiple third feature images with the same scale and the same number of channels can be fused into one feature image by point-by-point addition. This can make the size of the fused feature image closest to the size of the original fisheye image, which is convenient for subsequent segmentation tasks and obtains a segmentation mask image of the original image size.

[0122] Step S402: performing semantic segmentation on the fused feature image to obtain a segmentation mask image.

[0123] Semantic segmentation involves assigning a category label to each pixel in an image. For example, in a traffic image, the background, pedestrians, vehicles, and other traffic objects are labeled with different mask categories, such as 0, 1, and 2. Semantic segmentation helps us understand the semantic information of different regions in an image, thereby better understanding the image content.

[0124] The segmentation mask image is a grayscale image of the same size as the fisheye image. The segmentation mask image marks the mask category of each pixel, such as 0, 1, 2, etc.

[0125] The above is a further description of step S103 , and the following is a further description of step S104 .

[0126] In some implementations of the above step S104, the target detection result may be verified based on the segmentation result to obtain a target recognition result.

[0127] Specifically, the target detection box in the target recognition result can be verified based on the segmentation mask in the segmentation mask image.

[0128] Refer to FIG6 , which is a flowchart of the main steps of verifying the target detection result based on the segmentation result to obtain the target recognition result according to an embodiment of the present application. As shown in FIG6 , it mainly includes the following steps S601 to S604:

[0129] Step S601: Determine whether the mask category of the segmentation mask image is the same as the category of the target detection frame, and whether the intersection-over-union ratio of the mask area of ​​the category to the target detection frame area is greater than a preset threshold;

[0130] First, it is necessary to determine whether the mask category of each pixel in the segmentation mask image is the same as the category of the target detection box where the same pixel is located in the target recognition result. For example, if the mask category of pixel a in the segmentation mask image is pedestrian, and the category of the target detection box where pixel a is located in the target recognition result is also pedestrian, then the categories are determined to be the same; otherwise, the categories are determined to be different.

[0131] Furthermore, it is necessary to determine whether the intersection-over-union ratio of the mask area of ​​each category of target in the segmentation mask image and the area of ​​the detection box of the category of target in the target recognition result is greater than a preset threshold.

[0132] Specifically, we need to first calculate the mask area and target detection frame area of ​​each category of targets, where the mask area can be calculated based on the number of pixels of the category mask in the segmentation mask image, and the target detection frame area can be calculated based on the length and width of the category target detection frame. Then we need to calculate the intersection over union (IoU) of the mask area and the target detection frame area of ​​the category target, that is, the ratio of the intersection area of ​​the mask area and the target detection frame area to the union area. Finally, we determine whether the intersection over union (IoU) of the mask area and the target detection frame area of ​​each category of targets is greater than a preset threshold (such as 0.6). Generally speaking, the higher the IoU, the higher the match between the target detection frame and the segmentation mask, and the more accurate the target recognition result.

[0133] Furthermore, if the mask category of the segmentation mask image is the same as the category of the multiple target detection frames, and the intersection-over-union ratio of the mask area to the target detection frame area is greater than a preset threshold, step S602 is executed; otherwise, step S603 is executed.

[0134] Step S602: retain the target detection frame of the category;

[0135] Step S603: Eliminate the target detection frame of this category;

[0136] Furthermore, after each target detection frame in the target recognition result is verified based on the segmentation mask in the segmentation mask image, step S604 may be performed.

[0137] Step S604: Obtaining a target recognition result based on the retained target detection frame.

[0138] In some embodiments, referring to FIG7 , FIG7 is a schematic diagram of a target recognition result according to an embodiment of the present application. As shown in FIG7 , the target recognition result may include a target detection frame, a category of the detection frame, and a confidence level.

[0139] The above is a further explanation of step S104.

[0140] It should be pointed out that although the various steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effect of the present application, different steps do not have to be executed in such an order, and they can be executed simultaneously (in parallel) or in other orders.

[0141] For example, the above-mentioned steps S102 and S103 can be executed simultaneously, that is, target detection and segmentation are performed on the fisheye image to be identified at the same time; step S102 can also be executed first, and then step S103 is executed, that is, target detection is first performed on the fisheye image to be identified, and then the fisheye image to be identified is segmented; step S103 can also be executed first, and then step S102 is executed, that is, the fisheye image to be identified is segmented first, and then the target detection is performed on the fisheye image to be identified. These changes are all within the scope of protection of this application.

[0142] Furthermore, in some implementations, steps S102 to S104 may be performed based on a trained target recognition model.

[0143] The trained target recognition model has learned how to identify targets in images, has high accuracy and speed, can quickly and accurately identify traffic targets in fisheye images, and can be applied in different scenarios with good generalization ability and robustness.

[0144] Referring to FIG8 , FIG8 is a schematic diagram of the structure of a target recognition model according to an embodiment of the present application. As shown in FIG8 , the target recognition model includes multiple target detection modules, segmentation modules, verification modules, and feature extraction modules.

[0145] Multiple target detection modules utilize multiple CNN target detectors to detect targets in the fisheye image to be identified, generating target detection results. CNN target detectors possess powerful feature extraction capabilities and can perform operations such as convolution, pooling, and normalization on fisheye images to extract feature information. Classifiers are then used to classify objects and regress their positions, achieving target detection.

[0146] The segmentation module uses a CNN segmentation network to segment the fisheye image to obtain the segmentation result. CNN segmentation networks have powerful feature extraction capabilities and can achieve image segmentation. Some classic CNN segmentation networks include U-Net, Mask R-CNN, and DeepLab.

[0147] The verification module uses some CNN verification networks with detection capabilities, such as YOLOv3, Faster R-CNN, SSD, etc., to verify the target detection results based on the segmentation results and obtain the target recognition results.

[0148] The feature extraction module includes a lightweight backbone network, PeleeNet, and a bidirectional feature pyramid network, PAN. The lightweight backbone network, PeleeNet, uses depthwise separable convolution and a lightweight feature pyramid network to reduce model parameters and computational complexity, resulting in a smaller number of parameters and computation while maintaining high accuracy. The bidirectional feature pyramid network, PAN, adds a bottom-up branch to the feature pyramid network (FPN), enabling better feature fusion and target localization.

[0149] Both the lightweight backbone network and the bidirectional feature pyramid network include multiple branches, each capable of performing different feature extraction operations on the feature image. The lightweight backbone network can perform multiple first feature extraction operations on the fisheye image to be identified, generating multiple first feature images with varying scales and numbers of channels. The bidirectional feature pyramid network can perform second feature extraction operations on multiple first feature images with varying scales and numbers of channels, generating multiple second feature images with varying scales and the same number of channels.

[0150] The lightweight backbone network extracts features quickly and can meet the requirements of real-time recognition. Based on the features extracted by the lightweight backbone network, it continues to extract features through the bidirectional feature pyramid network, which can aggregate information in high-dimensional space and improve the expressive ability of the model.

[0151] It should be pointed out that the above examples of multiple target detection modules, segmentation modules, verification modules and feature extraction modules in the target recognition model are only schematic illustrations. In actual applications, those skilled in the art can make settings according to specific scenarios. For example, the target detection module and segmentation module can perform target detection and semantic segmentation based on other detectors and segmentation networks, and the feature extraction module can perform feature extraction based on other backbone networks and pyramid networks. There is no limitation here.

[0152] Furthermore, the above-mentioned target recognition model can be trained based on training samples, where the training samples are fisheye images marked with target detection labels and semantic segmentation labels.

[0153] In some embodiments, referring to FIG9 , FIG9 is a flow chart showing the main steps of training the target recognition model according to an embodiment of the present application. As shown in FIG9 , the training process mainly includes the following steps S901 to S903:

[0154] Step S901: Obtain training samples;

[0155] The training samples are fisheye images annotated with target detection labels and semantic segmentation labels. They can be obtained from public datasets or collected and annotated from actual scenes.

[0156] For the collected fisheye images, standard object detection labels and semantic segmentation labels can be used for annotation. Specifically, for the object detection task, bounding box annotation or keypoint annotation can be used to annotate the object position and shape information in the fisheye image; for the semantic segmentation task, pixel-level annotation can be used to assign each pixel in the image to a different category.

[0157] Step S902: training multiple target detection modules, segmentation modules, verification modules, and feature extraction modules in the target recognition model based on the training samples;

[0158] Specifically, it includes: training the feature extraction module to extract effective fisheye image features, training the target detection module to detect and locate different types of targets, training the semantic segmentation module to classify and label each pixel in the fisheye image, and training the verification module to verify the target recognition results based on the segmentation results.

[0159] Step S903: When the target recognition model converges to a preset error, the target recognition model training is completed.

[0160] The preset error is a stable low value, and those skilled in the art can set the preset error value according to actual application scenarios, which is not limited here.

[0161] The target recognition model trained by the above implementation method has instance-level target recognition capability, can identify multiple target objects of different categories in fisheye images, and has high generalization ability and robustness.

[0162] Furthermore, the target recognition model can be input into the trained target recognition model to obtain the image segmentation result.

[0163] Specifically, referring to FIG10 , FIG10 is a schematic diagram of the main flow of a method for object recognition in a fisheye image according to an embodiment of the present application.

[0164] As shown in Figure 10, the lightweight backbone network PeleeNet performs a first feature extraction operation on the input fisheye image to obtain multiple first feature images with different scales and different numbers of channels;

[0165] The bidirectional feature pyramid network PAN performs a second feature extraction operation on multiple first feature images with different scales and different numbers of channels to obtain multiple second feature images with different scales and the same number of channels;

[0166] Multiple CNN target detectors perform target detection on the multiple second feature images to obtain target detection results; at the same time, a fused feature image is obtained based on the multiple second feature images, and the CNN segmentation network segments the fused feature image to obtain a segmentation result;

[0167] Finally, the CNN verification network verifies the target detection results through the segmentation results to obtain the target recognition results.

[0168] Through the above implementation, it is possible to perform target detection and segmentation on fisheye images based on the trained dual-task target recognition model, and obtain the target recognition result of the fisheye image based on the target detection result and the segmentation result. Among them, the target detection task and the segmentation task can gain each other during model training, so that the model has better feature expression ability, improved generalization and robustness, and can quickly and accurately identify traffic targets in fisheye images, achieve the purpose of real-time perception of driving equipment environmental information, and enhance the user experience.

[0169] It will be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment of the present application can also be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium, etc. that can carry the computer program code.

[0170] Furthermore, the present application also provides an electronic device. Refer to Figure 11, which is a schematic diagram of the main structure of an electronic device according to an embodiment of the present application. As shown in Figure 11, the electronic device in the embodiment of the present application mainly includes a processor 1101 and a storage device 1102. The storage device 1102 can be configured to store a program for executing the target recognition method of the fisheye image of the above-mentioned method embodiment, and the processor 1101 can be configured to execute the program in the storage device 1102, which includes but is not limited to a program for executing the target recognition method of the fisheye image of the above-mentioned method embodiment. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application.

[0171] In some possible implementations of the present application, the electronic device may include multiple processors 1101 and multiple storage devices 1102. The program for executing the method for target recognition of fisheye images of the above-mentioned method embodiment may be divided into multiple subroutines, and each subroutine may be loaded and run by the processor 1101 to execute different steps of the method for target recognition of fisheye images of the above-mentioned method embodiment. Specifically, each subroutine may be stored in a different storage device 1102, and each processor 1101 may be configured to execute the program in one or more storage devices 1102 to jointly implement the method for target recognition of fisheye images of the above-mentioned method embodiment, that is, each processor 1101 executes different steps of the method for target recognition of fisheye images of the above-mentioned method embodiment to jointly implement the method for target recognition of fisheye images of the above-mentioned method embodiment.

[0172] The multiple processors 1101 may be processors deployed on the same device. For example, the electronic device may be a high-performance device composed of multiple processors, and the multiple processors 1101 may be processors configured on the high-performance device. Furthermore, the multiple processors 1101 may be processors deployed on different devices. For example, the electronic device may be a server cluster, and the multiple processors 1101 may be processors on different servers in the server cluster. Alternatively, the computer device may be a driving device cluster, and the multiple processors 1101 may be processors on different driving devices in the driving device cluster.

[0173] Furthermore, the present application also provides a computer-readable storage medium. In a computer-readable storage medium embodiment according to the present application, the computer-readable storage medium can be configured to store a program for executing the target recognition method of the fisheye image of the above-mentioned method embodiment, and the program can be loaded and run by the processor to implement the above-mentioned target recognition method of the fisheye image. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiment of the present application is a non-transitory computer-readable storage medium.

[0174] It should be noted that the relevant user personal information that may be involved in the various embodiments of this application is strictly in accordance with the requirements of laws and regulations, follows the principles of legality, legitimacy and necessity, and is based on the reasonable purposes of business scenarios to process personal information that users actively provide during the use of products / services or generated due to the use of products / services, as well as personal information obtained with the user's authorization.

[0175] The user personal information processed by this application will vary depending on the specific product / service scenario and must be based on the specific scenario in which the user uses the product / service. This may involve the user's account information, device information, driving information, vehicle information, or other related information. This application will treat the user's personal information and its processing with a high degree of diligence.

[0176] This application attaches great importance to the security of user personal information and has taken reasonable and feasible security protection measures that comply with industry standards to protect user information and prevent personal information from being accessed, disclosed, used, modified, damaged or lost without authorization.

[0177] Thus far, the technical solution of the present application has been described in conjunction with an embodiment shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of the present application is obviously not limited to these specific embodiments. Without departing from the principles of the present application, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present application.

Claims

1. A method for object recognition in fisheye images, characterized in that: The method comprises: S1, obtaining a fisheye image to be identified; S2, performing target detection on the fisheye image to be identified to obtain a target detection result; S3, segmenting the fisheye image to be identified to obtain a segmentation result; S4. Obtain a target recognition result based on the target detection result and the segmentation result.

2. The target recognition method of fisheye image according to claim 1, characterized in that: The performing target detection on the fisheye image to be identified to obtain a target detection result includes: Performing multiple feature extractions on the fisheye image to be identified to obtain multiple feature images; Performing target detection on the multiple feature images to obtain a target detection frame and attribute information of the target detection frame; The attribute information of the target detection frame includes at least one of the category, coordinates, size and confidence of the target detection frame.

3. The target recognition method of fisheye image according to claim 2, characterized in that: The step of segmenting the fisheye image to be identified to obtain a segmentation result includes: Processing the multiple feature images to obtain a fused feature image; Performing semantic segmentation on the fused feature image to obtain a segmentation mask image; The segmentation mask image marks the mask category of each pixel.

4. The method for object recognition of fisheye images according to claim 3, characterized in that: The obtaining of the target recognition result based on the target detection result and the segmentation result comprises: The target detection result is verified based on the segmentation result to obtain the target recognition result.

5. The method for object recognition of fisheye images according to claim 4, characterized in that: The verifying the target detection result based on the segmentation result to obtain the target recognition result includes: Determine whether the mask category of the segmentation mask image is the same as the category of the target detection frame, and whether the intersection-and-union ratio of the mask area of ​​the category to the area of ​​the target detection frame is greater than a preset threshold; If yes, retain the target detection frame of the category, and obtain the target recognition result based on the retained target detection frame; The target recognition result includes the target detection frame, the category of the detection frame and the confidence level.

6. The method for object recognition of fisheye images according to claim 3, characterized in that: The step of performing multiple feature extractions on the fisheye image to be identified to obtain multiple feature images includes: Performing a plurality of first feature extraction operations on the fisheye image to be identified to obtain a plurality of first feature images with different scales and different numbers of channels; A second feature extraction operation is performed on the plurality of first feature images with different scales and different numbers of channels to obtain a plurality of second feature images with different scales and the same number of channels.

7. The method for object recognition of fisheye images according to claim 6, characterized in that: The processing of the multiple feature images to obtain a fused feature image comprises: Based on the interpolation method, up-sampling operation is performed on all or part of the second feature images with different scales and the same number of channels to obtain multiple third feature images with the same scale and the same number of channels; The plurality of third feature images having the same scale and the same number of channels are fused to obtain the fused feature image.

8. The method for object recognition of fisheye images according to any one of claims 1 to 7, characterized in that: The steps S2-S4 are performed based on the trained target recognition model.

9. The method for object recognition of fisheye images according to claim 8, characterized in that: The target recognition model includes multiple target detection modules, segmentation modules and verification modules; The target detection module is configured to perform the target detection on the fisheye image to be identified to obtain the target detection result; The segmentation module is configured to perform the segmentation on the fisheye image to be identified to obtain the segmentation result; The verification module is configured to verify the target detection result based on the segmentation result to obtain the target recognition result.

10. The method for object recognition of fisheye images according to claim 8, characterized in that: The target recognition model also includes a feature extraction module, which includes a lightweight backbone network and a bidirectional feature pyramid network; The lightweight backbone network is configured to perform the multiple first feature extraction operations on the fisheye image to be identified to obtain the multiple first feature images with different scales and different numbers of channels; The bidirectional feature pyramid network is configured to perform the second feature extraction operation on the multiple first feature images with different scales and different numbers of channels to obtain the multiple second feature images with different scales and the same number of channels.

11. The method for object recognition of fisheye images according to claim 8, characterized in that: The target recognition model is trained based on training samples, and the training samples are fisheye images marked with target detection labels and semantic segmentation labels.

12. An electronic device comprising a processor and a storage device, wherein the storage device is suitable for storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by the processor to execute the object recognition method of the fisheye image according to any one of claims 1 to 11.

13. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by a processor to execute the object recognition method of a fisheye image according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Positioning element detection method and device, equipment and medium

    CN111274974A

  • Automatic driving vehicle control method and system and automatic driving vehicle

    CN114419603A

  • Parking space detection method and device, electronic equipment and vehicle

    CN116630834A

  • Target identification method of fisheye image, electronic equipment and storage medium

    CN117765493A

  • Method and system for image analysis using boundary detection

    DE102019129107A1

Cited By

  • Monitoring method and system for preventing external force damage, electronic equipment and storage medium

    CN121191091A

  • Structured document image preprocessing and defect detection method based on image recognition

    CN122473171A