Target perception method, device, system, electronic device and storage medium

By identifying and cropping the logo boxes in the panoramic bird's-eye bird's-eye bird's-eye image to remove distorted areas, the problem of low accuracy in the panoramic bird's-eye image is solved, and more accurate target position perception is achieved, providing a more accurate data basis for driverless driving and automatic parking.

CN118887643BActive Publication Date: 2025-05-13INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410800983.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2025-05-13
Estimated Expiration
2044-06-20

AI Technical Summary

Technical Problem

In the prior art, the accuracy of target perception based on panoramic bird's eye view images is not high, resulting in safety hazards in unmanned driving technology.

Method used

By obtaining a panoramic bird's-eye view of the target moving object, identifying and generating an identification box, cropping the identification box to remove distorted areas, and then obtaining the location information of the target in the real world.

Benefits of technology

Improves target perception accuracy based on panoramic bird's eye image, simplifies the elimination of distortion effects, and provides a more accurate data basis for unmanned driving and automatic parking technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118887643B_ABST
    Figure CN118887643B_ABST
Patent Text Reader

Abstract

The present invention provides a target perception method, device, system, electronic device and storage medium, the method comprising: identifying a first perceived target in a panoramic bird's-eye view image of a target moving body, and generating a first identification frame in the panoramic bird's-eye view image for marking the identified first perceived target; determining the point on the first identification frame in the panoramic bird's-eye view image that is closest to the geometric center of the target moving body as the target point corresponding to the first identification frame, cropping the first identification frame based on the target point corresponding to the first identification frame and a first border size threshold, and determining the cropped first identification frame as a second identification frame; based on the second identification frame, obtaining the position information of the first perceived target in the real world as the target perception result of the target moving body. The target perception method, device, system, electronic device and storage medium provided by the present invention can improve the accuracy of target perception based on panoramic bird's-eye view images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a target perception method, device, system, electronic device and storage medium. Background Art

[0002] With the rapid development of big data technology and artificial intelligence technology in recent years, driverless cars have become an important development trend in the automotive industry and even the entire transportation field.

[0003] The traditional target perception method in the related technology can perform target perception based on the panoramic bird's-eye view image of the vehicle's surrounding environment information, thereby providing a data basis for the vehicle's unmanned driving.

[0004] However, panoramic bird's-eye view images are usually collected by wide-angle or fisheye lenses, which usually have lens distortion, resulting in different degrees of distortion in panoramic bird's-eye view images. In addition, the degree of distortion of panoramic bird's-eye view images is easily affected by factors such as shooting height, terrain undulations, and the distance of the target object. The distortion in panoramic bird's-eye view images will seriously affect the accuracy of target perception and bring safety risks to the unmanned driving of vehicles.

[0005] Therefore, how to more accurately perceive targets based on panoramic bird's-eye view images is a technical problem that needs to be solved urgently in this field. Summary of the invention

[0006] The present invention provides a target perception method, device, system, electronic device and storage medium, which are used to solve the defect of low accuracy of target perception based on panoramic bird's-eye view images in the prior art, and realize more accurate target perception based on panoramic bird's-eye view images.

[0007] The present invention provides a target perception method, comprising the following steps.

[0008] Obtain a panoramic bird's-eye view image of the target moving object.

[0009] A first perceived target in the panoramic bird's-eye view image is identified, and a first identification frame for marking the identified first perceived target is generated in the panoramic bird's-eye view image, wherein the first perceived target includes at least one of an obstacle, a pedestrian, a ground sign, and a traffic sign.

[0010] The point on the first identification frame in the panoramic bird's-eye view image that is closest to the geometric center of the target moving body is determined as the target point corresponding to the first identification frame. Based on the target point corresponding to the first identification frame and the first border size threshold, the first identification frame is cropped, and the cropped first identification frame is determined as the second identification frame.

[0011] Based on the second identification frame, the position information of the first perceived target in the real world is obtained as the target perception result of the target moving body.

[0012] According to a target perception method provided by the present invention, after obtaining a panoramic bird's-eye view image of a target moving body, the method further includes: identifying a second perceived target in the panoramic bird's-eye view image, and generating a third identification frame in the panoramic bird's-eye view image for marking the identified second perceived target, wherein the second perceived target includes a parking space; determining the geometric center of the third identification frame in the panoramic bird's-eye view image as a target point corresponding to the third identification frame, cropping the third identification frame based on the target point corresponding to the third identification frame and a second border size threshold, and determining the cropped third identification frame as a fourth identification frame; based on the fourth identification frame, acquiring the position information of the second perceived target in the real world as the target perception result of the target moving body.

[0013] According to a target perception method provided by the present invention, after obtaining the position information of the first perceived target in the real world, the method further includes: in the case where there is a first perceived target tracking list, based on the Hungarian matching algorithm, matching the second identification box with the existing historical identification box in the first perceived target tracking list; in the case where the second identification box successfully matches the historical identification box, based on the position information of the first perceived target in the real world, the identification information of the first perceived target and the moment when the panoramic bird's-eye view image is obtained, updating the tracking information of the historical identification box in the first perceived target tracking list that successfully matches the second identification box; in the case where the second identification box successfully matches the historical identification box, When the identification frame fails to match successfully, the position information of the first perceived target in the real world, the identification information of the first perceived target and the moment of acquiring the panoramic bird's-eye view image are determined as the tracking information of the second identification frame, and the second identification frame and the tracking information of the second identification frame are added to the first perceived target tracking list. For the historical identification frames in the first perceived target tracking list that fail to match successfully with the second identification frame, when the duration of adding the historical identification frames that fail to match successfully with the second identification frame to the first perceived target tracking list exceeds a first duration threshold, the historical identification frames that fail to match successfully with the second identification frame are removed from the first perceived target tracking list.

[0014] According to a target perception method provided by the present invention, the first identification frame is cropped based on the target point corresponding to the first identification frame and the first frame size threshold, and the cropped first identification frame is determined as the second identification frame, including: determining the frame where the target point corresponding to the first identification frame is located as the target frame; when the lengths of both sides of the target point on the target frame are greater than half of the size threshold corresponding to the target frame in the first frame size threshold, taking the target frame as the boundary of the cropping area and the target point corresponding to the first identification frame as the midpoint of the boundary, determining the cropping area whose size meets the first frame size threshold in the first identification frame; when the lengths of either side of the target point on the target frame are less than half of the size threshold corresponding to the target frame in the first frame size threshold, taking the target frame as the boundary of the cropping area and taking the vertex on the target frame closer to the target point corresponding to the first identification frame as the starting point, determining the cropping area whose size meets the first frame size threshold in the first identification frame; and cropping the first identification frame along the boundary of the cropping area, and determining the cropped first identification frame as the second identification frame.

[0015] According to a target perception method provided by the present invention, the third identification frame is cropped based on the target point corresponding to the third identification frame and the second border size threshold, including: taking the target point corresponding to the third identification frame as the geometric center, determining a cropping area within the third identification frame whose size meets the second border size threshold; cropping the third identification frame along the boundary of the cropping area, and determining the cropped third identification frame as the fourth identification frame.

[0016] According to a target perception method provided by the present invention, after obtaining the position information of the second perceived target in the real world based on the fourth identification frame, the method further includes: in the case where there is a second perceived target tracking list, matching the fourth identification frame with a historical identification frame already in the second perceived target tracking list based on the Hungarian matching algorithm; in the case where the fourth identification frame successfully matches the historical identification frame, updating the tracking information of the historical identification frame in the second perceived target tracking list that successfully matches the fourth identification frame based on the position information of the second perceived target in the real world, the identification information of the second perceived target and the moment when the panoramic bird's-eye view image is obtained, When the historical identification frame fails to match successfully, the position information of the second perceived target in the real world, the identification information of the second perceived target and the moment of acquiring the panoramic bird's-eye view image are determined as the tracking information of the fourth identification frame, and the fourth identification frame and the tracking information of the fourth identification frame are added to the second perceived target tracking list. For the historical identification frame in the second perceived target tracking list that fails to match successfully with the fourth identification frame, when the duration of adding the historical identification frame that fails to match successfully with the fourth identification frame to the second perceived target tracking list exceeds a second duration threshold, the historical identification frame that fails to match successfully with the fourth identification frame is removed from the second perceived target tracking list.

[0017] According to a target perception method provided by the present invention, the first perception target in the panoramic bird's-eye view image is identified, and a first identification frame for marking the identified first perception target is generated in the panoramic bird's-eye view image, including: when the first perception target includes an obstacle, the panoramic bird's-eye view image is input into an obstacle recognition model, and an obstacle mask image of the panoramic bird's-eye view image output by the obstacle recognition model is obtained; when the first perception target includes a pedestrian, the panoramic bird's-eye view image is input into a pedestrian recognition model, and a pedestrian mask image of the panoramic bird's-eye view image output by the obstacle recognition model is obtained; and when the first perception target includes a pedestrian, the panoramic bird's-eye view image is input into a pedestrian recognition model, and a pedestrian mask image of the panoramic bird's-eye view image output by the obstacle recognition model is obtained; When the target includes a ground sign, the panoramic bird's-eye view image is input into a ground sign recognition model to obtain a ground sign mask image of the panoramic bird's-eye view image output by the obstacle recognition model; when the first perceived target includes a traffic sign, the panoramic bird's-eye view image is input into a traffic sign recognition model to obtain a traffic sign mask image of the panoramic bird's-eye view image output by the obstacle recognition model; contour detection is performed on a highlight area in the mask image of the panoramic bird's-eye view image that is not covered by a black layer to determine the contour of the first perceived target; the mask image of the panoramic bird's-eye view image includes an obstacle mask image of the panoramic bird's-eye view image, at least one of the pedestrian mask image of the panoramic bird's-eye view image, the ground sign mask image of the panoramic bird's-eye view image, and the traffic sign mask image of the panoramic bird's-eye view image; based on the outline of the first perceived target, removing the first perceived target whose size is smaller than a first outline perimeter threshold from the mask image; based on the outline of the first perceived target, generating a minimum circumscribed rectangle of the first perceived target in the mask image as the first identification frame, and then removing the black layer on the mask image; wherein the obstacle recognition model is obtained by training based on a sample panoramic bird's-eye view image and an obstacle mask image of the sample panoramic bird's-eye view image; the pedestrian mask image of the panoramic bird's-eye view image, the ground sign mask image of the panoramic bird's-eye view image, and the traffic sign mask image of the ... The human image recognition model is trained based on the sample panoramic bird's-eye view image and the pedestrian mask image of the sample panoramic bird's-eye view image; the ground sign image recognition model is trained based on the sample panoramic bird's-eye view image and the ground sign mask image of the sample panoramic bird's-eye view image; the traffic sign image recognition model is trained based on the sample panoramic bird's-eye view image and the traffic sign mask image of the sample panoramic bird's-eye view image; the obstacle mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the area in the sample panoramic bird's-eye view image except the area where the obstacle is located;The pedestrian mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the area of ​​the sample panoramic bird's-eye view image except the area where the pedestrian is located; the ground sign mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the area of ​​the sample panoramic bird's-eye view image except the area where the ground sign is located; the traffic sign mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the area of ​​the sample panoramic bird's-eye view image except the area where the traffic sign is located. ;

[0018] The present invention also provides a target sensing device, comprising the following modules:

[0019] The bird's-eye view image acquisition module is used to acquire a panoramic bird's-eye view image of the target moving object.

[0020] A perception target recognition module is used to identify a first perception target in the panoramic bird's-eye view image and generate a first identification frame in the panoramic bird's-eye view image for marking the first perception target identified, wherein the first perception target includes at least one of an obstacle, a pedestrian, a ground sign and a traffic sign.

[0021] A perception target correction module is used to determine the point on the first identification frame in the panoramic bird's-eye view image that is closest to the geometric center of the target moving body as the target point corresponding to the first identification frame, and to crop the first identification frame based on the target point corresponding to the first identification frame and a first border size threshold, and determine the cropped first identification frame as the second identification frame.

[0022] A perception result output module is used to obtain the position information of the first perceived target in the real world based on the second identification frame as the target perception result of the target moving body.

[0023] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the target perception methods described above is implemented.

[0024] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the target perception methods described above.

[0025] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the target perception methods described above.

[0026] The target perception method, device, system, electronic device and storage medium provided by the present invention identify a first perceived target in a panoramic bird's-eye view image of a target moving body, and generate a first identification frame for marking the identified first perceived target in the panoramic bird's-eye view image, then determine the point on the first identification frame in the panoramic bird's-eye view image that is closest to the geometric center of the target moving body as the target point corresponding to the first identification frame, crop the first identification frame based on the target point corresponding to the first identification frame and a first border size threshold, determine the cropped first identification frame as a second identification frame, and then based on the second identification frame, The position information of the first perceived target in the real world is obtained as the target perception result of the target moving body. The edges of the image based on the panoramic bird's-eye view image will show a stretching and distorted effect, so that objects far from the center will show a deformed appearance defect. By cropping the first identification frame, the distorted area in the first identification frame is cropped out, thereby improving the accuracy of target perception based on the panoramic bird's-eye view image, and can more simply and efficiently eliminate the influence of the distortion of the panoramic bird's-eye view image on the target perception accuracy, and can provide a more accurate data basis for unmanned driving technology and automatic parking technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0028] Figure 1 This is one of the flow charts of the target perception method provided by the present invention.

[0029] Figure 2 This is the second flow chart of the target perception method provided by the present invention.

[0030] Figure 3 It is a schematic diagram of the structure of the target sensing device provided by the present invention.

[0031] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0033] In the description of the invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0034] In the description of the present application, the terms "first", "second", etc. are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are usually a class, and the number of objects is not limited. For example, the first object can be one or more. In addition, in the description of the present application, "and / or" represents at least one of the connected objects, and the character " / " generally represents that the front and back associated objects are in an "or" relationship.

[0035] It should be noted that with the rapid development of the automobile industry, traffic congestion and safe driving issues have attracted more and more attention from residents and researchers. With the rapid development of big data technology and artificial intelligence technology in recent years, driverless driving has become an important development trend in the automobile industry and even the entire transportation field. As an important part of driverless driving, the development of autonomous parking technology can greatly improve parking safety, improve parking efficiency and enhance driving comfort.

[0036] Among the related technologies, laser radar and ultrasonic radar are usually used for target perception in unmanned driving technology and autonomous parking technology. Laser radar can obtain environmental information in the form of point clouds based on optical principles and the characteristics of laser emission. Ultrasonic radar can perceive objects in the environment based on acoustic principles and the characteristics of sound wave reflection.

[0037] However, due to the high value of LiDAR itself, the industrial production of LiDAR-based autonomous driving technology and autonomous parking technology is constrained by high production costs. Ultrasonic radar, on the other hand, has difficulty in obtaining the shape characteristics of the perceived target due to the single information it obtains, and thus it is difficult to accurately identify the specific type of the perceived target, which limits the role of ultrasonic radar in autonomous driving technology and autonomous parking technology.

[0038] In order to overcome the application defects of lidar and ultrasonic radar in unmanned driving technology and autonomous parking technology, related technologies can also perform target perception based on panoramic bird's-eye view images of the vehicle's surrounding environment, thereby providing a data basis for the vehicle's unmanned driving.

[0039] However, panoramic bird's-eye view images are usually collected by wide-angle or fisheye lenses, which usually have lens distortion, resulting in different degrees of distortion in panoramic bird's-eye view images. In addition, the degree of distortion of panoramic bird's-eye view images is easily affected by factors such as shooting height, terrain undulations, and the distance of the target object. The distortion in panoramic bird's-eye view images will seriously affect the accuracy of target perception and bring safety risks to the unmanned driving of vehicles.

[0040] In this regard, the present invention provides a target perception method, device and system. The target perception method provided by the present invention uses a fisheye camera as a sensor to obtain fisheye images for computer vision perception, thereby detecting and tracking objects such as ground signs, traffic signs, obstacles, parking spaces, etc. in a bird's-eye view. The target perception system provided by the present invention has a lower production cost than the autonomous valet parking method based on laser radar. At the same time, the target perception system provided by the present invention can obtain richer visual information than the autonomous valet parking method based solely on ultrasonic radar, thereby identifying the category of the detected object.

[0041] Combine the following Figure 1-Figure 2 The object perception method of the present invention is described.

[0042] Figure 1 is one of the flow charts of the target perception method provided by the present invention, such as Figure 1 As shown, the method includes the following steps: Step 101, obtaining a panoramic bird's-eye view image of a target moving object.

[0043] It should be noted that the execution subject of the embodiment of the present invention is a target sensing device. The target sensing device may be a vehicle-mounted computer of the target moving body; the target sensing device may also be other electronic devices such as a user terminal.

[0044] Specifically, the perceived target in the surrounding environment of the target moving body is the perceived object of the target perception method provided by the present invention. Based on the target perception method provided by the present invention, the perceived target in the surrounding environment of the target moving body can be perceived during the movement of the target moving body.

[0045] It is understandable that the target moving object in the embodiment of the present invention may be determined based on actual needs. The target moving object is not specifically limited in the embodiment of the present invention.

[0046] It should be noted that the mobile body in the embodiment of the present invention may include movable objects such as vehicles, robots, and drones. The specific type of the mobile body is not limited in the embodiment of the present invention.

[0047] Optionally, the moving object in the embodiment of the present invention may be a vehicle. The target moving object is taken as an example to illustrate the target sensing method provided by the present invention.

[0048] At least four fisheye cameras or wide-angle cameras are arranged around the body of the target moving body. Using the above-mentioned fisheye cameras or wide-angle cameras, multiple environmental images including the surrounding environment of the target moving body can be obtained, and the above-mentioned environmental images include environmental information of all angles around the target moving body.

[0049] Optionally, a fisheye camera or a wide-angle camera is respectively arranged at the front, rear, left and right positions of the target moving body.

[0050] After acquiring multiple environmental images of the target moving object's surrounding environment captured by the above-mentioned fisheye cameras or wide-angle cameras, operations such as dedistortion, projection and splicing can be performed on the above-mentioned environmental images to obtain a panoramic bird's-eye view image of the target moving object.

[0051] The following takes the example of setting a fisheye camera at the front, rear, left and right positions of the target moving body as an example to specifically describe the specific steps of obtaining a panoramic bird's-eye view image of the target moving body.

[0052] Figure 2 This is the second flow chart of the target perception method provided by the present invention. Figure 2 As shown, the acquired environment image set I includes multiple environment images of the target moving body's surrounding environment. fisheye By performing distortion correction, we can obtain the distortion-corrected environment image set I correction :

[0053] I correction =C orrecti0n (I fisheye )

[0054] Among them, I fisheyerepresents an environmental image set including multiple environmental images of the target moving object's surrounding environment, including the front original environmental image I obtained from the front fisheye camera fisheye_front , the rear original environment image I obtained from the rear fisheye camera fisheye_rear , the left original environment image I obtained from the left fisheye camera fisheye_left And obtain the right original environment image I from the right fisheye camera fisheye_right ; C orrection Represents the fisheye correction algorithm; I correction Represents the set of environment images after distortion correction, including the original environment image I fisheye_front Image I after distortion correction correction_front , post-original environment image I fishey_erar Image I after distortion correction correction_rear , left original environment image I fisheye_left Image I after distortion correction correction_left And the right original environment image I fisheye_right Image I after distortion correction correction_right .

[0055] The transformation matrix M obtained by image calibration transform Set I of environmental images after distortion correction correction Perform a bird's-eye view projection to obtain an environmental image set I from a bird's-eye view bev :

[0056] I bev =T ransfom (I correction , M transform )

[0057] Among them, the transformation matrix M transform The distortion-corrected environment image set I correction The calibration process is carried out using the fisheye camera internal reference and fisheye correction algorithm; T rannsform represents the perspective transformation function; I bev Represents a set of environmental images from a bird's-eye view, including I correction_front Represents the front environment image I from the bird's-eye view obtained by projection transformation bev_front , I correction_rear Represents the rear environment image I from the bird's-eye view obtained by projection transformation bevr_ear , I correction_left Represents the left-placed environment image I from the bird's-eye view obtained by projection transformation bev_left , I correction_right Represents the right-side environment image I from the bird's-eye view obtained by projection transformation bev_right .

[0058] By matching the key points in the environment images under different bird's-eye view perspectives, the positions of the environment images when the panoramic bird's-eye view of the target moving body is stitched can be confirmed, and then the transformation matrix M that projects the environment image under the bird's-eye view perspective to the panoramic bird's-eye view can be obtained. stitch .

[0059] According to the position of the environment images under the bird's-eye view when the panoramic bird's-eye view of the target moving body is stitched, the overlapping area between the environment images can be determined for weighted fuzzy processing calibration. For the non-overlapping area, its fuzzy weight is 1; for the overlapping area, its fuzzy processing weight is q. Finally, the fuzzy weight mask set M can be obtained. ambigious .

[0060]

[0061] Among them, q i represents the fuzzy weight of the i-th pixel in the overlapped area; d 1 represents the Euclidean distance between the ith pixel in the overlapped area and the center point of the environment image where the ith pixel in the overlapped area is located; d 2 Represents the Euclidean distance from the center point of the environment image it overlaps.

[0062] The fuzzy weight mask set M obtained by calibrating the environment image ambigious A collection of environmental images from a bird's eye view bev Perform fuzzy processing to obtain the weighted environmental image set I ambigious :

[0063] I ambigious =M ambigious ×I bev

[0064] Among them, × represents the matrix multiplication operator; I ambigious represents the set of environmental images from the bird's-eye view after weighted blur processing, including the front environmental image I after weighted blur processing ambigious_front , the post-environment image I after weighted blurring ambigious_rear , the left-placed environment image I after weighted blurring ambigious_left And the right-side environment image I after weighted blur processing ambigious_right .

[0065] Using the transformation matrix M stitch The environmental image set I from the bird's-eye view after weighted fuzzy processing ambigious Stitching is performed to obtain a panoramic bird's-eye view image I of the target moving object for perception;

[0066] I=T ransform (I ambigious , Mstitch )

[0067] Among them, T ransform represents the perspective transformation function. I represents the panoramic bird's-eye view image of the target moving object used for perception.

[0068] Step 102: identify a first perceived target in the panoramic bird's-eye view image, and generate a first identification frame in the panoramic bird's-eye view image for marking the identified first perceived target, wherein the first perceived target includes at least one of an obstacle, a pedestrian, a ground sign, and a traffic sign.

[0069] It should be noted that the ground signs in the embodiments of the present invention may include but are not limited to lane signs, parking signs, safety signs, guide signs, and prohibition signs, etc. Among them, lane signs include lane lines, and parking signs include parking space lines.

[0070] The traffic signs in the embodiments of the present invention are facilities used to provide road users with guidance, warnings, prohibitions or instructions about road traffic, including but not limited to warning signs, prohibition signs, instruction signs, guide signs, road construction safety signs and auxiliary signs.

[0071] Specifically, after obtaining the panoramic bird's-eye view image I of the target moving object, deep learning technology can be used to perform image recognition on the panoramic bird's-eye view image I to identify the first perceived target in the panoramic bird's-eye view image I, and then a first identification box can be generated in the panoramic bird's-eye view image I to mark the identified first perceived target.

[0072] It is understandable that the number of first perception targets recognized by image recognition in the panoramic bird's-eye view image I may be one or more. Accordingly, the number of first identification frames generated in the panoramic bird's-eye view image I may be one or more.

[0073] As an optional embodiment, identifying a first perceived target in a panoramic bird's-eye view image, and generating a first identification frame in the panoramic bird's-eye view image for marking the identified first perceived target, includes: when the first perceived target includes an obstacle, inputting the panoramic bird's-eye view image into an obstacle recognition model, and obtaining an obstacle mask image of the panoramic bird's-eye view image output by the obstacle recognition model; when the first perceived target includes a pedestrian, inputting the panoramic bird's-eye view image into a pedestrian recognition model, and obtaining a pedestrian mask image of the panoramic bird's-eye view image output by the obstacle recognition model; when the first perceived target includes a ground sign, inputting the panoramic bird's-eye view image into a ground sign recognition model, and obtaining a ground sign mask image of the panoramic bird's-eye view image output by the obstacle recognition model; when the first perceived target includes a traffic sign, inputting the panoramic bird's-eye view image into a traffic sign recognition model, and obtaining a traffic sign mask image of the panoramic bird's-eye view image output by the obstacle recognition model.

[0074] Among them, the obstacle recognition model is trained based on the sample panoramic bird's-eye view image and the obstacle mask image of the sample panoramic bird's-eye view image; the pedestrian image recognition model is trained based on the sample panoramic bird's-eye view image and the pedestrian mask image of the sample panoramic bird's-eye view image; the ground sign image recognition model is trained based on the sample panoramic bird's-eye view image and the ground sign mask image of the sample panoramic bird's-eye view image; the traffic sign image recognition model is trained based on the sample panoramic bird's-eye view image and the traffic sign mask image of the sample panoramic bird's-eye view image; the obstacle mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the sample panoramic bird's-eye view image. the area in the sample panoramic bird's-eye view image excluding the area where the obstacle is located; the pedestrian mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the area in the sample panoramic bird's-eye view image excluding the area where the pedestrian is located; the ground sign mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the area in the sample panoramic bird's-eye view image excluding the area where the ground sign is located; the traffic sign mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the area in the sample panoramic bird's-eye view image excluding the area where the traffic sign is located.

[0075] Perform contour detection on the highlight area not covered by the black layer in the mask image of the panoramic bird's-eye view image to determine the contour of the first perceived target. The mask image of the panoramic bird's-eye view image includes at least one of an obstacle mask image of the panoramic bird's-eye view image, a pedestrian mask image of the panoramic bird's-eye view image, a ground sign mask image of the panoramic bird's-eye view image, and a traffic sign mask image of the panoramic bird's-eye view image.

[0076] Based on the contour of the first perceived target, the first perceived target whose size is smaller than the first contour perimeter threshold is eliminated from the mask image.

[0077] Based on the outline of the first perceived target, after generating a minimum circumscribed rectangle of the first perceived target in the mask image as a first identification frame, the black layer on the mask image is removed.

[0078] Specifically, when the first perceived target includes an obstacle, the panoramic bird's-eye view image I of the target moving body is input into the obstacle recognition model M. 0bject Afterwards, the obstacle recognition model M object The obstacles in the panoramic bird's-eye view image I can be recognized by image recognition, and then based on the recognition result, a black layer can be covered in the area except the area where the obstacles are located in the panoramic bird's-eye view image to obtain and output the obstacle mask image of the panoramic bird's-eye view image.

[0079] It can be understood that the obstacle mask image S of the panoramic bird's-eye view image obiect is a panoramic bird's-eye view image I covered with a black layer, and the black layer is used to cover the area of ​​the panoramic bird's-eye view image I except the area where the obstacle is located. obiect The highlighted area in the figure is where the obstacle is located.

[0080] Optionally, the obstacle recognition model M in the embodiment of the present invention obiect Models can be learned for semantic segmentation.

[0081] It should be noted that the obstacle mask image S in the embodiment of the present invention obiect The identification information of obstacles in each highlighted area is marked in the figure.

[0082] Based on the mask image S obiect By using the type and identification information of the first perceived target corresponding to each highlighted area marked in , the category set of different types of first perceived targets can be obtained, for example, the category set C of the first perceived target of type obstacle object .

[0083] It should be noted that, in the case where the types of the first perception targets include pedestrians, ground signs, traffic signs or other first perception targets, the specific process of obtaining the pedestrian mask images, ground sign mask images, traffic sign mask images or other first perception target mask images of the panoramic bird's-eye view image and the pedestrian mask images, ground sign mask images, traffic sign mask images or other types of first perception target mask images of the panoramic bird's-eye view image can refer to the description of the above embodiments, and will not be repeated in the embodiments of the present invention. The following takes the example of the type of the first perception target including obstacles to explain the specific process of generating a first identification box for marking the identified obstacles in the panoramic bird's-eye view image.

[0084] For the obstacle mask image S object By performing contour detection on the highlighted area in the panoramic bird’s-eye view image I, the initial set of obstacle targets can be obtained.

[0085]

[0086] Among them, S i Represents the obstacle mask image S object The i-th highlighted area in C ontourS represents the contour detection function; represents the initial set of obstacle targets in the panoramic bird's-eye view image I, including the obstacle mask image S object The set of all obstacle contours identified in .

[0087] According to the first contour perimeter threshold set The initial set of obstacle targets in the panoramic bird's-eye view image I Filter and remove the initial set of obstacle targets The size of the middle contour is less than the first contour perimeter threshold set Obstacles, get the obstacle target set of the panoramic bird's-eye view image I

[0088]

[0089] Among them, F ilter represents an obstacle contour filter, which is based on the first contour perimeter threshold set S tandard For the candidate obstacle target set Make a judgment and select the first contour perimeter threshold set S tandard All obstacles with the first contour perimeter threshold in .

[0090] It should be noted that the first contour perimeter threshold set S tandardThe first contour perimeter threshold value may include multiple first contour perimeter threshold values. The first contour perimeter threshold value in the embodiment of the present invention may be determined based on prior knowledge and / or actual conditions. The first contour perimeter threshold value is not specifically limited in the embodiment of the present invention.

[0091] Obstacle target set based on panoramic bird's-eye view image I The outline of the obstacle can be found in the obstacle mask image S object Generate the minimum bounding rectangle of the obstacle as the first identification box for marking the obstacle

[0092]

[0093] In the obstacle mask image S object Generate a first identification box for marking obstacles Afterwards, the obstacle mask image S can be removed object On the black layer, get the first identification box Panoramic bird's-eye view image with obstacles annotated.

[0094] It should be noted that, when the types of the first perception targets include pedestrians, ground signs, traffic signs or other types, the specific steps for obtaining the panoramic bird's-eye view image with pedestrians, ground signs, traffic signs or other first perception targets marked with the first identification frame can refer to the steps for obtaining the panoramic bird's-eye view image with the first identification frame. The specific steps of marking the panoramic bird's-eye view image with obstacles will not be described in detail in the embodiment of the present invention.

[0095] Step 103: determine the point on the first identification frame in the panoramic bird's-eye view image that is closest to the geometric center of the target moving body as the target point corresponding to the first identification frame, crop the first identification frame based on the target point corresponding to the first identification frame and the first border size threshold, and determine the cropped first identification frame as the second identification frame.

[0096] It should be noted that when a fisheye camera or a wide-angle camera is capturing images, due to the very wide field of view (usually 180 degrees or greater), the edges of the captured images will appear stretched and distorted, causing objects far from the center to appear deformed. In addition, in images captured by fisheye cameras or wide-angle cameras, objects in the central area will appear larger and are prone to radial distortion (i.e., gradual deformation from the center to the edge).

[0097] Correspondingly, the panoramic bird's-eye view image obtained based on the images collected by the fisheye camera or the wide-angle camera also has the defect that the edges of the image will appear stretched and distorted, making the objects far away from the center appear deformed.

[0098] Therefore, in the embodiment of the present invention, the edges of the panoramic bird's-eye view image I of the target moving body may show a stretching and distorted effect, so that objects far from the center may show a deformed appearance defect. After the point on the first identification frame in the panoramic bird's-eye view image I that is closest to the geometric center of the target moving body is determined as the target point corresponding to the first identification frame, the first identification frame is cropped based on the target point corresponding to the first identification frame and the first border size threshold. On the basis of ensuring that the position of the identified first perceived target remains unchanged and the relative position between the identified first perceived target and the target moving body remains unchanged, the distorted area in the first identification frame can be cropped out, thereby improving the accuracy of target perception based on the panoramic bird's-eye view image.

[0099] It should be noted that the first border size threshold in the embodiment of the present invention may include a height threshold and a width threshold of the border. The first border size threshold in the embodiment of the present invention may be determined based on prior knowledge and / or actual conditions. The specific value of the first border size is not limited in the embodiment of the present invention.

[0100] As an optional embodiment, based on the target point corresponding to the first identification frame and the first border size threshold, the first identification frame is cropped and the cropped first identification frame is determined as the second identification frame, including: determining the border where the target point corresponding to the first identification frame is located as the target border.

[0101] When the lengths on both sides of the target point on the target border are greater than half of the size threshold corresponding to the target border in the first border size threshold, the target border is used as the boundary of the cropping area, the target point corresponding to the first identification frame is used as the midpoint of the boundary, and a cropping area whose size meets the first border size threshold is determined in the first identification frame. When the length on either side of the target point on the target border is less than half of the size threshold corresponding to the target border in the first border size threshold, the target border is used as the boundary of the cropping area, the vertex on the target border closer to the target point corresponding to the first identification frame is used as the starting point, and a cropping area whose size meets the first border size threshold is determined in the first identification frame.

[0102] The first identification frame is cropped along the boundary of the cropping area, and the cropped first identification frame is determined as the second identification frame.

[0103] Specifically, when the first perceived target includes an obstacle, the first identification frame may be The point closest to the geometric center of the target moving body is determined as the first identification frame The corresponding target point P, based on the first border size threshold S 1 , you can set the first identification frame Crop and place the cropped first identification frame Determine the second identification frame

[0104]

[0105]

[0106] Among them, Bbox min Represents the bounding box clipping function.

[0107] Bbox border clipping function min You can get the first identification frame The length of the target point P on both sides of the target border where the corresponding target point P is located can be calculated based on the given first border size threshold S 1 and the length of both sides of the target point P on the target border, in the first identification frame The cropping area is determined in the figure, and then the cropping area can be cropped to obtain a size that is the first border size threshold S 1 And the first identification frame The corresponding target point P is located in the second identification box on the border

[0108] It should be noted that, when the target border is a border in the length direction, the size threshold corresponding to the target border is the size threshold in the length direction in the first border size threshold; when the target border is a border in the width direction, the size threshold corresponding to the target border is the size threshold in the width direction in the first border size threshold.

[0109] It should be noted that, when the first perception target includes pedestrians, ground signs, traffic signs or other first perception targets, the specific steps of cropping the first identification frame used to mark pedestrians, ground signs, traffic signs or other first perception targets are the same as the steps of cropping the first identification frame used to mark obstacles. Please refer to the steps of cropping the first identification frame used to mark obstacles, which will not be repeated in the embodiments of the present invention.

[0110] Step 104: Based on the second identification frame, obtain the position information of the first perceived target in the real world as the target perception result of the target moving body.

[0111] Specifically, after determining the second identification frame in the panoramic bird's-eye view image I, the position information of the first perceived target in the real world can be obtained through numerical calculation based on the position of the second identification frame in the panoramic bird's-eye view image I and the position projection of the panoramic bird's-eye view image I in the real world, as the target perception result of the target mobile object.

[0112] For example, in the case where the first perception target includes an obstacle, the position information of the obstacle in the real world can be obtained as the target perception result of the target mobile body through numerical calculation based on the position of the second identification frame used to mark the obstacle in the panoramic bird's-eye view image I and the position projection of the panoramic bird's-eye view image I in the real world.

[0113] It should be noted that the position information of the first perceived target in the real world may include the coordinates of the four vertices of the second identification box used to mark the first perceived target in the world coordinate system, and may also include the coordinates of the target point corresponding to the first identification box on the border of the second identification box in the world coordinate system.

[0114] In the embodiment of the present invention, after identifying a first perceived target in a panoramic bird's-eye view image of a target moving body and generating a first identification frame for marking the identified first perceived target in the panoramic bird's-eye view image, a point on the first identification frame in the panoramic bird's-eye view image that is closest to the geometric center of the target moving body is determined as a target point corresponding to the first identification frame, and the first identification frame is cropped based on the target point corresponding to the first identification frame and a first frame size threshold, and the cropped first identification frame is determined as a second identification frame, and then based on the second identification frame, the position information of the first perceived target in the real world is obtained as a target perception result of the target moving body. Based on the fact that the edges of the image in the panoramic bird's-eye view image will present a stretching and distorted effect, so that the object far from the center will present a deformed appearance defect, by cropping the first identification frame, the distorted area in the first identification frame is cropped, thereby improving the accuracy of target perception based on the panoramic bird's-eye view image, and more simply and efficiently eliminating the influence of the distortion of the panoramic bird's-eye view image on the target perception accuracy, and providing a more accurate data basis for unmanned driving technology and automatic parking technology.

[0115] As an optional embodiment, after obtaining the position information of the first perception target in the real world, the method further includes: in the case where there is already a first perception target tracking list, matching the second identification box with the existing historical identification box in the first perception target tracking list based on the Hungarian matching algorithm.

[0116] When the second identification frame successfully matches the historical identification frame, the tracking information of the historical identification frame that successfully matches the second identification frame in the first perception target tracking list is updated based on the position information of the first perceived target in the real world, the identification information of the first perceived target and the moment of obtaining the panoramic bird's-eye view image. When the second identification frame fails to match the historical identification frame, the position information of the first perceived target in the real world, the identification information of the first perceived target and the moment of obtaining the panoramic bird's-eye view image are determined as the tracking information of the second identification frame, and the second identification frame and the tracking information of the second identification frame are added to the first perception target tracking list. For the historical identification frame in the first perception target tracking list that fails to match the second identification frame successfully, when the duration of the historical identification frame that fails to match the second identification frame being added to the first perception target tracking list exceeds the first duration threshold, the historical identification frame that fails to match the second identification frame is removed from the first perception target tracking list.

[0117] Specifically, the first perception target tracking list in the embodiment of the present invention may include at least one of an obstacle target tracking list, a pedestrian tracking list, a ground sign tracking list, and a traffic sign tracking list.

[0118] In the case where the first perceived target includes an obstacle, the obstacle target tracking list T can be obtained by using the Hungarian matching algorithm based on key point tracking. object The existing historical identification box Bbox obiect ′ and the second identification frame in the panoramic bird's-eye view image I to make a match.

[0119] It should be noted that the obstacle target tracking list T object Includes the existing historical identification box Bbox obiect ′ tracking information, Bbox obiect The tracking information of ′ can include the historical identification box Bbox obiect ′Location information in the real world, historical identification box Bbox object ′The location information of the target point on the border in the real world, the historical identification box Bbox object The acquisition time of the historical panoramic bird's-eye view of the target moving body and the historical identification box Bbox obiect ’ The identification information of the obstacle marked.

[0120] It can be understood that the historical panoramic bird's-eye view of the target moving object is acquired before the panoramic bird's-eye view image I of the target moving object.

[0121] Optionally, a preset time interval is provided between the moment of acquiring the panoramic bird's-eye view image I of the target moving body and the moment of acquiring the previous historical panoramic bird's-eye view image of the panoramic bird's-eye view image I of the target moving body. The preset time interval may be determined by prior knowledge and / or actual conditions, for example, the preset time interval may be 1 second.

[0122] The historical identification box Bbox in the embodiment of the present invention obiect ′ is generated by the method described in the above embodiments in the historical panoramic bird's-eye view of the target moving body, and the historical identification box Bbox obiect ′ is the second identification frame in the historical panoramic bird's-eye view of the target moving body, the historical identification frame Bbox obiect ′The target point on the border is the historical identification box Bbox obiect The target point corresponding to the first identification box on the border of ′.

[0123] It should be noted that the historical identification box Bbox obiect ′Location information in the real world, which may include the historical identification box Bbox object ′The coordinates of the four vertices in the world coordinate system; historical identification box Bbox object ′The location information of the target point on the border in the real world, which may include the historical identification box Bbox obiect ’The coordinates of the target point on the border in the world coordinate system.

[0124] If the obstacle target tracking list T object A historical identification box Bbox in obiect ' and any second identification frame If the match is successful, the second identification box The position information of the marked obstacle in the real world, the second identification frame The identification information of the marked obstacles and the moment of obtaining the panoramic bird's-eye view image I of the target moving body are updated to the above historical identification box Bbox object ′ tracking information.

[0125] If the obstacle target tracking list T object A historical identification box Bbox in object ' and any second identification frame If no match is successful, you can compare the above historical identification box Bbox object 'Add to obstacle target tracking list xobject The duration of the first duration threshold, if the above historical identification box Bbox object 'Add to obstacle target tracking list T object If the duration is greater than the first duration threshold, then the obstacle target tracking list Tobject Delete the above historical identification box Bbox object ′ and the above historical identification box Bbox object ′ tracking information.

[0126] If a second identification frame and obstacle target tracking list T object Any historical identification box Bbox in object ' and neither match successfully, then the second identification frame The position information of the marked obstacle in the real world, the second identification frame The identification information of the marked obstacles and the time of obtaining the panoramic bird's-eye view image I of the target moving body are used as the second identification frame Tracking information, the second identification box above and the second identification box The tracking information of is added to the obstacle target tracking list T object middle.

[0127] It should be noted that when the first perception target includes pedestrians, ground signs, traffic signs or other first perception targets, the specific steps for updating the obstacle target tracking list, pedestrian tracking list, ground sign tracking list or traffic sign tracking list can be referred to the contents in the above embodiments and will not be repeated in the embodiments of the present invention.

[0128] It should be noted that the first duration threshold in the embodiment of the present invention is determined based on prior knowledge and / or actual conditions. The specific value of the first duration threshold in the embodiment of the present invention is not limited.

[0129] It should be noted that, when the first perceived target tracking list has not yet been created, the first perceived target tracking list can be created after the position information of the first perceived target in the real world is obtained, and the position information of the first perceived target in the real world, the identification information of the first perceived target and the moment when the panoramic bird's-eye view image I of the target moving body is obtained can be used as the tracking information of the second identification box where the first perceived target is located, and the above-mentioned second identification box and the tracking information of the above-mentioned second identification box can be added to the above-mentioned first perceived target tracking list.

[0130] The embodiment of the present invention updates the first perception target tracking list after matching the second identification frame with the existing historical identification frame in the first perception target tracking list based on the Hungarian matching algorithm. This can achieve uninterrupted perception of the same first perception target, continuously obtain the latest position of the first perception target, and further improve the accuracy of target perception based on panoramic bird's-eye view images.

[0131] As an optional embodiment, after acquiring the panoramic bird's-eye view image of the target moving body, the method also includes: identifying a second perceived target in the panoramic bird's-eye view image, and generating a third identification frame in the panoramic bird's-eye view image for marking the identified second perceived target, wherein the second perceived target includes a parking space.

[0132] Specifically, the second perceived target in the panoramic bird's-eye view image is identified, and a third identification frame for marking the identified second perceived target is generated in the panoramic bird's-eye view image, including: inputting the panoramic bird's-eye view image I of the target moving body into the parking space recognition model, and obtaining the parking space mask image of the panoramic bird's-eye view image output by the parking space recognition model.

[0133] Among them, the obstacle recognition model is trained based on the sample panoramic bird's-eye view image and the parking space mask image of the sample panoramic bird's-eye view image; the parking space mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the area in the sample panoramic bird's-eye view image except the area where the parking space is located.

[0134] As an optional embodiment, a second perceived target in a panoramic bird's-eye view image is identified, and a third identification frame for marking the identified second perceived target is generated in the panoramic bird's-eye view image, including: inputting the panoramic bird's-eye view image into a parking space recognition model, and obtaining a parking space mask image of the panoramic bird's-eye view image output by the parking space recognition model.

[0135] Among them, the parking space recognition model is trained based on the sample panoramic bird's-eye view image and the parking space mask image of the sample panoramic bird's-eye view image; the parking space mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the area in the sample panoramic bird's-eye view image except the area where the parking space is located.

[0136] Contour detection is performed on a highlight area not covered by a black layer in a second perception target mask image of the panoramic bird's-eye view image to determine the contour of the second perception target, wherein the second perception target mask image includes a parking space mask image.

[0137] Based on the contour of the second perceived target, the second perceived target whose size is smaller than the second contour perimeter threshold is eliminated from the mask image.

[0138] Based on the contour of the second perceived target, after generating the minimum circumscribed rectangle of the second perceived target in the mask image as the third identification frame, the black layer on the mask image is removed.

[0139] Specifically, the panoramic bird's-eye view image I of the target moving object is input into the parking space recognition model M parkingAfterwards, the parking space recognition model M parking The parking space in the panoramic bird's-eye view image I can be image-recognized, and based on the recognition result, a black layer can be covered in the area except the area where the parking space is located in the panoramic bird's-eye view image to obtain and output the parking space mask image S of the panoramic bird's-eye view image. parking .

[0140] It can be understood that the parking space mask image S of the panoramic bird's-eye view image parking The parking space mask image S is a panoramic bird's-eye view image I covered with a black layer, and the black layer is used to cover the area of ​​the panoramic bird's-eye view image I except the area where the parking space is located. parking The highlighted area in the figure is where the obstacle is located.

[0141] Optionally, the parking space recognition model M in the embodiment of the present invention parking Models can be learned for semantic segmentation.

[0142] It should be noted that the parking space recognition model M in the embodiment of the present invention parking The identification information of the parking spaces in each highlighted area is marked.

[0143] For the parking space mask image S parking By performing contour detection on the highlighted area in the panoramic bird’s-eye view image I, the initial set of parking space targets can be obtained.

[0144]

[0145] Among them, S i Represents the parking space mask image S parking The i-th highlighted area in C ontours represents the contour detection function; represents the initial set of parking space targets in the panoramic bird's-eye view image I, including the parking space mask image S parking The set of all parking space contours identified in .

[0146] According to the second contour perimeter threshold set The initial set of parking space targets for the panoramic bird's-eye view image I Filter and remove the initial set of parking space targets The size of the middle contour is less than the second contour perimeter threshold set Obstacles, get the parking space target set of the panoramic bird's-eye view image I

[0147]

[0148] Among them, F ilterrepresents a parking space contour filter, which comprises a second contour perimeter threshold set For the candidate parking space target set Make a judgment and select the second contour perimeter threshold set All obstacles within the second contour perimeter threshold.

[0149] It should be noted that the second contour perimeter threshold set The second contour perimeter threshold may include multiple second contour perimeter thresholds. The second contour perimeter threshold in the embodiment of the present invention may be determined based on prior knowledge and / or actual conditions. The second contour perimeter threshold is not specifically limited in the embodiment of the present invention.

[0150] Parking space target set based on panoramic bird's-eye view image I The outline of the parking space can be obtained by masking the parking space image S parking Generate the minimum enclosing rectangle of the parking space as the third identification box for marking obstacles

[0151]

[0152] Mask image Sx in parking space arking Generate a third identification frame for marking parking spaces Afterwards, the obstacle parking space mask image S can be removed parking On the black layer, get the third identification box Panoramic bird's-eye view image with parking spaces annotated.

[0153] The geometric center of the third identification frame in the panoramic bird's-eye view image is determined as the target point corresponding to the third identification frame, and the third identification frame is cropped based on the target point corresponding to the third identification frame and the second border size threshold, and the cropped third identification frame is determined as the fourth identification frame.

[0154] Specifically, since the panoramic bird's-eye view image I of the target moving object obtained by the image captured by the fisheye camera or the wide-angle camera also has the defect that the edges of the image will show a stretching and distortion effect, so that objects far away from the center will appear deformed.

[0155] Therefore, in the embodiment of the present invention, the edges of the panoramic bird's-eye view image I based on the target moving body will show a stretching and distorted effect, so that objects far away from the center will show a deformed appearance defect. After the geometric center of the third identification box in the panoramic bird's-eye view image I is determined as the target point corresponding to the third identification box, the third identification box is cropped based on the target point corresponding to the third identification box and the second border size threshold. On the basis of ensuring that the position of the identified second perceived target remains unchanged and the relative position between the identified second perceived target and the target moving body remains unchanged, the distorted area in the third identification box can be cropped out, thereby improving the accuracy of target perception based on the panoramic bird's-eye view image.

[0156] It should be noted that the second border size threshold in the embodiment of the present invention may include a height threshold and a width threshold of the border. The second border size threshold in the embodiment of the present invention may be determined based on prior knowledge and / or actual conditions. The specific value of the second border size is not limited in the embodiment of the present invention.

[0157] It can be understood that the second border size threshold in the embodiment of the present invention is greater than or equal to the border size threshold.

[0158] As an optional embodiment, the third identification frame is cropped based on the target point corresponding to the third identification frame and the second border size threshold, including: taking the target point corresponding to the third identification frame as the geometric center, and determining a cropping area in the third identification frame whose size meets the second border size threshold.

[0159] The third identification frame is cropped along the boundary of the cropping area, and the cropped third identification frame is determined as the fourth identification frame.

[0160] Specifically, the third identification frame The geometric center is determined as the third identification frame The corresponding target point Q, based on the second border size threshold S 2 , you can set the third identification frame Crop and place the cropped third marker frame Determine the fourth identification frame

[0161]

[0162]

[0163] Among them, Bbox min Represents the bounding box clipping function.

[0164] Bbox border clipping function min You can use the third identification frame The corresponding target point Q is the geometric center, in the third identification frame The cropping area is determined in the figure, and then the cropping area can be cropped to obtain a size that is the second border size threshold S 2 And the third identification box The corresponding target point Q is the fourth identification box of the geometric center

[0165] Based on the fourth identification frame, the position information of the second perceived target in the real world is obtained as the target perception result of the target moving body.

[0166] Specifically, after determining the fourth identification frame in the panoramic bird's-eye view image I, the position information of the second perceived target in the real world can be obtained through numerical calculation based on the position of the fourth identification frame in the panoramic bird's-eye view image I and the position projection of the panoramic bird's-eye view image I in the real world, as the target perception result of the target mobile object.

[0167] In the embodiment of the present invention, after identifying the second perceived target in the panoramic bird's-eye view image of the target moving body and generating a third identification frame for marking the identified second perceived target in the panoramic bird's-eye view image, the geometric center of the third identification frame in the panoramic bird's-eye view image is determined as the target point corresponding to the third identification frame, the third identification frame is cropped based on the target point corresponding to the third identification frame and the second frame size threshold, and the cropped third identification frame is determined as a fourth identification frame, and then based on the fourth identification frame, the position information of the second perceived target in the real world is obtained as the target perception result of the target moving body. Based on the fact that the edges of the image in the panoramic bird's-eye view image will present a stretching and distorted effect, so that the object far from the center will present a deformed appearance defect, by cropping the third identification frame, the distorted area in the third identification frame is cropped, thereby improving the accuracy of target perception based on the panoramic bird's-eye view image, and more simply and efficiently eliminating the influence of the distortion of the panoramic bird's-eye view image on the target perception accuracy, and providing a more accurate data basis for unmanned driving technology and automatic parking technology.

[0168] As an optional embodiment, after obtaining the position information of the second perception target in the real world, the method also includes: in the case where there is already a second perception target tracking list, based on the Hungarian matching algorithm, matching the fourth identification box with the existing historical identification box in the second perception target tracking list.

[0169] When the fourth identification frame successfully matches the historical identification frame, the tracking information of the historical identification frame that successfully matches the fourth identification frame in the second perception target tracking list is updated based on the position information of the second perceived target in the real world, the identification information of the second perceived target and the moment of obtaining the panoramic bird's-eye view image. When the fourth identification frame fails to match the historical identification frame, the position information of the second perceived target in the real world, the identification information of the second perceived target and the moment of obtaining the panoramic bird's-eye view image are determined as the tracking information of the fourth identification frame, and the fourth identification frame and the tracking information of the fourth identification frame are added to the second perception target tracking list. For the historical identification frame in the second perception target tracking list that fails to match the fourth identification frame successfully, when the duration of adding the historical identification frame that fails to match the fourth identification frame successfully to the second perception target tracking list exceeds the second duration threshold, the historical identification frame that fails to match the fourth identification frame successfully is removed from the second perception target tracking list.

[0170] Specifically, since the second perception target includes a parking space, the second perception target tracking list includes a parking space target tracking list T parking .

[0171] The Hungarian matching algorithm based on key point tracking is used to track the parking space target list T parking The existing historical identification box Bbox parking ′ and the panoramic bird's-eye view image I are the fourth identification frames to make a match.

[0172] It should be noted that the parking space target tracking list T parking The existing historical identification box Bbox parking ′ tracking information, historical identification box Bbox parking ′Location information in the real world, historical identification box Bbox′ parking The real-world location information and historical identification box Bbox of the target point in parking The acquisition time of the historical panoramic bird's-eye view of the target moving body and the historical identification box Bbox parking ’The identification information of the marked parking space.

[0173] It can be understood that the historical panoramic bird's-eye view of the target moving object is acquired before the panoramic bird's-eye view image I of the target moving object.

[0174] Optionally, a preset time interval is provided between the moment of acquiring the panoramic bird's-eye view image I of the target moving body and the moment of acquiring the previous historical panoramic bird's-eye view image of the panoramic bird's-eye view image I of the target moving body. The preset time interval may be determined by prior knowledge and / or actual conditions, for example, the preset time interval may be 1 second.

[0175] The historical identification box Bbox in the embodiment of the present invention parking ′ is generated by the method described in the above embodiments in the historical panoramic bird's-eye view of the target moving body, and the historical identification box Bbox parking ′ is the fourth identification frame in the historical panoramic bird's-eye view of the target moving body, the historical identification frame Bbox parking ′ is the target point in the historical identification box Bbox parking The geometric center of ′.

[0176] It should be noted that the historical identification box Bbox parking ′Location information in the real world, which may include the historical identification box Bbox parking ′The coordinates of the four vertices in the world coordinate system; historical identification box Bbox parking ′The location information of the target point on the border in the real world, which may include the historical identification box Bbox parking The coordinates of the target point in ′ in the world coordinate system.

[0177] If the parking space target tracking list T parking中 A historical mark Bbox parking ' and any fourth identification frame If the match is successful, the fourth identification box The location information of the marked parking space in the real world, the fourth identification frame The marked parking space identification information and the moment of obtaining the panoramic bird's-eye view image I of the target moving body are updated to the above historical identification box Bbox parking ′ tracking information.

[0178] If the parking space target tracking list T parking A historical identification box Bbox in parking ' and any fourth identification frame If no match is successful, you can compare the above historical identification box Bbox parking 'Add to the parking space target tracking list T parking The duration of the second duration threshold, if the above historical identification box Bbox parking 'Add to the parking space target tracking list T parking If the duration is greater than the second duration threshold, then the parking space target tracking list T parking Delete the above historical identification box Bbox parking ′ and the above historical identification box Bbox parking ′ tracking information.

[0179] If a fourth identification frame Tracking list with parking space target Tparking Any historical identification box Bbox in parking ' and neither match successfully, then the fourth identification frame The location information of the marked obstacle in the real world, the fourth identification frame The marking information of the marked parking space and the time of obtaining the panoramic bird's-eye view image I of the target moving body are used as the fourth marking frame Tracking information, the fourth identification box above and the fourth identification box The tracking information of the parking space is added to the parking space target tracking list T parking middle.

[0180] It should be noted that the second duration threshold in the embodiment of the present invention is determined based on prior knowledge and / or actual conditions. The specific value of the second duration threshold in the embodiment of the present invention is not limited.

[0181] It should be noted that, when the second perception target tracking list has not yet been created, the second perception target tracking list can be created after the position information of the second perception target in the real world is obtained, and the position information of the second perception target in the real world, the identification information of the second perception target and the moment when the panoramic bird's-eye view image I of the target moving body is obtained can be used as the tracking information of the fourth identification box where the second perception target is located, and the above-mentioned fourth identification box and the tracking information of the above-mentioned fourth identification box can be added to the above-mentioned second perception target tracking list.

[0182] The embodiment of the present invention updates the second perception target tracking list after matching the fourth identification frame with the existing historical identification frame in the second perception target tracking list based on the Hungarian matching algorithm. This can achieve uninterrupted perception of the same second perception target, continuously obtain the latest position of the second perception target, and further improve the accuracy of target perception based on panoramic bird's-eye view images.

[0183] Figure 3 This is a schematic diagram of the structure of the target sensing device provided by the present invention. Figure 3 The target sensing device provided by the present invention is described. The target sensing device described below and the target sensing method described above can be referred to each other. Figure 3 As shown, the device includes: a bird's-eye view image acquisition module 301, a perception target recognition module 302, a perception target correction module 303 and a perception result output module 304.

[0184] The bird's-eye view image acquisition module 301 is used to acquire a panoramic bird's-eye view image of the target moving object.

[0185] The perception target recognition module 302 is used to identify the first perception target in the panoramic bird's-eye view image and generate a first identification frame in the panoramic bird's-eye view image for marking the identified first perception target. The first perception target includes at least one of an obstacle, a pedestrian, a ground sign and a traffic sign.

[0186] The perceived target correction module 303 is used to determine the point on the first identification frame in the panoramic bird's-eye view image that is closest to the geometric center of the target moving body as the target point corresponding to the first identification frame, and to crop the first identification frame based on the target point corresponding to the first identification frame and the first border size threshold, and determine the cropped first identification frame as the second identification frame.

[0187] The perception result output module 304 is used to obtain the position information of the first perceived target in the real world based on the second identification frame as the target perception result of the target moving body.

[0188] Specifically, the bird's-eye view image acquisition module 301, the perception target recognition module 302, the perception target correction module 303 and the perception result output module 304 are electrically connected.

[0189] The target perception device in the embodiment of the present invention identifies the first perceived target in the panoramic bird's-eye view image of the target moving body, generates a first identification frame for marking the identified first perceived target in the panoramic bird's-eye view image, determines the point on the first identification frame in the panoramic bird's-eye view image that is closest to the geometric center of the target moving body as the target point corresponding to the first identification frame, crops the first identification frame based on the target point corresponding to the first identification frame and the first border size threshold, determines the cropped first identification frame as the second identification frame, and then obtains the position information of the first perceived target in the real world based on the second identification frame as the target perception result of the target moving body. Based on the fact that the edges of the image in the panoramic bird's-eye view image will show the effect of stretching and distortion, so that the object far from the center will show the defect of deformed appearance, by cropping the first identification frame, the distorted area in the first identification frame is cropped, thereby improving the accuracy of target perception based on the panoramic bird's-eye view image, and can more simply and efficiently eliminate the influence of the distortion of the panoramic bird's-eye view image on the target perception accuracy, and can provide a more accurate data basis for unmanned driving technology and automatic parking technology.

[0190] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the target perception method, which includes: obtaining a panoramic bird's-eye view image of the target moving body; identifying a first perceived target in the panoramic bird's-eye view image, and generating a first identification frame for marking the identified first perceived target in the panoramic bird's-eye view image, wherein the first perceived target includes at least one of an obstacle, a pedestrian, a ground sign and a traffic sign; determining the point on the first identification frame in the panoramic bird's-eye view image that is closest to the geometric center of the target moving body as the target point corresponding to the first identification frame, cropping the first identification frame based on the target point corresponding to the first identification frame and the first frame size threshold, and determining the cropped first identification frame as the second identification frame; based on the second identification frame, obtaining the position information of the first perceived target in the real world as the target perception result of the target moving body.

[0191] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0192] Based on the contents of the above embodiments, a target perception system includes: the electronic device as described above and a plurality of fisheye cameras; each fisheye camera is electrically connected to the electronic device;

[0193] Each fisheye camera is arranged on the target moving body, and is used to obtain an environmental image of the surrounding environment of the target moving body, and send the environmental image to the electronic device;

[0194] The electronic device is used to obtain a panoramic bird's-eye view image of a target moving body based on the received environment image, and then obtain position information of a first perceived target in the real world based on the panoramic bird's-eye view image, wherein the type of the first perceived target includes at least one of an obstacle, a pedestrian, a ground sign and a traffic sign.

[0195] In order to solve the problem that the autonomous parking technology based on laser radar in the prior art is also subject to high production costs in industrial development, and that ultrasonic radar cannot reflect the shape characteristics of the detected object due to the single information obtained, and is therefore difficult to use for identifying the sensed object, the present invention provides a target perception method and a target perception system for implementing the above target perception method.

[0196] The target perception system provided by the present invention uses four fisheye images acquired by four fisheye cameras in front, behind, left and right of the vehicle body, and performs distortion correction, projection, blur processing and splicing operations on them to obtain a surround bird's-eye view; uses computer vision technology to perform semantic recognition on the acquired surround bird's-eye view, and detects objects such as ground signs, traffic signs, obstacles and parking spaces based on the semantic recognition results; and uses the Hungarian matching algorithm to track the detected obstacles and parking spaces through the obstacle tracking list and parking space tracking list maintained by the system. The present invention uses fisheye cameras as sensors to perceive the environment around the vehicle, and has a lower production cost than laser radar. The present invention uses a surround bird's-eye view to perceive the environment around the vehicle, detects and tracks objects such as ground signs, traffic signs, obstacles and parking spaces, and provides assistance for autonomous valet parking technology.

[0197] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the target perception method provided by the above-mentioned methods, which includes: obtaining a panoramic bird's-eye view image of a target moving body; identifying a first perceived target in the panoramic bird's-eye view image, and generating a first identification frame in the panoramic bird's-eye view image for marking the identified first perceived target, wherein the first perceived target includes at least one of an obstacle, a pedestrian, a ground sign and a traffic sign; determining the point on the first identification frame in the panoramic bird's-eye view image that is closest to the geometric center of the target moving body as the target point corresponding to the first identification frame, cropping the first identification frame based on the target point corresponding to the first identification frame and a first border size threshold, and determining the cropped first identification frame as a second identification frame; based on the second identification frame, obtaining the position information of the first perceived target in the real world as the target perception result of the target moving body.

[0198] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the target perception method provided by the above-mentioned methods, the method comprising: obtaining a panoramic bird's-eye view image of a target moving body; identifying a first perceived target in the panoramic bird's-eye view image, and generating a first identification frame in the panoramic bird's-eye view image for marking the identified first perceived target, the first perceived target including at least one of an obstacle, a pedestrian, a ground sign and a traffic sign; determining the point on the first identification frame in the panoramic bird's-eye view image that is closest to the geometric center of the target moving body as the target point corresponding to the first identification frame, cropping the first identification frame based on the target point corresponding to the first identification frame and a first border size threshold, and determining the cropped first identification frame as a second identification frame; based on the second identification frame, obtaining the position information of the first perceived target in the real world as the target perception result of the target moving body.

[0199] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0200] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A target perception method, characterized in that: include: Acquire a panoramic bird's-eye view image of the target moving object; Identify a first perceived target in the panoramic bird's-eye view image, and generate a first identification frame in the panoramic bird's-eye view image for marking the identified first perceived target, wherein the first perceived target includes at least one of an obstacle, a pedestrian, a ground sign, and a traffic sign; determining a point on the first identification frame in the panoramic bird's-eye view image that is closest to the geometric center of the target moving object as a target point corresponding to the first identification frame, cropping the first identification frame based on the target point corresponding to the first identification frame and a first border size threshold, and determining the cropped first identification frame as a second identification frame; Based on the second identification frame, acquiring position information of the first perceived target in the real world as a target perception result of the target moving object; The step of cropping the first identification frame based on the target point corresponding to the first identification frame and the first border size threshold, and determining the cropped first identification frame as the second identification frame includes: Determine the frame where the target point corresponding to the first identification frame is located as the target frame; In the case where the lengths of both sides of the target point on the target border are greater than half of the size threshold corresponding to the target border in the first border size threshold, the target border is used as the boundary of the cropping area, the target point corresponding to the first identification frame is used as the midpoint of the boundary, and the cropping area whose size meets the first border size threshold is determined in the first identification frame; in the case where the lengths of either side of the target point on the target border are less than half of the size threshold corresponding to the target border in the first border size threshold, the target border is used as the boundary of the cropping area, the vertex on the target border closer to the target point corresponding to the first identification frame is used as the starting point, and the cropping area whose size meets the first border size threshold is determined in the first identification frame; The first identification frame is cropped along the boundary of the cropping area, and the cropped first identification frame is determined as the second identification frame.

2. The target perception method according to claim 1, characterized in that: After acquiring the panoramic bird's-eye view image of the target moving object, the method further includes: Identifying a second perceived target in the panoramic bird's-eye view image, and generating a third identification frame in the panoramic bird's-eye view image for marking the identified second perceived target, wherein the second perceived target includes a parking space; determining the geometric center of the third identification frame in the panoramic bird's-eye view image as the target point corresponding to the third identification frame, cropping the third identification frame based on the target point corresponding to the third identification frame and a second border size threshold, and determining the cropped third identification frame as a fourth identification frame; Based on the fourth identification frame, the position information of the second perceived target in the real world is obtained as the target perception result of the target moving body.

3. The target perception method according to claim 1, characterized in that: After acquiring the position information of the first perception target in the real world, the method further includes: In the case where there is a first perception target tracking list, matching the second identification frame with the existing historical identification frames in the first perception target tracking list based on the Hungarian matching algorithm; In the case where the second identification frame successfully matches the historical identification frame, based on the position information of the first perceived target in the real world, the identification information of the first perceived target and the time of obtaining the panoramic bird's-eye view image, the tracking information of the historical identification frame that successfully matches the second identification frame in the first perceived target tracking list is updated; in the case where the second identification frame does not successfully match the historical identification frame, the position information of the first perceived target in the real world, the identification information of the first perceived target and the time of obtaining the panoramic bird's-eye view image are determined as the tracking information of the second identification frame, and the second identification frame and the tracking information of the second identification frame are added to the first perceived target tracking list; for the historical identification frame in the first perceived target tracking list that does not successfully match the second identification frame, when the duration of adding the historical identification frame that does not successfully match the second identification frame to the first perceived target tracking list exceeds a first duration threshold, the historical identification frame that does not successfully match the second identification frame is removed from the first perceived target tracking list.

4. The target perception method according to claim 2, characterized in that: The step of cropping the third identification frame based on the target point corresponding to the third identification frame and the second frame size threshold comprises: Taking the target point corresponding to the third identification frame as the geometric center, determining a cropping area within the third identification frame whose size meets the second frame size threshold; The third identification frame is cropped along the boundary of the cropping area, and the cropped third identification frame is determined as a fourth identification frame.

5. The target perception method according to claim 2, characterized in that: After acquiring the position information of the second perceived target in the real world based on the fourth identification frame, the method further includes: In the case where there is a second perception target tracking list, matching the fourth identification frame with the existing historical identification frames in the second perception target tracking list based on the Hungarian matching algorithm; In the case where the fourth identification frame successfully matches the historical identification frame, based on the position information of the second perceived target in the real world, the identification information of the second perceived target and the moment of acquiring the panoramic bird's-eye view image, the tracking information of the historical identification frame that successfully matches the fourth identification frame in the second perceived target tracking list is updated; in the case where the fourth identification frame fails to match the historical identification frame, the position information of the second perceived target in the real world, the identification information of the second perceived target and the moment of acquiring the panoramic bird's-eye view image are determined as the tracking information of the fourth identification frame, and the fourth identification frame and the tracking information of the fourth identification frame are added to the second perceived target tracking list; for the historical identification frame in the second perceived target tracking list that fails to match the fourth identification frame successfully, when the duration of adding the historical identification frame that fails to match the fourth identification frame successfully to the second perceived target tracking list exceeds a second duration threshold, the historical identification frame that fails to match the fourth identification frame successfully is removed from the second perceived target tracking list.

6. The target perception method according to any one of claims 1 to 5, characterized in that: The step of identifying a first perceived target in the panoramic bird's-eye view image and generating a first identification frame in the panoramic bird's-eye view image for marking the identified first perceived target includes: In the case where the first perception target includes an obstacle, the panoramic surround bird's-eye view image is input into an obstacle recognition model to obtain an obstacle mask image of the panoramic surround bird's-eye view image output by the obstacle recognition model; in the case where the first perception target includes a pedestrian, the panoramic surround bird's-eye view image is input into a pedestrian recognition model to obtain a pedestrian mask image of the panoramic surround bird's-eye view image output by the obstacle recognition model; in the case where the first perception target includes a ground sign, the panoramic surround bird's-eye view image is input into a ground sign recognition model to obtain a ground sign mask image of the panoramic surround bird's-eye view image output by the obstacle recognition model; in the case where the first perception target includes a traffic sign, the panoramic surround bird's-eye view image is input into a traffic sign recognition model to obtain a traffic sign mask image of the panoramic surround bird's-eye view image output by the obstacle recognition model; Performing contour detection on a highlight area not covered by a black layer in a mask image of the panoramic bird's-eye view image to determine a contour of the first perceived target, wherein the mask image of the panoramic bird's-eye view image includes at least one of an obstacle mask image of the panoramic bird's-eye view image, a pedestrian mask image of the panoramic bird's-eye view image, a ground sign mask image of the panoramic bird's-eye view image, and a traffic sign mask image of the panoramic bird's-eye view image; Based on the contour of the first perceived target, removing the first perceived target whose size is smaller than a first contour perimeter threshold from the mask image; Based on the outline of the first perceived target, after generating a minimum circumscribed rectangle of the first perceived target in the mask image as the first identification frame, removing a black layer on the mask image; Among them, the obstacle recognition model is trained based on the sample panoramic bird's-eye view image and the obstacle mask image of the sample panoramic bird's-eye view image; the pedestrian image recognition model is trained based on the sample panoramic bird's-eye view image and the pedestrian mask image of the sample panoramic bird's-eye view image; the ground sign image recognition model is trained based on the sample panoramic bird's-eye view image and the ground sign mask image of the sample panoramic bird's-eye view image; the traffic sign image recognition model is trained based on the sample panoramic bird's-eye view image and the traffic sign mask image of the sample panoramic bird's-eye view image; The obstacle mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the area of ​​the sample panoramic bird's-eye view image except the area where the obstacle is located; the pedestrian mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the area of ​​the sample panoramic bird's-eye view image except the area where the pedestrian is located; the ground sign mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the area of ​​the sample panoramic bird's-eye view image except the area where the ground sign is located; the traffic sign mask image of the sample panoramic bird's-eye view image is the sample panoramic bird's-eye view image covered with a black layer, and the black layer is used to cover the area of ​​the sample panoramic bird's-eye view image except the area where the traffic sign is located.

7. A target sensing device, characterized in that: include: A bird's-eye view image acquisition module is used to acquire a panoramic bird's-eye view image of the target moving object; a perception target recognition module, configured to recognize a first perception target in the panoramic bird's-eye view image, and generate a first identification frame in the panoramic bird's-eye view image for marking the recognized first perception target, wherein the first perception target includes at least one of an obstacle, a pedestrian, a ground sign, and a traffic sign; a perception target correction module, configured to determine the point on the first identification frame in the panoramic bird's-eye view image that is closest to the geometric center of the target moving body as a target point corresponding to the first identification frame, crop the first identification frame based on the target point corresponding to the first identification frame and a first border size threshold, and determine the cropped first identification frame as a second identification frame; A perception result output module, used for acquiring the position information of the first perception target in the real world based on the second identification frame as the target perception result of the target moving body; The perceived target correction module crops the first identification frame based on the target point corresponding to the first identification frame and a first border size threshold, and determines the cropped first identification frame as a second identification frame, including: Determine the frame where the target point corresponding to the first identification frame is located as the target frame; In the case where the lengths of both sides of the target point on the target border are greater than half of the size threshold corresponding to the target border in the first border size threshold, the target border is used as the boundary of the cropping area, the target point corresponding to the first identification frame is used as the midpoint of the boundary, and the cropping area whose size meets the first border size threshold is determined in the first identification frame; in the case where the lengths of either side of the target point on the target border are less than half of the size threshold corresponding to the target border in the first border size threshold, the target border is used as the boundary of the cropping area, the vertex on the target border closer to the target point corresponding to the first identification frame is used as the starting point, and the cropping area whose size meets the first border size threshold is determined in the first identification frame; The first identification frame is cropped along the boundary of the cropping area, and the cropped first identification frame is determined as the second identification frame.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the target perception method as described in any one of claims 1 to 6.

9. A target perception system, characterized in that: include: The electronic device as claimed in claim 8 and a plurality of fisheye cameras; each of the fisheye cameras is electrically connected to the electronic device; Each of the fisheye cameras is arranged on the target moving body, and is used to obtain an environmental image of the surrounding environment of the target moving body, and send the environmental image to the electronic device; The electronic device is used to obtain a panoramic bird's-eye view image of the target mobile body based on the received environmental image, and then obtain position information of a first perceived target in the real world based on the panoramic bird's-eye view image, wherein the type of the first perceived target includes at least one of an obstacle, a pedestrian, a ground sign and a traffic sign.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the target perception method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Multi-source fused aerial view sensing target detection method, device and equipment and medium

    CN117315424A

  • Parking space rapid detection method and system, electronic equipment and storage medium

    CN117456510A