Obstacle classification method, intelligent device and computer readable storage medium
Through intelligent devices, image acquisition results and visual object detection results from multiple camera perspectives are obtained, and obstacle classification is used to use 3D target frames and deep learning models to solve the problem of inaccurate obstacle distinction in the prior art, achieving high-precision obstacle classification and improving driving safety.
Patent Information
- Application Number
- CN202510173957.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to achieve accurate and effective distinction between obstacles, resulting in insufficient interaction between drivers and assisted driving systems with other vehicles and people on social roads, affecting safety.
The image acquisition results and visual object detection results from multiple camera perspectives are obtained through intelligent devices, and the target images from the camera perspectives are obtained using the 3D target frame and image acquisition results, and obstacle classification is performed based on the deep learning model.
The accurate classification of obstacles is achieved, driving safety is improved, the repeated identification problems of multiple perspectives and multiple cameras are avoided, and the accuracy and efficiency of obstacle classification are improved.
Smart Images

Figure CN120107930A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and in particular to an obstacle classification method, an intelligent device, and a computer-readable storage medium. Background Art
[0002] The popularity of advanced driver assistance functions in vehicles has led to more and more extensive interactions between intelligent vehicles and other vehicles and people on social roads. This requires drivers and driver assistance systems to effectively distinguish different obstacles on social roads and respond in a targeted manner, such as taking driving actions such as avoidance and braking, so as to effectively ensure the safety of drivers and surrounding vehicles and people. How to effectively distinguish obstacles more accurately and efficiently is a problem that needs to be solved in this field.
[0003] Accordingly, a new obstacle classification scheme is needed in this field to solve the above problems. Summary of the invention
[0004] In order to overcome the above-mentioned defects, the present application is proposed to solve or at least partially solve the technical problem of how to achieve accurate and effective distinction of obstacles.
[0005] In a first aspect, a method for classifying obstacles is provided, the method being applied to a smart device, the method comprising:
[0006] Obtain image acquisition results and visual target detection results of the smart device under multiple camera perspectives of the surrounding environment; wherein the visual target detection results include 3D target frames of environmental targets in the surrounding environment of the smart device;
[0007] Acquire a target image of the environmental target from a camera perspective according to the 3D target frame of the environmental target and the image acquisition result;
[0008] Obstacle classification results of the environmental target are obtained according to the target image.
[0009] In a technical solution of the above obstacle classification method, acquiring a target image of the environmental target under a camera perspective according to the 3D target frame of the environmental target and the image acquisition result includes:
[0010] Projecting the 3D target frame of the environmental target onto the image acquisition results under multiple camera viewing angles to obtain the 2D projection target frame of the environmental target under multiple camera viewing angles;
[0011] According to the 2D projection target frames of the environmental target under multiple camera viewing angles, a target image of the environmental target under one camera viewing angle is acquired.
[0012] In a technical solution of the above obstacle classification method, acquiring a target image of the environmental target under a camera perspective according to the 2D projected target frames under multiple camera perspectives of the environmental target includes:
[0013] For each environmental target to be classified, determining whether there is an image acquisition result of the environmental target under the perspective of the surround camera;
[0014] According to the judgment result and the 2D projection target frames of the environmental target under multiple camera viewing angles, a target image of the environmental target under one camera viewing angle is acquired.
[0015] In a technical solution of the above obstacle classification method, obtaining a target image of the environmental target under a camera perspective according to the judgment result and the 2D projected target frames of the environmental target under multiple camera perspectives includes:
[0016] If there is no image acquisition result of the environmental target under the perspective of the surround camera, a target image of the environmental target under the perspective of a camera is obtained according to the target cutoff rate and area of the 2D projection target frame;
[0017] If the environmental target has an image acquisition result under the perspective of a surround camera, a target image of the environmental target under the perspective of a camera is obtained according to the target cutoff rate of the 2D projection target frame under the perspective of the surround camera and other camera perspectives.
[0018] In a technical solution of the above obstacle classification method, obtaining a target image of the environmental target under a camera perspective according to the target cutoff rate and area of the 2D projected target frame includes:
[0019] Determine whether there is a 2D projection target frame whose target truncation rate of the environmental target is less than a preset first truncation rate among the 2D projection target frames under multiple camera perspectives;
[0020] If so, select an image corresponding to the 2D projection target frame with the largest area among the 2D projection target frames whose target truncation rate is less than the first truncation rate as the target image of the environmental target;
[0021] If it does not exist, select the image corresponding to the 2D projection target frame with the largest area among the 2D projection target frames whose target truncation rate is greater than or equal to the first truncation rate and less than a preset second truncation rate as the target image of the environmental target.
[0022] In a technical solution of the above obstacle classification method, obtaining a target image of the environmental target under a camera perspective according to the target cutoff rate of the 2D projected target frame under the perspective of the surround camera and other cameras comprises:
[0023] Determine whether there is a 2D projection target frame of the environmental target with a target truncation rate less than a preset first truncation rate under other camera perspectives except the surround camera perspective;
[0024] If so, select an image corresponding to the 2D projection target frame with the largest area among the 2D projection target frames whose target truncation rate is less than the first truncation rate under the other camera viewing angles as the target image of the environmental target;
[0025] If it does not exist, an image corresponding to the 2D projected target frame with the smallest target truncation rate of the environmental target under the perspective of the surround camera is selected as the target image of the environmental target.
[0026] In a technical solution of the above obstacle classification method, obtaining the obstacle classification result of the environmental target according to the target image includes:
[0027] Acquiring target features of the environmental target according to the target image;
[0028] Obtaining obstacle classification results of the environmental target according to the target features and based on a preset deep learning model;
[0029] Among them, the deep learning model is a classification model set according to preset classification requirements.
[0030] In a technical solution of the above obstacle classification method, acquiring the target feature of the environmental target according to the target image includes:
[0031] Perform feature extraction according to the target image to obtain image features of the environmental target;
[0032] According to the 3D target frame of the environmental target and the image feature, the target feature of the environmental target is acquired.
[0033] In a technical solution of the above obstacle classification method, the deep learning model includes a plurality of target detection heads, each of which is used to predict an obstacle category in the classification requirement;
[0034] The step of obtaining the obstacle classification result of the environmental target according to the target feature and based on a preset deep learning model includes:
[0035] Inputting the target feature into the plurality of target detection heads to obtain a prediction score of each target detection head;
[0036] Obtain an obstacle classification result of the environmental target according to the prediction score.
[0037] In a technical solution of the above obstacle classification method, the method further includes:
[0038] The environmental target is visualized according to the obstacle classification result and the 3D target frame of the environmental target.
[0039] In a technical solution of the above obstacle classification method, the obstacle classification result includes special obstacles;
[0040] The visual display of the environmental target according to the obstacle classification result and the 3D target frame of the environmental target includes:
[0041] If the obstacle classification result of the environmental target is a special obstacle, the environmental target is distinguished from other environmental targets and displayed visually.
[0042] In a second aspect, a smart device is provided, comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein a computer program is stored in the memory, and when the computer program is executed by the at least one processor, the method described in any one of the technical solutions of the above-mentioned obstacle classification method is implemented.
[0043] In a third aspect, a computer-readable storage medium is provided, wherein a plurality of program codes are stored therein, wherein the program codes are suitable for being loaded and run by a processor to execute the method described in any one of the technical solutions of the above-mentioned obstacle classification method.
[0044] Solution 1. An obstacle classification method, characterized in that the method is applied to a smart device, and the method comprises:
[0045] Obtain image acquisition results and visual target detection results of the smart device under multiple camera perspectives of the surrounding environment; wherein the visual target detection results include 3D target frames of environmental targets in the surrounding environment of the smart device;
[0046] Acquire a target image of the environmental target from a camera perspective according to the 3D target frame of the environmental target and the image acquisition result;
[0047] Obstacle classification results of the environmental target are obtained according to the target image.
[0048] Solution 2. The obstacle classification method according to Solution 1 is characterized in that:
[0049] The step of acquiring a target image of the environmental target under a camera perspective according to the 3D target frame of the environmental target and the image acquisition result includes:
[0050] Projecting the 3D target frame of the environmental target onto the image acquisition results under multiple camera viewing angles to obtain the 2D projection target frame of the environmental target under multiple camera viewing angles;
[0051] According to the 2D projection target frames of the environmental target under multiple camera viewing angles, a target image of the environmental target under one camera viewing angle is acquired.
[0052] Solution 3. The obstacle classification method according to Solution 2 is characterized in that:
[0053] The step of acquiring a target image of the environmental target under a camera perspective according to the 2D projected target frames under multiple camera perspectives of the environmental target comprises:
[0054] For each environmental target to be classified, determining whether there is an image acquisition result of the environmental target under the perspective of the surround camera;
[0055] According to the judgment result and the 2D projection target frames of the environmental target under multiple camera viewing angles, a target image of the environmental target under one camera viewing angle is acquired.
[0056] Solution 4. The obstacle classification method according to Solution 3 is characterized in that:
[0057] The step of acquiring a target image of the environmental target under a camera perspective according to the judgment result and the 2D projection target frames of the environmental target under multiple camera perspectives includes:
[0058] If there is no image acquisition result of the environmental target under the perspective of the surround camera, a target image of the environmental target under the perspective of a camera is obtained according to the target cutoff rate and area of the 2D projection target frame;
[0059] If the environmental target has an image acquisition result under the perspective of a surround camera, a target image of the environmental target under the perspective of a camera is obtained according to the target cutoff rate of the 2D projection target frame under the perspective of the surround camera and other camera perspectives.
[0060] Solution 5. The obstacle classification method according to Solution 4 is characterized in that:
[0061] The step of acquiring a target image of the environmental target under a camera viewing angle according to the target cutoff rate and area of the 2D projected target frame includes:
[0062] Determine whether there is a 2D projection target frame whose target truncation rate of the environmental target is less than a preset first truncation rate among the 2D projection target frames under multiple camera perspectives;
[0063] If so, select an image corresponding to the 2D projection target frame with the largest area among the 2D projection target frames whose target truncation rate is less than the first truncation rate as the target image of the environmental target;
[0064] If it does not exist, select the image corresponding to the 2D projection target frame with the largest area among the 2D projection target frames whose target truncation rate is greater than or equal to the first truncation rate and less than a preset second truncation rate as the target image of the environmental target.
[0065] Solution 6. The obstacle classification method according to Solution 4 is characterized in that:
[0066] The step of acquiring a target image of the environmental target under a camera perspective according to the target cutoff rate of the 2D projected target frame under the perspective of the surround camera and other cameras comprises:
[0067] Determine whether there is a 2D projection target frame of the environmental target with a target truncation rate less than a preset first truncation rate under other camera perspectives except the surround camera perspective;
[0068] If so, select an image corresponding to the 2D projection target frame with the largest area among the 2D projection target frames whose target truncation rate is less than the first truncation rate under the other camera viewing angles as the target image of the environmental target;
[0069] If it does not exist, an image corresponding to the 2D projected target frame with the smallest target truncation rate of the environmental target under the perspective of the surround camera is selected as the target image of the environmental target.
[0070] Solution 7. The obstacle classification method according to Solution 1 is characterized in that:
[0071] The step of obtaining the obstacle classification result of the environmental target according to the target image includes:
[0072] Acquiring target features of the environmental target according to the target image;
[0073] Obtaining obstacle classification results of the environmental target according to the target features and based on a preset deep learning model;
[0074] Among them, the deep learning model is a classification model set according to preset classification requirements.
[0075] Solution 8. The obstacle classification method according to Solution 7 is characterized in that:
[0076] The step of acquiring the target feature of the environmental target according to the target image includes:
[0077] Perform feature extraction according to the target image to obtain image features of the environmental target;
[0078] According to the 3D target frame of the environmental target and the image feature, the target feature of the environmental target is acquired.
[0079] Solution 9. The obstacle classification method according to Solution 7 is characterized in that the deep learning model includes a plurality of target detection heads, each of which is used to predict an obstacle category in the classification requirement;
[0080] The step of obtaining the obstacle classification result of the environmental target according to the target feature and based on a preset deep learning model includes:
[0081] Inputting the target feature into the plurality of target detection heads to obtain a prediction score of each target detection head;
[0082] Obtain an obstacle classification result of the environmental target according to the prediction score.
[0083] Solution 10. The obstacle classification method according to Solution 1, characterized in that the method further comprises:
[0084] The environmental target is visualized according to the obstacle classification result and the 3D target frame of the environmental target.
[0085] Solution 11. The obstacle classification method according to Solution 10, characterized in that the obstacle classification result includes special obstacles;
[0086] The visual display of the environmental target according to the obstacle classification result and the 3D target frame of the environmental target includes:
[0087] If the obstacle classification result of the environmental target is a special obstacle, the environmental target is distinguished from other environmental targets and displayed visually.
[0088] Solution 12. A smart device, comprising:
[0089] at least one processor;
[0090] and, a memory communicatively coupled to the at least one processor;
[0091] Wherein, a computer program is stored in the memory, and when the computer program is executed by the at least one processor, the obstacle classification method described in any one of Schemes 1 to 11 is implemented.
[0092] Solution 13. A computer-readable storage medium storing a plurality of program codes, characterized in that the program codes are suitable for being loaded and run by a processor to execute the obstacle classification method described in any one of Solutions 1 to 11.
[0093] The above one or more technical solutions of the present application have at least one or more of the following beneficial effects:
[0094] In implementing the technical solution of the obstacle classification method provided by the present application, the present application obtains the image acquisition results and visual target detection results of the surrounding environment under multiple camera perspectives of the intelligent device, obtains the target image of the environmental target under a camera perspective according to the 3D target frame and image acquisition results of the environmental target in the visual target detection results, and obtains the obstacle classification result of the environmental target according to the target image. Through the above configuration method, the present application can classify obstacles for environmental targets in combination with the visual target detection results of the intelligent device, and can achieve different obstacle classification processes for different obstacle classification requirements under the premise of effectively alleviating the computing power bottleneck, which is conducive to improving the expansion capability of the obstacle classification process. At the same time, based on the 3D target frame in the visual target detection result, a target image under the perspective of one camera is selected from the image acquisition results under multiple camera perspectives, which can effectively avoid the problem of repeated recognition of environmental targets from multiple perspectives and multiple cameras, and can effectively improve the accuracy and efficiency of the obstacle classification process for environmental targets.
[0095] Furthermore, the present application can realize the obstacle classification results and 3D target frames combined with the environmental targets, realize the visualization display of the environmental targets, and can effectively improve the visualization display capability of the driving safety of smart devices.
[0096] Furthermore, the present application can perform differentiated visual displays for special obstacles, enabling the driver and assisted driving system of the smart device to respond to special obstacles with higher priority, and implement coordinated driving behaviors such as avoidance and braking, which can effectively ensure the safety of the driver and surrounding vehicles and people, and effectively improve the assisted driving capabilities of smart devices that comply with social ethics and special safety requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0097] The disclosure of the present application will become easier to understand with reference to the accompanying drawings. It is easy for those skilled in the art to understand that these drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present application. Among them:
[0098] Figure 1 is a schematic flow chart of the main steps of an obstacle classification method according to an embodiment of the present application;
[0099] Figure 2It is a bird's-eye view of obstacles from different camera perspectives of a vehicle;
[0100] Figure 3 is a schematic diagram of a process for determining distorted pixels of a surround view camera according to an implementation of an embodiment of the present application;
[0101] Figure 4 is a schematic diagram of a model structure for obtaining obstacle classification results according to a target image of an environmental target according to an implementation of an embodiment of the present application;
[0102] Figure 5 It is a schematic diagram of the main steps of visually displaying environmental targets according to an implementation of an embodiment of the present application;
[0103] Figure 6 is a schematic diagram of a visual display example according to an example of an embodiment of the present application;
[0104] Figure 7 It is a schematic diagram of the main structure of a smart device according to an embodiment of the present application.
[0105] Reference numerals:
[0106] 11: memory; 12: processor. DETAILED DESCRIPTION
[0107] Some embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the protection scope of the present application.
[0108] In the description of this application, "module" and "processor" may include hardware, software or a combination of the two. A module may include hardware circuits, various suitable sensors, communication ports, memories, and may also include software parts, such as program codes, or a combination of software and hardware. The term "A and / or B" means all possible combinations of A and B, such as just A, just B, or A and B. The term "at least one A or B" or "at least one of A and B" has a similar meaning to "A and / or B" and may include just A, just B, or A and B. The singular terms "a" and "the" may also include plural forms.
[0109] The relevant user personal information that may be involved in the various embodiments of this application is strictly in accordance with the requirements of laws and regulations, following the principles of legality, legitimacy and necessity, based on the reasonable purposes of business scenarios, to process the personal information that users actively provide during the use of products / services or generated due to the use of products / services, as well as the personal information obtained with the user's authorization.
[0110] The user personal information processed by this application will vary depending on the specific product / service scenario, and shall be based on the specific scenario in which the user uses the product / service, and may involve the user's account information, device information, driving information, vehicle information or other related information. This application will treat the user's personal information and its processing with a high degree of diligence.
[0111] This application attaches great importance to the security of user personal information and has taken reasonable and feasible security protection measures that meet industry standards to protect user information and prevent personal information from being accessed, disclosed, used, modified, damaged or lost without authorization.
[0112] See attached Figure 1 , Figure 1 FIG. 1 is a flow chart of the main steps of the obstacle classification method according to an embodiment of the present application. Figure 1 As shown, the obstacle classification method in the embodiment of the present application mainly includes the following steps S101 to S103.
[0113] Step S101: Obtain image acquisition results and visual target detection results of the surrounding environment of the smart device under multiple camera perspectives; wherein the visual target detection results include 3D target frames of environmental targets in the surrounding environment of the smart device.
[0114] In this embodiment, the image acquisition results of the surrounding environment from multiple camera perspectives of the smart device and the visual target detection results obtained by the visual target detection model can be obtained. The visual target detection model uses at least one of the image acquisition results of the smart device and the point cloud data collected by the laser mine as input data, and outputs the coordinates of the 3D target box of the environmental target and the target label for identifying the environmental target as the visual target detection result.
[0115] In one embodiment, the smart device may be provided with multiple cameras, through which image acquisition results of the surrounding environment from multiple camera perspectives may be obtained. The 3D target frame of environmental targets such as vehicles, pedestrians, and bicycles may be located from multiple cameras through a visual target detection model.
[0116] In one implementation, the visual target detection model may be a BEV (Bird's Eye View) detection model.
[0117] Step S102: acquiring a target image of the environmental target under a camera perspective according to the 3D target frame of the environmental target and the image acquisition result.
[0118] In this embodiment, the 3D target frame of the environmental target and the image acquisition result can be combined to obtain a target image of the environmental target under the camera perspective.
[0119] In one implementation, according to preset classification requirements, a 3D target frame can be selected whose target label corresponds to the classification requirements, whose center point is within a given range and which complies with preset rules, and the selected 3D target frame can be projected onto the image acquisition results under multiple camera perspectives, and the image corresponding to the most easily recognizable 2D projected target frame can be selected as the target image of the environmental target under the final camera perspective.
[0120] For details, please refer to the attached Figure 2 , Figure 2 is a bird's-eye view of obstacles from different camera perspectives of a vehicle. Taking the smart device as a vehicle as an example, Figure 2 As shown in the figure, a vehicle (ego vehicle) is often equipped with multiple cameras, each of which collects images of the surrounding environment. The environmental targets in the ego vehicle's surrounding environment include fire trucks, ordinary cars, mounted police, police cars, police, ambulances, etc. Figure 2 Taking the police car in the video as an example, multiple camera perspectives can capture the image acquisition results of the police car. If obstacle classification reasoning is performed on the image acquisition results under each camera perspective, there will be a problem of repeated recognition from multiple camera perspectives. In the embodiment of the present application, based on the 3D target frame of the police car and the image acquisition results of the police car under multiple camera perspectives, a target image of the police car under one camera perspective (right front camera) can be selected for subsequent obstacle classification reasoning (model judgment), which can effectively avoid the repeated recognition and reasoning process of multiple camera perspectives and effectively improve the accuracy and efficiency of the reasoning process.
[0121] Step S103: Obtain obstacle classification results of the environmental target according to the target image.
[0122] In this embodiment, obstacle classification prediction can be performed on the environmental target according to the target image to obtain the obstacle classification result of the environmental target.
[0123] In one implementation, the obstacle classification result of the environmental target and the 3D target frame can be combined to visualize the environmental target, for example, on a vehicle screen.
[0124] In one embodiment, the obstacle classification result may include special obstacles. Special obstacles refer to obstacles with special markings. Special markings may include but are not limited to special vehicle license plates, sirens, identification indicator lights, etc. For example, environmental targets such as police cars, ambulances, fire trucks, police, and mounted police can all be special obstacles. If the obstacle classification result of an environmental target is a special obstacle, the environment can be distinguished from other environmental targets and visualized for warning of the driving environment and improving driving safety.
[0125] In one implementation, autonomous driving planning and control can be performed based on the obstacle classification results of the environmental target. For example, special obstacles such as police cars, ambulances, fire trucks, police, and mounted police can be effectively avoided, thereby achieving autonomous driving capabilities that meet social ethics and special safety requirements.
[0126] Based on the method described in steps S101 to S103 above, the embodiment of the present application obtains the image acquisition results and visual target detection results of the surrounding environment under multiple camera perspectives of the smart device, obtains the target image of the environmental target under a camera perspective according to the 3D target frame and image acquisition results of the environmental target in the visual target detection results, and obtains the obstacle classification result of the environmental target according to the target image. Through the above configuration method, the embodiment of the present application can classify obstacles for environmental targets in combination with the visual target detection results of the smart device, and can effectively alleviate the computing power bottleneck and realize different obstacle classification processes for different obstacle classification requirements, which is conducive to improving the expansion capability of the obstacle classification process. At the same time, based on the 3D target frame in the visual target detection result, a target image under a camera perspective is selected from the image acquisition results under multiple camera perspectives, which can effectively avoid the problem of repeated recognition of environmental targets from multiple perspectives and multiple cameras, and can effectively improve the accuracy and efficiency of the obstacle classification process for environmental targets.
[0127] Step S102 and step S03 are further described below.
[0128] In one implementation of the embodiment of the present application, step S102 may further include the following steps S1021 and S1022:
[0129] Step S1021: Project the 3D target frame of the environmental target onto the image acquisition results under multiple camera viewing angles to obtain a 2D projected target frame of the environmental target under multiple camera viewing angles.
[0130] In this embodiment, the 3D target frame of the environmental target that needs to be classified as an obstacle can be projected into the image acquisition results under multiple camera perspectives to obtain a 2D projection target frame of the environmental target under multiple camera perspectives.
[0131] Step S1022: acquiring a target image of the environmental target under a camera perspective according to the 2D projection target frames of the environmental target under multiple camera perspectives.
[0132] In this implementation, step S1022 may further include step S10221 and step S10222:
[0133] Step S10221: for each environmental target to be classified, determine whether there is an image acquisition result of the environmental target under the perspective of the surround camera.
[0134] Step S10222: acquiring a target image of the environmental target under a camera perspective according to the judgment result and the 2D projection target frames of the environmental target under multiple camera perspectives.
[0135] In this embodiment, a selection strategy of a 2D projection target frame may be determined based on an image acquisition result of an environmental target under the perspective of a surround view camera (SVC) to obtain a target image of the environmental target under the perspective of a camera.
[0136] Specifically:
[0137] (1) If the environmental target does not have an image acquisition result from the perspective of the surround camera, obtain a target image of the environmental target from the perspective of a camera based on the target truncation rate and area of the 2D projected target box:
[0138] That is, if the environmental target does not have an image acquisition result under the perspective of the surround camera, it can be determined whether there is a 2D projection target frame with a target truncation rate of the environmental target less than the first truncation rate in the 2D projection target frames under the perspectives of multiple cameras; if so, the image corresponding to the 2D projection target frame with the largest area in the 2D projection target frame with the target truncation rate less than the first truncation rate can be selected as the target image of the environmental target; if not, the image corresponding to the 2D projection target frame with the largest area in the 2D projection target frame with the target truncation rate greater than or equal to the first truncation rate and less than the preset second truncation rate can be selected as the target image of the environmental target. If the target truncation rates do not meet the requirements, the 2D projection target frame of the environmental target may not be selected.
[0139] The target truncation rate refers to the ratio of the truncated portion (such as the portion whose edge is clipped) in the image of the environmental target to the pixel points of the entire image. The number of truncated pixels and the total number of pixels in the image of the environmental target can be counted, and the number of truncated pixels is divided by the total number of pixels to obtain the target truncation rate. Those skilled in the art can set the values of the first truncation rate and the second truncation rate according to the needs of actual applications.
[0140] (2) If the environmental target has an image acquisition result from the perspective of the surround camera, the target image of the environmental target from the perspective of one camera is obtained according to the target truncation rate of the 2D projected target frame from the perspective of the surround camera and the perspective of other cameras:
[0141] That is, if there is an image acquisition result of an environmental target in the perspective of the surround-view camera, it can be determined whether there is a 2D projection target box with a target truncation rate less than a preset first truncation rate in other camera perspectives except the surround-view camera perspective; if so, the image corresponding to the 2D projection target box with the largest area among the 2D projection target boxes with a target truncation rate less than the first truncation rate in other camera perspectives can be selected as the target image of the environmental target; if not, the image corresponding to the 2D projection target box with the smallest target truncation rate of the environmental target in the surround-view camera perspective can be selected as the target image of the environmental target.
[0142] Specifically, reference can be made to the appendix Figure 3 , Figure 3 which is a schematic diagram of the process of determining the distorted pixels of the surround-view camera according to an embodiment of the present application. As Figure 3 shown, it can be determined whether there is a truncation in the 2D projection target box according to the coordinates of the 2D projection target box (bbox). For the upper left corner coordinates (x, y) of the bbox, if x is less than the truncation pixel x coordinate of the pinhole camera (χ < pinehole_trunc_pixel_χ), or y is less than the truncation pixel y coordinate of the pinhole camera (y < pinehole_trunc_pixel_y), it is considered that there is a truncation. Similarly, for the lower left corner coordinates (x, y) of the bbox, if χ > 1920 - pinehole_trunc_pixel_x, or y > 1080 - pinehole_trunc_pixel_y), it is considered that there is a truncation. Here, the pinhole camera is another camera set on the intelligent device except the surround-view camera.
[0143] In one embodiment, the maximum number of target boxes can be set, and the 2D projection target boxes can be retained from near to far. Those skilled in the art can set the maximum number of target boxes according to the actual application needs.
[0144] In one embodiment, the target category can be determined according to the target label of the 3D target box, and different expansion ratios can be set according to different categories to output the target image corresponding to the 2D projection target box.
[0145] In one embodiment of the embodiments of the present application, step S103 may further include the following steps S1031 and step S1032:
[0146] Step S1031: Obtain the target features of the environmental target according to the target image.
[0147] In this embodiment, step S1031 may further include step S10311 and step S10312:
[0148] Step S10311: Perform feature extraction based on the target image to obtain image features of the environmental target.
[0149] Step S10312: Acquire target features of the environmental target based on the 3D target frame and image features of the environmental target.
[0150] In this embodiment, please refer to the attached Figure 4 , Figure 4 FIG. 1 is a schematic diagram of a model structure for obtaining obstacle classification results based on a target image of an environmental target according to an implementation of an embodiment of the present application. Figure 4 As shown, for the target image of the obtained environmental target, the target image can be scaled to obtain input data with a data dimension of N×3×64×64; wherein N represents the batch size of the input data, 3 represents the number of color channels (ie, RGB), and 64×64 represents the uniform size of the target image after scaling.
[0151] Feature extraction can be performed based on the input data to obtain the image features of the environmental target. Specifically, the image can be scaled to 8 times based on five convolutional layers while performing feature extraction at the semantic level. Three residual blocks are used to obtain the image features of the environmental target. After obtaining the image features of the environmental target, the location information of the environmental target can be obtained based on the 3D target box of the environmental target. The location information is embedded (PositionEmbedding) into the image features and normalized using layernorm to obtain the target features of the environmental target, so that the target features help the deep learning model understand the spatial structure of the environmental target. Among them, the target feature dimension of the environmental target is N×64.
[0152] Step S1032: Obtain obstacle classification results of environmental targets according to target features and based on a preset deep learning model; wherein the deep learning model is a classification model set according to preset classification requirements.
[0153] In this embodiment, the deep learning model includes multiple target detection heads (Head), each target detection head is used to predict an obstacle category in the classification requirement; the target features can be input into the multiple target detection heads to obtain the prediction score of each target detection head, and the obstacle classification result of the environmental target is obtained according to the prediction score.
[0154] Specifically, Figure 4As shown in the figure, based on the classification requirements, multiple target detection heads (Head1, Head2, ..., Headn) can be set, and the multi-head attention mechanism can be used to extract the features of the environmental target at different angles. Classification prediction is performed based on the extracted features to obtain the obstacle classification results of the environmental target. The number of target detection heads represents the number of different classification tasks in the classification requirements. A score threshold can be set. If the prediction score of the target feature in the corresponding target detection head is higher than the score threshold, it can be considered that the category corresponding to the target detection head is the obstacle classification result of the environmental target.
[0155] In one embodiment, Figure 4 The deep learning model shown can be set on an independent DLA (deep learning acceleration) module of the Orin chip of the smart device, thereby optimizing the computing power used by the deep learning model.
[0156] In one embodiment, see the attached Figure 5 , Figure 5 It is a flow chart of the main steps for visually displaying environmental targets according to an implementation of an embodiment of the present application. The image acquisition results under multiple camera perspectives can be input into the BEV detection model deployed in the GPU (Graphics Processing Unit) to obtain the 3D target frame of the environmental target. The hpc (High performance computing) operator in the GPU can be used to project the 3D target frame into the image acquisition results under multiple camera perspectives to obtain a target image of the environmental target under a camera perspective. The target image of the environmental target is transmitted to the DLA module, and the obstacle category of the environmental target is predicted by the deep learning model deployed in the DLA module to obtain the obstacle classification result. The 3D target frame and obstacle classification results of the visual target detection can be combined for asynchronous transmission to realize model post-processing of the 3D target frame and obstacle classification results.
[0157] In one embodiment, see the attached Figure 6 , Figure 6 FIG. 1 is a schematic diagram of a visualization display example according to an example of an embodiment of the present application. Figure 6 As shown, the obstacle classification results of the environmental targets can be bound to the detection attributes of the 3D target box one by one. If it is a special obstacle (such as a police car, ambulance, fire truck, police, mounted police, etc.), these special obstacles can be replaced with animation models with warning functions. When the vehicle computer visualizes the obstacles in the surrounding environment, the vehicle computer screen can distinguish between ordinary obstacles (i.e., non-special obstacles) and special obstacles to warn the driving environment and improve driving safety, so as to achieve advanced assisted driving capabilities that meet social ethics and special safety requirements. Among them, Figure 6 The obstacles encircled by the dotted circle are special obstacles.
[0158] It should be pointed out that although the various steps in the above embodiments are described in a specific order, those skilled in the art can understand that in order to achieve the effect of the present application, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in other orders. These adjusted schemes are equivalent to the technical schemes described in this application, and therefore will also fall within the scope of protection of this application.
[0159] It is understood by those skilled in the art that all or part of the processes in the method for implementing the above-mentioned embodiment of the present application can also be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device, medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal and software distribution medium, etc. that can carry the computer program code.
[0160] Another aspect of the present application also provides a computer-readable storage medium.
[0161] In an embodiment of a computer-readable storage medium according to the present application, the computer-readable storage medium may be configured to store a program for executing the obstacle classification method of the above method embodiment, and the program may be loaded and run by a processor to implement the above obstacle classification method. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The computer-readable storage medium may be a storage device formed by various electronic devices, such as a disk, a hard disk, an optical disk, a flash memory, a read-only memory, a random access memory, etc. Optionally, the computer-readable storage medium in the embodiment of the present application is a non-temporary computer-readable storage medium.
[0162] Another aspect of the present application also provides a smart device.
[0163] In an embodiment of an intelligent device according to the present application, the intelligent device may include at least one processor; and a memory connected to the at least one processor in communication; wherein a computer program is stored in the memory, and when the computer program is executed by the at least one processor, the method described in any of the above embodiments is implemented. The intelligent device described in the present application may include a driving device, a smart car, a robot, and the like. Figure 7 , Figure 7 FIG. 4 exemplarily shows that the memory 11 and the processor 12 are communicatively connected via a bus.
[0164] In some embodiments of the present application, the smart device may further include at least one sensor, which is used to sense information. The sensor is connected to any type of processor mentioned in the present application. Optionally, the smart device may also include an autonomous driving system, which is used to guide the smart device to drive itself or assist driving. The processor communicates with the sensor and / or the autonomous driving system to complete the method described in any of the above embodiments. The processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. The processor has data and / or signal processing functions. The processor may be implemented in software, hardware, or a combination of the two.
[0165] So far, the technical solution of the present application has been described in conjunction with an embodiment shown in the accompanying drawings, but it is easy for those skilled in the art to understand that the protection scope of the present application is obviously not limited to these specific embodiments. Without departing from the principles of the present application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present application.
Claims
1. An obstacle classification method, characterized in that: The method is applied to a smart device, and the method comprises: Obtain image acquisition results and visual target detection results of the smart device under multiple camera perspectives of the surrounding environment; wherein the visual target detection results include 3D target frames of environmental targets in the surrounding environment of the smart device; Acquire a target image of the environmental target from a camera perspective according to the 3D target frame of the environmental target and the image acquisition result; Obstacle classification results of the environmental target are obtained according to the target image.
2. The obstacle classification method according to claim 1, characterized in that: The step of acquiring a target image of the environmental target under a camera perspective according to the 3D target frame of the environmental target and the image acquisition result includes: Projecting the 3D target frame of the environmental target onto the image acquisition results under multiple camera viewing angles to obtain the 2D projection target frame of the environmental target under multiple camera viewing angles; According to the 2D projection target frames of the environmental target under multiple camera viewing angles, a target image of the environmental target under one camera viewing angle is acquired.
3. The obstacle classification method according to claim 2, characterized in that: The step of acquiring a target image of the environmental target under a camera perspective according to the 2D projected target frames under multiple camera perspectives of the environmental target comprises: For each environmental target to be classified, determining whether there is an image acquisition result of the environmental target under the perspective of the surround camera; According to the judgment result and the 2D projection target frames of the environmental target under multiple camera viewing angles, a target image of the environmental target under one camera viewing angle is acquired.
4. The obstacle classification method according to claim 3, characterized in that: The step of acquiring a target image of the environmental target under a camera perspective according to the judgment result and the 2D projection target frames of the environmental target under multiple camera perspectives includes: If there is no image acquisition result of the environmental target under the perspective of the surround camera, a target image of the environmental target under the perspective of a camera is obtained according to the target cutoff rate and area of the 2D projection target frame; If the environmental target has an image acquisition result under the perspective of a surround camera, a target image of the environmental target under the perspective of a camera is obtained according to the target cutoff rate of the 2D projection target frame under the perspective of the surround camera and other camera perspectives.
5. The obstacle classification method according to claim 4, characterized in that: The step of acquiring a target image of the environmental target under a camera viewing angle according to the target cutoff rate and area of the 2D projected target frame includes: Determine whether there is a 2D projection target frame whose target truncation rate of the environmental target is less than a preset first truncation rate among the 2D projection target frames under multiple camera perspectives; If so, select an image corresponding to the 2D projection target frame with the largest area among the 2D projection target frames whose target truncation rate is less than the first truncation rate as the target image of the environmental target; If it does not exist, select the image corresponding to the 2D projection target frame with the largest area among the 2D projection target frames whose target truncation rate is greater than or equal to the first truncation rate and less than a preset second truncation rate as the target image of the environmental target.
6. The obstacle classification method according to claim 4, characterized in that: The step of acquiring a target image of the environmental target under a camera perspective according to the target cutoff rate of the 2D projected target frame under the perspective of the surround camera and other cameras comprises: Determine whether there is a 2D projection target frame of the environmental target with a target truncation rate less than a preset first truncation rate under other camera perspectives except the surround camera perspective; If so, select an image corresponding to the 2D projection target frame with the largest area among the 2D projection target frames whose target truncation rate is less than the first truncation rate under the other camera viewing angles as the target image of the environmental target; If it does not exist, an image corresponding to the 2D projected target frame with the smallest target truncation rate of the environmental target under the perspective of the surround camera is selected as the target image of the environmental target.
7. The obstacle classification method according to claim 1, characterized in that: The step of obtaining the obstacle classification result of the environmental target according to the target image includes: Acquiring target features of the environmental target according to the target image; Obtaining obstacle classification results of the environmental target according to the target features and based on a preset deep learning model; Among them, the deep learning model is a classification model set according to preset classification requirements.
8. The obstacle classification method according to claim 7, characterized in that: The step of acquiring the target feature of the environmental target according to the target image includes: Perform feature extraction according to the target image to obtain image features of the environmental target; According to the 3D target frame of the environmental target and the image feature, the target feature of the environmental target is acquired.
9. The obstacle classification method according to claim 7, characterized in that: The deep learning model includes a plurality of target detection heads, each of which is used to predict an obstacle category in the classification requirement; The step of obtaining the obstacle classification result of the environmental target according to the target feature and based on a preset deep learning model includes: Inputting the target feature into the plurality of target detection heads to obtain a prediction score of each target detection head; Obtain an obstacle classification result of the environmental target according to the prediction score.
10. The obstacle classification method according to claim 1, characterized in that: The method further comprises: The environmental target is visualized according to the obstacle classification result and the 3D target frame of the environmental target.