Region detection and visual perception model training method, electronic device, and storage medium
By performing boundary point update detection and visual perception model training on 2D environmental images of autonomous driving devices, the problem of unreliable predictions by deep learning models in unknown scenarios is solved, thereby improving obstacle detection rate and autonomous driving safety.
Patent Information
- Application Number
- CN202310512029.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-05
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-05-05
AI Technical Summary
Existing deep learning models based on visual perception are unreliable in predicting traffic scenarios and targets outside the training sample set for autonomous driving, leading to safety hazards.
By updating and detecting boundary points of drivable areas in two-dimensional environmental images collected by autonomous driving equipment, the stability and consistency of texture features of drivable areas are utilized to detect abnormal points of boundary changes, determine the target drivable area, and train the model by labeling boundary key points in the training samples of the visual perception model.
It improves the detection rate of unknown obstacle categories, enhances the driving safety of autonomous driving equipment, and avoids misjudgments caused by traditional models that have not learned or perceived unknown obstacle categories.
Smart Images

Figure CN116563821B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of automatic driving, and in particular to a region detection and visual perception model training method, an electronic device, and a computer storage medium. BACKGROUND
[0002] With the development of automatic driving technology, devices with automatic driving function (such as vehicles with full automatic driving function, or vehicles with incomplete automatic driving function, etc.) have been widely used in people's work and life. In order to achieve safe and stable automatic driving, devices with automatic driving function need to be trained based on deep learning models.
[0003] At present, the training of deep learning models based on visual perception, including image detection and image segmentation, is often data-driven, that is, it is obtained by using a limited training sample set for training the model. However, the deep learning model obtained by this training method is highly dependent on or even over-fitted to the distribution of the training sample set, which is manifested as: it can better handle traffic scenes and targets that have been learned through training, but its prediction results are unreliable for traffic scenes and targets outside the training sample set. However, in actual application, the automatic driving application scenario is highly open, and there are a large number of unpredictable and difficult-to-enumerate dangerous obstacles, a considerable part of which will be misjudged by the visual perception deep learning model as obstacles that do not affect driving, thereby causing serious hidden dangers to the safety of automatic driving. SUMMARY
[0004] Therefore, embodiments of the present application provide a region detection and visual perception model training scheme to at least partially solve the above problems.
[0005] According to a first aspect of embodiments of the present application, a region detection method is provided, comprising: performing boundary point update detection on a drivable region of a two-dimensional environment image collected by an automatic driving device to obtain a boundary point update detection result; determining a boundary change abnormal point according to the boundary point update detection result; and determining a target driving region of the automatic driving device according to the boundary change abnormal point.
[0006] According to a second aspect of the embodiments of the present application, a visual perception model training method is provided, comprising: obtaining a training sample for training a visual perception model, wherein the training sample comprises a two-dimensional environment image sample and boundary key point labeling of a drivable area in the two-dimensional environment image sample; performing boundary point detection of the drivable area on the two-dimensional environment image sample by the visual perception model to obtain corresponding sample boundary key points; and training the visual perception model according to the difference between the sample boundary key points and the boundary key point labeling, so that the trained visual perception model can perform boundary point detection and boundary point update detection on a two-dimensional environment image.
[0007] According to a third aspect of the embodiments of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface and a communication bus, the processor, the memory and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction makes the processor perform operations corresponding to the method of the first aspect or the second aspect.
[0008] According to a fourth aspect of the embodiments of the present application, a computer storage medium is provided, which stores a computer program, and the program is executed by a processor to implement the method of the first aspect or the second aspect.
[0009] According to the scheme provided by the embodiments of the present application, based on the characteristics that the texture features of the drivable area are relatively stable and consistent, boundary point update detection is performed on the real-time collected and constantly changing two-dimensional environment image, and boundary change abnormal points that may appear due to the approach of obstacles are detected from the two-dimensional environment image, so that the perception of the obstacles is realized. Therefore, whether the obstacles have been learned or perceived or not, whether the obstacles are unknown categories, the obstacles can be detected, and the target driving area that is relatively safe and drivable for the automatic driving device is determined. Therefore, the misjudgment with safety hazards caused by the deep learning model of the traditional visual perception due to the learning or perception of unknown categories of obstacles is avoided. Thus, the detection rate of the obstacles is improved, and the driving safety of the automatic driving device is improved. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the embodiments of the present application, and other drawings can also be obtained by those skilled in the art according to these drawings.
[0011] Figure 1 A schematic diagram of an exemplary system suitable for the embodiments of the present application;
[0012] Figure 2A A flow chart of steps of a region detection method according to an embodiment of the present application;
[0013] Figure 2B A process schematic diagram of a detection process in the embodiment shown in Figure 2A
[0014] Figure 2C A schematic diagram of a detection result in the example shown in Figure 2B
[0015] Figure 3A A flow chart of steps of a visual perception model training method according to an embodiment of the present application;
[0016] Figure 3B A process schematic diagram of a generation process of an extended training sample in the embodiment shown in Figure 3A
[0017] Figure 3C A process schematic diagram of another generation process of an extended training sample in the embodiment shown in Figure 3A
[0018] A structural schematic diagram of an electronic device according to an embodiment of the present application. Figure 4 DETAILED DESCRIPTION
[0019] In order to enable persons skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments of the present application. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by persons skilled in the art should belong to the scope of protection of the embodiments of the present application.
[0020] The specific implementation of the embodiments of the present application will be further described below with reference to the drawings in the embodiments of the present application.
[0021] Figure 1 An exemplary system to which the embodiments of the present application are applicable is shown. As shown in the figure, the system 100 can include a cloud server 102, a communication network 104 and / or one or more user devices 106, Figure 1 In the example, the user devices are multiple user devices. In this example, the user devices can be implemented as vehicle-side devices with automatic driving function. Figure 1
[0022] The cloud server 102 can be any appropriate device for storing information, data, programs, and / or any other suitable type of content, including but not limited to a distributed storage system device, a server cluster, a computing cloud server cluster, etc. In some embodiments, the cloud server 102 can store two-dimensional environment images collected by an autonomous driving device, and detection results of target drivable region detection on the two-dimensional environment images. In other embodiments, the cloud server 102 can perform any appropriate function. For example, in some embodiments, the cloud server 102 can be used to train a visual perception model for detecting drivable regions in two-dimensional environment images. As an optional example, in some embodiments, the cloud server 102 can detect boundary points of drivable regions in two-dimensional environment image samples by a visual perception model, obtain corresponding sample boundary key points, and train the visual perception model according to differences between the sample boundary key points and boundary key point labels corresponding to the two-dimensional environment image samples, so that the trained visual perception model can detect boundary points of drivable regions and update boundary points of drivable regions in two-dimensional environment images. As another example, in some embodiments, the cloud server 102 can deploy the trained visual perception model to a user device.
[0023] In some embodiments, the communication network 104 can be any appropriate combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the Internet, an intranet, a Wide Area Network (WAN), a Local Area Network (LAN), a wireless network, a Digital Subscriber Line (DSL) network, a frame relay network, an Asynchronous Transfer Mode (ATM) network, a Virtual Private Network (VPN), and / or any other suitable communication network. The user device 106 can connect to the communication network 104 through one or more communication links, such as the communication link 112, which can be linked to the cloud server 102 via one or more communication links, such as the communication link 114. The communication links can be any communication links suitable for communicating data among the user device 106 and the cloud server 102, such as network links, dial-up links, wireless links, hard-wired links, any other suitable communication links, or any suitable combination of such links.
[0024] The user device 106 can be any one or more user devices capable of drivable area detection on a two-dimensional environment image, such as a vehicle end device with automatic driving function, or a vehicle end console, or a vehicle data processing system, and / or any other suitable type of user device.
[0025] Based on the above system, the scheme provided by the embodiments of the present application is described below through a plurality of embodiments.
[0026] Embodiment One
[0027] Referring to Figure 2A , a step flowchart of a region detection method according to Embodiment One of the present application is shown.
[0028] The region detection method of the present embodiment includes the following steps:
[0029] Step S202: Perform boundary point update detection on the drivable area of the two-dimensional environment image collected by the automatic driving device to obtain the result of boundary point update detection.
[0030] In the process of driving, the automatic driving device will collect two-dimensional images of the environment it is in, i.e. two-dimensional environment images. In this environment, there are usually obstacles that affect traffic, such as motor or non-motor vehicles, pedestrians, and static obstacles such as road stakes, railings, temporary fences, etc. In this case, it is necessary to detect the drivable area of the automatic driving device. In the present embodiment, the drivable area means the closest drivable area to the automatic driving device. Because in the actual environment, there may be obstacles blocking within a certain range, but the road on the other side of the obstacle is still drivable, but the automatic driving device cannot cross the obstacle, so for the automatic driving device, the traffic area between the automatic driving device and the obstacle is the drivable area concerned in the present embodiment.
[0031] In addition, because the automatic driving device is dynamic in the process of driving, the two-dimensional environment images collected by the automatic driving device are also dynamically changing, therefore, it is necessary to dynamically update the boundary points of the drivable area, i.e. boundary point update detection, to obtain real-time boundary point update detection results, i.e. real-time boundary point information of the drivable area.
[0032] Although the corresponding detection result can be obtained through the boundary point update detection between adjacent single-frame two-dimensional environment images. In order to guarantee the privacy of the detection and effectively use the detection result of the historical two-dimensional environment image, in a feasible manner, the above boundary point update detection can be implemented as follows: according to the size of a preset detection window, the boundary point detection of the drivable area is performed on the multiple two-dimensional environment images in the detection window collected by the automatic driving device, to obtain the corresponding boundary key points; the multiple two-dimensional environment images in the detection window are updated, and the boundary point detection of the drivable area is performed on the updated multiple two-dimensional environment images in the detection window, and the result of the boundary point update detection is obtained based on the obtained boundary key points and the detection result. The size of the detection window can be appropriately set by a person skilled in the art according to actual needs, and the embodiments of the present application do not limit this. Exemplarily, the size of the detection window can be set to 10 frames. For example, the two-dimensional environment images existing in the current detection window are frames 1-10, and the boundary point update detection is performed on the two-dimensional environment images of the frames 1-10; further, the automatic driving device collects new two-dimensional environment images, such as frame 11, and the detection window eliminates the first frame 1 and becomes frames 2-11, and then the boundary point update detection is performed based on the two-dimensional environment images of the frames 2-11. Because there is less mutation between adjacent frames, the obtained boundary key points can be effectively used in the detection, and the changed boundary points are compared based on the detection result of the updated two-dimensional environment images. In this way, the changed boundary points can be found out, and the detection result of the historical two-dimensional environment image is fully utilized.
[0033] The update of the two-dimensional environment images in the detection window can be implemented by any appropriate manner by a person skilled in the art, such as every preset time, or every time a new two-dimensional environment image is detected, and preferably, the multiple two-dimensional environment images in the detection window can be updated according to the speed of the automatic driving device, so that the update is adapted to the speed of the automatic driving device, which can timely reflect the environmental change and avoid unnecessary waste of computing resources.
[0034] In the above process, both the boundary point update detection and the boundary point detection can be implemented based on the same method or model. The boundary point update detection can be understood as the boundary point detection of the updated two-dimensional environment image, which is implemented based on the boundary point detection of the drivable area.
[0035] In one feasible approach, boundary point detection of the drivable area includes: traversing boundary pixels of a two-dimensional environment image; and obtaining the boundary point of the drivable area closest to the autonomous driving device based on the traversal results. The pixel traversal can be performed row-by-row or column-by-column. Since the drivable area is the area close to the autonomous driving device, optionally, the traversal can be performed according to the driving direction of the autonomous driving device to determine the marker boundary pixels, which serve as the boundary points of the drivable area. For example, traversing the image in the width direction, i.e., traversing boundary pixels in row-by-row, starting from the direction of the autonomous driving device and moving away from it, and using the traversed non-road pixel closest to the autonomous driving device as the boundary point of the drivable area. Alternatively, traversing the image in the length direction, performing boundary pixel traversal in column-by-column, also starting from the direction of the autonomous driving device and moving away from it, and using the traversed non-road pixel closest to the autonomous driving device as the boundary point of the drivable area. By traversing pixels, we focus more on boundary pixels rather than obstacles and their categories. Regardless of the type of obstacle, it can be excluded from the driving area of the autonomous driving device by using boundary points. This can accurately determine the boundary of the driving area and avoid the computational burden caused by unnecessary obstacle classification and recognition.
[0036] In practical implementation, boundary point detection of drivable areas can be achieved through a drivable area keypoint regression model. This model takes a 2D image as input and outputs the boundary keypoints of the drivable area within the 2D image. The drivable area keypoint regression model can be implemented using any suitable deep learning model capable of image-based object detection, such as a convolutional neural network model based on feature pyramids. This drivable area keypoint regression model can perform multi-scale object detection on the input 2D image, especially 2D environment images, such as multi-scale drivable area boundary point detection. After feature fusion, it outputs the boundary point coordinates. To improve the model's detection speed during boundary point detection, key landmarks or change points on the boundary can be detected, i.e., boundary keypoint detection, finally outputting the point-like boundary of the drivable area. Using the drivable area keypoint regression model allows for more accurate detection.
[0037] By detecting and updating boundary points in the drivable area of a two-dimensional environment image, the corresponding detection results can be obtained, such as boundary points that have changed in the environment, and subsequent processing can be performed.
[0038] Step S204: Based on the results of the boundary point update detection, identify the boundary change anomaly points.
[0039] Based on the detection of the new and old two-dimensional environment images, such as the detection of two adjacent two-dimensional environment images, or based on the detection of multiple two-dimensional environment images in adjacent two window in the detection window, the boundary points of the changeable drivable region boundary can be determined, i.e. the boundary change abnormal points. Based on the boundary change abnormal points, the boundary of the new drivable region can be formed, and the new drivable region of the autonomous driving device is also defined.
[0040] In one possible manner, when determining the boundary change abnormal points, the boundary points close to the obstacle region of the autonomous driving device can be obtained according to the results of the boundary point update detection; and the boundary change abnormal points can be determined according to the obstacle region boundary points. In the actual driving environment, it is difficult to determine whether a certain obstacle will affect the driving of the autonomous driving device from a static perspective, because if the obstacle moves towards the autonomous driving device, it may affect the driving of the autonomous driving device; and if the obstacle moves away from or deviates from the autonomous driving device, it may not affect the driving of the autonomous driving device. Therefore, in this manner, whether the obstacle will affect the driving of the autonomous driving device and whether the obstacle region boundary point will become a boundary change abnormal point are determined according to the dynamic of the obstacle. Thus, more accurate determination of the boundary change abnormal points is achieved.
[0041] However, the manner of directly determining the changed boundary points as the boundary change abnormal points is also applicable to the scheme of the embodiments of the present application.
[0042] Step S206: determining the target driving region of the autonomous driving device according to the boundary change abnormal points.
[0043] As described above, after the boundary change abnormal points are determined, the boundary points of the drivable region can be determined according to the obtained other boundary points, and then the drivable region, i.e. the target driving region, can be determined. For example, the target driving region of the autonomous driving device can be determined according to the detected boundary key points of the drivable region and the boundary change abnormal points. On this basis, the autonomous driving device can automatically drive in the target driving region according to the preset driving plan.
[0044] But in order to further improve the accuracy of the determined target driving area and ensure the safe driving of the automatic driving device, in a feasible manner, other target driving area detection methods are also provided, which are combined with the method of the embodiment of the application for comprehensive judgment to ensure the accuracy of the target driving area determination and ensure the driving safety. In this way, image segmentation and / or target detection can also be performed on the two-dimensional environment image; according to the image segmentation result and / or the target detection result, the corresponding first drivable area and / or the second drivable area are obtained. On this basis, according to the boundary change abnormal point, the target driving area of the automatic driving device can be determined as follows: according to the boundary change abnormal point, the third drivable area is determined; according to the first drivable area and / or the second drivable area, and the third drivable area, and the confidence degree corresponding thereto, the target driving area of the automatic driving device is determined. Wherein, the specific implementation of image segmentation and target detection can be realized by any appropriate algorithm or model according to the actual needs of those skilled in the art, for example, a machine learning model such as a model based on convolutional neural network can be used, and the embodiment of the application does not limit this.
[0045] Through the embodiment, based on the characteristics that the texture features of the drivable area are relatively stable and consistent, the boundary point update detection is performed on the real-time collected and constantly changing two-dimensional environment image, the boundary change abnormal points that may appear due to the approach of obstacles are detected from the image, and the perception of the obstacles is realized. Therefore, whether the obstacles have been learned or perceived or not, whether the obstacles are unknown category obstacles, the obstacles can be detected, and the target driving area that is relatively safe and drivable for the automatic driving device is determined. Therefore, the misjudgment with safety hazards caused by the deep learning model of the traditional visual perception due to the unknown category obstacles that have not been learned or perceived is avoided. Thus, the detection rate of the obstacles is improved, and the driving safety of the automatic driving device is improved.
[0046] In the following, the above process is exemplarily explained in combination with Figure 2B and Figure 2C .
[0047] In this example, it is assumed that the automatic driving device is an automatic driving vehicle, and it is assumed that the size of the detection window is 10 frames, i.e. 10 frames of two-dimensional environment images can be accommodated in the detection window.
[0048] As Figure 2BAs shown, first, boundary point detection is performed on the current 1-10th frame of the two-dimensional environment image in the detection window by the drivable area key point regression model, also referred to as the FS (Free Space) key point regression model, to obtain a corresponding detection result, referred to as the first boundary key point of FS in this example. It should be understood by those skilled in the art that the first boundary key point does not refer to one boundary key point, but a set of boundary key points obtained after boundary point detection of the 1-10th frame of the two-dimensional environment image, which includes a plurality of boundary key points.
[0049] Further, the two-dimensional environment image in the detection window is updated according to the speed of the autonomous vehicle, assuming that the update is the 2-11th frame. Boundary point detection is then performed on the 2-11th frame of the two-dimensional environment image in the detection window by the FS key point regression model, i.e., boundary point update detection, to obtain a corresponding result to determine the boundary change abnormal points of FS, i.e., boundary change abnormal points.
[0050] Then, based on the boundary change abnormal points, the abnormal obstacle area can be determined, and the drivable area of the autonomous vehicle can be determined. At the same time, in this example, image segmentation and image-based target detection are also used to detect the drivable area to obtain a plurality of detection results. If the plurality of detection results are consistent with the result obtained by the FS key point regression model, the target driving area of the autonomous vehicle can be directly determined according to the FS key point regression model. If the plurality of detection results are inconsistent with the result obtained by the FS key point regression model, the final target driving area needs to be determined according to the confidence of each detection result. Generally, the drivable area indicated by the detection result with the highest confidence can be determined as the final target driving area.
[0051] In one example, the boundary points of the target driving area obtained are shown as dotted lines in FIG. 1. Figure 2C
[0052] Thus, by performing boundary point update detection on a plurality of consecutive two-dimensional environment images, perception of unknown category obstacles is achieved. Since the boundary of the drivable area can be a set of points in a two-dimensional environment image, such as an original two-dimensional environment image or an environment bird's eye view or an environment perspective view, it is easier to generate rich detection data compared to multi-frame image detection or tracking tasks, so that only the boundary point changes of the multi-frame two-dimensional environment images can be used to detect whether an obstacle is approaching the autonomous driving device, achieving classless obstacle perception and higher obstacle detection rate.
[0053] Embodiment Two
[0054] In this embodiment, the acquisition of the visual perception model for boundary point detection in the foregoing embodiment, i.e., the drivable area key point regression model, is described.
[0055] Referring to Figure 3A FIG. 3 shows a flowchart of steps of a method for training a visual perception model according to an embodiment of the present application.
[0056] The method for training a visual perception model according to the embodiment includes the following steps:
[0057] Step S302: Obtain training samples for training the visual perception model.
[0058] In the embodiment, the visual perception model is a drivable area key point regression model for boundary point detection of a two-dimensional environment image, and the training samples include two-dimensional environment image samples and boundary key point annotations of drivable areas in the two-dimensional environment image samples. Compared with the conventional method, in the embodiment, the boundaries of the drivable areas are annotated in the form of boundary key points, which has low annotation cost, fast data accumulation, low algorithm complexity, and easy simulation data generation.
[0059] Because the number of conventional training samples is limited and the describable environment scenes are also limited, in order to expand the description of the environment scenes and increase the number of training samples, in a feasible manner, before the step, a set of original training samples for training the visual perception model can also be obtained; a scrambling operation and / or a perspective transformation operation are performed on the training samples in the set of original training samples, and based on the operation results and the training samples in the set of original training samples, a set of extended training samples is generated. Then, obtaining the training samples for training the visual perception model can be implemented as follows: from the set of extended training samples, obtaining the training samples for training the visual perception model.
[0060] Optionally, the scrambling operation on the training samples in the set of original training samples can be implemented as follows: obtaining, by the visual perception model, associated multiple frames of two-dimensional environment image samples in the set of original training samples, performing boundary point detection of drivable areas on the multiple frames of two-dimensional environment image samples, and obtaining images of multiple frames of sample boundary key points; superimposing the images of the multiple frames of sample boundary key points to obtain a superimposed key point image; and performing boundary key point scrambling on the superimposed key point image to generate a sample image with noise. The associated multiple frames of two-dimensional environment image samples can be multiple frames of environment-associated images continuously collected.
[0061] For the visual perception model, whether it is the original two-dimensional environment image or the two-dimensional environment image represented by key points, it will be input into the model in a way acceptable to the visual perception model, such as image data vector, and after corresponding feature extraction and other operations, the corresponding feature map described by feature data is obtained. Therefore, although the noise-added sample images generated by scrambling are different in the visual field that can be perceived by the human eye, they can be perceived and learned by the visual perception model. Therefore, the noise-added sample images generated by scrambling the boundary key points can also be used as training samples to input the visual perception model for training. In this way, the training sample set is expanded, and the visual perception model obtained by training is more robust and has better robustness.
[0062] As shown in the example of FIG. 6, in one example, the same physical environment is continuously sampled to obtain multiple frames of two-dimensional environment images, which are exemplarily set to 10 frames of two-dimensional environment image samples. Since training samples are to be generated, the untrained visual perception model can be used to detect the boundary points of the 10 frames of two-dimensional environment image samples (since the accuracy of boundary point detection is not required to be too high when generating training samples), to obtain 10 frames of boundary key point images formed by the detected boundary key points. Then, the 10 frames of boundary key point images are superimposed to generate a superimposed key point image. Then, the superimposed key point image is scrambled to obtain a noise-added sample image. The scrambling process can be implemented by any appropriate scrambling method according to actual needs, such as Gaussian noise, Poisson noise, etc. Figure 3B
[0063] The visual perception model can be used to detect the boundary points of the associated multiple frames of two-dimensional environment image samples in the original training sample set to obtain multiple frames of sample boundary key point images. The multiple frames of sample boundary key point images are superimposed to obtain a superimposed key point image. The superimposed key point image is regionally spliced and / or coordinate transformed to generate a new sample image. As mentioned above, the associated multiple frames of two-dimensional environment image samples can be continuously collected environment-associated multiple frames of images.
[0064] Region stitching can be performed on the superimposed keypoint image. Alternatively, those skilled in the art can use preset rules to extract keypoints from the superimposed keypoint image and stitch them into it. These preset rules can be random rules, sequential extraction rules, etc. For example, keypoints in a region of size M*N (pixels) containing keypoints can be extracted from the superimposed keypoint image (e.g., the left side of the image) and stitched into it to obtain a new sample image.
[0065] Similarly, coordinate transformation operations can be performed on the superimposed keypoint image by those skilled in the art according to preset rules. This transformation can change the key points in the superimposed keypoint image from one viewpoint to another, for example, transforming a rear view image into a left view image; or, through coordinate transformation, moving the autonomous driving device closer to an obstacle to create a "collision" image, and so on. In this application, the specific coordinate transformation method used is not limited.
[0066] For example, such as Figure 3C As shown in (1), in one example, the same physical environment is continuously sampled to obtain multiple frames of two-dimensional environment images. For example, 10 frames of two-dimensional environment image samples are set. Since training samples are to be generated, an untrained visual perception model can be used to detect boundary points on these 10 frames of two-dimensional environment image samples (since the intention is to generate training samples, the accuracy of boundary point detection is not required to be too high) to obtain 10 frames of boundary key point images formed by the detected boundary key points. Then, the images of these 10 frames of boundary key points are overlaid to generate overlaid key point images. Subsequently, the overlaid key point images are stitched together to obtain new sample images after stitching.
[0067] For example Figure 3C As shown in (2), in one example, the same physical environment is continuously sampled to obtain multiple frames of two-dimensional environment images. For example, 10 frames of two-dimensional environment image samples are set. Since training samples are to be generated, an untrained visual perception model can be used to detect boundary points on these 10 frames of two-dimensional environment image samples (since the intention is to generate training samples, the accuracy of boundary point detection is not required to be too high) to obtain 10 frames of boundary key point images formed by the detected boundary key points. Then, the images of these 10 frames of boundary key points are superimposed to generate superimposed key point images. Subsequently, coordinate transformation is performed on the superimposed key point images so that the autonomous driving device "collides" with the obstacle to obtain new sample images.
[0068] Thus, the original training sample set is expanded by the sample images after the scrambling operation and / or the perspective transformation operation to form a new expanded training sample set. On this basis, training samples for training the visual perception model can be obtained from the expanded training sample set.
[0069] Step S304: The boundary point detection of the drivable area is performed on the two-dimensional environment image sample by the visual perception model to obtain the corresponding sample boundary key point.
[0070] As described above, the specific implementation of the visual perception model and its model structure are not limited in the embodiments of the present application, and can be implemented as a visual perception model based on a feature pyramid, or a convolutional neural network model based on an attention mechanism, etc.
[0071] After the two-dimensional environment image sample is input into the visual perception model, the corresponding sample boundary key point, i.e., the sample boundary point of the drivable area of the autonomous driving device in the two-dimensional environment image sample, is obtained through step-by-step feature extraction and boundary point prediction operations.
[0072] Step S306: The visual perception model is trained according to the difference between the sample boundary key point and the boundary key point annotation, so that the trained visual perception model can perform the boundary point detection and the boundary point update detection of the drivable area of the two-dimensional environment image.
[0073] The boundary key point annotation of the drivable area in the two-dimensional environment image sample can be used as a supervision condition. In the training process of the visual perception model, a preset loss function is used to calculate the difference, i.e., the loss value, between the sample boundary key point predicted by the model and the boundary key point annotation, and the visual perception model is trained based on the difference until a training termination condition is reached (such as reaching a preset training number or the loss value meeting a preset threshold, etc.). The loss function can be appropriately set by those skilled in the art according to actual needs, including but not limited to a cross-entropy loss function, a smooth loss function, etc.
[0074] Since the boundary point detection and the boundary point update detection are essentially the same, both are boundary point detection on the two-dimensional environment image, therefore, the trained visual perception model can perform the boundary point detection and the boundary point update detection of the drivable area of the two-dimensional environment image.
[0075] Through the embodiment, for the difficult-to-enumerate obstacles appearing in the driving environment, the texture characteristics of the FS are relatively stable and consistent, and the FS modeling is performed, and the visual perception model capable of end-to-end output of the FS boundary is obtained by regression training. Unlike image information, the FS is essentially fixed number of road boundary coordinate point information, and therefore is easy to batch model to generate rich simulation data for expanding the training sample set. The trained visual perception model can realize the perception of unknown category obstacles through FS boundary change detection of continuous multiple two-dimensional environment images, thereby improving the obstacle detection rate.
[0076] Embodiment three
[0077] Reference Figure 4 The embodiment of the application is not limited to the specific implementation of the electronic device.
[0078] As shown in Figure 4 , the electronic device can include a processor 402, a communications interface 404, a memory 406, and a communications bus 408.
[0079] Among them:
[0080] The processor 402, the communications interface 404, and the memory 406 complete mutual communication through the communications bus 408.
[0081] The communications interface 404 is configured to communicate with other electronic devices or servers.
[0082] The processor 402 is configured to execute the program 410, and specifically can execute the related steps in the above method embodiments.
[0083] Specifically, the program 410 can include program code including computer operation instructions.
[0084] The processor 402 can be a CPU, or a GPU (Graphic Processing Unit, graphic processing unit), or an ASIC (Application Specific Integrated Circuit, application specific integrated circuit), or one or more integrated circuits configured to implement one or more embodiments of the application. One or more processors included in the smart device can be the same type of processor, such as one or more CPUs; or can be different types of processors, such as one or more CPUs and one or more ASICs.
[0085] The memory 406 is configured to store a program 410. The memory 406 can include a high-speed RAM memory, and can further include a non-volatile memory such as at least one disk memory.
[0086] The program 410 can include a plurality of computer instructions, and the program 410 can specifically cause the processor 402 to perform operations corresponding to the method described in any of the foregoing method embodiments.
[0087] The specific implementation of each step in the program 410 can refer to the corresponding description in the corresponding steps and units in the foregoing method embodiments, and has corresponding beneficial effects, which will not be described here. Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the device and the module described above can refer to the corresponding process description in the foregoing method embodiments, which will not be described here.
[0088] The embodiments of the present application further provide a computer storage medium, which stores a computer program. The program is executed by a processor to implement the method described in any of the foregoing method embodiments. The computer storage medium includes but is not limited to a compact disc read-only memory (CD-ROM), a random access memory (RAM), a floppy disk, a hard disk, or a magneto-optical disk, etc.
[0089] The embodiments of the present application further provide a computer program product, which includes computer instructions. The computer instructions instruct a computing device to perform operations corresponding to any of the foregoing method embodiments.
[0090] In addition, it needs to be explained that the information related to the user (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to sample data used for training the model, data used for analysis, stored data, displayed data, user driving behavior data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of the related data need to comply with the relevant laws, regulations and standards of the country and region, and provide corresponding operation entrances for the user to choose authorization or refusal.
[0091] It needs to be pointed out that, according to the needs of implementation, each component / step described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or part of the operation of the components / steps can be combined into a new component / step, to achieve the purpose of the embodiments of the present application.
[0092] The methods according to the embodiments of the present application described above can be implemented in hardware, firmware, or software, or a combination thereof, and can be stored in a recording medium such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk, or be downloaded by a network from a remote recording medium or non-transitory machine-readable medium originally stored in a local recording medium and then stored in a local recording medium, so that the methods described herein can be processed by such software processing using a general-purpose computer, a special-purpose processor, or programmable or special-purpose hardware such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA). It can be understood that the computer, processor, microprocessor controller, or programmable hardware includes a storage component (for example, Random Access Memory (RAM), Read-Only Memory (ROM), flash memory, etc.) that can store or receive software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code for implementing the methods shown herein, the execution of the code converts the general-purpose computer into a special-purpose computer for executing the methods shown herein.
[0093] Those skilled in the art can appreciate that the units and method steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. A skilled person can use different methods to implement the described functions for a specific application, but such implementation should not be considered beyond the scope of the embodiments of the present application.
[0094] The above embodiments are only used to illustrate but not limit the embodiments of the present application, and a person of ordinary skill in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application, therefore all equivalent technical solutions belong to the scope of the embodiments of the present application, and the patent protection scope of the embodiments of the present application should be defined by the claims.
Claims
1. A region detection method, comprising: boundary point update detection of a drivable region of a two-dimensional environment image collected by an autonomous driving device, to obtain a boundary point update detection result, including: boundary point detection of a plurality of two-dimensional environment images collected by the autonomous driving device within a preset detection window, to obtain corresponding boundary key points; updating the plurality of two-dimensional environment images within the detection window, and performing boundary point detection of the updated plurality of two-dimensional environment images within the detection window, to obtain the boundary point update detection result based on the obtained boundary key points and the detection result; determining a boundary change abnormal point according to the boundary point update detection result; determining a target driving region of the autonomous driving device according to the boundary change abnormal point.
2. The method of claim 1, wherein, The updating of the plurality of two-dimensional environment images within the detection window comprises: updating the plurality of two-dimensional environment images within the detection window according to the speed of the autonomous driving device.
3. The method according to any one of claims 1-2, wherein, The boundary point update detection of the drivable region is based on boundary point detection of the drivable region; The boundary point detection of the drivable region comprises: boundary pixel point traversal of the two-dimensional environment image; obtaining the boundary point of the drivable region closest to the autonomous driving device according to the traversal result.
4. The method of claim 3, wherein, The boundary point detection of the drivable region is performed by a drivable region key point regression model; The drivable region key point regression model takes a two-dimensional image as input and outputs a boundary key point of a drivable region in the two-dimensional image.
5. The method of any one of claims 1-2, wherein, The determination of the boundary change abnormal point according to the boundary point update detection result comprises: obtaining an obstacle region boundary point gradually approaching the autonomous driving device according to the boundary point update detection result; determining the boundary change abnormal point according to the obstacle region boundary point.
6. The method of any one of claims 1-2, wherein, The determination of the target driving region of the autonomous driving device according to the boundary change abnormal point comprises: determining the target driving region of the autonomous driving device according to the detected boundary key points of the drivable region and the boundary change abnormal point.
7. The method of any one of claims 1-2, wherein: the method further comprises image segmentation and / or target detection of the two-dimensional environment image; and obtaining a first drivable region and / or a second drivable region corresponding thereto according to the image segmentation result and / or the target detection result; the determination of the target driving region of the autonomous driving device according to the boundary change abnormal point comprises: determining a third drivable region according to the boundary change abnormal point; and determining the target driving region of the autonomous driving device according to the first drivable region and / or the second drivable region, the third drivable region, and the confidence degrees corresponding thereto, respectively.
8. A visual perception model training method, comprising: obtaining training samples for training a visual perception model, wherein the training samples include two-dimensional environment image samples and boundary key point annotations of drivable regions in the two-dimensional environment image samples. The visual perception model is used for boundary point detection of the drivable area of the two-dimensional environment image sample, to obtain corresponding sample boundary key points; According to the difference between the sample boundary key points and the boundary key point labels, the visual perception model is trained, so that the trained visual perception model can perform boundary point detection and boundary point update detection on the two-dimensional environment image. The boundary point detection and boundary point update detection on the two-dimensional environment image include: performing boundary point detection on a plurality of two-dimensional environment images collected by the autonomous driving device within a preset detection window according to the size of the detection window, to obtain corresponding boundary key points; updating the plurality of two-dimensional environment images within the detection window, and performing boundary point detection on the updated plurality of two-dimensional environment images within the detection window, to obtain the result of boundary point update detection based on the obtained boundary key points and the detection result.
9. The method of claim 8, wherein, Before the training sample used for training the visual perception model is obtained, the method further comprises: obtaining an original training sample set used for training the visual perception model; performing a scrambling operation and / or a perspective transformation operation on the training samples in the original training sample set, and generating an extended training sample set based on the operation result and the training samples in the original training sample set; The training sample used for training the visual perception model is obtained from the extended training sample set.
10. The method of claim 9, wherein, The scrambling operation on the training samples in the original training sample set includes: obtaining the image of the plurality of sample boundary key points obtained by the visual perception model performing boundary point detection on the associated plurality of two-dimensional environment image samples in the original training sample set; superimposing the images of the plurality of sample boundary key points to obtain a superimposed key point image; performing boundary key point scrambling on the superimposed key point image to generate a noisy sample image.
11. The method of claim 9, wherein, The perspective transformation operation on the training samples in the original training sample set includes: obtaining the image of the plurality of sample boundary key points obtained by the visual perception model performing boundary point detection on the associated plurality of two-dimensional environment image samples in the original training sample set; superimposing the images of the plurality of sample boundary key points to obtain a superimposed key point image; performing region splicing and / or coordinate transformation on the superimposed key point image to generate a new sample image.
12. An electronic device comprising: The processor, the memory, the communication interface, and the communication bus complete communication among each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the corresponding operations of the method of any one of claims 1-7 or claims 8-11.
13. A computer storage medium having stored thereon a computer program which, when executed by a processor, implements the method of any one of claims 1-7, or claims 8-11.
Citation Information
Patent Citations
Detection method, virtual radar device, electronic equipment and storage medium
CN111552289A
Automatic driving method and device and storage medium
CN115320637A