Visual detection method and system, storage medium and electronic equipment

By splicing the panoramic surrounding image through multiple vehicle cameras and using detection models to detect parking spot corners, travelable areas and target objects, the problems of inaccurate information identification and low integration in the prior art are solved, and high-precision automatic parking detection is achieved.

CN119942480APending Publication Date: 2025-05-06SHANGHAI BAOLONG AUTOMOTIVE CORP (WUHAN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411800761.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has problems in the parking assist system with inaccurate information identification and low integration, which affects the safety and accuracy of automatic parking.

Method used

Multiple on-board cameras are used to capture images and splice them into a panoramic circumferential image. The parking space corner point detection, semantic segmentation and target detection are carried out through the detection model, and the detection results of the parking space corner point, the driving area and the target object are obtained, and these results are packaged with the CAN signal data of the vehicle body into a standardized data format to transmit to the upper computer.

Benefits of technology

Accurate identification of parking spot corners, driving areas and target objects is achieved, the safety and detection accuracy of automatic parking are improved, and the error rate during parking is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942480A_ABST
    Figure CN119942480A_ABST
Patent Text Reader

Abstract

The invention provides a visual detection method and system, a storage medium and electronic equipment. The visual detection method comprises the following steps: acquiring images shot by a plurality of vehicle-mounted cameras, and acquiring a panoramic look-around image according to the images; the image and the panoramic look-around image are input into a detection model, a detection result is obtained through the detection model, and the detection model comprises a parking space corner detection model, a drivable area detection model and a target object detection model; the detection result comprises a parking space corner detection result, a driving area detection result and a target object detection result; packaging the detection result, CAN signal data of the vehicle body and other to-be-transmitted data into a standardized data format to obtain a packaged file; and transmitting the packaged file to an upper computer. The visual detection method has relatively high recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of autonomous driving technology and relates to a visual detection method, and in particular to a visual detection method, system, storage medium and electronic device. Background Art

[0002] With the rapid development of society, the number of private cars has continued to increase, making it extremely difficult to find an empty parking space in a parking lot. In this context, the Park Assist System (PAS) came into being. The functions of the Park Assist System mainly include target position recognition, path planning, and parking guidance or path tracking. Among them, target position recognition, as a key link of the Park Assist System, is responsible for accurately detecting the location of the empty parking space.

[0003] At present, the use of image features of the panoramic surround view monitor (AVM) system to realize parking space detection has become the main development direction of parking space detection technology. With the help of the powerful feature extraction capability of the dynamic convolution neural network (DCNN), the detection method based on DCNN has gradually emerged in recent years, greatly improving the accuracy of free parking space detection. However, parking space detection not only needs to identify the location of the parking space, but also includes the extraction of other key features, such as the location of the limiter, the distribution of the ground lock, the height and position of the curb around the parking space, and the spatial layout of nearby pillars and corners. This information is of great value to the path planning module, which can significantly improve the safety of automatic parking and effectively reduce the error rate during parking. When planning the parking path, it is necessary to obtain the parking space corner point, driving area and obstacle information. However, the existing technology generally has problems such as inaccurate information recognition and low integration. Summary of the invention

[0004] The embodiments of the present application provide a visual inspection method, system, storage medium and electronic device for solving the problems of inaccurate information recognition and low integration in the prior art.

[0005] In a first aspect, an embodiment of the present application provides a visual detection method for autonomous driving parking, the visual detection method comprising: acquiring images captured by multiple vehicle-mounted cameras and acquiring a panoramic surround view image based on the images; inputting the images and the panoramic surround view image into a detection model, and using the detection model to obtain detection results, wherein the detection model comprises a parking corner point detection model, a drivable area detection model, and a target object detection model, and the detection results comprise parking corner point detection results, drivable area detection results, and target object detection results; encapsulating the detection results, CAN signal data of the vehicle body, and other data to be transmitted into a standardized data format to obtain an encapsulated file; and transmitting the encapsulated file to a host computer.

[0006] In an implementation of the first aspect, using the parking corner point detection model to obtain the parking corner point detection result includes: inputting the image and the panoramic surround image into the parking corner point detection model; according to the image and the panoramic surround image, using the backbone network of the parking corner point detection model to obtain a first feature map; according to the first feature map, using the neck network of the parking corner point detection model to obtain a second feature map; and obtaining the parking corner point detection result based on the second feature map.

[0007] In an implementation of the first aspect, the backbone network of the parking corner detection model extracts features of the target layer through a hybrid encoder, and uses a spatial pyramid pooling layer to downsample to obtain spatial features, and each pixel in the spatial features is mapped to a grid unit evenly distributed on the original image plane; the neck network of the parking corner detection model includes an autoencoder module and a channel mapping module, the autoencoder module is configured with multiple self-attention heads and uses GELU as an activation function, and the channel mapping module is used to process feature maps from the backbone network or the feature pyramid network.

[0008] In an implementation of the first aspect, obtaining the parking space corner point detection result based on the second feature map includes: processing the second feature map to obtain representative features; performing preliminary predictions based on the representative features to obtain preliminary prediction results; and dynamically adjusting the preliminary prediction results to obtain the parking space corner point detection result.

[0009] In an implementation of the first aspect, performing preliminary prediction based on the representative features to obtain preliminary prediction results includes: obtaining a bounding box of feature points based on the representative features; aligning the bounding boxes by dynamically allocating pixel cells; and performing parking space corner point prediction based on the pixel cells to obtain the preliminary prediction results.

[0010] In an implementation of the first aspect, the visual detection method includes: regressing the bounding box using a point-by-point convolutional layer; expanding the regressed bounding box by m times to cover all key points, where m is a preset multiple; dividing the expanded bounding box into a dynamic number of pixel cells in the horizontal and vertical directions; and performing dynamic pixel cell encoding based on the division result.

[0011] In an implementation of the first aspect, the drivable area detection model is a semantic segmentation model constructed based on a side adapter network, and using the drivable area detection model to obtain the drivable area detection result includes: inputting the image and the panoramic surround image into the drivable area detection model; the drivable area detection model uses a semantic segmentation-based detection method to classify pixels in the image and / or the panoramic surround image, and obtains the drivable area detection result based on the classification result.

[0012] In a second aspect, an embodiment of the present application provides a visual detection system for autonomous driving parking, the visual detection system comprising: a visual data acquisition module, used to acquire images taken by multiple vehicle-mounted cameras and acquire a panoramic surround view image based on the images; a detection module, used to input the images and the panoramic surround view image into a detection model, and use the detection model to obtain a detection result, wherein the detection model comprises a parking corner point detection model, a drivable area detection model and a target object detection model, and the detection result comprises a parking corner point detection result, a drivable area detection result and a target object detection result; a data encapsulation module, used to encapsulate the detection result, CAN signal data of the vehicle body and other data to be transmitted into a standardized data format to obtain a packaged file; a data sending module, used to transmit the packaged file to a host computer.

[0013] In a third aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the visual detection method described in any one of the first aspects of the embodiments of the present application is implemented.

[0014] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising: a memory storing a computer program; and a processor communicatively connected to the memory, for executing any visual detection method described in any one of the first aspects of the embodiment of the present application when the computer program is called.

[0015] As described above, the visual inspection method, system, storage medium and electronic device provided in the embodiments of the present application have the following features:

[0016] Beneficial effects:

[0017] The embodiments of the present application can use the detection model to perform parking corner detection, semantic segmentation and target detection. The detection model has been trained with a large amount of data and can accurately identify targets such as parking spaces, drivable areas, road signs and parking space frames. It has good generalization ability and high detection accuracy.

[0018] In the embodiment of the present application, images taken by multiple vehicle-mounted cameras can be stitched into a panoramic surround view image. The vehicle-mounted cameras can, for example, be four 3-megapixel cameras, which can provide clearer image quality while controlling costs and significantly improve detection accuracy.

[0019] In the embodiment of the present application, the image quality can be improved and the processing speed can be optimized through hardware acceleration, the image can be corrected and enhanced by using ISP (Image Signal Processor), and the image processing speed can be further improved by DSP (Digital Signal Processor) to ensure the real-time performance of the system.

[0020] The visual detection method provided in the embodiment of the present application can work stably under various lighting conditions to ensure the reliability and accuracy of detection.

[0021] In the embodiment of the present application, the protobuf protocol and other protocols can be used for data encapsulation to ensure that the data format is compact and easy to parse, and the TCP protocol can be used for data transmission to ensure the stability and real-time performance of data transmission.

[0022] The visual inspection system provided in the embodiment of the present application has good compatibility and scalability, is easy to integrate with other electronic systems of the vehicle, is easy to install and maintain, and is suitable for various vehicle models. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 Shown is a schematic diagram of an overall implementation environment of an embodiment of the present application.

[0024] Figure 2 Shown is a flow chart of a visual inspection method provided by an embodiment of the present application.

[0025] Figure 3 Shown is a flow chart for obtaining parking space corner point detection results in an embodiment of the present application.

[0026] Figure 4 Shown is a flow chart for obtaining parking space corner point detection results in an embodiment of the present application.

[0027] Figure 5 Shown is a flow chart for obtaining preliminary test results in an embodiment of the present application.

[0028] Figure 6Shown is a schematic diagram of the structure of a visual inspection system provided in an embodiment of the present application.

[0029] Figure 7 Shown is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0030] Component number description

[0031] 600 Vision Inspection System

[0032] 601 Visual Data Acquisition Module

[0033] 602 Detection Module

[0034] 603 Data Encapsulation Module

[0035] 604 Data sending module

[0036] 700 Electronic equipment

[0037] 701 Processor

[0038] 702 Memory

[0039] 7021 Operating System

[0040] 7022 Applications

[0041] 703 Network Interface

[0042] 704 bus system

[0043] 705 User Interface

[0044] Steps S21 to S24

[0045] Steps S31 to S34

[0046] Steps S41 to S43

[0047] Steps S51 to S53 DETAILED DESCRIPTION

[0048] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0049] It should be noted that the illustrations provided in the following embodiments are only used to illustrate the basic concept of the present application in a schematic manner, and therefore the illustrations only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0050] In the embodiments of the present application, the words "exemplary" or "for example" represent examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0051] In the embodiments of the present application, "at least one" refers to one or more, and "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc or abc, where a, b, c can be single or multiple.

[0052] Figure 1 A schematic diagram of an overall implementation environment of an embodiment of the present application is shown. Figure 1 The implementation environment shown includes M cameras, a lower computer and a host computer, where M is a positive integer, and its value is, for example, 4. The M cameras are vehicle-mounted cameras, which are connected to the lower computer for communication and are used to collect images around the vehicle. The lower computer is, for example, a Black Sesame Intelligent A1000 board, which is used to process the image to achieve visual detection, and send the detection result to the host computer together with other relevant data. The host computer is connected to the lower computer for further processing according to the received data, for example, the data can be parsed and displayed.

[0053] In some implementations, the camera may include a normal camera, a panoramic camera, a thermal imaging camera, etc., and these cameras may have different resolutions and field of view angles. Exemplarily, the camera is, for example, a 4-megapixel camera, but the present application is not limited thereto.

[0054] In some implementations, other sensors (such as ultrasonic radar) can also be integrated to enhance detection accuracy and environmental perception capabilities.

[0055] Figure 2 The flowchart of the visual detection method for automatic driving parking provided by the embodiment of the present application is shown. The visual detection method can be applied to Figure 1 The lower computer shown in the figure. Figure 2 As shown, the visual inspection method includes the following steps S21 to S24.

[0056] S21, acquiring images captured by multiple vehicle-mounted cameras and acquiring a panoramic surround view image based on the images.

[0057] S22, input the image and the panoramic surround image into a detection model, and use the detection model to obtain a detection result, wherein the detection model includes a parking corner point detection model, a drivable area detection model and a target object detection model, and the detection result includes a parking corner point detection result, a drivable area detection result and a target object detection result.

[0058] In some implementations, the parking space corner point detection result may include the coordinates of each corner point of the parking space. The drivable area detection result may include whether each pixel in the image belongs to the feasible area. The target object detection result may include the type, location, and bounding box coordinates of the target object.

[0059] Exemplarily, the target object includes, for example, a ground obstacle, a ground sign, a wheel stopper, etc., but the present application is not limited thereto.

[0060] S23, encapsulate the test results, the CAN signal data of the vehicle body and other data to be transmitted into a standardized data format to obtain a packaged file. Among them, the CAN signal may include vehicle speed, steering wheel information, gear information, etc. Other data to be transmitted can be selected according to actual needs, for example, the current location information of the vehicle, which includes GPS data and vehicle body posture (such as tilt angle, pitch angle, etc.); image acquisition timestamp, used to synchronize image data from different cameras; camera identification information, used to distinguish data from different cameras; system status information, including processor load, memory usage, etc., for system monitoring and debugging.

[0061] S24, transferring the packaged file to the host computer.

[0062] In some implementations, the host computer and the slave computer can be connected through TCP (Transmission Control Protocol) communication to ensure the stability of the transmission channel. The slave computer can use streaming to send the encapsulated file to the host computer to ensure the continuity and real-time nature of the transmission. In addition, error detection and retransmission functions can be added to the transmission code so that the system can perform error detection during the file transmission process to ensure the integrity of the data. When a data transmission error is detected, the slave computer can automatically retransmit the data to ensure the accuracy of the data received by the host computer.

[0063] In some implementations, obtaining a panoramic surround view image based on an image captured by a camera includes: performing a first preprocessing on the image captured by the camera, the first preprocessing including, for example, white balance, noise reduction, demosaicing, and color correction; performing a second preprocessing on the image after the first preprocessing, the second preprocessing including, for example, dedistortion, scaling, cropping, and image pyramid; converting the RGBA image after the second preprocessing into a format that can be processed by a GPU of a lower computer (for example, NV12 format), and stitching the converted images to obtain a panoramic surround view image.

[0064] See also Figure 3 In some implementations, obtaining parking space corner point detection results using a parking space corner point detection model includes the following steps S31 to S34.

[0065] S31, inputting the image and the panoramic surround image into a parking space corner point detection model.

[0066] S32, obtaining a first feature map using a backbone network of a parking corner point detection model according to the image and the panoramic surround view image.

[0067] S33: According to the first feature map, a second feature map is obtained by using the neck network of the parking space corner point detection model.

[0068] S34, obtaining a parking space corner point detection result based on the second feature map.

[0069] In some implementations, the parking corner detection model can be an RTMO model. The initialization model of the backbone network of the parking corner detection model adopts the S version of YOLOX, and extracts the features of the target layer through a hybrid encoder, the target layer is, for example, the 2nd layer, the 3rd layer and the 4th layer, and downsamples using the spatial pyramid pooling (SPP) layer with kernel sizes of (5, 9, 13) to generate spatial features P4 and P5. Each pixel in these two features is mapped to a uniformly distributed grid unit on the original image plane.

[0070] Exemplarily, the neck network of the parking corner detection model can be based on HybridEncoder as the core, integrating an autoencoder module and a channel mapping module. Among them, the autoencoder module is configured with multiple self-attention heads, GELU is used as the activation function, and the dropout rate of the feedforward neural network is set to 0. Among them, the number of self-attention heads is, for example, 8. The channel mapping module adopts the ChannelMapper type and does not use an activation function. Its main function is to process the feature map from the backbone network or the Feature Pyramid Networks (FPN) to adapt to the input requirements of the subsequent network. To ensure the real-time performance of the network, the width of the neck network can be designed to be 50% of the original width.

[0071] See also Figure 4 In some implementations, obtaining the parking space corner point detection result based on the second feature map includes the following steps S41 to S43.

[0072] S41, processing the second feature map to obtain a representative feature. For example, the second feature map may be processed by convolution, activation function, and pooling to obtain the representative feature.

[0073] S42, perform preliminary prediction based on representative features to obtain preliminary prediction results. The preliminary prediction results may include, for example, classification scores, bounding box predictions, key point offsets, key point visibility, posture vectors, and target heat maps. Among them, the classification score refers to the classification score of each feature point, which is used to determine whether the feature point belongs to the target category. Bounding box prediction refers to the prediction of the bounding box coordinates of each feature point, which is used to locate the target corner point. Key point offset refers to the offset of the key point relative to the feature point, which is used to further accurately locate the key point. Key point visibility refers to the visibility score of each key point, which is used to determine whether the key point is visible.

[0074] S43, dynamically adjusting the preliminary prediction result to obtain a parking space corner point detection result.

[0075] In some solutions, coordinates are assigned by dividing the entire input image into bins (similar cells). However, since each sub-target occupies only a small part of the image, this approach results in a large waste of bins. In other solutions, the assignment is optimized by setting a bin within a predefined range around each anchor point. However, in this approach, larger instances may lose keypoints, while smaller instances are prone to significant quantization errors. For the above issues, please refer to Figure 5 In the embodiment of the present application, performing preliminary prediction based on representative features to obtain preliminary prediction results includes the following steps S51 to S53.

[0076] S51, obtaining a bounding box of feature points based on representative features.

[0077] S52, dynamically allocating bins to align the bounding boxes.

[0078] Exemplarily, the visual detection method provided in the embodiment of the present application may also include: regressing the bounding box using a point-by-point convolution layer; expanding the regressed bounding box by m times to cover all key points, where m is a preset multiple, and its value is, for example, 1.25; dividing the expanded bounding box into a dynamic number of bins in the horizontal and vertical directions; and performing dynamic bin encoding based on the division results. The purpose of dynamically allocating bins is to adjust the bins in the x and y directions according to the given bounding box center and scale.

[0079] For example, assuming that the center of the bounding box is c=(c x ,c y ) scale is s = (s x ,s y ), the expanded bounding box is evenly divided into B along the horizontal and vertical axes x and B y In the embodiment of the present application, it can be expanded and adjusted to new bins in the x and y directions, as shown in the following formula:

[0080]

[0081] Among them, for each bin b x and b y , Among them (b x ,b y ) represents the coordinates of the corresponding bin.

[0082] In dynamic bin encoding, Sine Positional Encoding (SPE) can be used to encode bins, and the encoded features are converted to the required dimensions through a fully connected layer (FC). First, the position encoding of the bins in the x and y directions is calculated. For each position p and temperature parameter T, the position encoding formula is as follows:

[0083]

[0084] Among them, SPE is the sinusoidal position encoding function. d is the dimension after encoding. Then, the bins in the x and y directions are position encoded:

[0085]

[0086] Next, the position-encoded features are transformed into the required dimensions through a fully connected layer:

[0087]

[0088] Among them, FC x and FC y represents a fully connected layer.

[0089] S53, performing parking space corner point prediction based on the dynamically allocated bin to obtain a preliminary prediction result.

[0090] In some implementations, the loss functions used in the training of the parking corner detection model include classification loss, bounding box loss, key point visibility loss, key point loss, object key point similarity loss, auxiliary loss, and overall loss.

[0091] Classification loss (L cls ) is used to calculate the difference between the predicted classification score and the true category. The specific formula is as follows:

[0092]

[0093] Where N is the number of samples; C is the number of categories; y i,c is the true label of the i-th sample belonging to the c-th class; p i,c is the probability that the model predicts that the i-th sample belongs to the i-th class.

[0094] Bounding box loss (L bbox ) is used to calculate the difference between the predicted bounding box and the true bounding box. The specific formula is as follows:

[0095]

[0096] Where N is the number of samples; b i is the predicted bounding box of the i-th positive sample; is the true bounding box of the i-th positive sample;

[0097] The SmoothL1 loss is defined as follows:

[0098]

[0099] Keypoint visibility loss (L vis ) is used to calculate the difference between the predicted key point visibility score and the actual visibility. The specific formula is as follows:

[0100]

[0101] Where N is the number of samples; K is the number of key points; v i,k is the true visibility of the kth key point of the i-th sample; q i,kis the visibility score of the kth keypoint of the i-th sample predicted by the model.

[0102] Keypoint loss (L mle ) is used to calculate the difference between the predicted key point coordinates and the true key point coordinates. The specific formula is as follows:

[0103]

[0104] Where N is the number of samples; K is the number of key points; v i,k is the true visibility of the kth key point of the ith sample; x i,k is the predicted coordinate of the kth key point of the i-th sample; is the true coordinate of the kth key point of the ith sample.

[0105] Object keypoint similarity loss (L OKS ) is used to calculate the similarity difference between the predicted key points and the true key points. The specific formula is as follows:

[0106]

[0107] Where N is the number of samples; K is the number of key points; v i,k is the true visibility of the kth key point of the ith sample; x i,k is the predicted coordinate of the kth key point of the i-th sample; is the true coordinate of the kth key point of the ith sample; σ k is the scale factor of the kth key point; d i is the scale of the target object of the i-th sample.

[0108] Auxiliary loss (L aux ) is used to calculate the difference between the predicted auxiliary bounding box and the true auxiliary bounding box. The specific formula is as follows:

[0109]

[0110] Where, N is the number of samples; is the predicted auxiliary bounding box of the i-th positive sample; is the true auxiliary bounding box of the i-th positive sample.

[0111] The final loss function (overall loss) is the weighted sum of the above losses:

[0112] L=λ cls L cls +λ bbox L bbox +λ vis L vis +λ mle Lmle +λ OKS L OKS +λ aux L aux

[0113] Where λ represents the weight factor of each loss term. For example, we can set λ cls =2,λ bbox =λ mle =5,λ vis =λ aux =1,λ OKS =10.

[0114] In the training phase, the calculated loss can be passed through back propagation and the model parameters can be updated to minimize the loss. In the testing phase, the trained parking corner detection model can be used to generate the final parking corner detection result.

[0115] After obtaining the prediction results, the prediction results can be expanded and sigmoid activated. The high-confidence prediction results can be screened out according to the confidence threshold and NMS (Non-Maximum Suppression). The DCC (Dynamic Coordinate Classification) module is used to dynamically adjust the key point coordinates according to the posture vector and the target heat map to generate the final prediction results, which include key point coordinates, visibility scores, bounding boxes, etc.

[0116] In the process of deploying the parking corner detection model, the size of the input image (for example, 640×640) can be generated according to the configuration, and then the specific size of the feature map is calculated by combining the size of the input image and the stride of the feature map, and then the corresponding grid coordinates are generated and concatenated into a smooth grid coordinate tensor. According to the size and stride of the feature map, the smooth stride tensor is calculated and saved as the final image output result.

[0117] In some implementations, the drivable area detection model is a semantic segmentation model constructed based on a side adapter network. Using the drivable area detection model to obtain the drivable area detection result may include: inputting an image and a panoramic surround image into the drivable area detection model; the drivable area detection model uses a semantic segmentation-based detection method to classify pixels in the image and / or the panoramic surround image, and obtaining the drivable area detection result based on the classification result.

[0118] Specifically, the semantic segmentation model built based on the side adapter network freezes the features of the CLIP model and combines the end-to-end pipeline to maximize the use of the frozen CLIP model. The image and the panoramic surround image are predicted in the model, and each pixel in the image is classified by the semantic segmentation method, so as to achieve fine-grained pixel-level distinction of different categories such as roads, pavements, vehicles, pedestrians, etc., and accurately identify and segment the drivable area of ​​the car.

[0119] The protection scope of the visual inspection method provided by the embodiment of the present application is not limited to the execution order of the steps listed in this embodiment. All solutions implemented by adding, reducing or replacing steps in the prior art based on the principles of the present application are included in the protection scope of the present application.

[0120] An embodiment of the present application also provides a visual inspection system, which can implement the visual inspection method described in the present application. However, the implementation device of the visual inspection method described in the present application includes but is not limited to the structure of the visual inspection system listed in the present embodiment. All structural deformations and replacements of the prior art made according to the principles of the present application are included in the protection scope of the present application.

[0121] Figure 6 is a schematic block diagram of a visual inspection system provided in an embodiment of the present application. Figure 6 As shown, the visual inspection system 600 includes a visual data acquisition module 601 , a detection module 602 , a data encapsulation module 603 and a data sending module 604 .

[0122] The visual data acquisition module 601 is used to acquire images taken by multiple vehicle-mounted cameras and obtain a panoramic surround image based on the images. The detection module 602 is used to input the image and the panoramic surround image into the detection model, and obtain the detection result using the detection model, wherein the detection model includes a parking corner point detection model, a drivable area detection model and a target object detection model, and the detection result includes a parking corner point detection result, a drivable area detection result and a target object detection result. The data encapsulation module 603 is used to encapsulate the detection result, the CAN signal data of the vehicle body and other data to be transmitted into a standardized data format to obtain an encapsulated file. The data sending module 604 is used to transmit the encapsulated file to the host computer.

[0123] It should be understood that the specific process of each module executing the above corresponding steps has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.

[0124] The present application embodiment also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by the processor to implement the visual detection method provided by the present application embodiment. It can be understood by those of ordinary skill in the art that all or part of the steps in the method for implementing the above-mentioned embodiment can be completed by a program to instruct the processor, and the program can be stored in a computer-readable storage medium, and the storage medium is a non-transitory medium, such as a random access memory, a read-only memory, a flash memory, a hard disk, a solid-state hard disk, a magnetic tape, a floppy disk, an optical disc, and any combination thereof. The above-mentioned storage medium can be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital video disc (digital video disc, DVD)), or a semiconductor medium (for example, a solid-state hard disk (solid state disk, SSD)), etc.

[0125] An embodiment of the present application also provides an electronic device. Figure 7 is a schematic block diagram of an electronic device provided in an embodiment of the present application. Figure 7 As shown, the electronic device 700 includes: at least one processor 701, a memory 702, at least one network interface 703 and a user interface 705. The various components in the electronic device 700 are coupled together via a bus system 704. It can be understood that the bus system 704 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 704 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 704 is not described in detail. Figure 7 In the specification, various buses are labeled as bus systems.

[0126] The user interface 705 may include a display, a keyboard, a mouse, a trackball, a click gun, keys, buttons, a touch pad or a touch screen.

[0127] It is understood that the memory 702 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), which is used as an external cache. By way of exemplary but not limiting explanation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM). The memory described in the embodiments of the present application is intended to include but is not limited to these and any other suitable categories of memory.

[0128] The memory 702 in the embodiment of the present application is used to store various categories of data to support the operation of the electronic device 700. Examples of these data include: any executable program for operating on the electronic device 700, such as an operating system 7021 and an application 7022; the operating system 7021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application 7022 may include various applications, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The visual detection method provided in the embodiment of the present application may be included in the application 7022.

[0129] The method disclosed in the above embodiment of the present application can be applied to the processor 701, or implemented by the processor 701. The processor 701 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor 701 or an instruction in the form of software. The above processor 701 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 701 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor 701 may be a microprocessor or any conventional processor, etc. In combination with the steps of the accessory optimization method provided in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0130] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD) to execute the aforementioned method.

[0131] The terms "component", "module", "system", etc. used in this specification are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program and / or a computer. By way of illustration, both applications running on a computing device and a computing device can be components. One or more components may reside in a process and / or an execution thread, and a component may be located on a computer and / or distributed between two or more computers. In addition, these components may be executed from various computer-readable media having various data structures stored thereon. Components may, for example, communicate through local and / or remote processes according to signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system and / or a network, such as the Internet interacting with other systems through signals).

[0132] Those of ordinary skill in the art will appreciate that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0133] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0134] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0135] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0136] In the above embodiments, the functions of each functional unit can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When loading and executing computer program instructions (programs) on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. Computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.

[0137] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk or an optical disk.

[0138] The above embodiments are merely illustrative of the principles and effects of the present application and are not intended to limit the present application. Anyone familiar with the technology may modify or change the above embodiments without violating the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed in the present application shall still be covered by the claims of the present application.

Claims

1. A visual detection method for autonomous driving parking, characterized in that: The visual detection method comprises: Acquire images captured by multiple vehicle-mounted cameras and acquire a panoramic surround view image based on the images; Inputting the image and the panoramic surround image into a detection model, and obtaining a detection result using the detection model, wherein the detection model includes a parking corner point detection model, a drivable area detection model, and a target object detection model, and the detection result includes a parking corner point detection result, a drivable area detection result, and a target object detection result; Encapsulating the test results, the CAN signal data of the vehicle body and other data to be transmitted into a standardized data format to obtain a packaged file; The packaged file is transmitted to the host computer.

2. The visual inspection method according to claim 1, characterized in that: Using the parking space corner point detection model to obtain the parking space corner point detection result includes: Inputting the image and the panoramic surround image into the parking space corner point detection model; According to the image and the panoramic surround image, a first feature map is acquired using a backbone network of the parking corner point detection model; According to the first feature map, a second feature map is obtained by using a neck network of the parking space corner point detection model; The parking space corner point detection result is obtained based on the second feature map.

3. The visual inspection method according to claim 2, characterized in that: The backbone network of the parking corner detection model extracts the features of the target layer through a hybrid encoder, and downsamples using a spatial pyramid pooling layer to obtain spatial features, where each pixel in the spatial features is mapped to a grid unit evenly distributed on the original image plane; The neck network of the parking corner point detection model includes an autoencoder module and a channel mapping module. The autoencoder module is configured with multiple self-attention heads and adopts GELU as an activation function. The channel mapping module is used to process the feature map from the backbone network or the feature pyramid network.

4. The visual inspection method according to claim 2, characterized in that: Acquiring the parking space corner point detection result based on the second feature map includes: Processing the second feature map to obtain representative features; Perform preliminary prediction based on the representative features to obtain preliminary prediction results; The preliminary prediction result is dynamically adjusted to obtain the parking space corner point detection result.

5. The visual inspection method according to claim 4, characterized in that: A preliminary prediction is made based on the representative features to obtain preliminary prediction results including: Acquire a bounding box of feature points based on the representative features; The bounding boxes are aligned by dynamically allocating pixel cells; A parking space corner point prediction is performed based on the pixel cells to obtain the preliminary prediction result.

6. The visual inspection method according to claim 4, characterized in that: The visual detection method comprises: regressing the bounding box using a point-wise convolutional layer; Expand the regressed bounding box to m times to cover all key points, where m is a preset multiple; Dividing the expanded bounding box into a dynamic number of pixel cells in the horizontal direction and the vertical direction; Dynamic pixel cell encoding is performed based on the segmentation results.

7. The visual inspection method according to claim 1, characterized in that: The drivable area detection model is a semantic segmentation model constructed based on a side adapter network, and using the drivable area detection model to obtain the drivable area detection result includes: Inputting the image and the panoramic surround image into the drivable area detection model; The drivable area detection model uses a detection method based on semantic segmentation to classify pixels in the image and / or the panoramic surround image, and obtains the drivable area detection result according to the classification result.

8. A visual detection system for autonomous parking, characterized in that: The visual inspection system comprises: A visual data acquisition module, used to acquire images taken by multiple vehicle-mounted cameras and obtain a panoramic surround view image based on the images; a detection module, configured to input the image and the panoramic surround image into a detection model, and obtain a detection result using the detection model, wherein the detection model includes a parking corner point detection model, a drivable area detection model, and a target object detection model, and the detection result includes a parking corner point detection result, a drivable area detection result, and a target object detection result; A data encapsulation module is used to encapsulate the detection results, the CAN signal data of the vehicle body and other data to be transmitted into a standardized data format to obtain a packaged file; The data sending module is used to transmit the encapsulated file to the host computer.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the visual inspection method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: The electronic device comprises: A memory storing a computer program; A processor is communicatively connected to the memory, and executes the visual inspection method according to any one of claims 1 to 7 when calling the computer program.

Citation Information

Cited By

  • Target detection method and device, equipment and medium

    CN121213897A