Road obstacle recognition method and system and vehicle

By introducing attention mechanism and reward mechanism into the semantic segmentation model, combined with multi-scale feature extraction, the problem of unknown anomaly target recognition is solved, and the environmental adaptability and safety of obstacle recognition are improved.

CN120047920APending Publication Date: 2025-05-27CHERY AUTOMOBILE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510115770.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing semantic segmentation model is difficult to accurately identify when facing unknown abnormal targets, resulting in safety hazards and difficult to effectively identify all possible obstacle types, and has poor environmental adaptability.

Method used

A deep learning-based road obstacle recognition model is adopted, combining attention mechanisms and reward mechanisms, and through multi-scale feature extraction and attention map fusion, attention is focused on the uncertain areas of the road image to be identified.

Benefits of technology

It effectively improves the model's ability to identify unknown types of obstacles, enhances environmental adaptability, and reduces safety hazards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047920A_ABST
    Figure CN120047920A_ABST
Patent Text Reader

Abstract

The invention provides a road obstacle recognition method and system, and belongs to the technical field of vehicle intelligent auxiliary control, and the method comprises the steps: collecting a road image in a vehicle advancing direction in real time; based on the obtained road image, obtaining an obstacle recognition result through a pre-trained road obstacle recognition model based on deep learning; wherein in the training process of the road obstacle recognition model, an attention map obtained based on an attention mechanism is adopted, obtained prediction output is combined, and a correct recognition result and the prediction output are combined through a preset reward mechanism; therefore, the road obstacle recognition model focuses on the uncertain area of the to-be-recognized road image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vehicle intelligent auxiliary control, and particularly relates to a road obstacle recognition method, system and vehicle. Background Art

[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.

[0003] With the development of deep convolutional neural networks, semantic segmentation technology has been increasingly widely applied in the field of autonomous driving environment perception, especially in obstacle detection.

[0004] The inventors found that current research on semantic segmentation mainly focuses on improving the accuracy of segmentation performance. Semantic segmentation models can often only classify pre-defined categories in the dataset, which means that to achieve accurate and comprehensive obstacle detection, all possible objects need to be included in the training set. However, since the real world is open and unexpected things may happen at any time, in the face of unknown abnormal targets (for example, a small animal suddenly jumps onto the road), existing semantic segmentation models may not be able to accurately identify them, but classify them into the pre-defined categories in the training set. This processing method will pose serious safety hazards and greatly limit the application of deep learning algorithms in autonomous driving. At the same time, it is essentially infeasible to collect all types of obstacles that may appear on the road to construct a training set. Therefore, traditional solutions are difficult to effectively identify unknown type obstacles and have poor environmental adaptability. Summary of the Invention

[0005] The present invention provides a road obstacle recognition method, system and vehicle to solve the problem that traditional solutions are difficult to effectively identify unknown type obstacles and have poor environmental adaptability.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a road obstacle recognition method, including:

[0008] Real-time collecting road images in the traveling direction of the vehicle;

[0009] Based on the obtained road image, through a pre-trained deep learning-based road obstacle recognition model, an obstacle recognition result is obtained; the road obstacle recognition model specifically performs the following processing procedures: taking the road image as an input, based on a preset encoder based on a multi-scale feature extraction network, multi-scale fusion features are obtained; taking the preset scale features as the input of the attention module, an attention map is obtained; the multi-scale fusion features and the attention map are fused to obtain attention fusion features; based on the obtained attention fusion features, through a preset decoder, a prediction output is obtained.

[0010] Among them, during the training process of the road obstacle recognition model, an attention map obtained based on the attention mechanism is used, combined with the obtained prediction output, and the correct recognition result and the prediction output are combined through a preset reward mechanism to enable the road obstacle recognition model to focus its attention on the uncertain area of the road image to be recognized.

[0011] Furthermore, the acquisition of the attention fusion features is specifically represented as follows:

[0012]

[0013] Among them, Fa is the input feature of the attention module, Matt(Fa) is the attention map, F sc is the multi-scale fusion feature, α is a learnable parameter, represents element-wise multiplication.

[0014] Furthermore, in the attention module, the sum of the attention values is restricted to a fixed value through the softmax function.

[0015] Furthermore, the reward mechanism is specifically as follows: during the model training process, the correct classification result of the target object in the input image is obtained, and a reward map is constructed based on the obtained correct classification result; the obtained attention map is upsampled to the same size as the input image; the input image is predicted based on the prediction output of the encoder, the upsampled attention map, and the reward map.

[0016] Furthermore, the prediction of the input image based on the prediction output of the encoder, the upsampled attention map, and the reward map is specifically represented as follows:

[0017]

[0018] Among them, τ is a reward hyperparameter that controls the reward ratio, is the upsampled attention map, Pa is the prediction output of the encoder, G award is the reward map, represents element-wise multiplication.

[0019] Furthermore, during the model training process, an attention loss function term is introduced into its loss function, and the attention loss function term is specifically expressed as follows:

[0020]

[0021] where m h,w represents the attention value corresponding to the position coordinates (h, w) in the input image, β and ψ are adjustable penalty hyperparameters, and k h,w is the anomaly score corresponding to the position coordinates (h, w) in the input image.

[0022] In a second aspect, the present invention provides a road obstacle recognition system, including:

[0023] A data acquisition unit for real-time collecting road images in the traveling direction of the vehicle;

[0024] An obstacle recognition unit for obtaining an obstacle recognition result based on the obtained road image through a pre-trained deep learning-based road obstacle recognition model; the road obstacle recognition model specifically performs the following processing process: taking the road image as an input, obtaining multi-scale fusion features based on a preset encoder based on a multi-scale feature extraction network; taking the preset scale features as the input of an attention module to obtain an attention map; fusing the multi-scale fusion features and the attention map to obtain attention fusion features; and obtaining a prediction output based on the obtained attention fusion features through a preset decoder.

[0025] Wherein, during the training process of the road obstacle recognition model, an attention map obtained based on an attention mechanism is used, combined with the obtained prediction output, and the correct recognition result and the prediction output are combined through a preset reward mechanism to enable the road obstacle recognition model to focus its attention on the uncertain area of the road image to be recognized.

[0026] According to a third aspect of the embodiments of the present invention, there is provided an electronic device including a memory, a processor, and a computer program running on the memory, and when the processor executes the program, it implements the described road obstacle recognition method.

[0027] According to a fourth aspect of the embodiments of the present invention, there is provided a non-transitory computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, it implements the described road obstacle recognition method.

[0028] According to a fifth aspect of the embodiments of the present invention, there is provided a vehicle that uses the described road obstacle recognition method.

[0029] The above one or more technical solutions have the following beneficial effects:

[0030] The present invention provides a method and system for road obstacle recognition. Based on a proposed reward mechanism and combined with the attention map generated by the attention module, the model focuses its attention on the uncertain regions of the road image to be recognized, thereby effectively improving the model's effective recognition of unknown type obstacles and enhancing its environmental adaptability.

[0031] In the solution of the present invention during the model training process, an attention loss function term is introduced on the basis of the traditional loss function, enabling the model to focus its attention on regions with higher uncertainty to a certain extent, and further improving the model's effective recognition of unknown type obstacles.

[0032] Advantages of additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be learned through the practice of the present invention. Brief Description of the Drawings

[0033] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0034] Figure 1 It is a schematic diagram of the overall process of a method for road obstacle recognition according to an embodiment of the present invention;

[0035] Figure 2 It is a schematic diagram of the process for obtaining the attention fusion features according to an embodiment of the present invention;

[0036] Figure 3 It is a schematic diagram of the processing process of the reward mechanism according to an embodiment of the present invention;

[0037] Figure 4 It is a schematic diagram of the experimental results of a method for road obstacle recognition according to an embodiment of the present invention. Detailed Description of the Embodiments

[0038] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0039] Attention should be paid to a method and system for in-vehicle child monitoring based on the fusion of millimeter-wave radar and vision. The embodiments and features in the embodiments of the present invention can be combined with each other.

[0040] In one or more embodiments, as Figure 1 shown, a method for road obstacle recognition is presented, including:

[0041] Step 1: Real-time collect road images in the driving direction of the vehicle;

[0042] Step 2: Based on the obtained road images, through a pre-trained deep learning-based road obstacle recognition model, obtain the obstacle recognition results; the road obstacle recognition model specifically performs the following processing process: taking the road images as input, based on a preset encoder based on a multi-scale feature extraction network, obtain multi-scale fusion features; taking the preset scale features as the input of the attention module, obtain the attention map; fuse the multi-scale fusion features and the attention map to obtain the attention fusion features; based on the obtained attention fusion features, through a preset decoder, obtain the prediction output;

[0043] Among them, in the training process of the road obstacle recognition model, using the attention map obtained based on the attention mechanism, combined with the obtained prediction output, through a preset reward mechanism, combine the correct recognition results with the prediction output, so as to enable the road obstacle recognition model to focus its attention on the uncertain areas of the road images to be recognized.

[0044] In a specific implementation, the acquisition of the road images can be specifically obtained through an in-vehicle camera, and the acquired image forms can be color images and infrared images, or other types of images.

[0045] In a specific implementation, the deep learning-based road obstacle recognition model includes an encoder, a decoder, an attention module, and a reward mechanism module; specifically:

[0046] The encoder uses a multi-scale feature extraction network, such as: a pyramid network, a dilated convolution network;

[0047] The decoder is not set in the solution of this embodiment and can use a traditional network model;

[0048] As Figure 2 shown, the attention module specifically performs the following processing process:

[0049] Given an intermediate feature map Fa ∈ R C″×H′×W′ as the input, based on the attention module, output an attention map Matt ∈ R 1×H′×W′ . The higher the value on the map Matt, the lower the confidence of the model in the predicted pixel class.

[0050] Specifically, we first compress the features used by Fa through average pooling and max pooling operations along the channel axis to generate an efficient feature description similar to CBAM (Convolutional Block Attention Module). Then, we obtain an attention map of the same size as the input feature by concatenating the input feature Fa and feeding the feature into a standard convolutional layer. Specifically, the softmax function is used to limit the sum of attention values to a fixed value, generating our curiosity-driven attention map Matt. The arrangement of using softmax instead of the sigmoid function is to let different pixels compete for limited attention, thus achieving an optimal fit. In short, the attention map is calculated as follows:

[0051] Matt(F a ) = δ(f([Avgpool(F a );Maxpool(Fa)]))

[0052] where δ represents the softmax function and f represents the standard convolutional operation.

[0053] To obtain more detailed model inference features in the high-attention regions, we fuse the multi-scale features F low and F mid , which are sequentially selected from two different layers of the encoder. Concatenation and summation are two common feature fusion methods, and our experimental results show that the concatenation method is superior to the summation method. We use a convolutional layer with a kernel size of 1×1 and an upsampling layer to convert the features to the same size as Fa. The fused multi-scale feature is calculated as follows:

[0054] F sc = η(Concat(F low ,F mid ,Fa))

[0055] where η represents the feature processing operation using convolution and upsampling. The entire attention process can be summarized as follows:

[0056]

[0057] where represents element-wise multiplication, and α is set as a learnable parameter that is gradually modified during the training process.

[0058] After obtaining an attention module, inserting it at different positions in the semantic segmentation network may affect the overall performance. Empirically, the closer the attention is to the output, the greater the impact on the output result, thus obtaining a more accurate attention map. On the other hand, for high-dimensional features, the closer the attention is to the deeper position of the model, the more information is obtained, and the greater the impact on the segmentation result. In the layer close to the output, the attention weights extracted by the attention model are under-generalized, but the attention is sensitive, which often has a negative impact on the result. On the contrary, when the attention is placed deeper, the accurate attention map will enhance the features in the uncertain regions of the model, and this enhancement is fully fitted by the decoder, which can play the role of enhancing local features by the attention module for recognition. At the same time, this differentiates the recognition model described in this embodiment from the uncertainty prediction task.

[0059] In a specific implementation, as Figure 3 shown, the reward mechanism specifically includes the following processing procedures:

[0060] Let X ∈ r 3×H×W represent the input image, where H and W are the height and width of the input image respectively. Before normalization, the semantic segmentation network generates a prediction probability P c,h,w for each category at each position (h, w). Then, the output of the entire model can be denoted as Pa ∈ R C×H×W . The correct classification result of the input image is recorded in G ∈ Z H×W , and g h,w in G represents the class number C corresponding to the position (h, w). In the ground truth (GT) information, only normal categories are included, and out-of-distribution (OoD) samples are not covered.

[0061] To construct the reward map G award , the solution described in this embodiment expands the Grund Turth at each position, that is, g h,w from one-dimensional number of categories to a three-dimensional vector form G award ∈ Z C×H×W . In this vector representation, for the position (h, w), when c is the correct category at this position, the value of its vector element is 1, and for other categories it is 0. By using such an encoding method, the reward matrix is converted to the same size as the prediction output.

[0062] In addition, operations on the attention map are also involved. First, the attention map is upsampled to the same size as the input image. This attention map is denoted as Next, the reward feature is input into the model to help the network make predictions:

[0063]

[0064] Among them, τ is a reward hyperparameter that controls the reward ratio. When the model understands the input image, for each position (h, w) on the image, the higher its attention value, the more rewards the model obtains. This reward directly acts on improving the prediction probability of the correct class by the model. The principle of action of the GT (Ground Truth) reward value is that it can increase the prediction probability value of the correct class in the model output. When the model prediction value is close to the true class, the loss function of the classification problem will decrease accordingly, which helps the training and optimization of the model.

[0065] In addition, to prevent the addition of the reward mechanism from making the model lazily overly dependent on rewards and resulting in a decline in semantic segmentation performance, the solution described in this embodiment randomly selects whether to give the model a reward with a certain probability in each iteration, so that the model is continuously optimized in both the test and training scenarios. In the test scenario, the model is not allowed to obtain rewards. In contrast, the reward mechanism is used to motivate the model to recognize its own deficiencies in the training scenario.

[0066] According to the relevant theory of the uncertainty estimation method, different uncertainty estimation functions g(S) can be selected to calculate the final anomaly score, denoted as K ∈ R H×W . However, when only the reward mechanism is adopted, the difference in the attention values assigned by the model to abnormal and normal targets is not significant.

[0067] Based on the above problems, in one or more embodiments, in order to increase the gap between attention values in this embodiment, we add an attention loss function term on the basis of the original segmentation loss. The attention map can be obtained from the previous reward method. Then, the average value of the attention values at all correctly predicted positions is calculated. The correct positions are the positions that the model can correctly predict without using the GT reward calculated from Pa. On the contrary, the incorrect positions are those. The intuition behind this is that the model should have a lower attention value for the positions that it can easily and correctly predict. Specifically, the attention loss is calculated as follows:

[0068]

[0069] where m h,w represents the attention value corresponding to the position coordinates (h, w) in the input image, and β, ψ are adjustable penalty hyperparameters that balance the gap between the attention values of each position. The smaller β is, the greater the difference in attention values between abnormal and normal pixels. k h,w is the anomaly score corresponding to the position coordinates (h, w) in the input image. Assuming that the loss function of the original semantic segmentation network is L seg , then the total loss function is defined as:

[0070] L total = L seg + λL att

[0071] where λ is the weight parameter of L att Since the network has poor segmentation ability at the beginning of training, we hope that the network will focus on improving the semantic segmentation accuracy. As the number of training times increases, we hope that the network can focus on improving the semantic segmentation accuracy. As the number of training times increases, we expect the model to focus on learning how to improve the accuracy of anomaly detection by reasonably allocating attention. Therefore, the attention weight should gradually increase as the training accuracy improves. Based on this idea, a self-adjusting method for λ is required. λ is calculated as:

[0072]

[0073] So far, a curiosity-driven attention map M att and a probabilistic prediction output K can be obtained. The visualization result of the attention map is as shown in Figure 4 (b). K can be calculated using various uncertainty estimation functions. We choose the SML tool to generate the anomaly score and show the visualization result in Figure 4 (c). Finally, we integrate the probabilistic prediction output and the attention map using weighted summation:

[0074]

[0075] where σ represents the anomaly score weighting hyperparameter. Through this simple integration, the recognition model described in this embodiment can be combined with any anomaly detection method to obtain better performance, and its visualization result is as shown in Figure 4 (d). By combining the anomaly detection and semantic segmentation results, the final segmentation prediction map is as shown in Figure 4 (e).

[0076] In addition, in one or more embodiments, an attempt is made to combine the method described in this embodiment with some other uncertainty estimation methods, and the effectiveness of the solution is shown. It can be understood that training a model focused on anomaly obstacles will give rise to an interesting optimization problem. When the network can successfully focus its attention on areas with higher uncertainty, it can obtain rewards and thus reduce the overall loss.

[0077] In one or more embodiments, corresponding to the above method, this embodiment provides a road obstacle recognition system, including:

[0078] A data acquisition unit for real-time collecting road images in the vehicle traveling direction;

[0079] An obstacle recognition unit, which is used to obtain an obstacle recognition result based on the acquired road image through a pre-trained deep learning-based road obstacle recognition model; the road obstacle recognition model specifically performs the following processing process: taking the road image as an input, obtaining multi-scale fusion features based on a preset encoder based on a multi-scale feature extraction network; taking the preset scale features as the input of an attention module to obtain an attention map; fusing the multi-scale fusion features and the attention map to obtain attention fusion features; based on the obtained attention fusion features, obtaining a prediction output through a preset decoder;

[0080] Wherein, during the training process of the road obstacle recognition model, an attention map obtained based on an attention mechanism is used, combined with the obtained prediction output, and the correct recognition result and the prediction output are combined through a preset reward mechanism, so that the road obstacle recognition model focuses its attention on the uncertain area of the road image to be recognized.

[0081] It can be understood that the system in this embodiment corresponds to the above method embodiment, and its technical details have been described in detail in the method embodiment, so they will not be repeated here.

[0082] In more embodiments, there is also provided:

[0083] An electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the above embodiment is completed. For the sake of brevity, it will not be repeated here.

[0084] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0085] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0086] A computer-readable storage medium for storing computer instructions, which when executed by a processor, completes the method described in the above embodiment.

[0087] The methods in the above embodiments can be directly implemented by a hardware processor, or by a combination of hardware and software modules in the processor. The software modules can be located in mature storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage media is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above methods. To avoid repetition, it will not be described in detail here.

[0088] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0089] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0090] In more embodiments, a vehicle is also provided, which adopts the above road obstacle recognition method.

[0091] In specific implementation, the vehicle may further include components such as an RF (Radio Frequency) circuit, a memory including one or more computer-readable storage media, an input unit, a display unit, sensors, an audio circuit, a WiFi (Wireless Fidelity) module, a processor including one or more processing cores, and a power supply. Those skilled in the art can understand that the above components do not constitute a limitation to the vehicle, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Among them:

[0092] The RF circuit can be used for receiving and sending information during information reception and call processes. In particular, after receiving the downlink information of the base station, it is handed over to one or more processors for processing. Additionally, the data related to the uplink is sent to the base station. Generally, the RF circuit includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a subscriber identity module (SIM) card, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit can also communicate with the network and other devices through wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), LTE (Long Term Evolution), email, SMS (Short Messaging Service), etc.

[0093] The memory can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the vehicle (such as audio data, a phone book, etc.), etc. In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory can also include a memory controller to provide access to the memory for the processor and the input unit.

[0094] The input unit can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control. Specifically, the input unit can include a touch-sensitive surface and other input devices. The touch-sensitive surface, also known as a touch display screen or a touchpad, can collect touch operations of the user on or near it (such as operations of the user using any suitable object or accessory such as a finger or a stylus on or near the touch-sensitive surface), and drive the corresponding connection device according to a pre-set program. Optionally, the touch-sensitive surface can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor, and can receive and execute the commands sent by the processor. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch-sensitive surface. In addition to the touch-sensitive surface, the input unit can also include other input devices. Specifically, the other input devices can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), trackballs, mice, joysticks, etc.

[0095] The display unit can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the vehicle. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit can include a display panel. Optionally, the display panel can be configured in forms such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode). Further, the touch-sensitive surface can cover the display panel. When the touch-sensitive surface detects a touch operation on or near it, it is transmitted to the processor to determine the type of touch event. Subsequently, the processor provides a corresponding visual output on the display panel according to the type of touch event. The touch-sensitive surface and the display panel are implemented as two independent components to achieve input and input functions. However, in some embodiments, the touch-sensitive surface and the display panel can be integrated to achieve input and output functions.

[0096] The vehicle may also include at least one sensor, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel according to the brightness of the ambient light, and the proximity sensor can turn off the display panel and / or the backlight when the vehicle moves close to the ear. As a kind of motion sensor, the gravity acceleration sensor can detect the magnitude of the acceleration in each direction (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used in applications for identifying the posture of the vehicle (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. As for other sensors that the vehicle can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.

[0097] The audio circuit, speaker, and microphone can provide an audio interface between the user and the vehicle. The audio circuit can transmit the electrical signal converted from the received audio data to the speaker, and the speaker converts it into a sound signal for output. On the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit and then converted into audio data. After the audio data is output to the processor for processing, it is sent via the RF circuit to, for example, another vehicle, or the audio data is output to the memory for further processing. The audio circuit may also include an earphone jack to provide communication between the external earphone and the vehicle.

[0098] WiFi belongs to short - range wireless transmission technology. The vehicle can help users send and receive emails, browse the web, and access streaming media through the WiFi module, which provides users with wireless broadband Internet access. Although the WiFi module is shown, it can be understood that it does not belong to the essential components of the vehicle and can be omitted entirely according to needs within the scope of not changing the essence of the invention.

[0099] The processor is the control center of the vehicle, connecting all parts of the entire vehicle through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory, and calling the data stored in the memory, it executes various functions of the vehicle and processes data, thereby monitoring the vehicle as a whole. Optionally, the processor may include one or more processing cores. Preferably, the processor may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor either.

[0100] The vehicle further includes a power source (such as a battery) for supplying power to various components. Preferably, the power source can be logically connected to the processor through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power source can also include any components such as one or more DC or AC power sources, a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator, etc.

[0101] Although not shown, the vehicle may further include a camera, a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the display unit of the vehicle is a touch screen display, and the vehicle further includes a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include those for executing the methods shown in the above embodiments.

[0102] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A road obstacle recognition method, characterized in that: include: Collect road images in the direction of vehicle travel in real time; Obtain an obstacle recognition result based on the obtained road image through a pre-trained road obstacle recognition model based on deep learning; the road obstacle recognition model specifically performs the following processing process: taking the road image as input, obtaining multi-scale fusion features based on a preset encoder based on a multi-scale feature extraction network; Use the preset scale feature as the input of the attention module to obtain the attention map; Fusing the multi-scale fusion feature and the attention map to obtain an attention fusion feature; Based on the obtained attention fusion features, the prediction output is obtained through the preset decoder; Among them, during the training process of the road obstacle recognition model, an attention map obtained based on the attention mechanism is used, combined with the obtained prediction output, and the correct recognition result is combined with the prediction output through a preset reward mechanism, so as to enable the road obstacle recognition model to focus on the uncertain area of ​​the road image to be recognized.

2. A road obstacle recognition method as claimed in claim 1, characterized in that: The acquisition of the attention fusion feature is specifically expressed as follows: Among them, Fa is the input feature of the attention module, Matt(Fa) is the attention map, and F sc is a multi-scale fusion feature, α is a learnable parameter, Represents element-wise multiplication.

3. A road obstacle recognition method as claimed in claim 1, characterized in that: In the attention module, the sum of attention values ​​is limited to a fixed value through the softmax function.

4. A road obstacle recognition method as claimed in claim 1, characterized in that: The reward mechanism is specifically as follows: during the model training process, the correct classification result of the target object in the input image is obtained, and a reward map is constructed based on the correct classification result obtained; the obtained attention map is upsampled to the same size as the input image; and the input image is predicted based on the predicted output of the encoder, the upsampled attention map, and the reward map.

5. A road obstacle recognition method as claimed in claim 4, characterized in that: The prediction output of the encoder, the upsampled attention map, and the reward map are used to predict the input image, which is specifically expressed as follows: Among them, τ is the reward hyperparameter that controls the reward ratio, is the upsampled attention map, Pa is the predicted output of the encoder, and G award For the reward picture, Represents element-wise multiplication.

6. A road obstacle recognition method as claimed in claim 1, characterized in that: During the model training process, its loss function introduces an attention loss function term, which is specifically expressed as follows: Among them, m h,w represents the attention value corresponding to the position coordinate (h, w) in the input image, β and ψ are adjustable penalty hyperparameters, and k h,w is the anomaly score corresponding to the position coordinate (h, w) in the input image.

7. A road obstacle recognition system, characterized in that: include: A data acquisition unit, which is used to collect road images in the direction of vehicle travel in real time; The obstacle recognition unit is used to obtain an obstacle recognition result based on the obtained road image through a pre-trained road obstacle recognition model based on deep learning; the road obstacle recognition model specifically performs the following processing process: taking the road image as input, obtaining multi-scale fusion features based on a preset encoder based on a multi-scale feature extraction network; Use the preset scale feature as the input of the attention module to obtain the attention map; Fusing the multi-scale fusion feature and the attention map to obtain an attention fusion feature; Based on the obtained attention fusion features, the prediction output is obtained through the preset decoder; Among them, during the training process of the road obstacle recognition model, an attention map obtained based on the attention mechanism is used, combined with the obtained prediction output, and the correct recognition result is combined with the prediction output through a preset reward mechanism, so as to enable the road obstacle recognition model to focus on the uncertain area of ​​the road image to be recognized.

8. An electronic device comprising a memory, a processor and a computer program stored and running on the memory, characterized in that: When the processor executes the program, a road obstacle recognition method as described in any one of claims 1-6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a road obstacle recognition method as described in any one of claims 1 to 6 is implemented.

10. A vehicle, characterized in that: It adopts a road obstacle recognition method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Obstacle recognition method, obstacle recognition device and electronic equipment

    CN113128386A