Obstacle detection method, system, equipment and medium
By feature encoding and splicing of obstacle detection images, combined with the context information of the initial abnormal feature map, the problem of insufficient obstacle detection accuracy in the prior art is solved, and more accurate and robust obstacle recognition is achieved.
Patent Information
- Application Number
- CN202510237664.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-06
AI Technical Summary
The existing semantic segmentation model independently analyzes each pixel in obstacle detection, and does not fully consider the relationship between the pixel and the surrounding environment, resulting in the edges of the obstacles being damaged and small objects cannot be fully identified, affecting the accuracy and safety of the autonomous driving system.
The first image feature is obtained by encoding the image to be detected, and the initial abnormal feature map is encoded to obtain the second image feature, the two are spliced to obtain the fusion feature, and obstacle detection is performed based on the fusion feature.
By introducing context information of the initial anomaly feature map, the model can better understand the position, size, shape, and relationship with other elements of the obstacle in the scene, reduce misdetection and missed detection, and improve the accuracy and robustness of obstacle identification.
Smart Images

Figure CN120107931A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image target detection, and in particular to an obstacle detection method, system, equipment and medium. Background Art
[0002] With the development of deep learning, semantic segmentation models have been introduced into obstacle detection. By classifying each pixel in the image, it can accurately determine what object category each pixel belongs to (such as obstacle, road, background, etc.), thereby providing relatively detailed detection results and identifying obstacles more carefully than traditional methods.
[0003] However, when detecting obstacles, existing technologies often analyze each pixel independently without fully considering the relationship between the pixel and the surrounding environment. As a result, existing semantic segmentation models are insufficient in segmenting complete objects. For example, at the edge of an obstacle, due to the complex transition of pixel features, it is difficult for the model to accurately determine the boundary, resulting in incomplete edges of the segmented obstacle. For small objects, the model may not be able to fully outline its contours due to insufficient feature extraction, causing some abnormal objects to be unable to be fully identified. In autonomous driving scenarios, this will prevent the system from accurately perceiving the full picture of obstacles and making incorrect driving decisions, such as misjudging the size and position of obstacles, leading to dangers such as vehicle collisions, which seriously endanger driving safety. Summary of the invention
[0004] In view of this, in order to overcome at least one aspect of the above problems, an embodiment of the present invention provides an obstacle detection method, comprising the following steps: Encoding the image to be detected to obtain a first image feature, and performing obstacle detection on the image to be detected to obtain an initial abnormal feature map; Encoding the initial abnormal feature map to obtain a second image feature; The first image feature and the second image feature are concatenated to obtain a fusion feature; Obstacle detection is performed based on the fusion features to obtain a detection result.
[0005] In some embodiments, encoding the initial abnormal feature map to obtain a second image feature further comprises: The initial abnormal feature map is encoded based on the dimension and size of the first image feature to obtain the second image feature, wherein the dimension and size of the second image feature are the same as the dimension and size of the first image feature.
[0006] In some embodiments, encoding the initial abnormal feature map based on the dimension and size of the first image feature to obtain the second image feature further includes: The initial abnormal feature map is sequentially input into the first convolution layer, the first linear layer, the first activation function layer, the second convolution layer, the second linear layer, the second activation function layer, and the third convolution layer to obtain the second image feature, wherein the strides of the first convolution layer, the second convolution layer, and the third convolution layer and the number and size of the convolution kernels are determined according to the dimension and size of the first image feature.
[0007] In some embodiments, performing obstacle detection based on the fusion feature to obtain a detection result further includes: Obtaining a weight based on the fused features; Using the weight to weight the first image feature and the second image feature in the fusion feature to obtain the weighted fusion feature; Obstacle detection is performed based on the weighted fusion features to obtain a detection result.
[0008] In some embodiments, obtaining weights based on the fused features further includes: Inputting the fused features into a weight filter, wherein the weight filter comprises a multi-layer convolutional network and a softmax layer; The weight is obtained by calculating the fusion feature using the multi-layer convolutional network and the softmax layer in the weight filter.
[0009] In some embodiments, performing obstacle detection based on the fusion feature to obtain a detection result further includes: Obtaining multiple obstacle masks based on the fused features; Preprocessing the plurality of obstacle masks to obtain a final obstacle mask; The detection result is obtained based on the final obstacle mask.
[0010] In some embodiments, preprocessing the plurality of obstacle masks to obtain a final obstacle mask further comprises: Non-maximum suppression is used to perform a redundancy removal operation on the plurality of obstacle masks.
[0011] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention further provides an obstacle detection system, comprising: A calculation module is configured to encode the image to be detected to obtain a first image feature, and perform obstacle detection on the image to be detected to obtain an initial abnormal feature map; an alignment module, configured to encode the initial abnormal feature map to obtain a second image feature; A splicing module configured to splice the first image feature and the second image feature to obtain a fusion feature; The detection module is configured to perform obstacle detection based on the fusion features to obtain a detection result.
[0012] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention further provides a computer device, including: at least one processor; and A memory storing a computer program executable on the processor, wherein the processor executes the steps of any one of the obstacle detection methods described above when executing the program.
[0013] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any one of the obstacle detection methods described above are performed.
[0014] The present invention has one of the following beneficial technical effects: the solution proposed by the present invention introduces the initially obtained abnormal feature map into the coding features of the image, that is, introduces additional context information into the coding features of the image. In subsequent detection, the model can better understand the position, size, shape and relationship between the obstacle and other elements in the entire scene. Through such information integration and fusion, the problems of false detection and missed detection can be reduced, and the accuracy and robustness of identifying different types of obstacles can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can be obtained based on these drawings without paying creative work.
[0016] Figure 1 A schematic diagram of a flow chart of an obstacle detection method provided by an embodiment of the present invention; Figure 2 A flowchart of an obstacle detection method provided by an embodiment of the present invention; Figure 3 A schematic diagram of an image to be detected provided by an embodiment of the present invention; Figure 4 A schematic diagram of the true value of an obstacle area in an image to be detected provided by an embodiment of the present invention; Figure 5 A schematic diagram of an initial abnormal characteristic graph provided by an embodiment of the present invention; Figure 6 A schematic diagram of the final detection result provided by an embodiment of the present invention; Figure 7 A schematic diagram of the structure of a feature alignment module provided by an embodiment of the present invention; Figure 8 A schematic diagram of the structure of a feature fusion module provided by an embodiment of the present invention; Fig. 9 A schematic diagram of the structure of an obstacle detection system provided by an embodiment of the present invention; Fig.10 A schematic diagram of the structure of a computer device provided by an embodiment of the present invention; Fig.11 A schematic diagram of the structure of a computer-readable storage medium provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0017] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the embodiments of the present invention are further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0018] It should be noted that all expressions using "first" and "second" in the embodiments of the present invention are for distinguishing two non-identical entities with the same name or non-identical parameters. It can be seen that "first" and "second" are only for the convenience of expression and should not be understood as limitations on the embodiments of the present invention. The subsequent embodiments will not explain this one by one.
[0019] According to one aspect of the present invention, an embodiment of the present invention provides an obstacle detection method, such as Figure 1 As shown, it may include the steps of: S1, encoding the image to be detected to obtain a first image feature, and performing obstacle detection on the image to be detected to obtain an initial abnormal feature map; S2, encoding the initial abnormal feature map to obtain a second image feature; S3, concatenating the first image feature and the second image feature to obtain a fusion feature; S4, performing obstacle detection based on the fusion features to obtain a detection result.
[0020] In some embodiments, Figure 2 As shown, the image encoder of Sam (segment anything model, semantic segmentation model) can be used to encode the image to be detected to obtain the first image feature , use the pixel-level obstacle segmentation model to encode the image to be detected to obtain the initial abnormal feature map Then the feature alignment module is used to encode the initial abnormal feature map to obtain the second image feature Then, the feature fusion module is used to fusion the first image features and the second image feature Splice to get fusion features Finally, the mask decoder and mask filtering module are used to fusion feature Obstacle detection is performed to obtain detection results. The solution proposed by the present invention introduces the initially obtained abnormal feature map into the encoding features of the image, that is, introduces additional context information into the encoding features of the image. In subsequent detection, the model can better understand the position, size, shape and relationship between the obstacle and other elements in the entire scene. Through such information integration and fusion, the problems of false detection and missed detection can be reduced, and the accuracy and robustness of identifying different types of obstacles can be improved.
[0021] For example, Figure 3 The true value of the obstacle in the image to be detected can be Figure 4 As shown in the figure, the true value refers to the standard answer that is considered to be completely accurate. In the obstacle segmentation task, the true value image accurately marks the actual position, shape, boundary and other information of the obstacles in the image. Figure 3 After the image to be detected is processed by the pixel-level obstacle segmentation model, we get Figure 5 The initial abnormal feature map shown in Figure 1. This feature map contains the initial recognition information of the obstacles in the image by the pixel-level obstacle segmentation model, which is obtained from the difference between the pixel-level obstacle segmentation model and the real value image ( Figure 4 ) can be seen from the comparison between the initial anomaly feature map and the obstacle boundary that the obstacle boundary in the initial anomaly feature map is blurred. Figure 6 is the detection result obtained based on the solution proposed by the present invention. Compared with the initial abnormal feature map ( Figure 5 ), the boundary of the obstacle in this result becomes clear, which shows that in the scheme proposed in the present invention, additional context information (initial abnormal feature map) is introduced into the encoding features of the image, so that the final detection result can more accurately reflect the real shape and boundary of the obstacle, thereby improving the accuracy and robustness of the detection.
[0022] In some embodiments, encoding the initial abnormal feature map to obtain a second image feature further comprises: The initial abnormal feature map is encoded based on the dimension and size of the first image feature to obtain the second image feature, wherein the dimension and size of the second image feature are the same as the dimension and size of the first image feature.
[0023] Specifically, since the image to be detected is an RGB image, the RGB image becomes the first image feature of the patch-level after passing through the image encoder. , which no longer completely corresponds to the spatial position information of the original image. Therefore, it is necessary to use the feature alignment module to encode the initial abnormal feature map so that the initial abnormal feature map With the first image feature The spatial information of the feature alignment module is aligned, and the features encoded by the feature alignment module can be recorded as .
[0024] For example, if the image to be detected is an RGB image of (3, 1024, 1024), the path_size of the image encoder is 16, dim=256, and after encoding, the first image feature is (dim, 1024 / 16, 1024 / 16), that is, (256, 64, 64). After the image to be detected passes through the pixel-level obstacle segmentation model, the initial abnormal feature map of (1, 256, 256) is obtained. The feature alignment module can be used to align the initial abnormal feature map based on the dimension and size of the first image feature. Encode and obtain the second image feature which is also (256, 64, 64) .
[0025] In some embodiments, encoding the initial abnormal feature map based on the dimension and size of the first image feature to obtain the second image feature further includes: The initial abnormal feature map is sequentially input into the first convolution layer, the first linear layer, the first activation function layer, the second convolution layer, the second linear layer, the second activation function layer, and the third convolution layer to obtain the second image feature, wherein the strides of the first convolution layer, the second convolution layer, and the third convolution layer and the number and size of the convolution kernels are determined according to the dimension and size of the first image feature.
[0026] Specifically, in the feature alignment module, the initial abnormal feature map can be calibrated using the superposition operation features of the convolutional layer and the linear layer. Encoding, the number of convolutional layers and linear layers in the feature alignment module can be set according to the actual detection results. The more convolutional layers and linear layers, the more contextual information is retained, which is more conducive to subsequent detection, but it will lead to longer model training and inference time, and higher requirements on the computing power of hardware devices. The fewer the number of convolutional layers and linear layers, the less information is retained, resulting in low subsequent detection accuracy. Therefore, in practical applications, the number of convolutional layers and linear layers in the feature alignment module can be reasonably set according to the specific detection task requirements, hardware conditions, and the trade-off between detection accuracy and computing efficiency.
[0027] For example, Figure 7As shown, the feature alignment module proposed in the present invention may include a first convolutional layer, a first linear layer, a first activation function layer, a second convolutional layer, a second linear layer, a second activation function layer, and a third convolutional layer. The strides of the first convolutional layer, the second convolutional layer, and the third convolutional layer, as well as the number and size of the convolution kernels, are determined according to the dimension and size of the first image feature. If the initial abnormal feature map of (1, 256, 256) The second image feature of (256, 64, 64) is obtained by encoding , the stride of the first convolutional layer is (2, 2), the convolution kernel size is (2, 2), the number is 4, and the initial abnormal feature map of (1, 256, 256) After the first convolutional layer, the feature of (4, 128, 128) is obtained. The stride of the second convolutional layer is (2, 2), the convolution kernel size is (2, 2), and the number is 16. After the second convolutional layer, the feature of (4, 128, 128) is obtained. The feature of (16, 64, 64) is obtained. After the third convolutional layer, the stride of (2, 2), the convolution kernel size is (2, 2), and the number is 256. After the third convolutional layer, the feature of (16, 64, 64) is obtained. The feature of (256, 64, 64) is obtained.
[0028] In some embodiments, performing obstacle detection based on the fusion feature to obtain a detection result further includes: Obtaining a weight based on the fused features; Using the weight to weight the first image feature and the second image feature in the fusion feature to obtain the weighted fusion feature; Obstacle detection is performed based on the weighted fusion features to obtain a detection result.
[0029] Specifically, in order to better utilize the fused features for obstacle detection, weights can be calculated based on the fused features, and the weights can be used to highlight the more important parts of the fused features for obstacle detection, so that the more important feature parts are enhanced, while the relatively unimportant feature parts are weakened, thereby making the fused features more conducive to subsequent obstacle detection. For example, if the second image feature contains some background information, and the weights show that the background information is less important for obstacle detection, then after weighting, the impact of the background information will be reduced; on the contrary, if the first image feature contains key obstacle details, weighting can make it more prominent in the fused feature.
[0030] For example, if the weight obtained based on the fusion feature is W, then W can be used to Perform block-by-block weighted averaging to generate multimodal features , the specific formula is as follows:
[0031] In some embodiments, obtaining weights based on the fused features further includes: Inputting the fused features into a weight filter, wherein the weight filter comprises a multi-layer convolutional network and a softmax layer; The weight is obtained by calculating the fusion feature using the multi-layer convolutional network and the softmax layer in the weight filter.
[0032] Specifically, Figure 8 As shown, the weight filter may include a multi-layer convolutional subnetwork and a softmax layer, and the multi-layer convolutional subnetwork may include a convolutional layer, an activation function layer, and a convolutional layer. After the first image feature and the second image feature are spliced and input into the weight filter, the weight is output after being calculated by the convolutional network and the softmax layer in the weight filter. In this way, through the feature extraction of the multi-layer convolutional subnetwork and the nonlinear processing of the activation function, the fused feature is converted into a more discriminative feature representation, and then converted into a set of weights through the softmax layer, which provides a weight allocation scheme with a certain adaptive ability for subsequent processing steps (such as weighting the fused features, etc.), so that the information in the fused features can be better utilized to improve the performance of the entire system and the accuracy of task completion.
[0033] In some embodiments, performing obstacle detection based on the fusion feature to obtain a detection result further includes: Obtaining multiple obstacle masks based on the fused features; Preprocessing the plurality of obstacle masks to obtain a final obstacle mask; The detection result is obtained based on the final obstacle mask.
[0034] Specifically, since the road obstacles in the image may be complex, such as different shapes, sizes, occlusions, etc., the mask decoder will not directly obtain a unique and accurate obstacle mask, but will generate multiple possible candidate masks that meet certain conditions. These candidate masks may represent the area where the real road obstacles are located in the image, so the final obstacle mask can be obtained through preprocessing, and finally the detection result can be obtained based on the final obstacle mask. In order to observe the obstacle distribution more intuitively, the color map can be used to render the mask to obtain a color image.
[0035] In some embodiments, preprocessing the plurality of obstacle masks to obtain a final obstacle mask further comprises: Non-maximum suppression is used to perform a redundancy removal operation on the plurality of obstacle masks.
[0036] Specifically, the core idea of non-maximum suppression is to retain only the best one among multiple candidate results, and suppress other overlapping and low-scoring results. Each obstacle mask can be regarded as a prediction of the possibility of obstacles in the image, so the most representative and accurate masks are selected, and those redundant masks that highly overlap with other masks and have low confidence are removed. Removing redundant obstacle masks can avoid repeated detection of the same obstacle, so that the detection results more accurately reflect the number and location of obstacles actually existing in the image, and reduce false detections.
[0037] The solution proposed in this invention introduces the initially obtained abnormal feature map into the coded features of the image, that is, introduces additional context information into the coded features of the image. In subsequent detection, the model can better understand the position, size, shape and relationship between the obstacle and other elements in the entire scene. Through such information integration and fusion, the problems of false detection and missed detection can be reduced, and the accuracy and robustness of identifying different types of obstacles can be improved.
[0038] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention further provides an obstacle detection system 400, such as Fig. 9 As shown, including: The calculation module 401 is configured to encode the image to be detected to obtain a first image feature, and perform obstacle detection on the image to be detected to obtain an initial abnormal feature map; An alignment module 401 is configured to encode the initial abnormal feature map to obtain a second image feature; A splicing module 401 is configured to splice the first image feature and the second image feature to obtain a fusion feature; The detection module 401 is configured to perform obstacle detection based on the fusion features to obtain a detection result.
[0039] In some embodiments, encoding the initial abnormal feature map to obtain a second image feature further comprises: The initial abnormal feature map is encoded based on the dimension and size of the first image feature to obtain the second image feature, wherein the dimension and size of the second image feature are the same as the dimension and size of the first image feature.
[0040] In some embodiments, encoding the initial abnormal feature map based on the dimension and size of the first image feature to obtain the second image feature further includes: The initial abnormal feature map is sequentially input into the first convolution layer, the first linear layer, the first activation function layer, the second convolution layer, the second linear layer, the second activation function layer, and the third convolution layer to obtain the second image feature, wherein the strides of the first convolution layer, the second convolution layer, and the third convolution layer and the number and size of the convolution kernels are determined according to the dimension and size of the first image feature.
[0041] In some embodiments, performing obstacle detection based on the fusion feature to obtain a detection result further includes: Obtaining a weight based on the fused features; Using the weight to weight the first image feature and the second image feature in the fusion feature to obtain the weighted fusion feature; Obstacle detection is performed based on the weighted fusion features to obtain a detection result.
[0042] In some embodiments, obtaining weights based on the fused features further includes: Inputting the fused features into a weight filter, wherein the weight filter comprises a multi-layer convolutional network and a softmax layer; The weight is obtained by calculating the fusion feature using the multi-layer convolutional network and the softmax layer in the weight filter.
[0043] In some embodiments, performing obstacle detection based on the fusion feature to obtain a detection result further includes: Obtaining multiple obstacle masks based on the fused features; Preprocessing the plurality of obstacle masks to obtain a final obstacle mask; The detection result is obtained based on the final obstacle mask.
[0044] In some embodiments, preprocessing the plurality of obstacle masks to obtain a final obstacle mask further comprises: Non-maximum suppression is used to perform a redundancy removal operation on the plurality of obstacle masks.
[0045] Based on the same inventive concept, according to another aspect of the present invention, Fig.10 As shown, an embodiment of the present invention further provides a computer device 501, including: at least one processor 520; and The memory 510 stores a computer program 511 that can be run on the processor. When the processor 520 executes the program, the steps of any one of the above obstacle detection methods are performed.
[0046] Based on the same inventive concept, according to another aspect of the present invention, Fig.11 As shown, an embodiment of the present invention further provides a computer-readable storage medium 601, which stores a computer program 610. When the computer program 610 is executed by a processor, the steps of any of the above obstacle detection methods are performed.
[0047] Finally, it should be noted that a person skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods.
[0048] Furthermore, it should be appreciated that the computer-readable storage medium (eg, memory) herein may be either volatile memory or nonvolatile memory, or may include both volatile memory and nonvolatile memory.
[0049] It will also be appreciated by those skilled in the art that various exemplary logic blocks, modules, circuits and algorithm steps described in conjunction with the disclosure herein can be implemented as electronic hardware, computer software or a combination of the two. In order to clearly illustrate this interchangeability of hardware and software, a general description has been given to the functions of various schematic components, blocks, modules, circuits and steps. Whether this function is implemented as software or hardware depends on specific applications and the design constraints imposed on the entire system. Those skilled in the art can implement the function in various ways for each specific application, but this implementation decision should not be interpreted as causing a departure from the disclosed scope of the embodiments of the present invention.
[0050] The above are exemplary embodiments disclosed in the present invention, but it should be noted that various changes and modifications may be made without departing from the scope disclosed in the embodiments of the present invention as defined in the claims. The functions, steps and / or actions of the method claims according to the disclosed embodiments described herein do not need to be performed in any particular order. In addition, although the elements disclosed in the embodiments of the present invention may be described or required in individual form, they may also be understood as multiple unless explicitly limited to the singular.
[0051] It should be understood that, as used herein, the singular forms "a", "an" are intended to include the plural forms as well, unless the context clearly supports an exception. It should also be understood that, as used herein, "and / or" refers to any and all possible combinations including one or more of the associated listed items.
[0052] The serial numbers of the embodiments disclosed in the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0053] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0054] A person skilled in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the disclosure of the embodiments of the present invention (including the claims) is limited to these examples; under the concept of the embodiments of the present invention, the technical features in the above embodiments or different embodiments can also be combined, and there are many other changes in different aspects of the above embodiments of the present invention, which are not provided in detail for the sake of simplicity. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention should be included in the protection scope of the embodiments of the present invention.
Claims
1. An obstacle detection method, characterized in that: The following steps are involved: Encoding the image to be detected to obtain a first image feature, and performing obstacle detection on the image to be detected to obtain an initial abnormal feature map; Encoding the initial abnormal feature map to obtain a second image feature; The first image feature and the second image feature are concatenated to obtain a fusion feature; Obstacle detection is performed based on the fusion features to obtain a detection result.
2. The method according to claim 1, characterized in that Encoding the initial abnormal feature map to obtain a second image feature further includes: The initial abnormal feature map is encoded based on the dimension and size of the first image feature to obtain the second image feature, wherein the dimension and size of the second image feature are the same as the dimension and size of the first image feature.
3. The method according to claim 2, characterized in that Encoding the initial abnormal feature map based on the dimension and size of the first image feature to obtain the second image feature further includes: The initial abnormal feature map is sequentially input into the first convolution layer, the first linear layer, the first activation function layer, the second convolution layer, the second linear layer, the second activation function layer, and the third convolution layer to obtain the second image feature, wherein the strides of the first convolution layer, the second convolution layer, and the third convolution layer and the number and size of the convolution kernels are determined according to the dimension and size of the first image feature.
4. The method according to claim 1, characterized in that Performing obstacle detection based on the fusion feature to obtain a detection result further includes: Obtaining a weight based on the fused features; Using the weight to weight the first image feature and the second image feature in the fusion feature to obtain the weighted fusion feature; Obstacle detection is performed based on the weighted fusion features to obtain a detection result.
5. The method according to claim 4, characterized in that Obtaining weights based on the fusion features further includes: Inputting the fused features into a weight filter, wherein the weight filter comprises a multi-layer convolutional network and a softmax layer; The weight is obtained by calculating the fusion feature using the multi-layer convolutional network and the softmax layer in the weight filter.
6. The method according to claim 1, characterized in that Performing obstacle detection based on the fusion feature to obtain a detection result further includes: Obtaining multiple obstacle masks based on the fused features; Preprocessing the plurality of obstacle masks to obtain a final obstacle mask; The detection result is obtained based on the final obstacle mask.
7. The method according to claim 6, characterized in that Preprocessing the plurality of obstacle masks to obtain a final obstacle mask further includes: Non-maximum suppression is used to perform a redundancy removal operation on the plurality of obstacle masks.
8. An obstacle detection system, characterized in that: include: A calculation module is configured to encode the image to be detected to obtain a first image feature, and perform obstacle detection on the image to be detected to obtain an initial abnormal feature map; an alignment module, configured to encode the initial abnormal feature map to obtain a second image feature; A splicing module configured to splice the first image feature and the second image feature to obtain a fusion feature; The detection module is configured to perform obstacle detection based on the fusion features to obtain a detection result.
9. A computer device comprising: at least one processor; as well as A memory storing a computer program executable on the processor, wherein the processor executes the steps of the method according to any one of claims 1 to 7 when executing the program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are performed.