Obstacle detection and model training method and device, equipment, chip and medium

By using obstacle detection and model training methods in an autonomous driving environment, combining multi-view image, historical detection information and global semantic information, the obstacle detection model is trained, and the problem of insufficient accuracy of dynamic obstacle detection is solved, achieving higher detection accuracy and robustness.

CN120088760APending Publication Date: 2025-06-03XIAOMI EV TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510238149.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In autonomous driving environments, it is difficult for the prior art to effectively detect and identify dynamically changing obstacles, resulting in insufficient accuracy and robustness of obstacle detection.

Method used

An obstacle detection and model training method is proposed. By obtaining images collected by multiple sample cameras, obstacle detection information and image global semantic information at historical moments, the obstacle detection model is trained so that it can learn the dynamic change characteristics of obstacles.

Benefits of technology

It significantly improves the accuracy, robustness and generalization capabilities of the obstacle detection model, improves the accuracy and reliability of obstacle detection, and thus enhances driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088760A_ABST
    Figure CN120088760A_ABST
Patent Text Reader

Abstract

The invention provides an obstacle detection and model training method and device, equipment, a chip and a medium, and the method comprises the steps: obtaining training data; wherein the training data comprises first images collected by a plurality of sample cameras at a first moment, first detection information and global semantic information; the first detection information is obtained by performing obstacle detection on second images acquired by the plurality of sample cameras at a second moment before the first moment; the global semantic information is obtained by performing semantic extraction on the plurality of first images; inputting the training data into an initial obstacle detection model for obstacle detection to obtain second detection information; and according to the lane line marking information associated with the plurality of first images and the second detection information, the obstacle detection model is trained, so that the accuracy, robustness and generalization ability of the obstacle detection model are significantly improved, and the accuracy of obstacle detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of autonomous driving technology, and particularly to an obstacle detection and model training method, apparatus, device, chip and medium. Background Art

[0002] In an autonomous driving environment, a vehicle must have the ability to accurately perceive the surrounding environment and be able to accurately detect all static or dynamic obstacles that may affect the driving path, such as pedestrians, cyclists or other vehicles. By real-time and accurately identifying these obstacles, the vehicle can discover potential risks and dangerous situations in advance, make a quick response, and take necessary obstacle avoidance measures, thereby effectively avoiding accidents and ensuring the safety of passengers and other road users. Summary of the Invention

[0003] The present disclosure aims to solve one of the technical problems in the related art to a certain extent.

[0004] To this end, the present disclosure provides an obstacle detection and model training method, apparatus, device, chip and medium to enable an obstacle detection model to learn the dynamic change characteristics of obstacles, improve the accuracy, robustness and generalization ability of the obstacle detection model, and thus improve the accuracy of obstacle detection.

[0005] An embodiment of one aspect of the present disclosure provides a method for training an obstacle detection model, including:

[0006] Obtaining training data; wherein, the training data includes: first images, first detection information and global semantic information collected by a plurality of sample cameras at a first moment; the first detection information is obtained by performing obstacle detection on second images collected by the plurality of sample cameras at a second moment before the first moment; the global semantic information is obtained by performing semantic extraction on a plurality of the first images;

[0007] Inputting the training data into an initial obstacle detection model for obstacle detection to obtain second detection information;

[0008] Training the obstacle detection model according to lane line annotation information associated with a plurality of the first images and the second detection information.

[0009] An embodiment of another aspect of the present disclosure provides an obstacle detection method, including:

[0010] Obtaining target images collected by a plurality of vehicle-mounted cameras at a target moment;

[0011] Obtaining historical detection information; wherein, the historical detection information is obtained by performing obstacle detection on historical images collected by the plurality of vehicle-mounted cameras at a historical moment before the target moment;

[0012] Obtain the global semantic information of the image; wherein, the global semantic information of the image is obtained by performing semantic extraction on a plurality of the target images;

[0013] Input a plurality of the target images, the historical detection information, and the global semantic information of the image into an obstacle detection model trained by the method described in an embodiment of one aspect to perform obstacle detection, so as to obtain target detection information.

[0014] An embodiment of another aspect of the present disclosure provides a training device for an obstacle detection model, including:

[0015] An acquisition module, configured to acquire training data; wherein, the training data includes: first images collected by a plurality of sample cameras at a first moment, first detection information, and global semantic information; the first detection information is obtained by performing obstacle detection on second images collected by the plurality of sample cameras at a second moment before the first moment; the global semantic information is obtained by performing semantic extraction on a plurality of the first images;

[0016] An input module, configured to input the training data into an initial obstacle detection model to perform obstacle detection, so as to obtain second detection information;

[0017] A training module, configured to train the obstacle detection model according to lane line annotation information associated with a plurality of the first images and the second detection information.

[0018] An embodiment of still another aspect of the present disclosure provides an obstacle detection device, including:

[0019] A first acquisition module, configured to acquire target images collected by a plurality of vehicle-mounted cameras at a target moment;

[0020] A second acquisition module, configured to acquire historical detection information; wherein, the historical detection information is obtained by performing obstacle detection on historical images collected by the plurality of vehicle-mounted cameras at a historical moment before the target moment;

[0021] A third acquisition module, configured to acquire global semantic information of the image; wherein, the global semantic information of the image is obtained by performing semantic extraction on a plurality of the target images;

[0022] A detection module, configured to input a plurality of the target images, the historical detection information, and the global semantic information of the image into an obstacle detection model trained by the device described in an embodiment of another aspect to perform obstacle detection, so as to obtain target detection information.

[0023] In another aspect of the present disclosure, an embodiment provides an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method as described in an embodiment of one aspect, or implement the method as described in an embodiment of another aspect.

[0024] In another aspect of the present disclosure, an embodiment provides a chip, the chip includes a processing circuit configured to execute the method as described in an embodiment of one aspect, or execute the method as described in an embodiment of another aspect.

[0025] In another aspect of the present disclosure, an embodiment provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the method as described in an embodiment of one aspect, or implement the method as described in an embodiment of another aspect when executed by a processor.

[0026] In another aspect of the present disclosure, an embodiment provides a computer program product having a computer program stored thereon, and the program implements the method as described in an embodiment of one aspect, or implements the method as described in an embodiment of another aspect when executed by a processor.

[0027] The obstacle detection and model training method, apparatus, device, chip and medium proposed by the present disclosure collect training data, where the training data includes first images collected by multiple sample cameras at a first moment, obstacle detection information (i.e., first detection information) corresponding to second images collected by multiple cameras at a second moment (i.e., before the first moment), and global semantic information obtained by semantic extraction of multiple first images. This realizes that the training data not only contains image information at the first moment, but also incorporates the obstacle detection results at historical moments, which helps the obstacle detection model learn the dynamic change characteristics of obstacles. The global semantic information helps the obstacle detection model more accurately identify and classify obstacles in a complex road environment. Then, during the training process, the training data is input into an initial obstacle detection model for obstacle detection to obtain second detection information. Subsequently, the model is iteratively optimized using the lane line annotation information associated with multiple first images and the second detection information, significantly improving the accuracy, robustness and generalization ability of the obstacle detection model, thereby improving the accuracy of obstacle detection.

[0028] Additional aspects and advantages of the present disclosure will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present disclosure. Description of the Drawings

[0029] The above and / or additional aspects and advantages of the present disclosure will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, where:

[0030] Figure 1 It is a schematic flowchart of a method for training an obstacle detection model provided by an embodiment of the present disclosure;

[0031] Figure 2 It is a schematic flowchart of another method for training an obstacle detection model provided by an embodiment of the present disclosure;

[0032] Figure 3 It is a schematic flowchart of another method for training an obstacle detection model provided by an embodiment of the present disclosure;

[0033] Figure 4 It is a schematic flowchart of another method for training an obstacle detection model provided by an embodiment of the present disclosure;

[0034] Figure 5 It is a schematic diagram of the principle of a method for training an obstacle detection model provided by an embodiment of the present disclosure;

[0035] Figure 6 It is a schematic flowchart of an obstacle detection method provided by an embodiment of the present disclosure;

[0036] Figure 7 It is a schematic structural diagram of a device for training an obstacle detection model provided by an embodiment of the present disclosure;

[0037] Figure 8 It is a schematic structural diagram of an obstacle detection device provided by an embodiment of the present disclosure;

[0038] Figure 9 It is a block diagram of an electronic device provided by an embodiment of the present disclosure;

[0039] Figure 10 It is a schematic structural diagram of a chip proposed by an embodiment of the present disclosure. Detailed Description of the Embodiment

[0040] The embodiments of the present disclosure will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, but should not be construed as limiting the present disclosure.

[0041] In the related art, the obstacle recognition algorithm is sparse obstacle recognition based on query information. Through the learning of the model, the query contains the attributes, location information, etc. of the obstacles. Furthermore, the similarity between the query of the obstacles in the T-th frame and the query of the T-1-th frame can be calculated as the similarity for Hungarian matching, which is often used in 2D detection algorithms. However, in the sparse 3D detection algorithm based on query, the query represents the attributes and location of the obstacles. However, the location information in the query changes with the ego vehicle, and the relative position of the query to the ego vehicle is also different, which may lead to a decrease in the similarity of the queries of the obstacles between different frames (e.g., the T-th frame and the T-1-th frame).

[0042] In view of the above problems, the present disclosure proposes an obstacle detection and model training method, apparatus, device, chip and medium.

[0043] The following describes the obstacle detection and model training method, apparatus, device, chip and medium according to the embodiments of the present disclosure with reference to the accompanying drawings.

[0044] Figure 1 It is a schematic flowchart of a training method for an obstacle detection model provided by an embodiment of the present disclosure.

[0045] In this embodiment, the training method of the obstacle detection model is exemplified as being configured in a training apparatus for the obstacle detection model, and the training apparatus for the obstacle detection model can be set on an electronic device.

[0046] Among them, the electronic device can be a mobile terminal, an Internet of Things (IOT) device, a vehicle, etc. The mobile terminal is, for example, a hardware device with various operating systems such as an in-vehicle device, a mobile phone, a watch, a wearable device, a tablet computer, a personal digital assistant, etc.

[0047] As Figure 1 shown, the training method of the obstacle detection model may include the following steps:

[0048] Step 101, obtain training data.

[0049] Among them, the training data includes: first images, first detection information, and global semantic information collected by a plurality of sample cameras at a first moment; the first detection information is obtained by performing obstacle detection on second images collected by the plurality of sample cameras at a second moment before the first moment; the global semantic information is obtained by performing semantic extraction on the plurality of first images.

[0050] In order to train an efficient and accurate obstacle detection model, in the embodiments of the present disclosure, training data including multi-view images, obstacle detection results at historical moments, and global semantic information is used to train the obstacle detection model.

[0051] As an example, first images, first detection information, and global semantic information collected by multiple sample cameras at a first moment are obtained. Among them, the first images can be images captured by multiple cameras (i.e., multiple sample cameras) installed on a vehicle at the first moment (e.g., a specified moment); each sample camera covers a different field of view, and these first images provide a multi-view view of the surrounding environment; the first detection information is the result obtained by performing obstacle detection on the images captured by the multiple sample cameras at a certain historical moment (e.g., a second moment before the first moment) before the first moment.

[0052] It should be noted that if the first moment is the first moment, the first detection information can be generated by the perception device of the vehicle according to the perception information at the first moment; in addition, training the obstacle detection model based on the first detection information can help the obstacle detection model understand the historical behavior patterns of obstacles. Based on the data of the historical behavior patterns, the obstacle detection model can establish stronger correlations between consecutive frames, so as to more accurately re-identify and continuously track the same obstacle in different frames, which can enhance the re-identification ability of the obstacle detection model, thereby improving the similarity of obstacle detection results between different frames; the global semantic information is obtained by performing semantic extraction from multiple first images, and the global semantic information helps the obstacle detection model to more comprehensively understand the image content, thereby improving the accuracy and robustness of obstacle recognition.

[0053] Step 102: Input the training data into an initial obstacle detection model for obstacle detection to obtain second detection information.

[0054] In the embodiments of the present disclosure, the initial obstacle detection model can be an untrained obstacle detection model, and the purpose of this obstacle detection model is to identify and classify obstacles in images and predict their attributes such as position and speed.

[0055] Furthermore, input the training data into the initial obstacle detection model. Inside the obstacle detection model, the self-attention mechanism and the cross-attention mechanism are used to process the input data, so as to effectively integrate the features of the training data and generate accurate obstacle detection results.

[0056] Step 103: Train the obstacle detection model according to the lane line annotation information associated with multiple first images and the second detection information.

[0057] In order to improve the understanding ability and prediction accuracy of the obstacle detection model, in the embodiments of the present disclosure, the obstacle detection model is trained by combining lane line annotation information and second detection information.

[0058] It should be noted that, in order to improve the accuracy of the obstacle detection model, the lane line annotation information is generated based on multiple first images.

[0059] In summary, by collecting training data, which includes first images collected by multiple sample cameras at a first moment, obstacle detection information (i.e., first detection information) corresponding to second images collected by multiple cameras at a second moment (i.e., before the first moment), and global semantic information obtained by semantic extraction of multiple first images, it is achieved that the training data not only contains the image information at the first moment but also incorporates the obstacle detection results at historical moments, which helps the obstacle detection model learn the dynamic change characteristics of obstacles, and the global semantic information helps the obstacle detection model more accurately identify and classify obstacles in a complex road environment; then, during the training process, the training data is input into an initial obstacle detection model for obstacle detection to obtain second detection information; subsequently, using the lane line annotation information associated with multiple first images and the second detection information to iteratively optimize the model, significantly improving the accuracy, robustness, and generalization ability of the obstacle detection model, thereby improving the accuracy of obstacle detection.

[0060] To clearly illustrate how the obstacle detection model is trained according to the lane line annotation information associated with multiple first images and the second detection information in the above embodiments, the present disclosure proposes another training method for the obstacle detection model.

[0061] Figure 2 It is a schematic flowchart of another training method for the obstacle detection model provided by the embodiments of the present disclosure.

[0062] As Figure 2 shown, the training of the obstacle detection model may include the following steps:

[0063] Step 201, obtain training data.

[0064] Among them, the training data includes: first images collected by multiple sample cameras at a first moment, first detection information, and global semantic information; the first detection information is obtained by performing obstacle detection on second images collected by multiple sample cameras at a second moment before the first moment; the global semantic information is obtained by performing semantic extraction on multiple first images.

[0065] Step 202, input the training data into an initial obstacle detection model for obstacle detection to obtain second detection information.

[0066] Step 203: Predict the driving trajectories of obstacles based on the lane line annotation information and the second detection information, so as to obtain the predicted driving trajectories of at least one first obstacle.

[0067] In the embodiments of the present disclosure, the lane line annotation information refers to the lines that mark different lanes on the road, which are used to indicate the driving direction of the vehicle and the lane division. The lane line annotation information helps the vehicle understand its position on the road and the drivable paths. The second detection information is the result obtained by detecting obstacles based on the first images collected by multiple sample cameras at the first moment. The second detection information includes information such as the position, category (such as pedestrians, cyclists, other vehicles), speed, and direction of the obstacles. Furthermore, by combining the lane line annotation information associated with the multiple first images and the states of the obstacles in the second detection information, the predicted driving trajectories of at least one first obstacle within a future period of time can be predicted.

[0068] Among them, in order to improve the accuracy of the lane line annotation information, the lane line annotation information is obtained by the following steps:

[0069] 1. Perform spatial alignment on multiple first images to obtain multiple aligned first images;

[0070] In the embodiments of the present disclosure, by performing spatial alignment on multiple first images, the spatial consistency of the first images taken from different perspectives is ensured, which can reduce the errors caused by perspective differences, thereby improving the accuracy of lane line annotation.

[0071] 2. Extract features from each aligned first image to obtain the key features of each aligned first image;

[0072] In order to extract features from the image that can describe key information such as lane lines, in the embodiments of the present disclosure, features are extracted from each aligned first image to obtain the key features of each aligned first image.

[0073] 3. Fuse multiple key features to obtain a second fused feature.

[0074] In order to obtain a more complete and comprehensive lane line description information, in the embodiments of the present disclosure, multiple key features are fused to obtain a second fused feature with more comprehensive information.

[0075] 4. Perform lane line annotation according to the second fused feature to obtain lane line annotation information.

[0076] In order to accurately generate lane line annotation information, in the embodiments of the present disclosure, lane line annotation is performed based on the second fused feature to obtain lane line annotation information.

[0077] As an example, obtain a trained trajectory prediction model; input lane line annotation information and second detection information into the trajectory prediction model to obtain the predicted driving trajectories of at least one first obstacle output by the trajectory prediction model.

[0078] That is to say, in order to improve the accuracy of obstacle driving trajectory prediction, in the embodiments of the present disclosure, lane line annotation information and second detection information are input into the trained trajectory prediction model to obtain the predicted driving trajectories of at least one first obstacle output by the trajectory prediction model.

[0079] Step 204, train the obstacle detection model according to the predicted driving trajectories of at least one first obstacle.

[0080] In order to effectively train the obstacle detection model, in the embodiments of the present disclosure, by comparing the predicted driving trajectories of at least one first obstacle output by the trajectory prediction model with the annotated driving trajectories of at least one second obstacle actually annotated in a plurality of first images, the difference between the predicted driving trajectories of at least one first obstacle and the annotated driving trajectories of at least one second obstacle is obtained, and the obstacle detection model is optimized based on this difference; that is, by comparing the predicted driving trajectories and the annotated driving trajectories, the errors of the obstacle detection model in identifying obstacles and key parameters such as the position and speed of obstacles can be found, and these feedback information are used to adjust the model parameters to realize the training of the obstacle detection model.

[0081] As an example, obtain the annotated driving trajectories of at least one second obstacle in a plurality of first images; generate a loss function value according to the difference between the predicted driving trajectories of at least one first obstacle and the annotated driving trajectories of at least one second obstacle; train the obstacle detection model according to the loss function value.

[0082] That is to say, compare the predicted driving trajectories of at least one first obstacle with the known annotated driving trajectories of at least one second obstacle. Furthermore, based on the difference between the predicted driving trajectories of at least one first obstacle and the annotated driving trajectories of at least one second obstacle, generate a loss function value. Furthermore, based on the loss function value, train the obstacle detection model.

[0083] Among them, the differences between the predicted driving trajectories of at least one first obstacle and the labeled driving trajectories of at least one second obstacle may include at least one of the following: the differences between the predicted driving trajectories and the labeled driving trajectories of the same obstacle among at least one first obstacle and at least one second obstacle, the obstacle trajectories that exist in the predicted driving trajectories but do not exist in the labeled driving trajectories, and the obstacle trajectories that exist in the labeled driving trajectories but do not exist in the predicted driving trajectories. Among them, for the differences between the predicted driving trajectories and the labeled driving trajectories of the same obstacle, they can be determined by directly comparing the trajectory shapes, speeds, directions, etc. between the two; for the obstacle trajectories that exist in the predicted driving trajectories but do not exist in the labeled driving trajectories, it means that the obstacle detection model may have produced a false alarm, that is, predicted an obstacle that actually does not exist, and the difference is reflected in the inconsistency between the predicted driving trajectory and the actual situation (the labeled driving trajectory without the false alarm obstacle); for the obstacle trajectories that exist in the labeled driving trajectories but do not exist in the predicted driving trajectories, it means that the obstacle detection model may have produced a missed detection, that is, failed to detect an actually existing obstacle, and the difference is reflected in the inconsistency between the labeled driving trajectory and the predicted driving trajectory (the predicted driving trajectory without the missed detection obstacle).

[0084] It should be noted that the execution processes of steps 201 to 202 can be implemented in any one of the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be elaborated further.

[0085] In summary, based on the lane line annotation information and the second detection information, the driving trajectories of obstacles are predicted to obtain the predicted driving trajectories of at least one first obstacle; based on the predicted driving trajectories of at least one first obstacle, the obstacle detection model is trained. Thus, by combining the lane line annotation information and the second detection information to predict the driving trajectories of obstacles and training the obstacle detection model according to the predicted driving trajectories of obstacles, the accuracy of the obstacle detection model is improved, enabling the obstacle detection model to learn more complex environmental features and enhancing the generalization ability of the obstacle detection model in different driving scenarios.

[0086] To clearly illustrate how the second detection information is generated in the above embodiments, the present disclosure proposes another training method for an obstacle detection model.

[0087] Figure 3 It is a schematic flowchart of another training method for an obstacle detection model provided by an embodiment of the present disclosure.

[0088] As Figure 3 shown, the training method of the obstacle detection model may include the following steps:

[0089] Step 301, obtain training data.

[0090] Among them, the training data includes: first images collected by multiple sample cameras at a first moment, first detection information, and global semantic information; the first detection information is obtained by performing obstacle detection on second images collected by multiple sample cameras at a second moment before the first moment; the global semantic information is obtained by performing semantic extraction on multiple first images.

[0091] Step 302: Input the training data into an initial obstacle detection model, and use the self-attention layer in the obstacle detection model to extract features from the training data to obtain multiple sample features; use the cross-attention layer in the obstacle detection model to perform feature fusion on the multiple sample features to obtain a first fused feature; use the prediction layer in the obstacle detection model to perform obstacle detection on the first fused feature to obtain second detection information.

[0092] In order to improve the accuracy and robustness of the obstacle detection model, in the embodiments of the present disclosure, the obstacle detection model includes at least a self-attention layer, a cross-attention layer, and a prediction layer. Among them, the self-attention layer is used to extract key features from the input data, the cross-attention layer is used to integrate key features from different sources, and the prediction layer is used to perform obstacle detection based on the fused features to identify and classify obstacles in the image.

[0093] As an example, the self-attention layer in the obstacle detection model extracts features from the training data to obtain multiple sample features. Further, the cross-attention layer in the obstacle detection model performs feature fusion on the multiple sample features to obtain a first fused feature. Finally, the prediction layer in the obstacle detection model performs obstacle detection on the first fused feature to obtain second detection information.

[0094] Step 303: Train the obstacle detection model according to the lane line annotation information associated with the multiple first images and the second detection information.

[0095] It should be noted that the execution process of step 303 can be implemented in any one of the embodiments of the present disclosure. The embodiments of the present disclosure do not limit this and will not be elaborated further.

[0096] In summary, through the self-attention layer in the obstacle detection model, key features are effectively extracted from the training data, and noise interference is effectively reduced. The cross-attention layer further integrates the information between different features, enhancing the model's recognition ability for complex scenarios. Finally, based on these fused feature information, the prediction layer can more accurately detect obstacles, improving the accuracy and robustness of obstacle detection.

[0097] To clearly illustrate how the above embodiments generate global semantic information, the present disclosure provides a schematic flowchart of another method for training an obstacle detection model.

[0098] Figure 4 Schematic flowchart of another method for training an obstacle detection model provided by an embodiment of the present disclosure.

[0099] As Figure 4 shown, the method for training the obstacle detection model may include the following steps:

[0100] Step 401, obtain first images and first detection information collected by a plurality of sample cameras at a first moment.

[0101] Step 402, perform semantic feature extraction on each first image to obtain the global semantic feature of each first image.

[0102] In order to accurately obtain the global semantic information associated with a plurality of first images, in an embodiment of the present disclosure, by performing semantic feature extraction on each first image and integrating its global semantic feature, global semantic information associated with a plurality of first images is finally generated.

[0103] As an example, for any one of the first images, perform global semantic feature extraction on the any one of the first images to obtain the global semantic feature of the any one of the first images.

[0104] Step 403, generate a semantic representation for each first image according to the global semantic feature of each first image.

[0105] In order to facilitate the obstacle detection model to accurately perform context understanding, in an embodiment of the present disclosure, based on the global semantic feature of each first image, a semantic representation for each first image is generated. For example, based on a set mapping mechanism, the global semantic feature of each first image is transformed into a semantic representation; wherein, the semantic representation can be in text form (such as labels, descriptions, paragraphs, etc.), or in other forms (such as vectors, graph structures, etc.), and the present disclosure does not make specific limitations.

[0106] Step 404, integrate the semantic representations of each first image to obtain global semantic information associated with a plurality of first images.

[0107] In order to obtain more comprehensive and coherent semantic information, in an embodiment of the present disclosure, the semantic representations of each first image are integrated to obtain global semantic information.

[0108] Step 405, generate training data according to a plurality of first images, first detection information, and global semantic information.

[0109] Step 406: Input the training data into the initial obstacle detection model for obstacle detection to obtain the second detection information.

[0110] Step 407: Train the obstacle detection model according to the lane line annotation information associated with multiple first images and the second detection information.

[0111] It should be noted that the execution processes of Step 401, Steps 405 to 407 can be implemented in any one of the embodiments of the present disclosure respectively. The embodiments of the present disclosure do not make any limitations in this regard and will not be elaborated further.

[0112] In summary, by extracting semantic features from each first image to obtain the global semantic features of each first image; generating semantic representations of each first image according to the global semantic features of each first image; and integrating the semantic representations of each first image to obtain the global semantic information associated with multiple first images. Thus, by extracting and integrating semantic features from the first images obtained from different cameras, a comprehensive global semantic information can be obtained. The global semantic information contains rich environmental information. Training the obstacle detection model based on the global semantic information can improve the accuracy of obstacle model detection and classification.

[0113] In any embodiment of the present disclosure, as Figure 5 shown, the training method of the obstacle detection model in the embodiment of the present disclosure can also be implemented based on the following steps:

[0114] 1. Obtain multi-view images (multiple first images), historical detection information (i.e., the first detection information), and global semantic information;

[0115] 2. Input the multi-view images, historical detection information, and global semantic information into the initial obstacle detection model to obtain obstacle detection information (i.e., the second detection information); wherein, the obstacle detection information includes: obstacle attributes, positions, speeds, etc.;

[0116] 3. Input the obstacle detection information and the lane line annotation information into the trained trajectory prediction model to obtain the predicted driving trajectories of at least one first obstacle;

[0117] 4. Train the obstacle detection model according to the differences between the predicted driving trajectories of at least one first obstacle and the annotated driving trajectories of at least one second obstacle.

[0118] The training method of the obstacle detection model in the embodiment of the present disclosure enables the obstacle detection model to learn the dynamic change characteristics of obstacles, significantly improves the accuracy, robustness, and generalization ability of the obstacle detection model, and thus improves the accuracy of obstacle detection.

[0119] Based on any of the above embodiments, the present disclosure proposes an obstacle detection method.

[0120] Figure 6 FIG. is a schematic flowchart of an obstacle detection method provided by an embodiment of the present disclosure. It should be noted that, for the purpose of illustration, the obstacle detection method of the embodiment of the present disclosure is configured in an obstacle detection device, and the obstacle detection device can be arranged in a vehicle.

[0121] As Figure 6 shown, the obstacle detection method includes the following steps:

[0122] Step 601, obtain target images collected by multiple vehicle-mounted cameras at a target moment.

[0123] In order to obtain multi-view images, in an embodiment of the present disclosure, target images collected by multiple vehicle-mounted cameras at a target moment can be obtained. Each vehicle-mounted camera covers a different field of view and jointly provides 360-degree visual perception ability for the vehicle.

[0124] Step 602, obtain historical detection information.

[0125] The historical detection information is obtained by performing obstacle detection on historical images collected by multiple vehicle-mounted cameras at a historical moment before the target moment.

[0126] In an embodiment of the present disclosure, the historical detection information can be the result of performing obstacle detection on multiple images collected by the vehicle-mounted camera before the target moment. For example, the historical detection information can be obtained by performing obstacle detection on historical images collected by multiple vehicle-mounted cameras at the previous moment of the target moment.

[0127] Step 603, obtain global semantic information of the images.

[0128] The global semantic information of the images is obtained by performing semantic extraction on multiple target images.

[0129] In order to accurately obtain the global semantic information of the images, in an embodiment of the present disclosure, semantic feature extraction is performed on each target image to obtain the global semantic features of each target image; according to the global semantic features of each target image, a semantic representation of each target image is generated; and the semantic representations of each target image are integrated to obtain the global semantic information of the images associated with multiple target images.

[0130] Step 604, input the multiple target images, the historical detection information, and the global semantic information of the images into a trained obstacle detection model for obstacle detection to obtain target detection information.

[0131] Furthermore, input multiple target images, historical detection information, and image global semantic information into a trained obstacle detection model for obstacle detection to obtain target detection information, where the obstacle detection model can be Figures 1 to 5 a model trained by the training method of the obstacle detection model described in any embodiment.

[0132] In the obstacle detection method of the embodiments of the present disclosure, target images collected by multiple vehicle-mounted cameras at a target moment are obtained; historical detection information is obtained; where the historical detection information is obtained by performing obstacle detection on historical images collected by multiple vehicle-mounted cameras at a historical moment before the target moment; image global semantic information is obtained; where the image global semantic information is obtained by performing semantic extraction on multiple target images; multiple target images, historical detection information, and image global semantic information are input into a trained obstacle detection model for obstacle detection to obtain target detection information. Thus, by comprehensively considering the images, historical detection information, and image global semantic information collected by multiple vehicle-mounted cameras at the target moment and inputting them into a trained obstacle detection model for obstacle detection, the accuracy and reliability of obstacle detection can be significantly improved, thereby enhancing driving safety.

[0133] To implement the above Figures 1 to 5 embodiment, the present disclosure proposes a training device for an obstacle detection model.

[0134] Figure 7 FIG. is a schematic structural diagram of a training device for an obstacle detection model provided by an embodiment of the present disclosure.

[0135] As Figure 7 shown, the training device 700 for the obstacle detection model includes: an acquisition module 710, an input module 720, and a training module 730.

[0136] Among them, the acquisition module 710 is used to acquire training data; where the training data includes: first images collected by multiple sample cameras at a first moment, first detection information, and global semantic information; the first detection information is obtained by performing obstacle detection on second images collected by multiple sample cameras at a second moment before the first moment; the global semantic information is obtained by performing semantic extraction on multiple first images; the input module 720 is used to input the training data into an initial obstacle detection model for obstacle detection to obtain second detection information; the training module 730 is used to train the obstacle detection model according to lane line annotation information associated with multiple first images and the second detection information.

[0137] As a possible implementation, the training module 730 is configured to predict the driving trajectories of obstacles based on lane line annotation information and second detection information to obtain the predicted driving trajectories of at least one first obstacle; and train the obstacle detection model according to the predicted driving trajectories of at least one first obstacle.

[0138] As a possible implementation, the training module 730 is configured to obtain a trained trajectory prediction model; input the lane line annotation information and the second detection information into the trajectory prediction model to obtain the predicted driving trajectories of at least one first obstacle output by the trajectory prediction model.

[0139] As a possible implementation, the training module 730 is configured to obtain the annotated driving trajectories of at least one second obstacle in multiple first images; generate a loss function value according to the difference between the predicted driving trajectories of at least one first obstacle and the annotated driving trajectories of at least one second obstacle; and train the obstacle detection model according to the loss function value.

[0140] As a possible implementation, the second detection information is generated by the following modules: a first extraction module, a first fusion module, and a prediction module.

[0141] Among them, the first extraction module is configured to extract features from the training data by using the self-attention layer in the obstacle detection model to obtain multiple sample features; the first fusion module is configured to perform feature fusion on the multiple sample features by using the cross-attention layer in the obstacle detection model to obtain a first fusion feature; the prediction module is configured to perform obstacle detection on the first fusion feature by using the prediction layer in the obstacle detection model to obtain the second detection information.

[0142] As a possible implementation, the global semantic information is generated by the following modules: a second extraction module, a generation module, and a second fusion module.

[0143] Among them, the second extraction module is configured to perform semantic feature extraction on each first image to obtain the global semantic features of each first image; the generation module is configured to generate the semantic representations of each first image according to the global semantic features of each first image; the second fusion module is configured to integrate the semantic representations of each first image to obtain the global semantic information associated with multiple first images.

[0144] As a possible implementation, the lane line annotation information is obtained by the following modules: an alignment module, a third extraction module, a third fusion module, and an annotation module.

[0145] Among them, the alignment module is used to spatially align multiple first images to obtain multiple aligned first images; the third extraction module is used to extract features from each aligned first image to obtain the key features of each aligned first image; the third fusion module is used to fuse multiple key features to obtain a second fusion feature; the annotation module is used to perform lane line annotation based on the second fusion feature to obtain the lane line annotation information.

[0146] The training device for the obstacle detection model implemented in this disclosure obtains training data, where the training data includes first images collected by multiple sample cameras at a first moment, obstacle detection information (i.e., first detection information) corresponding to second images collected by multiple cameras at a second moment (i.e., before the first moment), and global semantic information obtained by performing semantic extraction on multiple first images. This realizes that the training data not only contains the image information at the first moment but also incorporates the obstacle detection results at historical moments, which helps the obstacle detection model learn the dynamic change characteristics of obstacles. The global semantic information helps the obstacle detection model more accurately identify and classify obstacles in a complex road environment. Then, during the training process, the training data is input into the initial obstacle detection model for obstacle detection to obtain second detection information. Subsequently, using the lane line annotation information associated with multiple first images and the second detection information to iteratively optimize the model significantly improves the accuracy, robustness, and generalization ability of the obstacle detection model, thereby improving the accuracy of obstacle detection.

[0147] To implement the above Figure 6 embodiment, this disclosure proposes an obstacle detection device.

[0148] Figure 8 It is a schematic structural diagram of an obstacle detection device provided by an embodiment of this disclosure.

[0149] As Figure 8 shown, the obstacle detection device 800 includes: a first acquisition module 810, a second acquisition module 820, a third acquisition module 830, and a detection module 840.

[0150] Among them, the first acquisition module 810 is used to acquire target images collected by multiple vehicle-mounted cameras at a target moment; the second acquisition module 820 is used to acquire historical detection information, where the historical detection information is obtained by performing obstacle detection on historical images collected by multiple vehicle-mounted cameras at a historical moment before the target moment; the third acquisition module 830 is used to acquire image global semantic information, where the image global semantic information is obtained by performing semantic extraction on multiple target images; the detection module 840 is used to input multiple target images, historical detection information, and image global semantic information into Figure 7The obstacle detection model trained by the device of the embodiment is used to detect obstacles to obtain target detection information.

[0151] The obstacle detection device according to the embodiment of the present disclosure obtains target images collected by a plurality of vehicle-mounted cameras at a target time; obtains historical detection information, where the historical detection information is obtained by detecting obstacles in historical images collected by the plurality of vehicle-mounted cameras at a historical time before the target time; obtains global image semantic information, where the global image semantic information is obtained by semantic extraction of the plurality of target images; inputs the plurality of target images, the historical detection information, and the global image semantic information into a trained obstacle detection model for obstacle detection to obtain target detection information. Thus, by comprehensively integrating the images, historical detection information, and global image semantic information collected by the plurality of vehicle-mounted cameras at the target time and inputting them into the trained obstacle detection model for obstacle detection, the accuracy and reliability of obstacle detection can be significantly improved, thereby enhancing driving safety.

[0152] To implement the above embodiment, the present disclosure also proposes an electronic device, including a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the Figures 1 to 5 training method of the obstacle detection model as described in the foregoing Figure 6 embodiment, or implement the

[0153] obstacle detection method as described in the foregoing Figures 1 to 5 embodiment. Figure 6 To implement the above embodiment, the present disclosure also proposes a chip, where the chip includes a processing circuit configured to execute the

[0154] training method of the obstacle detection model as described in the foregoing Figures 1 to 5 embodiment, or execute the Figure 6 obstacle detection method as described in the foregoing

[0155] embodiment. Figures 1 to 5 To implement the above embodiment, the present disclosure also proposes a computer-readable storage medium storing computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the Figure 6 training method of the obstacle detection model as described in the foregoing

[0156] Figure 9 A block diagram of an electronic device provided by an embodiment of the present disclosure. For example, the electronic device 900 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0157] Referring to Figure 9 , the electronic device 900 may include one or more of the following components: a processing component 902, a memory 904, a power component 906, a multimedia component 908, an audio component 910, an input / output (I / O) interface 912, a sensor component 914, and a communication component 916.

[0158] The processing component 902 generally controls the overall operation of the electronic device 900, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 902 may include one or more processors 920 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 902 may include one or more modules to facilitate the interaction between the processing component 902 and other components. For example, the processing component 902 may include a multimedia module to facilitate the interaction between the multimedia component 908 and the processing component 902.

[0159] The memory 904 is configured to store various types of data to support the operation of the electronic device 900. Examples of these data include instructions for any application or method operating on the electronic device 900, contact data, phone book data, messages, pictures, videos, etc. The memory 904 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0160] The power component 906 provides power to various components of the electronic device 900. The power component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 900.

[0161] The multimedia component 908 includes a screen that provides an output interface between the electronic device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 908 includes a front camera and / or a rear camera. When the electronic device 900 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0162] The audio component 910 is configured to output and / or input audio signals. For example, the audio component 910 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 900 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 904 or transmitted via the communication component 916. In some embodiments, the audio component 910 further includes a speaker for outputting audio signals.

[0163] The I / O interface 912 provides an interface between the processing component 902 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.

[0164] The sensor component 914 includes one or more sensors for providing status assessments of various aspects of the electronic device 900. For example, the sensor component 914 can detect the on / off state of the electronic device 900, the relative positioning of components, such as the display and the keypad of the electronic device 900. The sensor component 914 can also detect a change in the position of the electronic device 900 or a component of the electronic device 900, the presence or absence of user contact with the electronic device 900, the orientation or acceleration / deceleration of the electronic device 900, and a change in the temperature of the electronic device 900. The sensor component 914 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 914 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 914 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0165] The communication component 916 is configured to facilitate communication between the electronic device 900 and other devices in a wired or wireless manner. The electronic device 900 can access a communication standard-based wireless network, such as WiFi, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 916 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0166] In an exemplary embodiment, the electronic device 900 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above methods.

[0167] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 904 including instructions, and the above instructions can be executed by a processor 920 of the electronic device 900 to complete the above methods. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0168] Figure 10 is a schematic structural diagram of a chip proposed by an embodiment of the present disclosure. Reference can be made to Figure 10 the schematic structural diagram of the chip 1000 shown, but not limited thereto.

[0169] The chip 1000 includes a processing circuit 1001, and the processing circuit 1001 is configured to execute any of the above methods.

[0170] In some embodiments, the chip 1000 further includes one or more interface circuits 1002. Optionally, the interface circuit 1002 is connected to the memory 1003. The interface circuit 1002 can be used to receive signals from the memory 1003 or other devices, and the interface circuit 1002 can be used to send signals to the memory 1003 or other devices. For example, the interface circuit 1002 can read instructions stored in the memory 1003 and send the instructions to the processing circuit 1001.

[0171] In some embodiments, the interface circuit 1002 executes at least one of the communication steps such as sending and / or receiving in the above methods, and the processing circuit 1001 executes other steps.

[0172] In some embodiments, terms such as interface circuit, interface, transceiver pin, transceiver, etc. may be used interchangeably.

[0173] In some embodiments, chip 1000 further includes one or more memories 1003 for storing instructions. Optionally, all or part of memories 1003 may be outside chip 1000.

[0174] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0175] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0176] Any process or method description in a flowchart or described in other ways herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present disclosure pertain.

[0177] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.

[0178] It should be understood that various parts of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0179] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0180] In addition, in each embodiment of the present disclosure, each functional unit may be integrated into a processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0181] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present disclosure.

Claims

1. A training method for an obstacle detection model, characterized in that: include: Acquire training data; wherein the training data includes: a first image captured by a plurality of sample cameras at a first moment, first detection information, and global semantic information; the first detection information is obtained by performing obstacle detection on a second image captured by the plurality of sample cameras at a second moment before the first moment; the global semantic information is obtained by performing semantic extraction on the plurality of first images; Inputting the training data into an initial obstacle detection model to perform obstacle detection to obtain second detection information; The obstacle detection model is trained according to the lane line marking information associated with the plurality of the first images and the second detection information.

2. The method according to claim 1, characterized in that The training of the obstacle detection model according to the lane line marking information associated with the plurality of the first images and the second detection information includes: Predicting a driving trajectory of an obstacle according to the lane line marking information and the second detection information to obtain a predicted driving trajectory of at least one first obstacle; The obstacle detection model is trained according to the predicted driving trajectory of the at least one first obstacle.

3. The method according to claim 2, characterized in that The predicting of the driving trajectory of the obstacle according to the lane line marking information and the second detection information to obtain a predicted driving trajectory of at least one first obstacle includes: Get the trained trajectory prediction model; The lane line marking information and the second detection information are input into the trajectory prediction model to obtain a predicted driving trajectory of at least one first obstacle output by the trajectory prediction model.

4. The method according to claim 2, characterized in that: The step of training the obstacle detection model according to the predicted driving trajectory of the at least one first obstacle comprises: Acquire a marked driving trajectory of at least one second obstacle in a plurality of the first images; generating a loss function value according to a difference between a predicted driving trajectory of the at least one first obstacle and a labeled driving trajectory of the at least one second obstacle; The obstacle detection model is trained according to the loss function value.

5. The method according to claim 1, characterized in that The second detection information is generated by the following steps: Using the self-attention layer in the obstacle detection model to extract features from the training data to obtain multiple sample features; Using a cross attention layer in the obstacle detection model to perform feature fusion on the multiple sample features to obtain a first fused feature; The prediction layer in the obstacle detection model is used to perform obstacle detection on the first fusion feature to obtain the second detection information.

6. The method according to any one of claims 1 to 5, characterized in that The global semantic information is generated by the following steps: Extracting semantic features from each of the first images to obtain global semantic features of each of the first images; generating a semantic representation of each of the first images according to the global semantic features of each of the first images; The semantic representations of the first images are integrated to obtain global semantic information associated with the plurality of the first images.

7. The method according to any one of claims 1 to 5, characterized in that The lane marking information is obtained by the following steps: spatially aligning the plurality of first images to obtain a plurality of aligned first images; Performing feature extraction on each of the aligned first images to obtain key features of each of the aligned first images; Fusing the plurality of key features to obtain a second fused feature; Lane line marking is performed according to the second fusion feature to obtain the lane line marking information.

8. An obstacle detection method, characterized in that: include: Obtain target images captured by multiple vehicle-mounted cameras at a target time; Acquire historical detection information; wherein the historical detection information is obtained by performing obstacle detection on historical images captured by the multiple vehicle-mounted cameras at historical moments before the target moment; Acquire global semantic information of the image; wherein the global semantic information of the image is obtained by semantically extracting a plurality of the target images; Input the plurality of target images, the historical detection information and the global semantic information of the image into an obstacle detection model trained by the method according to any one of claims 1 to 7 to perform obstacle detection to obtain target detection information.

9. A training device for an obstacle detection model, characterized in that: include: An acquisition module is used to acquire training data; wherein the training data includes: a first image captured by a plurality of sample cameras at a first moment, first detection information, and global semantic information; the first detection information is obtained by performing obstacle detection on a second image captured by the plurality of sample cameras at a second moment before the first moment; and the global semantic information is obtained by performing semantic extraction on the plurality of first images; An input module, used for inputting the training data into an initial obstacle detection model to perform obstacle detection to obtain second detection information; A training module is used to train the obstacle detection model according to the lane line marking information associated with the plurality of the first images and the second detection information.

10. An obstacle detection device, characterized in that: include: A first acquisition module is used to acquire target images captured by multiple vehicle-mounted cameras at a target time; A second acquisition module is used to acquire historical detection information; wherein the historical detection information is obtained by performing obstacle detection on historical images captured by the multiple vehicle-mounted cameras at historical moments before the target moment; A third acquisition module is used to acquire global semantic information of an image; wherein the global semantic information of an image is obtained by semantically extracting a plurality of the target images; A detection module is used to input the plurality of target images, the historical detection information and the image global semantic information into the obstacle detection model trained by the device as claimed in claim 9 to perform obstacle detection to obtain target detection information.

11. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7, or the method according to claim 8.

12. A chip, characterized in that: The chip comprises a processing circuit configured to execute the method according to any one of claims 1 to 7 or to execute the method according to claim 8.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 7, or the method according to claim 8.