Obstacle detection method, electronic equipment and storage medium

By generating the first text information for describing the image to be detected and detecting obstacles based on the text information, the problem of low detection accuracy of obstacles in complex scenarios in the prior art is solved, thereby improving the accuracy of detection and vehicle driving safety.

CN119942492APending Publication Date: 2025-05-06ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311464434.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When detecting obstacles around vehicles, the prior art has low accuracy when facing complex scenes, occlusion and light changes.

Method used

By acquiring the image to be detected, first text information for describing the image to be detected is generated, and whether there is an obstacle in the scene corresponding to the image is detected based on the text information.

Benefits of technology

It improves the accuracy and comprehensiveness of obstacle detection, and enhances vehicle driving safety and driving experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942492A_ABST
    Figure CN119942492A_ABST
Patent Text Reader

Abstract

The invention provides an obstacle detection method, electronic equipment and a storage medium, relates to the field of target detection, and is used for at least improving the accuracy of obstacle detection. The method comprises the steps of obtaining a to-be-detected image, and generating first text information used for describing the to-be-detected image based on the to-be-detected image; and detecting whether an obstacle exists in a scene corresponding to the to-be-detected image based on the first text information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of target detection, and in particular to an obstacle detection method, electronic equipment, and storage medium. Background Art

[0002] With the development of social economy, the issue of vehicle driving safety has attracted much attention. In order to ensure the safety of vehicles during driving, it is particularly important to detect obstacles around the vehicle in a timely manner and warn the user. When detecting obstacles around the vehicle, related technologies are usually based on image segmentation algorithms.

[0003] However, image segmentation algorithms face certain challenges when faced with complex scenes, occlusions, and lighting changes. For example, in a complex scene, there may be multiple overlapping, touching, or partially occluded obstacles, which makes it difficult for the segmentation algorithm to accurately identify obstacles; or, lighting changes may cause image quality to degrade, making the boundaries of obstacles unclear. Therefore, image segmentation algorithms have low accuracy in obstacle detection. Summary of the invention

[0004] The embodiments of the present disclosure provide an obstacle detection method, an electronic device, and a storage medium, which are used to at least improve the accuracy of obstacle detection.

[0005] In a first aspect, a method for obstacle detection is provided, the method comprising:

[0006] Acquire an image to be detected, and generate first text information for describing the image to be detected based on the image to be detected;

[0007] Based on the first text information, it is detected whether there is an obstacle in the scene corresponding to the image to be detected.

[0008] Based on the obstacle detection method provided by the embodiment of the present disclosure, by acquiring the image to be detected and generating the first text information for describing the image to be detected (that is, the first text information can describe the image features in the image to be detected in natural language), the comprehensiveness and accuracy of the understanding of the image features of the image to be detected can be enhanced, so that when detecting whether there are obstacles in the scene corresponding to the image to be detected based on the first text information, the accuracy and comprehensiveness of the obstacle detection in the scene corresponding to the image to be detected are improved, thereby improving the driving safety and driving experience of the vehicle.

[0009] In a second aspect, an obstacle detection device is provided, comprising:

[0010] An acquisition module, used for acquiring an image to be detected, and generating first text information for describing the image to be detected based on the image to be detected;

[0011] A detection module is used to detect whether there is an obstacle in the scene corresponding to the image to be detected based on the first text information.

[0012] In a third aspect, an electronic device is provided, comprising: a memory and a processor; the memory and the processor are coupled; the memory is used to store a computer program; and the processor implements the obstacle detection method of any of the above embodiments when executing the computer program.

[0013] In a fourth aspect, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the obstacle detection method of any of the above embodiments is implemented.

[0014] In a fifth aspect, a computer program product is provided, which includes computer program instructions, and when the computer program instructions are executed by a processor, the obstacle detection method of any of the above embodiments is implemented.

[0015] For the specific description of the second to fifth aspects and their various implementations in the present disclosure, reference can be made to the detailed description of the first aspect and its various implementations; and for the beneficial effects of the second to fifth aspects and their various implementations, reference can be made to the analysis of the beneficial effects in the first aspect and its various implementations, which will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present disclosure, the drawings required for use in some embodiments of the present disclosure will be briefly introduced below. Obviously, the drawings described below are only drawings of some embodiments of the present disclosure, and a person skilled in the art can also obtain other drawings based on these drawings.

[0017] Figure 1 A schematic diagram of the structure of a system provided in some embodiments of the present disclosure;

[0018] Figure 2 A flowchart of an obstacle detection method provided in some embodiments of the present disclosure;

[0019] Figure 3 A schematic diagram of an image to be detected provided in some embodiments of the present disclosure;

[0020] Figure 4 A flowchart of another obstacle detection method provided in some embodiments of the present disclosure;

[0021] Figure 5 A flowchart of another obstacle detection method provided in some embodiments of the present disclosure;

[0022] Figure 6A flowchart of another obstacle detection method provided in some embodiments of the present disclosure;

[0023] Figure 7 A flowchart of another obstacle detection method provided in some embodiments of the present disclosure;

[0024] Figure 8 A schematic diagram of the structure of an obstacle detection device provided in some embodiments of the present disclosure;

[0025] Fig. 9 A schematic diagram of the structure of an electronic device provided for some embodiments of the present disclosure. DETAILED DESCRIPTION

[0026] The following will be combined with the drawings in the present disclosure to clearly and completely describe the technical solutions in the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0027] It should be noted that, in the present disclosure, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present disclosure should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0028] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.

[0029] In the description of the present disclosure, unless otherwise specified, " / " means "or", for example, A / B can mean A or B. "And / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, "at least one" means one or more, and "a plurality" means two or more.

[0030] For ease of understanding, relevant concepts involved in the embodiments of the present disclosure are first briefly introduced.

[0031] Vehicle-to-everything (V2X) wireless communication technology refers to the communication and interaction technology between vehicles and everything. It is an important technology in intelligent transportation systems, aiming to improve road safety, traffic efficiency and passenger comfort. V2X technology enables vehicles to communicate with other vehicles, infrastructure and traffic participants such as pedestrians in real time. This communication can be carried out through wireless communication technologies such as wireless-fidelity (Wi-Fi) networks and cellular networks to transmit information such as vehicle location, speed, direction, etc.

[0032] Cellular-V2X (Cv2x) is used for direct wireless communication between vehicles.

[0033] PC5 method, in CV2X, PC5 method is a specific physical layer and scheduling method for direct communication between vehicles. It can realize Sidelink communication between vehicles. Sidelink is an air interface technology in the CV2X system, which allows vehicles to establish communication connections directly without going through base stations or network relays. This direct communication can be carried out without network coverage or network congestion, with lower latency and higher reliability.

[0034] The object detection (you only look once, YOLO) algorithm is used to detect and locate multiple objects in real time in an image or video. Compared with traditional object detection algorithms, the YOLO algorithm has faster speed and higher accuracy. The core idea of ​​the YOLO algorithm is to transform the object detection problem into a regression problem. It divides the input image into a grid of fixed size and predicts the bounding box and category probability of a certain object in each grid cell. This means that the YOLO algorithm only needs one forward pass to complete the detection and classification of the object at the same time, so it is called "you only look once".

[0035] The above is an introduction to some concepts involved in the embodiments of the present disclosure, which will not be repeated below.

[0036] In the field of autonomous driving and V2X vehicle-road collaborative perception, traditional computer vision (CV) algorithms have limited types of detection targets, and the types of detection targets supported may be missed under different lighting and angles. Currently, there is no ideal visible light technology that can detect obstacles that affect vehicle driving under all circumstances. For example, millimeter waves can detect obstacles, but can only detect moving targets; LiDAR can also detect obstacles, but LiDAR has high vehicle cost requirements.

[0037] In response to the above problems, an embodiment of the present disclosure provides an obstacle detection method, the idea of ​​which is: by acquiring an image to be detected and generating a first text information for describing the image to be detected (the first text information can describe the image features in the image to be detected in natural language), the comprehensiveness and accuracy of the understanding of the image features of the image to be detected can be enhanced, so that when detecting whether there is an obstacle in the scene corresponding to the image to be detected based on the first text information, the accuracy and comprehensiveness of the obstacle detection in the scene corresponding to the image to be detected are improved, thereby improving the driving safety and driving experience of the vehicle.

[0038] At the same time, the method provided by the embodiment of the present disclosure detects and identifies the image to be detected based on a large neural network model, can carry more detection categories, and realize efficient detection and identification of various obstacles in different scenes, avoiding the long-tail effect brought by small models and improving the user experience.

[0039] See also Figure 1 , is a schematic diagram of the structure of the system involved in the obstacle detection method provided in the embodiment of the present disclosure. Figure 1 The system comprises: an image acquisition device 100 and a detection device 200, and the image acquisition device 100 and the detection device 200 are communicatively connected.

[0040] The image acquisition device 100 is used to capture and record the image to be detected. After capturing the image to be detected, the image acquisition device 100 sends the image to be detected to the detection device 200 so that the detection device 200 can detect whether there is an obstacle in the scene corresponding to the image to be detected.

[0041] As an example, the image acquisition device 100 can directly capture an image in the target scene, and use the captured image as the image to be detected. Optionally, the target scene can be a road scene, a parking lot scene, or a gas station scene.

[0042] As another example, the image acquisition device 100 can capture video information in the target scene, wherein the video information includes continuous image frames acquired by the image acquisition device 100 at a certain frame rate, and each image frame can be an independent static image. The image acquisition device 100 can capture an image frame at a certain time point (i.e., a screenshot operation), or use a video frame extraction technology to extract an image frame at a certain time point from the video information, and use the extracted image frame as the image to be detected.

[0043] In some embodiments, the image acquisition device 100 may be a device including a visible light sensor, such as a camera, a photoresistor, etc.

[0044] It should be noted that there may be one or more image acquisition devices 100. The embodiment of the present disclosure does not limit the number of image acquisition devices 100.

[0045] It should be noted that the embodiment of the present disclosure mainly uses the image acquisition device 100 to capture the image to be detected, but in actual implementation, the image to be detected can also be captured by other methods, for example, by creating a three-dimensional point cloud map through a laser radar, etc. The embodiment of the present disclosure does not limit the method of capturing the image to be detected.

[0046] The detection device 200 is used to obtain the image to be detected sent by the image acquisition device 100, and detect whether there is an obstacle in the scene corresponding to the image to be detected.

[0047] In some embodiments, the detection device 200 generates first text information for describing the image to be detected based on the image to be detected, and detects whether there is an obstacle in the scene corresponding to the image to be detected based on the first text information.

[0048] In some embodiments, the detection device 200 may be a processor. Optionally, the processor may be a central processing unit (CPU), a general network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. Alternatively, the detection device 200 may also be other devices with processing functions, such as circuits, devices, or software modules, which are not limited in the embodiments of the present disclosure.

[0049] In some embodiments, the above system also includes: any one or more of a domain controller, an edge computing (mobile edge computing MEC) unit, and a roadside computing unit (roadside computing unit, RCU).

[0050] In the automatic driving system, the domain controller can coordinate the driving of the vehicle according to the detection result of the detection device 200 on the detection image, for example, path planning, speed control, coordination between vehicles, etc. Optionally, the domain controller can be integrated into the detection device 200 as a part of the detection device 200; or, the domain controller can exist independently of the detection device 200.

[0051] In the vehicle-road cooperative perception system, the MEC unit and / or RCU can receive the image to be detected captured by the image acquisition device 100, and use algorithms to process and analyze the image to be detected in real time, and perform decision-making and control tasks at the edge, such as target detection, tracking, and path planning. Optionally, the MEC unit and / or RCU can be integrated into the detection device 200 as part of the detection device 200; or, the MEC unit or RCU can exist independently of the detection device 200.

[0052] It should be noted that the system architecture and application scenarios described in the embodiments of the present disclosure are intended to more clearly illustrate the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. A person of ordinary skill in the art can appreciate that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0053] An obstacle detection method provided by an embodiment of the present disclosure is described below in conjunction with the accompanying drawings.

[0054] See also Figure 2 , is a flow chart of an obstacle detection method provided by an embodiment of the present disclosure. Figure 2 As shown, the obstacle detection method provided in the embodiment of the present disclosure is applied to a detection device, and can be specifically implemented as follows:

[0055] S101: Acquire an image to be detected, and generate first text information for describing the image to be detected based on the image to be detected.

[0056] As an example, when the detection device obtains the image to be detected from the image acquisition device, multiple image acquisition devices can be allocated to the detection area, and each image acquisition device performs independent detection. For example, the detection device divides the image acquisition device A and the image acquisition device B to detect different positions of the detection area and capture the image to be detected in the detection area.

[0057] As another example, when the detection device obtains the image to be detected from the image acquisition device, the picture captured by a certain image acquisition device can be divided into different sub-pictures, each sub-picture corresponding to a different sub-area in the detection area, and then the sub-pictures in each sub-area (that is, the image to be detected in each sub-area) are obtained in units of sub-areas.

[0058] In some embodiments, the first text information is used to describe the type, shape, size, etc. of the object in the image to be detected, as well as the lighting conditions, weather conditions, etc. of the image to be detected. For example, the first text information may be "The sky is blue. There is a dog on the road.".

[0059] It should be noted that the language form of the first text information can be Chinese, English, etc., and the language form of the first text information can also be different according to different requirements during specific implementation. The embodiment of the present disclosure does not limit the language form of the first text information.

[0060] It can be understood that compared with the problem of low accuracy in obstacle detection based on image segmentation algorithms in related technologies, the method provided in the embodiments of the present invention acquires an image to be detected and generates first text information for describing the image to be detected. The image features of the image to be detected can be expressed more comprehensively and accurately using the first text information, so that when obstacles are subsequently detected based on the first text information, the detection of obstacles can be more comprehensive and accurate.

[0061] S102: Based on the first text information, detect whether there is an obstacle in the scene corresponding to the image to be detected.

[0062] Exemplarily, if the first text information describes that there is a pedestrian and a conical object in the scene corresponding to the image to be detected, then there is an obstacle, the conical object, in the scene corresponding to the image to be detected.

[0063] It can be understood that the obstacle detection method provided by the embodiment of the present disclosure, by acquiring the image to be detected and generating the first text information for describing the image to be detected (the first text information can describe the image features in the image to be detected in natural language), can enhance the comprehensiveness and accuracy of the understanding of the image features of the image to be detected, so that when detecting whether there are obstacles in the scene corresponding to the image to be detected based on the first text information, the accuracy and comprehensiveness of the obstacle detection in the scene corresponding to the image to be detected are improved, thereby improving the driving safety and driving experience of the vehicle.

[0064] In some embodiments, when it is detected that there is an obstacle in the scene corresponding to the image to be detected, the detection device may also send an alarm message for prompting the existence of the obstacle.

[0065] In some embodiments, when an obstacle is detected in the scene corresponding to the image to be detected, the detection device sends a warning message to the electronic device involved in the scene corresponding to the image to be detected. Exemplarily, the detection device can transmit the warning message to the traffic participants on the roadside or the traffic manager in the cloud through the communication link through the RCU, roadside unit (RSU), MEC and other devices.

[0066] In some embodiments, when no obstacles are detected in the scene corresponding to the image to be detected, the detection device may also send an alarm message to the electronic device involved in the scene corresponding to the image to be detected, and the content of the alarm message may be determined based on whether there are objects in the scene. Exemplarily, if there are no obstacles in the scene, but there are known common traffic participants, such as pedestrians, buses, etc., the alarm message is used to prompt the traffic participants or the traffic managers in the cloud to track and detect the traffic participants; if there are no obstacles or other objects in the scene, the alarm message is used to prompt the traffic participants or the traffic managers in the cloud that there is nothing abnormal in the scene.

[0067] In some embodiments, traffic participants include: motor vehicles, non-motor vehicles, pedestrians, etc. As an example, the detection device can send warning information to traffic participants through CV2X technology. Optionally, the warning information can be a roadside infrastructure / road condition event (RSI / RTE) message; CV2X uses PC5 to broadcast RSI / RTE messages, the eventID of the message uses the national standard traffic event encoding, and the description field of the message is filled with the first text information obtained by the graphic generation model, and the first text information can be the original text information or the processed text information.

[0068] As another example, the detection device can send warning information to traffic participants through 5G technology. For example, the detection device can transmit the warning information through a base station, or through the cloud, server, etc. At the same time, traffic participants can also obtain warning information from the server through mobile phones and car software (such as navigation software).

[0069] It can be understood that the method provided by the embodiment of the present disclosure sends a warning message for indicating the existence of an obstacle when an obstacle is detected in the scene corresponding to the image to be detected, which can promptly remind the user to pay attention to safety, ensure the user's driving safety, and improve the user's driving experience.

[0070] In some embodiments, the first text information can be generated by some image-text conversion models. Based on the image to be detected, generating the first text information for describing the image to be detected can be implemented as follows: inputting the image to be detected into the image-text generation model to obtain the first text information output by the image-text generation model.

[0071] In some embodiments, the detection device inputs the image to be detected into the image-text generation model, and the image-text generation model performs artificial intelligence forward reasoning on the image to be detected to obtain the first text information. Figure 3 As shown, Figure 3The image to be detected is input into the image-text generation model, and the image-text generation model will output the first text information corresponding to the image: "The sky is blue. There is a cone shaped object on the road".

[0072] In some embodiments, the detection device also collects the first text information of each image to be detected, marks the region (region) corresponding to each image to be detected, generates a region identification code (regionID), and associates it with a key value dictionary of prompt (taking the first text information as the value and associating it with the key corresponding to the regionID) to facilitate subsequent processing and identification of each image to be detected.

[0073] Optionally, the detection device may determine a unique regionID according to the content of each image to be detected, ensuring that each region corresponding to each image to be detected has a different identifier.

[0074] It can be understood that compared with the related art based on traditional convolutional neural networks, which use semantic segmentation algorithms to identify obstacles in scenes and have a high false alarm rate, the embodiment of the present disclosure is based on a graphic generation model and can accurately determine the image features of the image to be detected, thereby improving the accuracy of obstacle detection.

[0075] In some embodiments, Figure 4 As shown, based on the first text information, detecting whether there is an obstacle in the scene corresponding to the image to be detected can be specifically implemented as follows: steps S1021-S1023.

[0076] S1021. Divide the first text information into at least one sentence.

[0077] As an example, the first text information can be divided based on punctuation marks (such as periods, commas, etc.) of the first text information to obtain at least one sentence. For example, if the first text information is "The sky is blue. There is a cone shaped object on the road", the first text information is divided according to the punctuation marks to obtain the sentence "The sky is blue" and the sentence "There is a cone shaped object on the road".

[0078] As another example, the first text information can be divided based on a machine learning model (such as a sentence boundary detection model). The sentence boundary detection model can continue to detect the input first text information to find a suitable division position for the first text information to obtain at least one sentence.

[0079] It should be noted that the above are only some examples of dividing the first text information given in the embodiments of the present disclosure. In actual implementation, different methods can be selected according to different actual needs; the embodiments of the present disclosure do not limit the method for dividing the first text information.

[0080] S1022: Identify whether the textual meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected.

[0081] In some embodiments, the above step S1022 may be specifically implemented as follows: inputting the sentence into a text classification model to obtain a text classification result output by the text classification model.

[0082] In some embodiments, the text classification result includes a first classification result or a second classification result, the first classification result is used to indicate that the text meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected, and the second classification result is used to indicate that the text meaning of the sentence is not used to describe the existence of an object in the scene corresponding to the image to be detected.

[0083] Exemplarily, if the sentence "The sky is blue" is input into the text classification model and the text classification result is the second classification result, it means that the text meaning of the sentence "The sky is blue" is not used to describe the existence of an object in the scene corresponding to the image to be detected. If the sentence "There is a cone shaped object on the road" is input into the text classification model and the text classification result is the first classification result, it means that the text meaning of the sentence "There is a cone shaped object on the road" is used to describe the existence of an object in the scene corresponding to the image to be detected.

[0084] In some embodiments, the following methods may also be considered to identify whether the textual meaning of a sentence is used to describe the existence of an object in the scene corresponding to the image to be detected:

[0085] Keyword matching method: create a list of keywords that contain as many objects as possible, and then perform keyword matching on the sentences to see if the sentences contain keywords corresponding to the objects. If the sentences contain any keyword corresponding to the objects, it is determined that there is an object in the scene corresponding to the image to be detected.

[0086] Natural language processing method uses natural language processing technology to analyze each sentence, such as using a bag of words model or a word embedding model (such as Word2Vec, BERT model, etc.) to understand the semantics of the sentence. For example, if the sentence contains the description "there is a bag in the middle of the road" or "an unknown animal is crossing the road", then this information may indicate that there is an object in the scene corresponding to the image to be detected.

[0087] The rule engine method builds a rule engine, writes a series of rules and pattern matching, and determines whether there is an obstacle based on the text description of each sentence. For example, when a sentence contains descriptions such as "there are spilled objects on the roadside" or "the road is icy", it can be determined that there is an object in the scene corresponding to the image to be detected.

[0088] It should be noted that whether the textual meanings of some recognition sentences provided in the above embodiments of the present disclosure are used to describe examples of objects existing in the scene corresponding to the image to be detected, different methods can be flexibly selected according to different actual conditions during specific implementation, and the embodiments of the present disclosure do not limit this.

[0089] It is understandable that the method provided by the embodiment of the present disclosure can classify the text more finely by dividing the first text information to obtain at least one sentence. Different sentences may emphasize different information or have different descriptions. Inputting the sentences into the text classification model respectively to obtain the text classification results output by the text classification model can enable the text classification model to more accurately identify the key information in the sentences, thereby improving the accuracy of classification.

[0090] S1023. When the textual meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected, detect whether there is an obstacle in the scene corresponding to the image to be detected.

[0091] In some embodiments, the above step S1023 can be specifically implemented as: steps a1-a2.

[0092] a1. Extract keywords corresponding to objects from sentences.

[0093] For example, if the sentence is "There is a cone shaped object on the road", the keyword corresponding to the extracted object is: cone. If the sentence is "A car is driving on the road", the keyword corresponding to the extracted object is: car.

[0094] a2. Based on the keywords corresponding to the object, determine whether the object is an obstacle in the scene corresponding to the image to be detected.

[0095] As a possible implementation of the above step a2: searching whether there is a keyword corresponding to the object in the first keyword list. If a keyword corresponding to the object is found in the first keyword list, the object is determined to be an obstacle; or, if a keyword corresponding to the object is not found in the first keyword list, the object is determined to be a non-obstacle.

[0096] The first keyword list includes at least one keyword corresponding to an obstacle, for example, the first keyword list includes keywords corresponding to uncommon building materials, keywords corresponding to wild animals, keywords corresponding to cubes of different shapes, and the like.

[0097] For example, if the keyword corresponding to the object is cone, then the keyword cone is searched in the first keyword list to see if it exists. If the keyword corresponding to the object is car, then the keyword car is searched in the first keyword list to see if it exists. If cone is found in the first keyword list, then the object is determined to be an obstacle. If cone is not found in the first keyword list, then the object is determined to be a non-obstacle.

[0098] It can be understood that the method provided in the embodiment of the present disclosure is based on the keywords corresponding to the objects, and can identify the obstacles in the image to be detected by matching with the preset obstacle keywords, thereby avoiding manual identification of obstacles and saving manpower, material resources and time costs. At the same time, for obstacle detection tasks in some specific scenarios, such as autonomous driving, robot navigation and other scenarios, the obstacle detection technology based on object keywords can improve the accuracy of obstacle identification.

[0099] In some embodiments, when the keyword corresponding to the object is not found in the first keyword list, the detection device sends the keyword of the object to the server. Based on the keyword of the object, the server further identifies the keyword of the object using an intelligent algorithm or manually and generates a first recognition result. The first recognition result is used to indicate whether the object is an obstacle. The detection device receives the first recognition result sent by the server, and when the first recognition result is used to indicate that the object is an obstacle, the detection device adds the keyword of the object to the first keyword list. Exemplarily, if the keyword of the object is cone, and cone is not found in the first keyword list, the detection device sends the keyword cone of the object to the server, the server determines that the object is an obstacle and generates a first recognition result, the detection device receives the first recognition result sent by the server, and updates the keyword cone into the first keyword list.

[0100] It is understandable that the method provided by the embodiment of the present disclosure sends the keyword of the object to the server when the keyword corresponding to the object is not found in the first keyword list, and can further improve the accuracy of obstacle detection by intelligent algorithm or manual recognition of whether the object is an obstacle. At the same time, the method provided by the embodiment of the present disclosure adds the keyword of the object to the first keyword list when the first recognition result indicates that the object is an obstacle, which can expand the keyword list and improve the accuracy of obstacle detection.

[0101] As another possible implementation of the above step a2: searching for keywords corresponding to the object in the second keyword list. If the keywords corresponding to the object are not found in the second keyword list, the object is determined to be an obstacle; or, if the keywords corresponding to the object are found in the second keyword list, the object is determined to be a non-obstacle.

[0102] The second keyword list includes at least one keyword corresponding to a non-obstacle.

[0103] For example, the second keyword list includes keywords corresponding to pedestrians, vehicles, zebra crossings, etc. Exemplarily, if the keyword corresponding to the object is car, then the second keyword list is searched for the presence of the keyword car. Exemplarily, assuming that the keyword corresponding to the object is cone, if cone is not found in the second keyword list, the object is determined to be an obstacle; if cone is found in the second keyword list, the object is determined to be a non-obstacle.

[0104] It should be noted that non-obstacles refer to other objects or subjects that are not related to the task goal or situation in a specific task or scenario. Traffic targets such as pedestrians, electric vehicles, bicycles, buses, zebra crossings, etc. are usually non-obstacles, which are common traffic participants in the traffic environment. In actual implementation, the definition of non-obstacles varies depending on the specific tasks or scenarios, so the disclosed embodiments do not limit the content of non-obstacles.

[0105] In some embodiments, when the keyword corresponding to the object is not found in the second keyword list, the detection device sends the keyword of the object to the server, and the server further identifies the keyword of the object using an intelligent algorithm or manually based on the keyword of the object and generates a second recognition result. The second recognition result is used to indicate whether the object is a non-obstacle. The detection device receives the second recognition result sent by the server, and when the second recognition result is used to indicate that the object is a non-obstacle, the detection device adds the keyword of the object to the second keyword list. Exemplarily, assuming that the corresponding keyword of the object is bus, and bus is not found in the second keyword list, the keyword bus of the object is sent to the server. The server determines that the object is a non-obstacle and generates a second recognition result. The detection device receives the second recognition result sent by the server and updates the keyword bus into the second keyword list.

[0106] It should be noted that if the second recognition result is used to indicate that the object is not an obstacle, it means that the keywords in the second keyword list are not comprehensive enough. At this time, the keywords of the object need to be added to the second keyword list to expand and update the second keyword list.

[0107] It is understandable that the method provided by the embodiment of the present disclosure sends the keyword of the object to the server when the keyword corresponding to the object is not found in the second keyword list, and can further improve the accuracy of obstacle detection by intelligent algorithm or manual recognition of whether the object is a non-obstacle. At the same time, the method provided by the embodiment of the present disclosure adds the keyword of the object to the second keyword list when the second recognition result indicates that the object is a non-obstacle, which can expand the keyword list and improve the accuracy of obstacle detection.

[0108] As another possible implementation of the above step a2: respectively search for keywords corresponding to the object in the first keyword list and the second keyword list. In the case where the keywords corresponding to the object are not found in the first keyword list and the keywords corresponding to the object are not found in the second keyword list, the detection device sends the keywords of the object to the server. The server further identifies the keywords of the object using an intelligent algorithm or manually based on the keywords of the object and generates a third recognition result. The third recognition result is used to indicate whether the object is an obstacle. The detection device receives the third recognition result sent by the server, and adds the keywords of the object to the first keyword list when the third recognition result is used to indicate that the object is an obstacle; or, when the third recognition result is used to indicate that the object is a non-obstruction, the keywords of the object are added to the second keyword list. Exemplarily, assuming that the corresponding keyword of the object is bus, and bus is not found in the first keyword list, and bus is not found in the second keyword list, the keyword bus of the object is sent to the server. The server determines that the object is a non-obstruction and generates a third recognition result, and the detection device receives the third recognition result sent by the server and adds the keyword bus to the second keyword list.

[0109] It is understandable that in the method provided by the embodiment of the present disclosure, the detection device can search for keywords corresponding to the object in the first keyword list and the second keyword list respectively, which can make the search method more flexible and the search range wider; at the same time, the method provided by the embodiment of the present disclosure can further improve the accuracy of obstacle detection by using intelligent algorithms or manual identification to identify whether the object is an obstacle without searching for keywords corresponding to the object in the first keyword list and the second keyword list. In addition, the method provided by the embodiment of the present disclosure can add the keywords of the object to the first keyword list or the second keyword list based on the third recognition result, which can expand the keyword list and improve the accuracy of obstacle detection.

[0110] As another possible implementation of the above step a2: search whether there is a keyword corresponding to the object in the third keyword list. When the keyword corresponding to the object is found in the third keyword list, determine the label of the keyword corresponding to the object. The label is used to indicate whether the object is an obstacle. Finally, determine whether the object is an obstacle based on the label. Optionally, the label can be 0 or 1; 0 indicates that the object is not an obstacle; 1 indicates that the object is an obstacle.

[0111] The third keyword list includes at least one keyword corresponding to an obstacle, at least one keyword corresponding to a non-obstacle, and a label of each keyword. Exemplarily, the third keyword list includes: vehicle (label 0), cone-shaped object (label 1).

[0112] It should be noted that the content of the third keyword list is only an example given in the embodiment of the present disclosure. In specific implementation, the storage format of each keyword and each keyword label in the third keyword list can be flexibly determined, and the embodiment of the present disclosure does not limit this.

[0113] In some embodiments, when the keyword corresponding to the object is not found in the third keyword list, the detection device sends the keyword of the object to the server. The server further identifies the keyword of the object using an intelligent algorithm or manually based on the keyword of the object and generates a fourth recognition result. The fourth recognition result is used to indicate whether the object is an obstacle. The detection device receives the fourth recognition result sent by the server, and based on the fourth recognition result, after adding a label to the keyword of the object, the keyword of the object and its label are added to the third keyword list. Exemplarily, if the keyword of the object is cone, and cone is not found in the third keyword list, the detection device sends the keyword cone of the object to the server, the server determines that the object is an obstacle and generates a fourth recognition result, and after the detection device receives the fourth recognition result sent by the server, since the object is an obstacle, the keyword cone is added with label 1, and cone (label 1) is added to the third keyword list.

[0114] It is understandable that in the method provided by the embodiment of the present disclosure, the third keyword list may include both at least one keyword corresponding to an obstacle and at least one keyword corresponding to a non-obstacle, which may make the keywords in the third keyword list wider and more diverse, thereby improving the keywords in the third keyword list; at the same time, the method provided by the embodiment of the present disclosure may further improve the accuracy of obstacle detection by using an intelligent algorithm or manually identifying whether the object is an obstacle, without searching for keywords corresponding to the object from the third keyword list. In addition, the method provided by the embodiment of the present disclosure determines the label corresponding to the keyword of the object based on the fourth recognition result, and adds it to the third keyword list, which may expand the keyword list and improve the accuracy of obstacle detection.

[0115] In some embodiments, based on the first text information, it is detected whether there is an obstacle in the scene corresponding to the image to be detected, and the first text information may not be segmented, but the text classification result of the first text information is directly obtained, and based on the text classification result, it is determined whether there is an obstacle in the scene corresponding to the image to be detected. Exemplarily, the first text information is input into a text classification model to obtain a text classification result output by the text classification model.

[0116] In some embodiments, the text classification results include a third classification result and a fourth classification result. The third classification result is used to indicate that the text meaning of the first text information is used to describe the existence of obstacles in the scene corresponding to the image to be detected, and the fourth classification result is used to indicate that the text meaning of the first text information is not used to describe the existence of obstacles in the scene corresponding to the image to be detected.

[0117] Exemplarily, the first text information "The sky is blue. There is a cone shaped object on the road" is input into the text classification model, and the classification result thereof is the third classification result, which means that the text meaning of the first text information is used to describe the existence of obstacles in the scene corresponding to the image to be detected.

[0118] In some embodiments, the training process of the text classification model may refer to the following steps S301-S302, which will not be described in detail in the embodiments of the present disclosure.

[0119] It can be understood that the method provided by the embodiment of the present disclosure can not only divide the first text information into sentences and input the sentences into a text classification model, but also directly input the first text information into the text classification model to obtain the text classification results output by the text classification model. While obtaining the text classification results is more flexible, it also improves the efficiency of obstacle detection and improves the user experience.

[0120] In some embodiments, Figure 5 As shown in the figure, the above-mentioned image-text generation model is trained through the following steps:

[0121] S201: Obtain a second training sample set.

[0122] In some embodiments, the second training sample set includes a plurality of second sample images and third text information corresponding to each of the plurality of second sample images.

[0123] In some embodiments, when acquiring the second training sample set, multiple second sample images may be collected, and corresponding text descriptions may be added to the second sample images to ensure that the second training sample set contains various types of obstacles and detection scenarios. Optionally, the second sample images may be images under different road conditions and scenarios, for example, the second sample images may be images under scenes such as urban roads, rural roads, highways, gas stations, etc. The second sample images may be captured by a camera on a vehicle, a drone, a ground camera, etc.

[0124] In some embodiments, the second sample images should cover different weather conditions, lighting conditions and obstacle types. Before acquiring and generating the second training sample set, the second sample images need to be resized, color balanced, and denoised.

[0125] In some embodiments, after obtaining the second training sample set, the second sample image and the third text information in the second training sample set can be preprocessed so that the data format of the second sample image and the third text information matches the parameter format of the image-text generation model. Exemplarily, multiple second sample images can be processed so that the format of the second sample image matches the input format of the image-text generation model; for example, the multiple second sample images can be converted into multiple feature vectors as inputs to the image-text generation model. Exemplarily, the third text information corresponding to each of the multiple sample images can be processed so that the third text information matches the input format of the image-text generation model; for example, the third text information can be encoded into a format understandable to the image-text generation model.

[0126] It can be understood that the method provided by the embodiment of the present disclosure can comprehensively consider the impact of different road conditions, weather conditions and other conditions on images by training the image-text generation model based on images under different road conditions and scenes, thereby improving the adaptability of the image-text generation model to images to be detected in complex scenes, and thereby improving the accuracy of the first text information output by the image-text generation model.

[0127] S202: Based on the second training sample set, the initial image-text generation model is trained to obtain a trained image-text generation model.

[0128] In some embodiments, the initial image-text generation model can be a model for natural language processing, and the initial image-text generation model can be extended for multimodal tasks, combining text and images for parameter input. Exemplarily, the initial image-text generation model can be a large open source model such as a chat generative pre-trained transformer (chatGPT), a deepbooru model, or a stable diffusion joint training (contrastivelanguage-image pre-training, CLIP).

[0129] In some embodiments, the parameter size of the initial image and text generation model needs to be determined in combination with the computing power allocation of the obstacle detection device. For example, when a large model with 6 billion parameters generates text information of the image to be detected on the roadside sensing MEC / RCU (for example, the orin nx detection device), it takes 20 seconds to complete the model output once. In addition, other AI models that are detecting the image to be detected in real time may take 30 seconds to complete the output of the large model once; if the computing power of the obstacle detection device is limited and 30 seconds is too long, the parameters of the large model can be reduced; for example, a large model with 1 billion parameters can complete the model output once in 5 seconds, and if the large model is quantized, the speed will be even faster.

[0130] It should be noted that the parameter size of the image-text generation model needs to be selected in combination with different actual scenarios and requirements, and the embodiments of the present disclosure do not limit this.

[0131] In some embodiments, after preprocessing the second sample image and the third text information in the second training sample set, a multimodal fusion method can be designed to fuse the second sample image and the third text information corresponding to the second sample image, so that the initial image-text generation model can understand the relationship between the second sample image and the third text information corresponding to the second sample image. Optionally, the architecture of the image-text generation model can also be modified so that the initial image-text generation model can accept multiple inputs and generate corresponding outputs.

[0132] In some embodiments, the initial image-text generation model is trained using a second training sample set after multimodal fusion (in this case, the data in the second training sample set is multimodal data). During the training, the image-text generation model will attempt to learn the association between the second sample image and the third text information corresponding to the second sample image to obtain a trained image-text generation model, so that after inputting the image to be detected, the image-text generation model can output the first text information used to describe the image to be detected.

[0133] In some embodiments, after obtaining the trained image-text generation model, appropriate evaluation indicators can be defined to measure the performance of the image-text generation model in obstacle detection, and the performance of the image-text generation model can be continuously improved according to the evaluation results. Once the image-text generation model performs well, it can be deployed in practical applications (for example, the image-text generation model is deployed in a traffic system) so that the image-text generation model generates the first text information of the image to be detected. Exemplarily, evaluation indicators such as the correlation between the second sample image and the third text information corresponding to the second sample image, and the output accuracy of the image-text generation model can be defined.

[0134] In some embodiments, improving the performance of the image-text generation model includes: adjusting the architecture of the image-text generation model; expanding the second training sample set to increase the diversity and quantity of data; adjusting the values ​​of hyperparameters of the image-text generation model; adjusting the training strategy of the image-text generation model, etc.

[0135] In some embodiments, considering real-time and efficiency issues, it may be necessary to quantize, prune, and optimize the image and text generation model to adapt to actual computing resources. According to the real-time requirements of the detection, select an appropriate sampling rate, such as performing a detection every 10-30 seconds. In order to provide more efficient AI computing power, it may be necessary to configure a dedicated embedded AI chip or high-performance computing unit for the detection device.

[0136] It is understandable that traditional image-text generation models, for example, the yolov5 model generally uses the coco dataset to detect 80 types of objects, the obj365 model has more classifications than the yolov5 model, and the imagenet dataset model has 1000 categories. However, the above models can carry fewer categories, and the parameters are generally at the level of about 30MB. The parameter level is small, and the model cannot detect and identify obstacles for images to be detected that exceed the categories that the model can carry, and the limitations are large. The disclosed embodiments are based on large neural network models, such as the chatgpt4 model, the deepbooru model, the stablediffusion CLIP model, etc., to detect and identify images to be detected (large neural network models have been trained in large quantities for parameters at the billion level), which can carry more detection categories, realize efficient detection and identification of various obstacles in different scenes, and improve the user experience.

[0137] In some embodiments, Figure 6 As shown in the figure, the above text classification model is trained through the following steps:

[0138] S301: Obtain a first training sample set.

[0139] In some embodiments, the first training sample set includes multiple first samples and respective labels of the multiple first samples, the first sample is second text information used to describe the first sample image, the label of the first sample is used to indicate whether the text meaning of a sentence of the second text information is used to describe whether there is an object in the scene corresponding to the first sample image, and the scene corresponding to the first sample image is the same as the scene corresponding to the image to be detected.

[0140] Optionally, the label of the first sample can be identified by 0 or 1; 0 means that the text meaning of the second text information is not used to describe the existence of an object in the scene corresponding to the first sample image; 1 means that the text meaning of the second text information is used to describe the existence of an object in the scene corresponding to the first sample image. Exemplarily, if the first sample (that is, the second text information corresponding to the first sample image) is "The sky is blue", the label of the first sample is 0.

[0141] S302: Based on the first training sample set, train an initial text classification model to obtain a trained text classification model.

[0142] As an example, the first sample in the first training sample set is input into the initial text classification model, and the text classification model is trained using a label corresponding to each first sample to obtain a trained text classification model.

[0143] As another example, the first sample is segmented into sentences, and a label corresponding to each sentence is determined. Then, each sentence of the first sample is input into an initial text classification model, and the text classification model is trained using the label corresponding to each sentence to obtain a trained text classification model.

[0144] It should be noted that the above-mentioned training method of the text classification model is only an example given in the embodiment of the present disclosure. In the specific implementation, the embodiment of the present disclosure does not limit the training method of the text classification model.

[0145] For ease of understanding, the method provided in the embodiment of the present disclosure is further described below in the form of examples.

[0146] Example 1: A user drives a vehicle and detects obstacles in a road scene.

[0147] For example, Figure 7 As shown, Example 1 can be implemented as the following steps:

[0148] Step b1: Collect pictures to pre-train the large model.

[0149] Step b2: Download the pre-trained large model.

[0150] Step b3: collect the image to be detected.

[0151] Step b4: input the image to be detected into the image-text generation model.

[0152] Step b5: Generate text information of the image to be detected based on the image-text generation model.

[0153] Step b6, parse the text information to determine whether there is an object; if so, execute step b7; if not, execute step b12.

[0154] Step b7, determine whether the object is a common traffic target; if not, execute step b8; if yes, execute step b9.

[0155] Step b8: Identify that an abnormal object (ie, obstacle) is occupying the road surface.

[0156] Step b9: Identify common traffic targets on the road surface.

[0157] Step b10: perform target detection and classification on common traffic targets.

[0158] Step b11: Structuring the classification results.

[0159] Step b12: No abnormality, this test is finished.

[0160] It can be understood that the obstacle detection method provided by the embodiment of the present disclosure, by acquiring the image to be detected and generating the first text information for describing the image to be detected (the first text information can describe the image features in the image to be detected in natural language), can enhance the comprehensiveness and accuracy of the understanding of the image features of the image to be detected, so that when detecting whether there are obstacles in the scene corresponding to the image to be detected based on the first text information, the accuracy and comprehensiveness of the obstacle detection in the scene corresponding to the image to be detected are improved, thereby improving the driving safety and driving experience of the vehicle.

[0161] At the same time, the method provided by the embodiment of the present disclosure detects and identifies the image to be detected based on a large neural network model, can carry more detection categories, and realize efficient detection and identification of various obstacles in different scenes, avoiding the long-tail effect brought by small models and improving the user experience.

[0162] The above mainly introduces the scheme of the embodiment of the present disclosure from the perspective of the method. It is understandable that in order to realize the above functions, the obstacle detection device includes at least one of the hardware structure and software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiment disclosed herein, the embodiment of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiment of the present disclosure.

[0163] It is understandable that, in order to realize the above functions, the obstacle detection device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments of the present disclosure, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present disclosure.

[0164] The embodiments of the present disclosure may divide the functional modules of the obstacle detection device according to the above method embodiments. For example, each functional module may be divided corresponding to each function, or two or more functions may be integrated into one functional module. The above integrated modules may be implemented in the form of hardware or software. It should be noted that the division of modules in the embodiments of the present disclosure is schematic and is only a logical functional division. There may be other division methods in actual implementation. The following is an example of dividing each functional module corresponding to each function.

[0165] Figure 8 is a schematic diagram of the structure of an obstacle detection device provided by an embodiment of the present disclosure, which can execute the obstacle detection method provided by the above method embodiment. Figure 8 As shown, the obstacle detection device 300 includes: an acquisition module 301 and a detection module 302 ; in some embodiments, the obstacle detection device 300 also includes: a judgment module 303 and a sending module 304 .

[0166] The acquisition module 301 is used to acquire an image to be detected, and generate first text information for describing the image to be detected based on the image to be detected.

[0167] The detection module 302 is used to detect whether there is an obstacle in the scene corresponding to the image to be detected based on the first text information.

[0168] In some embodiments, the detection module 302 is specifically used to divide the first text information to obtain at least one sentence; identify whether the text meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected; when the text meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected, detect whether there is an obstacle in the scene corresponding to the image to be detected.

[0169] In some embodiments, the detection module 302 is specifically used to extract keywords corresponding to the object from the sentence; the judgment module 303 is used to judge whether the object is an obstacle in the scene corresponding to the image to be detected based on the keywords corresponding to the object.

[0170] In some embodiments, the judgment module 303 is specifically used to search whether there are keywords corresponding to the object in the first keyword list, and the first keyword list includes at least one keyword corresponding to an obstacle; when the keyword corresponding to the object is found in the first keyword list, the object is determined to be an obstacle; or, when the keyword corresponding to the object is not found in the first keyword list, the object is determined to be a non-obstacle.

[0171] In some embodiments, the judgment module 303 is further used to send the keyword of the object to the server if the keyword corresponding to the object is not found in the first keyword list; receive the first recognition result sent by the server, and the first recognition result is used to indicate whether the object is an obstacle.

[0172] In some embodiments, the judgment module 303 is further configured to add a keyword of the object to the first keyword list when the first recognition result indicates that the object is an obstacle.

[0173] In some embodiments, the judgment module 303 is specifically used to search for keywords corresponding to the object in the second keyword list, and the second keyword list includes at least one keyword corresponding to a non-obstacle; if the keyword corresponding to the object is not found in the second keyword list, the object is determined to be an obstacle; or, if the keyword corresponding to the object is found in the second keyword list, the object is determined to be a non-obstacle.

[0174] In some embodiments, the judgment module 303 is also used to send the keyword of the object to the server if the keyword corresponding to the object is not found in the second keyword list; receive the second recognition result sent by the server, and the second recognition result is used to indicate whether the object is a non-obstacle.

[0175] In some embodiments, the judgment module 303 is further configured to add the keyword of the object to the second keyword list when the second recognition result indicates that the object is a non-obstacle.

[0176] In some embodiments, the detection module 302 is also used to input the sentence into a text classification model to obtain a text classification result output by the text classification model, wherein the text classification result includes a first classification result or a second classification result, wherein the first classification result is used to indicate that the text meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected, and the second classification result is used to indicate that the text meaning of the sentence is not used to describe the existence of an object in the scene corresponding to the image to be detected.

[0177] In some embodiments, a text classification model is trained in the following manner: a first training sample set is obtained, the first training sample set includes multiple first samples and respective labels of the multiple first samples, the first sample is second text information used to describe the first sample image, the label of the first sample is used to indicate whether the text meaning of a sentence of the second text information is used to describe whether there is an object in a scene corresponding to the first sample image, and the scene corresponding to the first sample image is the same as the scene corresponding to the image to be detected; based on the first training sample set, an initial text classification model is trained to obtain a trained text classification model.

[0178] In some embodiments, the detection module 302 is specifically used to input the first text information into a text classification model to obtain a text classification result output by the text classification model, wherein the text classification result includes a third classification result and a fourth classification result, wherein the third classification result is used to indicate that the text meaning of the first text information is used to describe the existence of obstacles in the scene corresponding to the image to be detected, and the fourth classification result is used to indicate that the text meaning of the first text information is not used to describe the existence of obstacles in the scene corresponding to the image to be detected.

[0179] In some embodiments, the acquisition module 301 is specifically used to input the image to be detected into the image-text generation model to obtain the first text information output by the image-text generation model.

[0180] In some embodiments, the image-text generation model is trained in the following manner: obtaining a second training sample set, the second training sample set including multiple second sample images and third text information corresponding to each of the multiple second sample images; based on the second training sample set, training the initial image-text generation model to obtain a trained image-text generation model.

[0181] In some embodiments, the sending module 304 is used to send warning information for prompting the existence of obstacles when it is detected that there are obstacles in the scene corresponding to the image to be detected.

[0182] In the case of implementing the functions of the above-mentioned integrated modules in the form of hardware, the embodiments of the present disclosure provide a possible structure of the electronic device involved in the above-mentioned embodiments. Fig. 9 As shown, the electronic device 400 includes: a processor 402 and a bus 404. Optionally, the electronic device 400 may further include a memory 401; optionally, the electronic device 400 may further include a communication interface 403.

[0183] The processor 402 may be a processor that implements or executes various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of the present disclosure. The processor 402 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of the present disclosure. The processor 402 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0184] The communication interface 403 is used to connect with other devices via a communication network, such as Ethernet, wireless access network, wireless local area network (WLAN), etc.

[0185] The memory 401 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program codes in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0186] As a possible implementation, the memory 401 may exist independently of the processor 402, and the memory 401 may be connected to the processor 402 via a bus 404 to store instructions or program codes. When the processor 402 calls and executes the instructions or program codes stored in the memory 401, the obstacle detection method provided in the embodiment of the present disclosure can be implemented.

[0187] In another possible implementation, the memory 401 may also be integrated with the processor 402 .

[0188] The bus 404 may be an extended industry standard architecture (EISA) bus, etc. The bus 404 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig. 9 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0189] Some embodiments of the present disclosure provide a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium), which stores computer program instructions. When the computer program instructions are executed on a computer, the computer executes the obstacle detection method as described in any of the above embodiments.

[0190] Exemplarily, the above-mentioned computer-readable storage media may include, but are not limited to: magnetic storage devices (e.g., hard disks, floppy disks or magnetic tapes, etc.), optical disks (e.g., compact disks (CD), digital versatile disks (DVD), etc.), smart cards and flash memory devices (e.g., erasable programmable read-only memory (EPROM), cards, sticks or key drives, etc.). The various computer-readable storage media described in the present disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing and / or carrying instructions and / or data.

[0191] An embodiment of the present disclosure provides a computer program product including instructions. When the computer program product is run on a computer, the computer is enabled to execute the obstacle detection method described in any one of the above embodiments.

[0192] The above is only a specific implementation of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present disclosure should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.

Claims

1. An obstacle detection method, characterized in that: The method comprises: Acquire an image to be detected, and generate first text information for describing the image to be detected based on the image to be detected; Based on the first text information, it is detected whether there is an obstacle in the scene corresponding to the image to be detected.

2. The method according to claim 1, characterized in that: The detecting, based on the first text information, whether there is an obstacle in the scene corresponding to the image to be detected includes: Dividing the first text information into at least one sentence; Identifying whether the textual meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected; In the case where the textual meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected, it is detected whether there is an obstacle in the scene corresponding to the image to be detected.

3. The method according to claim 2, characterized in that In the case where the textual meaning of the sentence is used to describe that there is an object in the scene corresponding to the image to be detected, detecting whether there is an obstacle in the scene corresponding to the image to be detected includes: Extracting keywords corresponding to the object from the sentence; Based on the keyword corresponding to the object, it is determined whether the object is an obstacle in the scene corresponding to the image to be detected.

4. The method according to claim 3, characterized in that The determining, based on the keyword corresponding to the object, whether the object is an obstacle in the scene corresponding to the image to be detected includes: Searching whether there is a keyword corresponding to the object in a first keyword list, wherein the first keyword list includes at least one keyword corresponding to an obstacle; In the case where a keyword corresponding to the object is found in the first keyword list, determining that the object is an obstacle; or, When the keyword corresponding to the object is not found in the first keyword list, the object is determined to be a non-obstacle.

5. The method according to claim 4, characterized in that The determining, based on the keyword corresponding to the object, whether the object is an obstacle in the scene corresponding to the image to be detected further includes: If the keyword corresponding to the object is not found in the first keyword list, sending the keyword of the object to the server; A first recognition result sent by the server is received, where the first recognition result is used to indicate whether the object is an obstacle.

6. The method according to claim 5, characterized in that The method further comprises: In a case where the first recognition result is used to indicate that the object is an obstacle, a keyword of the object is added to the first keyword list.

7. The method according to claim 3, characterized in that The determining, based on the keyword corresponding to the object, whether the object is an obstacle in the scene corresponding to the image to be detected includes: Searching for a keyword corresponding to the object in a second keyword list, where the second keyword list includes at least one keyword corresponding to a non-obstacle; If no keyword corresponding to the object is found in the second keyword list, determining that the object is an obstacle; or, When the keyword corresponding to the object is found in the second keyword list, the object is determined to be a non-obstacle.

8. The method according to claim 7, characterized in that The determining, based on the keyword corresponding to the object, whether the object is an obstacle in the scene corresponding to the image to be detected further includes: If no keyword corresponding to the object is found in the second keyword list, sending the keyword of the object to the server; A second recognition result sent by the server is received, where the second recognition result is used to indicate whether the object is a non-obstacle.

9. The method according to claim 8, characterized in that The method further comprises: In a case where the second recognition result is used to indicate that the object is a non-obstacle, a keyword of the object is added to the second keyword list.

10. The method according to claim 2, characterized in that The identifying whether the textual meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected includes: The sentence is input into a text classification model to obtain a text classification result output by the text classification model, wherein the text classification result includes a first classification result or a second classification result, wherein the first classification result is used to indicate that the text meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected, and the second classification result is used to indicate that the text meaning of the sentence is not used to describe the existence of an object in the scene corresponding to the image to be detected.

11. The method according to claim 10, characterized in that The text classification model is trained in the following way: Acquire a first training sample set, the first training sample set comprising a plurality of first samples and respective labels of the plurality of first samples, the first sample being second text information used to describe a first sample image, the label of the first sample being used to indicate whether a textual meaning of a sentence of the second text information is used to describe whether there is an object in a scene corresponding to the first sample image, and the scene corresponding to the first sample image is the same as the scene corresponding to the image to be detected; Based on the first training sample set, an initial text classification model is trained to obtain a trained text classification model.

12. The method according to claim 1, characterized in that The detecting, based on the first text information, whether there is an obstacle in the scene corresponding to the image to be detected includes: The first text information is input into a text classification model to obtain a text classification result output by the text classification model, wherein the text classification result includes a third classification result and a fourth classification result, wherein the third classification result is used to indicate that the text meaning of the first text information is used to describe the existence of an obstacle in the scene corresponding to the image to be detected, and the fourth classification result is used to indicate that the text meaning of the first text information is not used to describe the existence of an obstacle in the scene corresponding to the image to be detected.

13. The method according to claim 1, characterized in that The step of generating first text information for describing the image to be detected based on the image to be detected includes: The image to be detected is input into a picture-text generation model to obtain the first text information output by the picture-text generation model.

14. The method according to claim 13, characterized in that The image-text generation model is trained in the following way: Acquire a second training sample set, where the second training sample set includes a plurality of second sample images and third text information corresponding to each of the plurality of second sample images; Based on the second training sample set, the initial image-text generation model is trained to obtain the trained image-text generation model.

15. The method according to claim 1, characterized in that The method further comprises: When it is detected that an obstacle exists in the scene corresponding to the image to be detected, an alarm message for prompting the existence of the obstacle is sent.

16. An electronic device, characterized in that: include: a processor and a memory for storing instructions executable by the processor; The processor is configured to execute the instructions so that the electronic device performs the method as claimed in any one of claims 1 to 15.

17. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on an electronic device, the electronic device executes the method according to any one of claims 1 to 15.