Obstacle detection method, electronic device and storage medium

By generating the first text information to describe the characteristics of the image to be detected, the problem of low accuracy of obstacle detection in the prior art is solved, and higher accuracy and comprehensiveness of obstacle detection are achieved, and vehicle driving safety and driving experience are improved.

WO2025092466A1PCT designated stage expired Publication Date: 2025-05-08ZTE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/125877
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-03
Filing Date
2024-10-18
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The prior art has low accuracy when detecting obstacles around vehicles, especially under different light and angle conditions, and lacks an ideal visible light technology to detect obstacles that affect vehicle driving under any circumstances.

Method used

By acquiring the image to be detected, first text information for describing the image to be detected is generated, and understanding of image features is enhanced, thereby detecting obstacles based on the first text information, and improving the accuracy and comprehensiveness of the detection.

Benefits of technology

It improves vehicle driving safety and driving experience, enhances the accuracy and comprehensiveness of obstacle detection, can carry more detection categories, and achieves efficient detection and identification of various obstacles in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024125877_08052025_PF_FP_ABST
    Figure CN2024125877_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an obstacle detection method, an electronic device and a storage medium. The obstacle detection method comprises: acquiring an image under detection, and on the basis of said image, generating first text information used for describing said image; and on the basis of the first text information, detecting whether an obstacle is present in a scenario corresponding to said image.
Need to check novelty before this filing date? Find Prior Art

Description

Obstacle detection method, electronic device and storage medium

[0001] This disclosure claims priority to Chinese patent application No. 202311464434.2, filed on November 03, 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present disclosure relates to the field of target detection, and in particular to an obstacle detection method, electronic equipment, and storage medium. Background Art

[0003] With the development of society and economy, vehicle safety has attracted much attention. To ensure vehicle safety during driving, timely detection of obstacles around the vehicle and providing warnings to users are particularly important. Some technologies for detecting obstacles around vehicles are usually based on image segmentation algorithms.

[0004] Summary of the Invention

[0005] In a first aspect, an embodiment of the present disclosure provides an obstacle detection method. The obstacle detection method includes:

[0006] Acquire an image to be detected, and generate first text information for describing the image to be detected based on the image to be detected;

[0007] Based on the first text information, it is detected whether there is an obstacle in the scene corresponding to the image to be detected.

[0008] In a second aspect, an embodiment of the present disclosure provides an obstacle detection device. The obstacle detection device includes: an acquisition module and a detection module;

[0009] The acquisition module is configured to acquire an image to be detected and generate first text information describing the image to be detected based on the image to be detected;

[0010] The detection module is configured to detect whether there is an obstacle in the scene corresponding to the image to be detected based on the first text information.

[0011] In a third aspect, embodiments of the present disclosure provide an electronic device comprising: a memory and a processor; the memory and the processor are coupled; the memory is configured to store a computer program; and the processor implements the obstacle detection method of the first aspect when executing the computer program.

[0012] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the obstacle detection method of the first aspect described above is implemented.

[0013] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, which includes computer program instructions, and when the computer program instructions are executed by a processor, the obstacle detection method of the first aspect is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] To more clearly illustrate the technical solutions of the present disclosure, the following briefly introduces the drawings required for use in some embodiments of the present disclosure. Obviously, the drawings described below are only drawings of some embodiments of the present disclosure, and those skilled in the art can also derive other drawings based on these drawings.

[0015] FIG1 is a schematic structural diagram of a system according to some embodiments.

[0016] FIG2 is a flowchart of an obstacle detection method according to some embodiments.

[0017] FIG3 is a schematic diagram of an image to be detected according to some embodiments.

[0018] FIG4 is a flowchart of another obstacle detection method according to some embodiments.

[0019] FIG5 is a flowchart of yet another obstacle detection method according to some embodiments.

[0020] FIG6 is a flowchart of yet another obstacle detection method according to some embodiments.

[0021] FIG7 is a flowchart of yet another obstacle detection method according to some embodiments.

[0022] FIG8 is a schematic structural diagram of an obstacle detection device according to some embodiments.

[0023] FIG9 is a schematic structural diagram of an electronic device according to some embodiments. DETAILED DESCRIPTION

[0024] The following will clearly and completely describe the technical solutions of this disclosure in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of this disclosure, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0025] It should be noted that in this disclosure, expressions such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described in this disclosure as "exemplarily" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of expressions such as "exemplarily" or "for example" is intended to present the relevant concepts in a detailed manner.

[0026] In the following, the terms "first," "second," etc. are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the quantity of the technical features indicated. Therefore, a feature specified as "first," "second," etc. may explicitly or implicitly include one or more of the features.

[0027] In the description of this disclosure, unless otherwise specified, " / " means "or." For example, A / B can mean A or B. "And / or" herein is simply a description of an association relationship between associated objects, indicating that three possible relationships exist. For example, A and / or B can mean: only A, only B, and A and B. Furthermore, "at least one" means one or more, and "a plurality" means two or more.

[0028] To facilitate understanding, relevant concepts involved in the embodiments of the present disclosure are first briefly introduced.

[0029] Vehicle-to-everything (V2X) wireless communication technology refers to the communication and interaction between vehicles and everything. It is a key technology in intelligent transportation systems, aiming to improve road safety, traffic efficiency, and passenger comfort. V2X technology enables vehicles to communicate in real time with other vehicles, infrastructure, and other traffic participants such as pedestrians. This communication can be carried out through wireless communication technologies such as wireless fidelity (Wi-Fi) networks and cellular networks, transmitting information such as vehicle location, speed, and direction.

[0030] Cellular-V2X (Cv2x) is used for direct wireless communication between vehicles.

[0031] In CV2X, the PC5 method is a specific physical layer and scheduling method for direct communication between vehicles. It enables Sidelink communication between vehicles. Sidelink is an air interface technology in the CV2X system that allows vehicles to establish direct communication connections without going through a base station or network relay. This direct communication can be carried out in the absence of network coverage or network congestion, with lower latency and higher reliability.

[0032] The You Only Look Once (YOLO) object detection algorithm is used to detect and localize multiple objects in an image or video in real time. Compared to traditional object detection algorithms, the YOLO algorithm is faster and more accurate. The core idea of ​​the YOLO algorithm is to transform the object detection problem into a regression problem. It divides the input image into a fixed-size grid and predicts the bounding box and class probability of an object in each grid cell. This means that the YOLO algorithm only needs a single forward pass to simultaneously detect and classify objects, hence the name "You Only Look Once."

[0033] The above is an introduction to some concepts involved in the embodiments of the present disclosure, which will not be repeated below.

[0034] In the fields of autonomous driving and V2X vehicle-to-infrastructure collaborative perception, traditional computer vision (CV) algorithms have limited target detection capabilities and can miss detections under varying lighting conditions and angles. Currently, no ideal visible light technology can detect obstacles that impede vehicle movement in all conditions. For example, millimeter-wave sensors can detect obstacles, but only in motion. LiDAR can also detect obstacles, but it places a high premium on vehicle costs.

[0035] In response to the above problems, an embodiment of the present disclosure provides an obstacle detection method, the idea of ​​which is: by acquiring an image to be detected and generating first text information for describing the image to be detected (the first text information can describe the image features in the image to be detected in natural language), the comprehensiveness and accuracy of the understanding of the image features of the image to be detected can be enhanced, so that when detecting whether there is an obstacle in the scene corresponding to the image to be detected based on the first text information, the accuracy and comprehensiveness of the obstacle detection in the scene corresponding to the image to be detected are improved, thereby improving the vehicle's driving safety and driving experience.

[0036] At the same time, the method provided by the embodiment of the present disclosure detects and identifies the images to be detected based on a large neural network model, can carry more detection categories, and realize efficient detection and identification of various obstacles in different scenes, avoiding the long-tail effect brought by small models and improving the user experience.

[0037] Referring to Figure 1, which is a schematic diagram of the structure of a system involved in the obstacle detection method provided by an embodiment of the present disclosure, the system includes an image acquisition device 100 and a detection device 200, which are communicatively connected.

[0038] The image acquisition device 100 is used to capture and record the image to be detected. After capturing the image to be detected, the image acquisition device 100 sends the image to be detected to the detection device 200 so that the detection device 200 can detect whether there is an obstacle in the scene corresponding to the image to be detected.

[0039] As an example, the image acquisition device 100 can directly capture an image in the target scene and use the captured image as the image to be detected. In some embodiments, the target scene can be a road scene, a parking lot scene, or a gas station scene.

[0040] As another example, the image acquisition device 100 can capture video information of a target scene. The video information includes continuous image frames acquired by the image acquisition device 100 at a certain frame rate, each of which can be an independent static image. The image acquisition device 100 can intercept an image frame at a certain point in time (i.e., a screenshot operation), or use video frame extraction technology to extract an image frame at a certain point in time from the video information and use the extracted image frame as the image to be detected.

[0041] In some embodiments, the image acquisition device 100 may be a device including a visible light sensor, such as a camera, a photoresistor, etc.

[0042] It should be noted that there may be one or more image acquisition devices 100. The embodiment of the present disclosure does not limit the number of image acquisition devices 100.

[0043] It should be noted that the embodiment of the present disclosure mainly uses the image acquisition device 100 to capture the image to be detected, but in actual implementation, the image to be detected can also be captured by other methods, for example, by creating a three-dimensional point cloud map through a lidar, etc. The embodiment of the present disclosure does not limit the method of capturing the image to be detected.

[0044] The detection device 200 is used to obtain the image to be detected sent by the image acquisition device 100, and detect whether there is an obstacle in the scene corresponding to the image to be detected.

[0045] In some embodiments, the detection device 200 generates first text information for describing the image to be detected based on the image to be detected, and detects whether there is an obstacle in the scene corresponding to the image to be detected based on the first text information.

[0046] In some embodiments, the detection device 200 may be a processor. In some embodiments, the processor may be a central processing unit (CPU), a general-purpose network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. Alternatively, the detection device 200 may also be other devices having processing functions, such as circuits, devices, or software modules, which are not limited in the present disclosure.

[0047] In some embodiments, the above system further includes: any one or more of a domain controller, a mobile edge computing (MEC) unit, and a roadside computing unit (RCU).

[0048] In an autonomous driving system, the domain controller can coordinate vehicle driving based on the detection results of the detection device 200 on the image to be detected, for example, path planning, speed control, and inter-vehicle coordination. In some embodiments, the domain controller can be integrated into the detection device 200 as part of the detection device 200; alternatively, the domain controller can exist independently of the detection device 200.

[0049] In the vehicle-road cooperative perception system, the MEC unit and / or RCU can receive the images to be detected captured by the image acquisition device 100 and use algorithms to process and analyze the images in real time, performing decision-making and control tasks at the edge, such as target detection, tracking, and path planning. In some embodiments, the MEC unit and / or RCU can be integrated into the detection device 200 as part of the detection device 200; alternatively, the MEC unit or RCU can exist independently of the detection device 200.

[0050] It should be noted that the system architecture and application scenarios described in the embodiments of the present disclosure are intended to more clearly illustrate the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present disclosure are also applicable to similar technical problems.

[0051] An obstacle detection method provided by an embodiment of the present disclosure is described below with reference to the accompanying drawings.

[0052] Referring to Figure 2, which is a flow chart of an obstacle detection method provided by an embodiment of the present disclosure, as shown in Figure 2, the obstacle detection method provided by an embodiment of the present disclosure is applied to a detection device and can be implemented as follows: S101 to S102.

[0053] S101: Acquire an image to be detected, and generate first text information for describing the image to be detected based on the image to be detected.

[0054] As an example, when a detection device acquires images to be detected from an image acquisition device, multiple image acquisition devices can be assigned to the detection area, with each of the multiple image acquisition devices performing independent detection. For example, the detection device can be divided into image acquisition device A and image acquisition device B to detect different locations in the detection area and capture images to be detected in the detection area.

[0055] As another example, when the detection device obtains the image to be detected from the image acquisition device, the picture captured by a certain image acquisition device can be divided into different sub-pictures, each sub-picture corresponding to a different sub-area within the detection area, and then the sub-picture within each sub-area (that is, the image to be detected within each sub-area) is obtained in units of sub-areas.

[0056] In some embodiments, the first text information is used to describe the type, shape, size, etc. of the object in the image to be detected, as well as the lighting conditions, weather conditions, etc. of the image to be detected. For example, the first text information may be "The sky is blue. There is a dog on the road."

[0057] It should be noted that the language of the first text information can be Chinese, English, etc. Depending on different implementation requirements, the language of the first text information can also be different. The embodiment of the present disclosure does not limit the language of the first text information.

[0058] It can be understood that compared to the problem of low accuracy in obstacle detection based on image segmentation algorithms in some technologies, the method provided by the embodiments of the present disclosure obtains an image to be detected and generates first text information for describing the image to be detected. The image features of the image to be detected can be expressed more comprehensively and accurately using the first text information, so that when obstacles are subsequently detected based on the first text information, the detection of obstacles is more comprehensive and accurate.

[0059] S102: Based on the first text information, detect whether there is an obstacle in the scene corresponding to the image to be detected.

[0060] Exemplarily, if the first text information describes that there is a pedestrian and a conical object in the scene corresponding to the image to be detected, then the conical object is the obstacle in the scene corresponding to the image to be detected.

[0061] It can be understood that the obstacle detection method provided by the embodiment of the present disclosure, by acquiring the image to be detected and generating first text information for describing the image to be detected (the first text information can describe the image features in the image to be detected in natural language), can enhance the comprehensiveness and accuracy of the understanding of the image features of the image to be detected, so that when detecting whether there is an obstacle in the scene corresponding to the image to be detected based on the first text information, the accuracy and comprehensiveness of the obstacle detection in the scene corresponding to the image to be detected are improved, thereby improving the vehicle's driving safety and driving experience.

[0062] In some embodiments, when an obstacle is detected in the scene corresponding to the image to be detected, the detection device may further send an alarm message for prompting the existence of the obstacle.

[0063] In some embodiments, when an obstacle is detected in a scene corresponding to an image to be detected, the detection device sends an alert to an electronic device associated with the scene corresponding to the image to be detected. For example, the detection device can transmit the alert to roadside traffic participants or cloud-based traffic managers via a communication link using a device such as an RCU (roadside unit) or MEC.

[0064] In some embodiments, when no obstacles are detected in the scene corresponding to the image to be detected, the detection device may also send an alarm message to the electronic device involved in the scene corresponding to the image to be detected. The content of the alarm message may be determined based on whether there are objects in the scene. For example, if there are no obstacles in the scene, but there are known common traffic participants, such as pedestrians, buses, etc., the alarm message is used to prompt the traffic participants or the traffic manager in the cloud to track and detect the traffic participants; if there are no obstacles or other objects in the scene, the alarm message is used to prompt the traffic participants or the traffic manager in the cloud that there is nothing abnormal in the scene.

[0065] In some embodiments, traffic participants include motor vehicles, non-motor vehicles, pedestrians, and so on. As an example, the detection device can use CV2X technology to send warning information to traffic participants. In some embodiments, the warning information can be a roadside infrastructure / road condition event (RSI / RTE) message. CV2X uses PC5 to broadcast RSI / RTE messages. The event ID of the message uses the national standard traffic event code, and the description field of the message is filled with the first text information generated by the image-text generation model. This first text information can be the original text information or the processed text information.

[0066] As another example, a detection device can use 5G technology to send warning information to traffic participants. For example, the detection device can transmit warning information through a base station or a cloud server. Traffic participants can also obtain warning information from the server through their mobile phones or in-vehicle software (such as navigation software).

[0067] It can be understood that the method provided by the embodiment of the present disclosure sends a warning message to indicate the existence of an obstacle when an obstacle is detected in the scene corresponding to the image to be detected, which can promptly remind the user to pay attention to safety, ensure the user's driving safety, and improve the user's driving experience.

[0068] In some embodiments, the first text information can be generated by some image-text conversion models. Based on the image to be detected, generating the first text information for describing the image to be detected can be implemented as follows: inputting the image to be detected into the image-text generation model to obtain the first text information output by the image-text generation model.

[0069] In some embodiments, the detection device inputs the image to be detected into the image-text generation model, which performs artificial intelligence (AI) forward reasoning on the image to be detected to obtain the first text information. For example, as shown in FIG3 , if FIG3 is input as the image to be detected into the image-text generation model, the image-text generation model will output the first text information corresponding to the image: "The sky is blue. There is a cone-shaped object on the road."

[0070] In some embodiments, the detection device also collects the first text information of each image to be detected, marks the region (region) corresponding to each image to be detected, generates a region identification code (regionID), and associates it with a key-value dictionary of prompt (taking the first text information as the value and associating it with the key corresponding to the regionID) to facilitate subsequent processing and identification of each image to be detected.

[0071] In some embodiments, the detection device may determine a unique regionID based on the content of each image to be detected, ensuring that each region corresponding to the image to be detected has a different identifier.

[0072] It can be understood that compared with some technologies based on traditional convolutional neural networks, which use semantic segmentation algorithms to identify obstacles in the scene and have a high false alarm rate, the embodiment of the present disclosure is based on a graphic generation model, which can accurately determine the image features of the image to be detected, thereby improving the accuracy of obstacle detection.

[0073] In some embodiments, as shown in FIG4 , based on the first text information, detecting whether there is an obstacle in the scene corresponding to the image to be detected can be implemented as follows: S1021 to S1023 .

[0074] S1021. Divide the first text information into at least one sentence.

[0075] As an example, the first text information can be divided based on punctuation marks (e.g., periods, commas, etc.) in the first text information to obtain at least one sentence. For example, if the first text information is "The sky is blue. There is a cone-shaped object on the road," the first text information can be divided based on punctuation marks to obtain the sentence "The sky is blue" and the sentence "There is a cone-shaped object on the road."

[0076] As another example, the first text information can be divided based on a machine learning model (such as a sentence boundary detection model). The sentence boundary detection model can continue to detect the input first text information and find a suitable division position for the first text information to obtain at least one sentence.

[0077] It should be noted that the above are only some examples of dividing the first text information given in the embodiments of the present disclosure. In actual implementation, different methods can be selected according to different actual needs; the embodiments of the present disclosure do not limit the method of dividing the first text information.

[0078] S1022: Identify whether the textual meaning of each clause in the at least one clause is used to describe the existence of an object in the scene corresponding to the image to be detected.

[0079] In some embodiments, the above S1022 can be implemented as: inputting each sentence of the at least one sentence into a text classification model to obtain a text classification result output by the text classification model.

[0080] In some embodiments, the text classification result includes a first classification result or a second classification result. The first classification result is used to indicate that the text meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected, and the second classification result is used to indicate that the text meaning of the sentence is not used to describe the existence of an object in the scene corresponding to the image to be detected.

[0081] For example, if the sentence "The sky is blue" is input into the text classification model and the text classification result is the second classification result, it means that the text meaning of the sentence "The sky is blue" is not used to describe the existence of an object in the scene corresponding to the image to be detected. If the sentence "There is a cone-shaped object on the road" is input into the text classification model and the text classification result is the first classification result, it means that the text meaning of the sentence "There is a cone-shaped object on the road" is used to describe the existence of an object in the scene corresponding to the image to be detected.

[0082] In some embodiments, to determine whether the textual meaning of a sentence is used to describe the presence of an object in the scene corresponding to the image to be detected, the following methods may also be considered:

[0083] The keyword matching method creates a list of keywords that contain as many objects as possible. Then, it performs keyword matching on the sentences to see if any of the sentences contain keywords corresponding to the objects. If any of the sentences contain keywords corresponding to the objects, it is determined that the object exists in the scene corresponding to the image to be detected.

[0084] Natural language processing (NLP) analyzes each sentence using techniques such as bag-of-words models or word embedding models (such as Word2Vec and BERT) to understand the semantics of the sentence. For example, if a sentence contains descriptions such as "a bag is in the middle of the road" or "an unidentified animal is crossing the road," these information may indicate the presence of an object in the scene corresponding to the image being detected.

[0085] The rule engine approach involves building a rule engine that uses a series of rules and pattern matching to determine the presence of obstacles based on the text description of each sentence. For example, if a sentence contains descriptions such as "there are scattered objects on the roadside" or "the road is icy," it can be determined that an object exists in the scene corresponding to the image being detected.

[0086] It should be noted that whether the textual meanings of some recognition sentences provided in the above embodiments of the present disclosure are used to describe examples of objects existing in the scene corresponding to the image to be detected, different methods can be flexibly selected according to different actual conditions during implementation, and the embodiments of the present disclosure do not limit this.

[0087] It is understood that the method provided in the embodiments of the present disclosure, by dividing the first text information into at least one sentence, can more precisely classify the text. Different sentences may emphasize different information or have different descriptions. Inputting each of the at least one sentence into a text classification model and obtaining a text classification result output by the text classification model can enable the text classification model to more accurately identify key information in the sentence, thereby improving classification accuracy.

[0088] S1023: When the textual meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected, detect whether there is an obstacle in the scene corresponding to the image to be detected.

[0089] In some embodiments, the above S1023 can be implemented as: a1 to a2.

[0090] a1. Extract keywords corresponding to objects from sentences.

[0091] For example, if the sentence is "There is a cone-shaped object on the road", the keyword corresponding to the object to be extracted is: cone. If the sentence is "A car is driving on the road", the keyword corresponding to the object to be extracted is: car.

[0092] a2. Based on the keywords corresponding to the object, determine whether the object is an obstacle in the scene corresponding to the image to be detected.

[0093] As an implementation of step a2 above, the first keyword list is searched for keywords corresponding to the object. If a keyword corresponding to the object is found in the first keyword list, the object is determined to be an obstacle; or if a keyword corresponding to the object is not found in the first keyword list, the object is determined to be non-obstacle.

[0094] The first keyword list includes at least one keyword corresponding to an obstacle. For example, the first keyword list includes keywords corresponding to uncommon building materials, keywords corresponding to wild animals, keywords corresponding to cubes of different shapes, and the like.

[0095] For example, if the keyword corresponding to the object is "cone," the first keyword list is searched for the keyword "cone." If the keyword corresponding to the object is "car," the first keyword list is searched for the keyword "car." If "cone" is found in the first keyword list, the object is determined to be an obstacle. If "cone" is not found in the first keyword list, the object is determined to be a non-obstacle.

[0096] It can be understood that the method provided in the embodiments of the present disclosure is based on keywords corresponding to objects, and can identify obstacles in the image to be detected by matching them with preset obstacle keywords, thereby avoiding manual obstacle identification and saving manpower, material resources and time costs. At the same time, for obstacle detection tasks in some specific scenarios, such as autonomous driving, robot navigation and other scenarios, the obstacle detection technology based on object keywords can improve the accuracy of obstacle identification.

[0097] In some embodiments, when the keyword corresponding to the object is not found in the first keyword list, the detection device sends the keyword of the object to the server. Based on the keyword of the object, the server further identifies the keyword of the object using an intelligent algorithm or manually and generates a first recognition result. The first recognition result is used to indicate whether the object is an obstacle. The detection device receives the first recognition result sent by the server, and when the first recognition result is used to indicate that the object is an obstacle, the detection device adds the keyword of the object to the first keyword list. Exemplarily, if the keyword of the object is cone, and cone is not found in the first keyword list, the detection device sends the keyword cone of the object to the server, the server determines that the object is an obstacle and generates a first recognition result, the detection device receives the first recognition result sent by the server, and updates the keyword cone into the first keyword list.

[0098] It is understood that the method provided by the embodiments of the present disclosure, when no keywords corresponding to an object are found in the first keyword list, sends the object's keywords to the server, further enabling intelligent algorithm or manual identification of whether the object is an obstacle, thereby improving the accuracy of obstacle detection. Furthermore, when the first recognition result indicates that the object is an obstacle, the method provided by the embodiments of the present disclosure adds the object's keywords to the first keyword list, thereby expanding the keyword list and improving the accuracy of obstacle detection.

[0099] As another implementation of step a2 above: searching for keywords corresponding to the object in the second keyword list. If no keywords corresponding to the object are found in the second keyword list, the object is determined to be an obstacle; or, if the keywords corresponding to the object are found in the second keyword list, the object is determined to be non-obstacle.

[0100] The second keyword list includes at least one keyword corresponding to a non-obstacle.

[0101] For example, the second keyword list includes keywords corresponding to pedestrians, vehicles, and zebra crossings. For example, if the keyword corresponding to the object is "car," the second keyword list is searched for the presence of the keyword "car." For example, if the keyword corresponding to the object is "cone," if "cone" is not found in the second keyword list, the object is determined to be an obstacle; if "cone" is found in the second keyword list, the object is determined to be a non-obstacle.

[0102] It should be noted that non-obstructions refer to other objects or subjects that are unrelated to the task objective or context within a specific task or scenario. Traffic targets such as pedestrians, electric vehicles, bicycles, buses, and zebra crossings are typically non-obstructions and are common traffic participants in the traffic environment. In actual implementation, the definition of non-obstructions varies depending on the task or scenario, so the present disclosed embodiments do not limit the content of non-obstructions.

[0103] In some embodiments, when the keyword corresponding to the object is not found in the second keyword list, the detection device sends the keyword of the object to the server. The server further identifies the keyword of the object using an intelligent algorithm or manually based on the keyword of the object and generates a second recognition result. The second recognition result is used to indicate whether the object is a non-obstacle. The detection device receives the second recognition result sent by the server, and when the second recognition result is used to indicate that the object is a non-obstacle, the detection device adds the keyword of the object to the second keyword list. Exemplarily, assuming that the corresponding keyword of the object is bus, and bus is not found in the second keyword list, the keyword bus of the object is sent to the server. The server determines that the object is a non-obstacle and generates a second recognition result. The detection device receives the second recognition result sent by the server and updates the keyword bus into the second keyword list.

[0104] It should be noted that if the second recognition result is used to indicate that the object is not an obstacle, it means that the keywords in the second keyword list are not comprehensive enough. At this time, the keywords of the object need to be added to the second keyword list to expand and update the second keyword list.

[0105] It is understood that the method provided by the embodiments of the present disclosure, when no keywords corresponding to an object are found in the second keyword list, sends the object's keywords to the server, further enabling intelligent algorithm or manual identification of whether the object is a non-obstacle, thereby improving the accuracy of obstacle detection. Furthermore, when the second recognition result indicates that the object is a non-obstacle, the method provided by the embodiments of the present disclosure adds the object's keywords to the second keyword list, thereby expanding the keyword list and improving the accuracy of obstacle detection.

[0106] As another implementation of item a2 above, the detection device searches both the first and second keyword lists for keywords corresponding to the object. If no keywords corresponding to the object are found in the first keyword list or the second keyword list, the detection device sends the object's keywords to the server. The server further identifies the object's keywords based on the object's keywords using an intelligent algorithm or manual recognition, and generates a third recognition result. The third recognition result indicates whether the object is an obstacle. The detection device receives the third recognition result sent by the server and, if the third recognition result indicates that the object is an obstacle, adds the object's keywords to the first keyword list; or, if the third recognition result indicates that the object is not an obstacle, adds the object's keywords to the second keyword list. For example, assuming the keyword corresponding to the object is "bus," and "bus" is not found in the first keyword list or the second keyword list, the detection device sends the object's keyword "bus" to the server. The server determines that the object is not an obstacle and generates a third recognition result. The detection device receives the third recognition result sent by the server and adds the keyword "bus" to the second keyword list.

[0107] It is understood that through the method provided by the embodiments of the present disclosure, the detection device can search for keywords corresponding to an object in the first keyword list and the second keyword list, respectively, making the search method more flexible and extending the search scope. At the same time, the method provided by the embodiments of the present disclosure can further improve the accuracy of obstacle detection by using intelligent algorithms or manual identification to determine whether the object is an obstacle, without searching for keywords corresponding to the object in the first keyword list and the second keyword list. In addition, based on the third recognition result, the method provided by the embodiments of the present disclosure can expand the keyword list and improve the accuracy of obstacle detection.

[0108] As another implementation of step a2 above: a search is performed to determine whether a keyword corresponding to the object exists in the third keyword list. If a keyword corresponding to the object is found in the third keyword list, a tag for the keyword corresponding to the object is determined. The tag indicates whether the object is an obstacle. Finally, a determination is made based on the tag whether the object is an obstacle. In some embodiments, the tag can be 0 or 1; 0 indicates that the object is not an obstacle; 1 indicates that the object is an obstacle.

[0109] The third keyword list includes at least one keyword corresponding to an obstacle, at least one keyword corresponding to a non-obstacle, and a label for each keyword. For example, the third keyword list includes: vehicle (label 0), cone-shaped object (label 1).

[0110] It should be noted that the content of the above-mentioned third keyword list is only an example given in the embodiment of the present disclosure. During implementation, the storage form of each keyword and each keyword label in the third keyword list can be flexibly determined, and the embodiment of the present disclosure does not limit this.

[0111] In some embodiments, when the keyword corresponding to the object is not found in the third keyword list, the detection device sends the keyword of the object to the server. The server further identifies the keyword of the object using an intelligent algorithm or manually based on the keyword of the object and generates a fourth recognition result. The fourth recognition result is used to indicate whether the object is an obstacle. The detection device receives the fourth recognition result sent by the server, and based on the fourth recognition result, adds a label to the keyword of the object, and then adds the keyword of the object and its label to the third keyword list. Exemplarily, if the keyword of the object is cone, and cone is not found in the third keyword list, the detection device sends the keyword cone of the object to the server, and the server determines that the object is an obstacle and generates a fourth recognition result. After the detection device receives the fourth recognition result sent by the server, since the object is an obstacle, the label 1 is added to the keyword cone, and cone (label 1) is added to the third keyword list.

[0112] It is understood that in the method provided by the embodiments of the present disclosure, the third keyword list can include both at least one keyword corresponding to an obstacle and at least one keyword corresponding to a non-obstacle. This can extend the scope and variety of keywords in the third keyword list, thereby improving the accuracy of the keywords in the third keyword list. Furthermore, without searching the third keyword list for keywords corresponding to an object, the method provided by the embodiments of the present disclosure can further improve the accuracy of obstacle detection by using intelligent algorithms or manual methods to identify whether the object is an obstacle. Furthermore, based on the fourth recognition result, the method provided by the embodiments of the present disclosure determines the label corresponding to the keyword of the object and adds it to the third keyword list, thereby expanding the keyword list and improving the accuracy of obstacle detection.

[0113] In some embodiments, based on the first text information, detecting whether there are obstacles in the scene corresponding to the image to be detected may also not involve sentence segmentation of the first text information, but instead directly obtaining a text classification result for the first text information, and then determining whether there are obstacles in the scene corresponding to the image to be detected based on the text classification result. Exemplarily, the first text information is input into a text classification model, and the text classification result output by the text classification model is obtained.

[0114] In some embodiments, the text classification results include a third classification result and a fourth classification result. The third classification result is used to indicate that the text meaning of the first text information is used to describe the presence of obstacles in the scene corresponding to the image to be detected, and the fourth classification result is used to indicate that the text meaning of the first text information is not used to describe the presence of obstacles in the scene corresponding to the image to be detected.

[0115] For example, the first text information "The sky is blue. There is a cone-shaped object on the road" is input into the text classification model, and the classification result is the third classification result, which means that the text meaning of the first text information is used to describe the existence of an obstacle in the scene corresponding to the image to be detected.

[0116] In some embodiments, the training process of the text classification model may refer to S301 to S302 below, which will not be described in detail in the embodiments of the present disclosure.

[0117] It can be understood that the method provided by the embodiment of the present disclosure can not only divide the first text information into sentences and input the sentences into a text classification model, but also directly input the first text information into the text classification model to obtain the text classification results output by the text classification model. While obtaining the text classification results more flexibly, it also improves the efficiency of obstacle detection and improves the user experience.

[0118] In some embodiments, as shown in FIG5 , the above-mentioned image-text generation model is obtained through training from S201 to S202 .

[0119] S201: Obtain a second training sample set.

[0120] In some embodiments, the second training sample set includes a plurality of second sample images and third text information corresponding to each of the plurality of second sample images.

[0121] In some embodiments, when acquiring the second training sample set, multiple second sample images may be collected and corresponding text descriptions may be added to each of the second sample images to ensure that the second training sample set includes a variety of obstacle types and detection scenarios. In some embodiments, the second sample images may be images of different road conditions and scenarios, such as urban roads, rural roads, highways, and gas stations. The second sample images may be captured by a vehicle camera, a drone, a ground camera, or the like.

[0122] In some embodiments, the second sample images should cover different weather conditions, lighting conditions, and obstacle types. Before obtaining and generating the second training sample set, the second sample images need to be subjected to operations such as size cropping, color balancing, and denoising.

[0123] In some embodiments, after obtaining the second training sample set, the second sample images and the third text information in the second training sample set can be preprocessed so that the data format of the second sample images and the third text information matches the parameter format of the image-text generation model. Exemplarily, multiple second sample images can be processed so that the format of the second sample images matches the input format of the image-text generation model; for example, the multiple second sample images can be converted into multiple feature vectors as input to the image-text generation model. Exemplarily, the third text information corresponding to each of the multiple sample images can be processed so that the third text information matches the input format of the image-text generation model; for example, the third text information can be encoded into a format that the image-text generation model can understand.

[0124] It can be understood that the method provided by the embodiment of the present disclosure can comprehensively consider the impact of different road conditions, weather conditions and other conditions on images by training the image-text generation model based on images under different road conditions and scenes, thereby improving the adaptability of the image-text generation model to images to be detected in complex scenes, and thereby improving the accuracy of the first text information output by the image-text generation model.

[0125] S202: Based on the second training sample set, the initial image-text generation model is trained to obtain a trained image-text generation model.

[0126] In some embodiments, the initial image-text generation model can be a model for natural language processing. It can also be extended for multimodal tasks, combining text and images as parameter inputs. For example, the initial image-text generation model can be a large-scale open-source model such as the Chat Generative Pre-trained Transformer (ChatGPT), the DeepBooru model, or the Contrastive Language-Image Pre-training (CLIP) model.

[0127] In some embodiments, the parameter size of the initial image-text generation model needs to be determined in conjunction with the computing power allocation of the obstacle detection equipment. For example, when a large model with 6 billion parameters generates text information for an image to be detected on a roadside sensing MEC / RCU (for example, the Orin NX detection device), it takes 20 seconds to complete the model output. Combined with other AI models that are detecting the image to be detected in real time, it may take 30 seconds to complete the large model output. If the computing power of the obstacle detection equipment is limited and 30 seconds is too long, the parameters of the large model can be reduced. For example, a large model with 1 billion parameters can complete the model output in 5 seconds. If the large model is quantized, the speed will be even faster.

[0128] It should be noted that the parameter size of the image-text generation model needs to be selected in combination with the actual scenario and different requirements, and the embodiments of the present disclosure do not limit this.

[0129] In some embodiments, after preprocessing the second sample image and the third text information in the second training sample set, a multimodal fusion method can be designed to fuse the second sample image and the third text information corresponding to the second sample image, so that the initial image-text generation model can understand the relationship between the second sample image and the third text information corresponding to the second sample image. In some embodiments, the architecture of the image-text generation model can also be modified so that the initial image-text generation model can accept multiple inputs and generate corresponding outputs.

[0130] In some embodiments, the initial image-text generation model is trained using a second training sample set after multimodal fusion (the data in the second training sample set is multimodal data at this time). During the training, the image-text generation model will attempt to learn the association between the second sample image and the third text information corresponding to the second sample image to obtain a trained image-text generation model, so that after the image to be detected is input, the image-text generation model can output the first text information used to describe the image to be detected.

[0131] In some embodiments, after obtaining a trained image-text generation model, appropriate evaluation metrics can be defined to measure the performance of the image-text generation model in obstacle detection. Based on the evaluation results, the performance of the image-text generation model can be continuously improved. Once the image-text generation model performs well, it can be deployed in practical applications (for example, deploying the image-text generation model in a traffic system) so that the image-text generation model generates the first text information of the image to be detected. For example, evaluation metrics such as the correlation between the second sample image and the third text information corresponding to the second sample image, and the output accuracy of the image-text generation model can be defined.

[0132] In some embodiments, improving the performance of the image-text generation model includes: adjusting the architecture of the image-text generation model; expanding the second training sample set to increase the diversity and quantity of data; adjusting the values ​​of hyperparameters of the image-text generation model; adjusting the training strategy of the image-text generation model, etc.

[0133] In some embodiments, considering real-time and efficiency issues, it may be necessary to optimize the image and text generation model by quantizing and pruning it to accommodate actual computing resources. Based on the real-time requirements of the detection, an appropriate sampling rate is selected, such as performing a detection every 10-30 seconds. To provide more efficient AI computing power, it may be necessary to configure the detection device with a dedicated embedded AI chip or high-performance computing unit.

[0134] It is understandable that traditional image-text generation models, for example, the yolov5 model generally uses the coco dataset to detect 80 types of object types, the obj365 model has more classifications than the yolov5 model, and the imageNet dataset model has 1,000 categories of classifications. However, the above models can only carry a small number of categories, and the parameters are generally at the level of about 30MB. The parameter level is small, and the model cannot detect and identify obstacles for images to be detected that exceed the categories that the model can carry, and the limitations are relatively large. The disclosed embodiments are based on large-scale neural network models, such as the chatgpt4 model, the deepbooru model, the stablediffusion CLIP model, etc., to detect and identify images to be detected (large-scale neural network models have been trained extensively for parameters at the billion level), which can carry more detection categories, achieve efficient detection and identification of various obstacles in different scenarios, and improve the user experience.

[0135] In some embodiments, as shown in FIG6 , the text classification model is obtained through training from S301 to S302 .

[0136] S301: Obtain a first training sample set.

[0137] In some embodiments, the first training sample set includes multiple first samples and respective labels of the multiple first samples, the first sample is second text information used to describe the first sample image, the label of the first sample is used to indicate whether the textual meaning of the sentence of the second text information is used to describe whether there is an object in the scene corresponding to the first sample image, and the scene corresponding to the first sample image is the same as the scene corresponding to the image to be detected.

[0138] In some embodiments, the label of the first sample can be denoted by 0 or 1; 0 indicates that the textual meaning of the second text information is not used to describe the existence of an object in the scene corresponding to the first sample image; 1 indicates that the textual meaning of the second text information is used to describe the existence of an object in the scene corresponding to the first sample image. For example, if the first sample (i.e., the second text information corresponding to the first sample image) is "The sky is blue", the label of the first sample is 0.

[0139] S302: Based on the first training sample set, train an initial text classification model to obtain a trained text classification model.

[0140] As an example, the first sample in the first training sample set is input into the initial text classification model, and the text classification model is trained using a label corresponding to each first sample to obtain a trained text classification model.

[0141] As another example, the first sample is segmented into sentences, and a label corresponding to each sentence is determined. Then, each sentence of the first sample is input into an initial text classification model, and the text classification model is trained using the label corresponding to each sentence to obtain a trained text classification model.

[0142] It should be noted that the above-mentioned training method of the text classification model is only an example given in the embodiment of the present disclosure. During implementation, the embodiment of the present disclosure does not limit the training method of the text classification model.

[0143] For ease of understanding, the method provided in the embodiment of the present disclosure is further described below in the form of examples.

[0144] Example 1: A user drives a vehicle and detects obstacles in a road scene.

[0145] Exemplarily, as shown in FIG7 , Example 1 may be implemented as steps b1 to b12 .

[0146] Step b1: Collect images to pre-train the large model.

[0147] Step b2: Download the pre-trained large model.

[0148] Step b3: Acquire the image to be detected.

[0149] Step b4: input the image to be detected into the image-text generation model.

[0150] Step b5: Generate text information of the image to be detected based on the image-text generation model.

[0151] Step b6: parse the text information to determine whether there is an object; if so, execute step b7; if not, execute step b12.

[0152] Step b7: Determine whether the object is a common traffic target; if not, proceed to step b8; if so, proceed to step b9.

[0153] Step b8: Identify that an abnormal object (ie, obstacle) is occupying the road surface.

[0154] Step b9: Identify common traffic targets on the road surface.

[0155] Step b10: perform target detection and classification on common traffic targets.

[0156] Step b11: Structuring the classification results.

[0157] Step b12: No abnormality is found, and the test is finished.

[0158] It can be understood that the obstacle detection method provided by the embodiment of the present disclosure, by acquiring the image to be detected and generating first text information for describing the image to be detected (the first text information can describe the image features in the image to be detected in natural language), can enhance the comprehensiveness and accuracy of the understanding of the image features of the image to be detected, so that when detecting whether there is an obstacle in the scene corresponding to the image to be detected based on the first text information, the accuracy and comprehensiveness of the obstacle detection in the scene corresponding to the image to be detected are improved, thereby improving the vehicle's driving safety and driving experience.

[0159] At the same time, the method provided by the embodiment of the present disclosure detects and identifies the images to be detected based on a large neural network model, can carry more detection categories, and realize efficient detection and identification of various obstacles in different scenes, avoiding the long-tail effect brought by small models and improving the user experience.

[0160] The above mainly introduces the solution of the embodiment of the present disclosure from the perspective of method. It is understandable that in order to realize the above functions, the obstacle detection device includes at least one of the hardware structure and software modules corresponding to the execution of each function. It should be easily appreciated by those skilled in the art that, in combination with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present disclosure.

[0161] It is understandable that, in order to achieve the above functions, the obstacle detection device includes hardware structures and / or software modules corresponding to the execution of each function. It should be readily apparent to those skilled in the art that, in combination with the algorithmic steps of the various examples described in the embodiments of the present disclosure, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in hardware or in a computer software-driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present disclosure.

[0162] The embodiments of the present disclosure can divide the functional modules of the obstacle detection device according to the above-mentioned method embodiments. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one functional module. The above-mentioned integrated modules can be implemented in the form of hardware or software. It should be noted that the division of modules in the embodiments of the present disclosure is schematic and is only a logical functional division. In actual implementation, other division methods can be used. The following is an example of dividing each functional module according to each function.

[0163] Figure 8 is a schematic diagram of the structure of an obstacle detection device provided in an embodiment of the present disclosure, which can implement the obstacle detection method provided in the above method embodiment. As shown in Figure 8, obstacle detection device 300 includes: an acquisition module 301 and a detection module 302. In some embodiments, obstacle detection device 300 also includes: a determination module 303 and a sending module 304.

[0164] The acquisition module 301 is configured to acquire an image to be detected and generate first text information for describing the image to be detected based on the image to be detected.

[0165] The detection module 302 is configured to detect whether there is an obstacle in the scene corresponding to the image to be detected based on the first text information.

[0166] In some embodiments, the detection module 302 is used to divide the first text information to obtain at least one sentence; identify whether the text meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected; if the text meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected, detect whether there is an obstacle in the scene corresponding to the image to be detected.

[0167] In some embodiments, the detection module 302 is used to extract keywords corresponding to the object from the sentence; the judgment module 303 is used to judge whether the object is an obstacle in the scene corresponding to the image to be detected based on the keywords corresponding to the object.

[0168] In some embodiments, the judgment module 303 is used to search whether there are keywords corresponding to the object in the first keyword list, and the first keyword list includes at least one keyword corresponding to an obstacle; when the keyword corresponding to the object is found in the first keyword list, the object is determined to be an obstacle; or, when the keyword corresponding to the object is not found in the first keyword list, the object is determined to be a non-obstacle.

[0169] In some embodiments, the judgment module 303 is further used to send the keyword of the object to the server if the keyword corresponding to the object is not found in the first keyword list; receive the first recognition result sent by the server, and the first recognition result is used to indicate whether the object is an obstacle.

[0170] In some embodiments, the judgment module 303 is further configured to add a keyword of the object to the first keyword list when the first recognition result indicates that the object is an obstacle.

[0171] In some embodiments, the judgment module 303 is used to search for keywords corresponding to the object in a second keyword list, where the second keyword list includes at least one keyword corresponding to a non-obstacle; if the keyword corresponding to the object is not found in the second keyword list, the object is determined to be an obstacle; or, if the keyword corresponding to the object is found in the second keyword list, the object is determined to be a non-obstacle.

[0172] In some embodiments, the judgment module 303 is further used to send the keyword of the object to the server if the keyword corresponding to the object is not found in the second keyword list; receive the second recognition result sent by the server, and the second recognition result is used to indicate whether the object is a non-obstacle.

[0173] In some embodiments, the judgment module 303 is further configured to add the keyword of the object to the second keyword list when the second recognition result indicates that the object is a non-obstacle.

[0174] In some embodiments, the detection module 302 is also used to input the sentence into a text classification model to obtain a text classification result output by the text classification model. The text classification result includes a first classification result or a second classification result. The first classification result is used to indicate that the text meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected, and the second classification result is used to indicate that the text meaning of the sentence is not used to describe the existence of an object in the scene corresponding to the image to be detected.

[0175] In some embodiments, a text classification model is trained in the following manner: obtaining a first training sample set, the first training sample set including multiple first samples and respective labels of the multiple first samples, the first sample being second text information used to describe the first sample image, the label of the first sample being used to indicate whether the textual meaning of a sentence of the second text information is used to describe whether there is an object in the scene corresponding to the first sample image, and the scene corresponding to the first sample image is the same as the scene corresponding to the image to be detected; based on the first training sample set, training an initial text classification model to obtain a trained text classification model.

[0176] In some embodiments, the detection module 302 is used to input the first text information into a text classification model to obtain a text classification result output by the text classification model, wherein the text classification result includes a third classification result and a fourth classification result. The third classification result is used to indicate that the text meaning of the first text information is used to describe the existence of obstacles in the scene corresponding to the image to be detected, and the fourth classification result is used to indicate that the text meaning of the first text information is not used to describe the existence of obstacles in the scene corresponding to the image to be detected.

[0177] In some embodiments, the acquisition module 301 is used to input the image to be detected into the image-text generation model to obtain the first text information output by the image-text generation model.

[0178] In some embodiments, the image-text generation model is trained in the following manner: obtaining a second training sample set, the second training sample set including multiple second sample images and third text information corresponding to each of the multiple second sample images; based on the second training sample set, training the initial image-text generation model to obtain a trained image-text generation model.

[0179] In some embodiments, the sending module 304 is configured to send warning information indicating the presence of an obstacle when an obstacle is detected in the scene corresponding to the image to be detected.

[0180] In the case of implementing the functions of the above-mentioned integrated modules in hardware, the embodiments of the present disclosure provide a possible structure of the electronic device involved in the above-mentioned embodiments. As shown in Figure 9, the electronic device 400 includes: a processor 402 and a bus 404. In some embodiments, the electronic device 400 may also include a memory 401; in some embodiments, the electronic device 400 may also include a communication interface 403.

[0181] Processor 402 may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of the present disclosure. Processor 402 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. Processor 402 may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of the present disclosure. Processor 402 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0182] The communication interface 403 is used to connect to other devices via a communication network, such as Ethernet, wireless access network, wireless local area network (WLAN), etc.

[0183] The memory 401 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0184] As an implementation, memory 401 may exist independently of processor 402 and may be connected to processor 402 via bus 404 to store instructions or program code. When processor 402 calls and executes the instructions or program code stored in memory 401, the obstacle detection method provided in the embodiments of the present disclosure can be implemented.

[0185] In another implementation, the memory 401 may also be integrated with the processor 402 .

[0186] Bus 404 can be an Extended Industry Standard Architecture (EISA) bus, etc. Bus 404 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, FIG9 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.

[0187] Some embodiments of the present disclosure provide a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium), which stores computer program instructions. When the computer program instructions are executed on a computer, the computer executes the obstacle detection method as described in any of the above embodiments.

[0188] Exemplarily, the above-mentioned computer-readable storage media may include, but are not limited to: magnetic storage devices (e.g., hard disks, floppy disks, or magnetic tapes, etc.), optical disks (e.g., compact disks (CDs), digital versatile disks (DVDs), etc.), smart cards, and flash memory devices (e.g., erasable programmable read-only memories (EPROMs), cards, sticks, or key drives, etc.). The various computer-readable storage media described in the present disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.

[0189] An embodiment of the present disclosure provides a computer program product comprising instructions. When the computer program product is run on a computer, the computer is enabled to execute the obstacle detection method described in any one of the above embodiments.

[0190] The above is only a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or replacements within the technical scope disclosed in the present disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

Claims

1. An obstacle detection method, comprising: Acquire an image to be detected, and generate first text information for describing the image to be detected based on the image to be detected; Based on the first text information, it is detected whether there is an obstacle in the scene corresponding to the image to be detected.

2. The method according to claim 1, wherein: The detecting, based on the first text information, whether there is an obstacle in the scene corresponding to the image to be detected includes: Dividing the first text information into at least one sentence; Identifying whether the textual meaning of each sentence in the at least one sentence is used to describe the existence of an object in the scene corresponding to the image to be detected; In the case where the textual meaning of each sentence is used to describe the existence of an object in the scene corresponding to the image to be detected, it is detected whether there is an obstacle in the scene corresponding to the image to be detected.

3. The method according to claim 2, wherein: In the case where the textual meaning of each sentence is used to describe the existence of an object in the scene corresponding to the image to be detected, detecting whether there is an obstacle in the scene corresponding to the image to be detected includes: Extracting keywords corresponding to the object from each sentence; Based on the keyword corresponding to the object, it is determined whether the object is an obstacle in the scene corresponding to the image to be detected.

4. The method according to claim 3, wherein: The determining, based on the keyword corresponding to the object, whether the object is an obstacle in the scene corresponding to the image to be detected includes: Searching whether there is a keyword corresponding to the object in a first keyword list, wherein the first keyword list includes at least one keyword corresponding to an obstacle; In the case where a keyword corresponding to the object is found in the first keyword list, determining that the object is an obstacle; or When the keyword corresponding to the object is not found in the first keyword list, the object is determined to be a non-obstacle.

5. The method according to claim 4, wherein: The determining, based on the keyword corresponding to the object, whether the object is an obstacle in the scene corresponding to the image to be detected further includes: If the keyword corresponding to the object is not found in the first keyword list, sending the keyword of the object to the server; A first recognition result sent by the server is received, where the first recognition result is used to indicate whether the object is an obstacle.

6. The method according to claim 5, further comprising: In a case where the first recognition result is used to indicate that the object is an obstacle, a keyword of the object is added to the first keyword list.

7. The method according to claim 3, wherein: The determining, based on the keyword corresponding to the object, whether the object is an obstacle in the scene corresponding to the image to be detected includes: Searching for a keyword corresponding to the object in a second keyword list, where the second keyword list includes at least one keyword corresponding to a non-obstacle; If no keyword corresponding to the object is found in the second keyword list, determining that the object is an obstacle; or When the keyword corresponding to the object is found in the second keyword list, the object is determined to be a non-obstacle.

8. The method according to claim 7, wherein: The determining, based on the keyword corresponding to the object, whether the object is an obstacle in the scene corresponding to the image to be detected further includes: If no keyword corresponding to the object is found in the second keyword list, sending the keyword of the object to the server; A second recognition result sent by the server is received, where the second recognition result is used to indicate whether the object is a non-obstacle.

9. The method according to claim 8, further comprising: In a case where the second recognition result is used to indicate that the object is a non-obstacle, a keyword of the object is added to the second keyword list.

10. The method according to claim 2, wherein: The step of identifying whether the textual meaning of each sentence is used to describe the existence of an object in the scene corresponding to the image to be detected includes: Each of the sentences is input into a text classification model to obtain a text classification result output by the text classification model, wherein the text classification result includes a first classification result or a second classification result, wherein the first classification result is used to indicate that the text meaning of the sentence is used to describe the existence of an object in the scene corresponding to the image to be detected, and the second classification result is used to indicate that the text meaning of the sentence is not used to describe the existence of an object in the scene corresponding to the image to be detected.

11. The method according to claim 10, wherein: The text classification model is trained in the following way: Acquire a first training sample set, the first training sample set comprising a plurality of first samples and respective labels of the plurality of first samples, each of the plurality of first samples being used to describe second text information of a first sample image, the label of each of the plurality of first samples being used to indicate whether a textual meaning of a sentence of the second text information is used to describe whether there is an object in a scene corresponding to the first sample image, and the scene corresponding to the first sample image is the same as the scene corresponding to the image to be detected; Based on the first training sample set, an initial text classification model is trained to obtain a trained text classification model.

12. The method according to claim 1, wherein: The detecting, based on the first text information, whether there is an obstacle in the scene corresponding to the image to be detected includes: The first text information is input into a text classification model to obtain a text classification result output by the text classification model, wherein the text classification result includes a third classification result and a fourth classification result, wherein the third classification result is used to indicate that the text meaning of the first text information is used to describe the existence of an obstacle in the scene corresponding to the image to be detected, and the fourth classification result is used to indicate that the text meaning of the first text information is not used to describe the existence of an obstacle in the scene corresponding to the image to be detected.

13. The method according to claim 1, wherein: The step of generating first text information for describing the image to be detected based on the image to be detected includes: The image to be detected is input into a picture-text generation model to obtain the first text information output by the picture-text generation model.

14. The method according to claim 13, wherein: The image-text generation model is trained in the following way: Acquire a second training sample set, where the second training sample set includes a plurality of second sample images and third text information corresponding to each of the plurality of second sample images; Based on the second training sample set, the initial image-text generation model is trained to obtain the trained image-text generation model.

15. The method according to claim 1, further comprising: When it is detected that an obstacle exists in the scene corresponding to the image to be detected, an alarm message for prompting the existence of the obstacle is sent.

16. An electronic device, comprising: a processor and a memory for storing executable instructions for the processor; The processor is configured to execute the instructions so that the electronic device performs the method according to any one of claims 1 to 15.

17. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on an electronic device, the electronic device executes the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Abnormal event detection method generated based on image description and detection system thereof

    CN112732965A

  • Road surface obstacle detection and early warning method

    CN116597420A

  • Functional porridge having the activities of antioxidation, antidiabetes and antithrombosis comprising hempseed, Glehnia littoralis and yam

    KR1020220143525A

  • Image processing method and apparatus, and storage medium

    US20220058332A1