Traffic scene throwing object detection and classification method and system based on visual language model

Through a method based on the visual language model, the task of spilled object detection is decomposed into road segmentation, location detection and classification. Combined with the background difference method and non-maximum suppression technology, the problems of poor robustness of existing methods in complex environments and high demand for manual labeling are solved, and efficient spilled object detection and classification are achieved in different scenarios.

CN120673161APending Publication Date: 2025-09-19NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510805876.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing methods for detecting and classifying spilled objects have poor robustness in complex traffic environments and are greatly affected by environmental factors. In addition, fully supervised learning methods require a large amount of manual labeling and are difficult to generalize in different scenarios.

Method used

A method based on the visual language model is used to decouple the scattered object detection task into three parts: road segmentation, location detection, and classification. The visual language model CLIP is used for classification, combined with background difference method and non-maximum suppression technology to achieve scattered object detection and classification.

Benefits of technology

It can detect a variety of spilled objects in different scenarios, reducing the need for manual labeling, improving the robustness and generalization of detection, and providing more reliable automated detection support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673161A_ABST
    Figure CN120673161A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic scene throwing object detection and classification method and system based on a visual language model, and the method comprises the steps: building a vocabulary based on the category of a traffic scene throwing object; based on the vocabulary, constructing a throwing object sample; segmenting the road information to obtain road segmentation information; acquiring a thrown object detection area based on the road segmentation information; based on the thrown object sample and the thrown object detection area, detection and classification of the thrown object are completed. Based on the strong generalization of the visual language model, the method can achieve the detection of the thrown objects on the premise of few samples or even no samples, provides more reliable automatic detection support for a traffic management department, helps the traffic management department to better manage and control the traffic, and effectively prevents the traffic accidents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and intelligent transportation, and in particular to a method and system for detecting and classifying spilled objects in traffic scenes based on a visual language model. Background Art

[0002] With the acceleration of urbanization and the increasing complexity of road traffic networks, traffic safety has gradually become a key issue in urban management. During traffic operations, spilled debris caused by various reasons can not only cause traffic accidents and increase traffic congestion, but can also lead to serious casualties and property damage. Therefore, timely and accurate detection and handling of such spilled debris is crucial to ensuring traffic safety. However, since spilled debris often comes in a wide variety of shapes and colors, and is limited by the angle and position of traffic cameras, it is often small, making it more difficult to identify.

[0003] Typically, the detection and recognition of spilled objects relies on traditional visual methods such as background modeling technology, or on fully supervised learning methods. However, these methods often have significant limitations. On the one hand, although traditional detection methods such as background modeling have low requirements for computing resources, they perform poorly in complex and changing traffic environments and are easily affected by factors such as changes in lighting and weather conditions, resulting in high false detection and missed detection rates. On the other hand, although fully supervised learning methods can achieve high accuracy under specific conditions, this high accuracy comes at the cost of a large amount of time-consuming and labor-intensive manual data labeling. In addition, these two methods have weak generalization capabilities in different scenarios and are difficult to adapt to diverse practical application needs. Summary of the Invention

[0004] Through the above analysis, the problems and defects of the existing technology are as follows:

[0005] Traditional computer vision algorithms have poor robustness and are significantly affected by the environment; fully supervised methods require a large amount of manual labeling, and the models are difficult to generalize to different application scenarios.

[0006] In view of the shortcomings of current methods for detecting and classifying scattered objects in applications, the present invention proposes a method for detecting and classifying scattered objects in traffic scenes based on a visual language model. In simple terms, the method decouples the task of detecting scattered objects into three parts: road segmentation, position detection, and classification. The method first extracts the road surface information that needs attention through the line drawing method or the semantic segmentation method, which can effectively suppress complex environmental information. Then, the target detector detects the position of the vehicle on the road and laterally divides the area to be detected according to the position of the vehicle. In the area to be detected, the frame difference method is used to obtain a relatively loose foreground and obtain the area proposal of the area to be detected. The method will segment the object position from the original image based on the area proposal and transmit it to the visual language model CLIP for classification, obtain the confidence of each area proposal, and then perform post-processing operations such as non-maximum suppression on the proposed area. The output result is used as the final scattered object detection result.

[0007] To achieve the above object, the beneficial effects of the present invention are as follows:

[0008] A method for detecting and classifying spilled objects in traffic scenes based on a visual language model, comprising the following steps:

[0009] Build a vocabulary based on the categories of spilled objects in traffic scenes;

[0010] constructing a spill sample based on the vocabulary;

[0011] Segmenting the road information to obtain road segmentation information;

[0012] Based on the road segmentation information, obtaining a spilled object detection area;

[0013] Based on the spilled object sample and the spilled object detection area, detection and classification of the spilled objects are completed.

[0014] Preferably, for the constructed vocabulary, the text is encoded in advance by a text encoder of CLIP and the encoded text vector is stored in a memory.

[0015] Preferably, based on the road segmentation information, the position information of the vehicle on the road is detected by the yolov5 model to obtain the driving position; if the upper left and lower right coordinates of the vehicle position anchor frame are (x1, y1) and (x2, y2) respectively, the upper left and lower right coordinates of the selected area (x3, y3) and (x4, y4) are calculated as follows:

[0016]

[0017] Preferably, the steps of obtaining the spilled object detection area according to the driving position include: using a background difference method to obtain an area requiring attention in the area to be detected:

[0018] Mask(i,j)={1||Img(i,j)-BG(i,j)|>T}

[0019] Among them, Mask(i, j) represents the pixel information of the mask image position (i, j), and the pixel information of the position is set to zero by default; Img(i, j) represents the pixel information of the current frame image position (i, j); BG(i, j) represents the pixel information of the background image BG(i, j) position, and T represents the division threshold.

[0020] Preferably, after the spilled object detection area is obtained, it is cropped, adjusted to 224x224 and input into the image encoder of the visual language model CLIP for encoding; the encoded image vector is multiplied with the text vector respectively, and the similarity is calculated with the vector of the spilled object sample; the two sets of similarity results are fused in a weighted manner to obtain the confidence of the current detection area.

[0021] Preferably, after the detection is completed, non-maximum suppression processing is performed, the steps of which include: extracting detection boxes in descending order of confidence, and deleting detection targets whose GIOU with the detection box is higher than 0.5 until no detection box is deleted; the GIOU calculation formula is as follows:

[0022]

[0023] Among them, A and B represent detection boxes; A x 、A y 、B x 、B y Represents the center of each of the two detection boxes; w and h represent the width and height of the detection area.

[0024] The present invention also provides a system for detecting and classifying spilled objects in traffic scenes based on a visual language model. The system is used to implement the above method and includes: an acquisition module, a construction module, a segmentation module, an acquisition module, and a detection module;

[0025] The collection module is used to build a vocabulary based on the categories of spilled objects in traffic scenes;

[0026] The construction module is used to construct a spilled object sample based on the vocabulary;

[0027] The segmentation module is used to segment the road information to obtain road segmentation information;

[0028] The acquisition module is used to acquire a spilled object detection area based on the road segmentation information;

[0029] The detection module is used to complete the detection and classification of the spilled objects based on the spilled object samples and the spilled object detection area.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] The present invention provides a method and system for detecting and classifying spilled objects in traffic scenes based on a visual language model, which is used to detect a variety of spilled objects in different scenarios. Based on the strong generalization of the visual language model, the detection of spilled objects can be achieved with few or even no samples, providing more reliable automated detection support for traffic management departments, helping them to better control traffic and effectively prevent traffic accidents. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention;

[0034] Figure 2 Schematic diagram of segmentation results according to an embodiment of the present invention;

[0035] Figure 3 Schematic diagram of the segmentation process according to an embodiment of the present invention;

[0036] Figure 4 Schematic diagram of the improved U-Net model structure according to an embodiment of the present invention;

[0037] Figure 5 Schematic diagram of the channel attention convolution structure in the improved U-Net model according to an embodiment of the present invention;

[0038] Figure 6 Schematic diagram of an efficient spatial pyramid structure in an improved U-Net model according to an embodiment of the present invention;

[0039] Figure 7 A schematic diagram of a random frame generation area according to an embodiment of the present invention;

[0040] Figure 8 This is a schematic diagram of the effect of random frame generation area according to an embodiment of the present invention;

[0041] Figure 9 The overall process framework of an embodiment of the present invention is a schematic diagram. DETAILED DESCRIPTION

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0043] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0044] Example 1

[0045] like Figure 1 FIG. 1 is a flow chart of the method of this embodiment, and the steps include:

[0046] S1. Build a vocabulary based on the categories of spilled objects in traffic scenes.

[0047] The vocabulary is constructed based on the categories of scattered objects commonly seen in traffic scenes, as well as categories that are about to appear and easily confused categories.

[0048] Furthermore, in S1, the specific information also includes:

[0049] A1. Search for similar keywords based on publicly reported spillage incidents and spillage datasets at locations such as tunnels and highway intersections.

[0050] A2: Classify and group similar-looking spilled objects.

[0051] A3: Use a large language model to automatically generate feature descriptions for different categories of scattered objects.

[0052] A4: The initial categories used during the test included: garbage (cardboard boxes, wine bottles, plastic bottles, cans), animals (carcasses), and large spills (cloth rolls, tires, and cargo boxes).

[0053] S2. Construct a sample of spilled objects based on the vocabulary;

[0054] Sample construction and data vectorization. For the vocabulary built in S1, the text is encoded in advance through CLIP's text encoder and the encoded text vector is stored in memory for fast recall when used.

[0055] At the same time, a small amount of data on common categories in traffic scenes is collected, including not only the category of scattered objects, but also road background, vehicles, etc. This data is encoded through the CLIP image encoder and generated into a feature vector for reference.

[0056] S3. Segment the road information to obtain road segmentation information.

[0057] First, a fixed camera or a pan-tilt camera is used to obtain a road image and segment it. The segmentation results are as follows: Figure 2 As shown. This embodiment uses both manual marking and semantic segmentation. Manual marking involves manually marking out the road locations of interest. The semantic segmentation method uses deep learning (U-Net) to segment road locations, taking the maximum closure of the segmented foreground after erosion and expansion as the road location result. Since the camera position is essentially fixed, focusing on the current road location can be achieved with a single operation.

[0058] Specifically, in the preparation stage, road monitoring data from different angles is obtained from the pan-tilt camera, and the road parts are marked to obtain the segmentation data of the road surface. The constructed data is sent to the improved deep learning semantic segmentation model for training and learning, the semantic segmentation weights are learned, and the semantic segmentation model is used to preliminarily realize the division of road areas. In the actual application stage, if there is no manual marking, multiple key frames within 30 seconds are selected and sent to the semantic segmentation model to obtain multiple binary images. The maximum closure after corrosion and expansion of the segmented foreground is used as the road position result, and then the union area of ​​the binary image is taken. Finally, the maximum connected domain of each area is obtained, and the small area connected domain is deleted to obtain the final automatic segmentation of the road surface result. The flow of this process is as follows Figure 3 As shown, the improved U-Net model structure is as follows Figure 4 、 Figure 5 、 Figure 6 shown.

[0059] S4. Based on the road segmentation information, obtain the spilled object detection area.

[0060] S401. Obtain the driving position from the road segmentation map obtained in step S3.

[0061] Based on the road segmentation information, the position information of the vehicle on the road is detected by the yolov5 model to obtain the driving position; if the upper left and lower right coordinates of the vehicle position anchor box are (x1, y1) and (x2, y2) respectively, then the upper left and lower right coordinates of the selected area (x3, y3) and (x4, y4) are calculated as follows:

[0062] x3=2×x1-x2

[0063] x4=2×x2-x1

[0064] y3=1.5×y1-0.5×y2

[0065] y4=1.5×y2-0.5×y1

[0066] S402. Obtain the spilled object detection area based on the driving position.

[0067] The detection area for the spilled objects is selected based on the vehicle's position. In this embodiment, the background difference method is used to obtain the area that needs attention in the detection area to obtain a mask image of the foreground; the vehicle position in the mask image (the vehicle position is obtained by step S3) is uniformly set as the background.

[0068] At the same time, the minimum length and width of detection are set according to the size of common spilled objects, and multiple groups of random detection frames are generated at each randomly generated center point, thereby generating multiple groups of detection frame areas.

[0069] Specifically, for the selection of the detection frame, the background difference method will randomly generate a center point in the corresponding area of ​​the mask. For the generation of the detection frame, the length and width of the detection frame will not be less than the minimum set length, and the maximum will not exceed the length and width of the vehicle. At the same time, the foreground area in the detection frame must not be less than 25% of the detection frame area; otherwise, it will be treated as an invalid frame, and the current length and width will be adjusted as the maximum value to regenerate the detection frame. If the conditions cannot be met for three consecutive times or the current length and width have reached the minimum value, the point will be abandoned. The random frame generation area is as follows: Figure 7 The effect diagram is shown as Figure 8 shown.

[0070] The calculation method of the background difference method is as follows:

[0071] Mask(i,j)={1||Img(i,j)-BG(i,j)|>T}

[0072] Among them, Mask(i, j) represents the pixel information of the mask image position (i, j), and the pixel information of the position is set to zero by default; Img(i, j) represents the pixel information of the current frame image position (i, j); BG(i, j) represents the pixel information of the background image BG(i, j) position, and T represents the division threshold.

[0073] For background subtraction, the background modeling depends on the initial frame in which the algorithm is run. If no relevant targets (such as vehicles, pedestrians, or spilled objects) are detected on the current road for a period of time, the algorithm will use the current frame to update the background and continue detection.

[0074] S5. Complete detection and classification of the spilled objects based on the spilled object samples and the spilled object detection area.

[0075] The detection area corresponding to the detection box is cropped, resized to 224x224, and fed into the CLIP visual language model's image encoder for encoding. The encoded image vector is multiplied with the text vector and its similarity is calculated with the object sample vector. The two sets of similarity results are weighted and fused to obtain the confidence score for the current detection area. This allows all randomly generated detection boxes to obtain a confidence score for each classification. Non-maximum suppression is applied to all detection boxes to determine the threshold for detecting objects on the road.

[0076] Specifically, the cropped area is resized to an image with a length and width of 224 pixels and then input into the CLIP visual encoder for encoding. The image code i1 of the image to be classified is multiplied with the text vector (the text vector includes: the category text information t1 of the object to be classified, and the description text information t2 of the category of the object to be classified) and the similarity is calculated with the image code i0 of the template image. The calculation method is as follows:

[0077] L i1 =λ L i1t1+β L i1t2

[0078] C i1 =λ C D KL (L i0 ,L i1 )+β C D KL (i0,i1)

[0079] S i1 =Softmax(λL i1 +βC i1 )

[0080] Among them, D KL Represents the KL divergence between the calculated vectors. The obtained similarity result categories are divided and merged according to the grouping of A2, which can represent the confidence of the current detection area; T represents the output result of the text encoder; C i1 represents the image encoding of the image and the divergence of the template image; L i1 represents the probability of the image being classified as a scattered object combined with text information; S i1 Represents the total classification probability of the image; λ L and β L For L i1 The hyperparameters in this embodiment are set to 0.5 and 0.5; λ C and β C C i1 The hyperparameters in this embodiment are set to 0.8 and 0.2; S and β SFor S i1 The hyperparameters in this example are set to 1 and -0.3.

[0081] After the detection is completed, non-maximum suppression processing is performed, and the steps include:

[0082] Extract detection boxes in descending order of confidence and delete detection targets with a GIOU greater than 0.5 with the detection box until no detection box is deleted. The GIOU calculation formula is as follows:

[0083]

[0084] Among them, A and B represent detection boxes; A x 、A y 、B x 、B y Indicates the center of each of the two detection frames; w and h indicate the width and height of the detection area. The overall flow chart of this embodiment is as follows: Figure 9 shown.

[0085] Example 2

[0086] This embodiment also provides a traffic scene spilled object detection and classification system based on a visual language model, including: an acquisition module, a construction module, a segmentation module, an acquisition module and a detection module; the acquisition module is used to construct a vocabulary based on the traffic scene spilled object category; the construction module is used to construct spilled object samples based on the vocabulary; the segmentation module is used to segment road information and obtain road segmentation information; the acquisition module is used to obtain spilled object detection areas based on the road segmentation information; the detection module is used to complete the detection and classification of spilled objects based on the spilled object samples and the spilled object detection areas.

[0087] The following will describe in detail how the present invention solves technical problems in practical work in conjunction with this embodiment.

[0088] First, the collection module is used to build a vocabulary based on the categories of spilled objects in traffic scenes.

[0089] The vocabulary is constructed based on the categories of scattered objects commonly seen in traffic scenes, as well as categories that are about to appear and easily confused categories.

[0090] Further specific information also includes:

[0091] A1. Search for similar keywords based on publicly reported spillage incidents and spillage datasets at locations such as tunnels and highway intersections.

[0092] A2: Classify and group similar-looking spilled objects.

[0093] A3: Use a large language model to automatically generate feature descriptions for different categories of scattered objects.

[0094] A4: The initial categories used during the test included: garbage (cardboard boxes, wine bottles, plastic bottles, cans), animals (carcasses), large spills (cloth rolls, tires, cargo boxes),

[0095] The building module constructs the spilled object samples based on the vocabulary;

[0096] Sample construction and data vectorization: For the vocabulary built by the acquisition module, the text is encoded in advance through the CLIP text encoder and the encoded text vector is stored in memory for fast recall when used.

[0097] At the same time, a small amount of data on common categories in traffic scenes is collected, including not only the category of scattered objects, but also road background, vehicles, etc. This data is encoded through the CLIP image encoder and generated into a feature vector for reference.

[0098] The segmentation module segments the road information to obtain road segmentation information.

[0099] First, a fixed camera or a pan-tilt camera is used to obtain a road image and segment it. The segmentation results are as follows: Figure 2 As shown. This embodiment uses both manual marking and semantic segmentation. Manual marking involves manually marking out the road locations of interest. The semantic segmentation method uses deep learning (U-Net) to segment road locations, taking the maximum closure of the segmented foreground after erosion and expansion as the road location result. Since the camera position is essentially fixed, focusing on the current road location can be achieved with a single operation.

[0100] Specifically, in the preparation stage, road monitoring data from different angles is obtained from the pan-tilt camera, and the road parts are marked to obtain the segmentation data of the road surface. The constructed data is sent to the improved deep learning semantic segmentation model for training and learning, the semantic segmentation weights are learned, and the semantic segmentation model is used to preliminarily realize the division of road areas. In the actual application stage, if there is no manual marking, multiple key frames within 30 seconds are selected and sent to the semantic segmentation model to obtain multiple binary images. The maximum closure after corrosion and expansion of the segmented foreground is used as the road position result, and then the union area of ​​the binary image is taken. Finally, the maximum connected domain of each area is obtained, and the small area connected domain is deleted to obtain the final automatic segmentation of the road surface result. The flow of this process is as follows Figure 3 As shown, the improved U-Net model structure is as follows Figure 4 、 Figure 5 、 Figure 6 shown.

[0101] The acquisition module obtains the spilled object detection area based on the road segmentation information.

[0102] Get the vehicle position from the road segmentation map obtained from the segmentation module.

[0103] Based on the road segmentation information, the position information of the vehicle on the road is detected by the yolov5 model to obtain the driving position; if the upper left and lower right coordinates of the vehicle position anchor box are (x1, y1) and (x2, y2) respectively, then the upper left and lower right coordinates of the selected area (x3, y3) and (x4, y4) are calculated as follows:

[0104] x3=2×x1-x2

[0105] x4=2×x2-x1

[0106] y3=1.5×y1-0.5×y2

[0107] y4=1.5×y2-0.5×y1

[0108] Obtain the spilled object detection area based on the driving position.

[0109] The detection area for spilled objects is selected based on the position of the vehicle. In this embodiment, the background difference method is used to obtain the area that needs attention in the detection area to obtain a mask image of the foreground; the vehicle position in the mask image (the vehicle position is obtained by the segmentation module) is uniformly set as the background.

[0110] At the same time, the minimum length and width of detection are set according to the size of common spilled objects, and multiple groups of random detection frames are generated at each randomly generated center point, thereby generating multiple groups of detection frame areas.

[0111] Specifically, for the selection of the detection frame, the background difference method will randomly generate a center point in the corresponding area of ​​the mask. For the generation of the detection frame, the length and width of the detection frame will not be less than the minimum set length, and the maximum will not exceed the length and width of the vehicle. At the same time, the foreground area in the detection frame must not be less than 25% of the detection frame area; otherwise, it will be treated as an invalid frame, and the current length and width will be adjusted as the maximum value to regenerate the detection frame. If the conditions cannot be met for three consecutive times or the current length and width have reached the minimum value, the point will be abandoned. The random frame generation area is as follows: Figure 3 The effect diagram is shown as Figure 4 shown.

[0112] The calculation method of the background difference method is as follows:

[0113] Mask(i,j)={1||Img(i,j)-BG(i,j)|>T}

[0114] Among them, Mask(i, j) represents the pixel information of the mask image position (i, j), and the pixel information of the position is set to zero by default; Img(i, j) represents the pixel information of the current frame image position (i, j); BG(i, j) represents the pixel information of the background image BG(i, j) position, and T represents the division threshold.

[0115] For background subtraction, the background modeling depends on the initial frame in which the algorithm is run. If no relevant targets (such as vehicles, pedestrians, or spilled objects) are detected on the current road for a period of time, the algorithm will use the current frame to update the background and continue detection.

[0116] Finally, the detection module completes the detection and classification of spilled objects based on the spilled object samples and the spilled object detection area.

[0117] The detection area corresponding to the detection box is cropped, resized to 224x224, and fed into the image encoder of the visual language model CLIP for encoding. The encoded image vector is multiplied by the text vector and its similarity is calculated with the vector of the spilled object sample from the construction module. The two sets of similarity results are fused using a weighted method to obtain the confidence score for the current detection area. This allows all randomly generated detection boxes to obtain a confidence score for each classification. Non-maximum suppression is applied to all detection boxes to determine the threshold for detecting spilled objects on the road.

[0118] Specifically, the cropped area is resized to an image with a length and width of 224 pixels and then input into the CLIP visual encoder for encoding. The image code i1 of the image to be classified is multiplied with the text vector (the text vector includes: the category text information t1 of the object to be classified, and the description text information t2 of the category of the object to be classified) and the similarity is calculated with the image code i0 of the template image. The calculation method is as follows:

[0119] L i1 =λ L i1t1+β L i1t2

[0120] C i1 =λ C D KL (L i0 ,L i1 )+β C D KL (i0,i1)

[0121] S i1 =Softmax(λL ii +βC i1 )

[0122] Among them, D KLRepresents the KL divergence between the calculated vectors. The obtained similarity result categories are divided and merged according to the grouping of A2, which can represent the confidence of the current detection area; T represents the output result of the text encoder; C i1 represents the image encoding of the image and the divergence of the template image; L i1 represents the probability of the image being classified as a scattered object combined with text information; S i1 Represents the total classification probability of the image; λ L and β L For L i1 The hyperparameters in this embodiment are set to 0.5 and 0.5; λ C and β C C i1 The hyperparameters in this embodiment are set to 0.8 and 0.2; S and β S For S i1 The hyperparameters in this example are set to 1 and -0.3.

[0123] After the detection is completed, non-maximum suppression processing is performed, and the steps include:

[0124] Extract detection boxes in descending order of confidence and delete detection targets with a GIOU greater than 0.5 with the detection box until no detection box is deleted. The GIOU calculation formula is as follows:

[0125]

[0126] Among them, A and B represent detection boxes; A x 、A y 、B x 、B y Represents the center of each of the two detection boxes; w and h represent the width and height of the detection area.

[0127] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A method for detecting and classifying spilled objects in traffic scenes based on a visual language model, characterized in that the steps include: Build a vocabulary based on the categories of spilled objects in traffic scenes; constructing a spill sample based on the vocabulary; Segmenting the road information to obtain road segmentation information; Based on the road segmentation information, obtaining a spilled object detection area; Based on the spilled object sample and the spilled object detection area, detection and classification of the spilled objects are completed.

2. The method for detecting and classifying spilled objects in traffic scenes based on a visual language model according to claim 1, characterized in that: For the constructed vocabulary, the text is encoded in advance by the CLIP text encoder and the encoded text vector is stored in the memory.

3. The method for detecting and classifying spilled objects in traffic scenes based on a visual language model according to claim 1, characterized in that: Based on the road segmentation information, the position information of the vehicle on the road is detected by the yolov5 model to obtain the driving position; if the upper left and lower right coordinates of the vehicle position anchor box are (x1, y1) and (x2, y2) respectively, then the upper left and lower right coordinates of the selected area (x3, y3) and (x4, y4) are calculated as follows:

4. The method for detecting and classifying spilled objects in traffic scenes based on a visual language model according to claim 3 is characterized in that: According to the driving position, the spilled object detection area is obtained, and the steps include: using the background difference method to obtain the area that needs attention in the area to be detected: Mask(i,j)={1||Img(i,j)-BG(i,j)|>T} Among them, Mask(i, j) represents the pixel information of the mask image position (i, j), and the pixel information of the position is set to zero by default; Img(i, j) represents the pixel information of the current frame image position (i, j); BG(i, j) represents the pixel information of the background image BG(i, j) position, and T represents the division threshold.

5. The method for detecting and classifying spilled objects in traffic scenes based on a visual language model according to claim 4, characterized in that: After the spilled object detection area is acquired, it is cropped, resized to 224x224, and input into the image encoder of the visual language model CLIP for encoding; the encoded image vector is multiplied with the text vector, and the similarity is calculated with the vector of the spilled object sample; the two sets of similarity results are fused in a weighted manner to obtain the confidence of the current detection area.

6. The method for detecting and classifying spilled objects in traffic scenes based on a visual language model according to claim 5, characterized in that: After the detection is completed, non-maximum suppression processing is performed. The steps include: extracting detection boxes in descending order of confidence, and deleting detection targets with a GIOU greater than 0.5 with the detection box until no detection box is deleted; the GIOU calculation formula is as follows: Among them, A and B represent detection boxes; A x 、A y 、B x 、B y Represents the center of each of the two detection boxes; w and h represent the width and height of the detection area.

7. A system for detecting and classifying spilled objects in traffic scenes based on a visual language model, the system being used to implement the method according to any one of claims 1 to 6, characterized in that: include: Acquisition module, construction module, segmentation module, acquisition module and detection module; The collection module is used to build a vocabulary based on the categories of spilled objects in traffic scenes; The construction module is used to construct a spilled object sample based on the vocabulary; The segmentation module is used to segment the road information to obtain road segmentation information; The acquisition module is used to acquire a spilled object detection area based on the road segmentation information; The detection module is used to complete the detection and classification of the spilled objects based on the spilled object samples and the spilled object detection area.