Road throwing object detection method and device and electronic equipment

The road spilled object detection method that combines the Gaussian mixture model and the neural network model solves the problems of low detection efficiency and poor accuracy in the existing technology, realizes intelligent automatic detection, and improves detection efficiency and accuracy.

CN120635800APending Publication Date: 2025-09-12QINGDAO HISENSE TRANS TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510570392.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The existing technology for detecting spilled objects on roads has low efficiency and poor accuracy, and is prone to missed detection due to human negligence.

Method used

The Gaussian mixture model (GMM) is combined with grayscale histogram processing and neural network model to achieve automatic detection by performing block processing, feature extraction and enhancement on road images.

Benefits of technology

It improves the efficiency and accuracy of road spillage detection, avoids missed detection due to human negligence, and realizes intelligent automatic detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635800A_ABST
    Figure CN120635800A_ABST
Patent Text Reader

Abstract

The invention discloses a road throwing object detection method and device and electronic equipment, and the method comprises the steps: firstly dividing a to-be-detected road image into a plurality of first sub-images after obtaining the to-be-detected road image, and then determining a target gray histogram corresponding to the first sub-images; and feature extraction is carried out on the target gray level histogram based on a Gaussian mixture model (GMM), and whether the corresponding first sub-image contains a throwing object is determined. And if the road image has the first sub-image of which the detection result is that the throwing object is contained, determining that the throwing object exists in the road image. Therefore, intelligent automatic detection of the road spilled objects is realized. Compared with a detection method that personnel watch video images in the related technology, the method improves the detection efficiency of the road throwing objects, avoids the problem of missing detection of the throwing objects due to personnel negligence, and improves the detection accuracy of the road throwing objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer image technology, and in particular to a method, device and electronic equipment for detecting road spilled objects. Background Art

[0002] Road spills pose a significant threat to highway safety and require prompt detection and resolution. Traditional technologies rely on manual inspections or video polling and comparison to detect spills, making it difficult for management to detect incidents promptly and implement swift resolution. This "man-on-screen" approach to detecting spills consumes significant human resources and is inefficient. Furthermore, spills can be missed due to negligence, resulting in poor detection accuracy. Summary of the Invention

[0003] The present application provides a method, device and electronic equipment for detecting spilled objects on the road, which are used to solve the problems of low efficiency and poor accuracy in detecting spilled objects in related technologies.

[0004] In a first aspect, the present application provides a method for detecting spilled objects on a road, the method comprising:

[0005] Acquire a road image to be detected, and divide the road image into a plurality of first sub-images;

[0006] For the multiple first sub-images, determine a target grayscale histogram corresponding to the first sub-image, input the target grayscale histogram into a pre-trained Gaussian mixture model (GMM), and determine whether the first sub-image contains a detection result of a scattered object based on the GMM;

[0007] If the road image contains a first sub-image whose detection result indicates that the sub-image contains scattered objects, it is determined that the road image contains scattered objects.

[0008] The above technical solution has the following advantages or beneficial effects:

[0009] In the present application, after obtaining the road image to be detected, the road image is first divided into multiple first sub-images, and then the target grayscale histogram corresponding to the first sub-image is determined; then, based on the Gaussian mixture model (GMM), feature extraction is performed on the target grayscale histogram, and it is determined whether the corresponding first sub-image contains spilled objects. If there is a first sub-image in the road image that contains spilled objects, it is determined that there are spilled objects in the road image. This achieves intelligent automatic detection of road spilled objects. Compared with the method of personnel watching video image detection in the related art, the efficiency of road spilled object detection is improved, and the problem of missed detection of spilled objects due to personnel negligence is avoided, thereby improving the accuracy of road spilled object detection.

[0010] In an optional implementation, determining the target grayscale histogram corresponding to the first sub-image includes:

[0011] Counting a first grayscale histogram corresponding to the first sub-image, and performing equalization processing on the first grayscale histogram using an equalization algorithm to obtain a second grayscale histogram;

[0012] The number of pixels corresponding to grayscales less than a preset first grayscale threshold and greater than a preset second grayscale threshold in the second grayscale histogram is set to 0 to obtain a target grayscale histogram; wherein the preset first grayscale threshold is less than the preset second grayscale threshold.

[0013] The above technical solution has the following advantages or beneficial effects:

[0014] In the present application, the first sub-image is enhanced by combining histogram equalization with contrast limitation. Histogram equalization refers to statistically analyzing the first grayscale histogram corresponding to the first sub-image, and performing equalization processing on the first grayscale histogram using an equalization algorithm to obtain a second grayscale histogram. Contrast limitation refers to setting the number of pixels in the second grayscale histogram corresponding to grayscales less than a preset first grayscale threshold and greater than a preset second grayscale threshold to 0 to obtain a target grayscale histogram. The target grayscale histogram obtained in this way is more accurate, thereby making the detection of spilled objects based on the target grayscale histogram more accurate.

[0015] In an optional implementation, before performing equalization processing on the first grayscale histogram using an equalization algorithm to obtain a second grayscale histogram, the method further includes:

[0016] Acquiring first brightness information and first contrast information of the road image;

[0017] Determining target type information corresponding to the road image based on the first brightness information and the first contrast information, and the brightness information ranges and contrast information ranges corresponding to images of different types of information, wherein the type information includes type information requiring enhancement processing and type information not requiring enhancement processing;

[0018] If the target type information is information of a type requiring enhancement processing, performing an equalization process on the first grayscale histogram using an equalization algorithm to obtain a second grayscale histogram;

[0019] If the target type information is type information that does not require enhanced processing, the method further includes:

[0020] The first grayscale histogram is used as a target grayscale histogram corresponding to the first sub-image.

[0021] The above technical solution has the following advantages or beneficial effects:

[0022] In the present application, an equalization algorithm is used to perform equalization processing on the first grayscale histogram. Before obtaining the second grayscale histogram, the target type information corresponding to the road image is first determined based on the first brightness information and the first contrast information of the road image. If the target type information is information of a type that does not require enhanced processing, the first grayscale histogram is used as the target grayscale histogram corresponding to the first sub-image. The subsequent step of performing equalization processing on the first grayscale histogram using an equalization algorithm to obtain the second grayscale histogram is not performed, thereby taking into account both the efficiency and accuracy of scattered object detection. If the target type information is information of a type that requires enhanced processing, the step of performing equalization processing on the first grayscale histogram using an equalization algorithm to obtain the second grayscale histogram is performed; thereby ensuring the accuracy of scattered object detection.

[0023] In an optional implementation, before dividing the image into a plurality of first sub-images, the method further includes:

[0024] The road image is input into a pre-trained neural network model for detecting spilled objects. If it is determined based on the neural network model that there are spilled objects in the road image, a subsequent step of dividing the image into multiple first sub-images is performed.

[0025] The above technical solution has the following advantages or beneficial effects:

[0026] In the present application, before dividing the image into a plurality of first sub-images, the road image is first input into a pre-trained neural network model for spilled object detection. If the neural network model determines that no spilled objects are present in the road image, then the road image is determined to be free of spilled objects, and subsequent spilled object detection based on the GMM model is unnecessary. If the neural network model determines that spilled objects are present in the road image, a secondary detection of spilled objects is then performed based on the GMM model, thereby ensuring the accuracy of spilled object detection.

[0027] In an optional embodiment, before inputting the road image into a pre-trained neural network model for spilled object detection, the method further includes:

[0028] The road image is binarized, a preset morphological operation window is used to perform morphological processing on the binarized road image, and the morphologically processed road image is input into a pre-trained neural network model for spilled object detection.

[0029] The above technical solution has the following advantages or beneficial effects:

[0030] In this application, the road image is binarized, and then a preset morphological operation window is used to perform morphological processing on the binarized road image, thereby achieving denoising and enhancement of the road image. The morphologically processed road image is then input into a pre-trained neural network model for spilled object detection for spilled object detection, thereby improving the accuracy of spilled object detection based on the neural network model.

[0031] In an optional implementation, after acquiring the road image to be detected and before dividing the road image into a plurality of first sub-images, the method further includes:

[0032] Obtaining the tilt angle and rotation angle of the camera that shoots the road image; constructing a rotation matrix according to the tilt angle and rotation angle; and performing rotation correction on the road image using an affine transformation algorithm according to the rotation matrix.

[0033] The above technical solution has the following advantages or beneficial effects:

[0034] In this application, considering the adverse effects of image distortion on scattered object detection, after acquiring a road image to be detected, rotation correction is performed on the road image based on the tilt and rotation angles of the camera capturing the road image. The subsequent step of dividing the rotation-corrected road image into multiple first sub-images is then performed. This improves the accuracy of scattered object detection.

[0035] In an optional embodiment, the training process of the GMM model includes:

[0036] For each first sample image in the first training set, the first sample image is divided into multiple sample sub-images. For the multiple sample sub-images, a sample grayscale histogram corresponding to the sample sub-image is determined, and the sample grayscale histogram and the first label information corresponding to the sample sub-image are input into the GMM model to be trained, wherein the first label information is label information of whether the sample sub-image contains spilled objects or not; based on each Gaussian distribution initialized in the GMM model to be trained, a sample feature vector corresponding to the sample sub-image is determined, and a first prediction result of whether the sample sub-image contains spilled objects is determined based on the sample feature vector; a first loss value is determined according to the first prediction result and the first label information; and the GMM model to be trained is trained according to the first loss value.

[0037] In an optional embodiment, the training process of the neural network model includes:

[0038] For each second sample image in the second training set, the second sample image and the corresponding second label information are input into the neural network model to be trained, wherein the second label information is label information of whether the second sample image contains spilled objects or not; based on the neural network model to be trained, a second prediction result of whether the second sample image contains spilled objects is determined; according to the second prediction result and the second label information, a second loss value is determined; and according to the second loss value, the neural network model to be trained is trained.

[0039] In a second aspect, the present application provides a road spilled object detection device, the device comprising:

[0040] an acquisition module, configured to acquire a road image to be detected and divide the road image into a plurality of first sub-images;

[0041] a determination module, configured to determine, for each of the plurality of first sub-images, a target grayscale histogram corresponding to each of the first sub-images, input the target grayscale histogram into a pre-trained Gaussian mixture model (GMM), and determine, based on the GMM, whether the first sub-image contains a detection result of a scattered object;

[0042] The detection module is configured to determine whether there are scattered objects in the road image if a first sub-image detected as containing scattered objects exists in the road image.

[0043] In an optional embodiment, the determination module is specifically used to count the first grayscale histogram corresponding to the first sub-image, and use an equalization algorithm to equalize the first grayscale histogram to obtain a second grayscale histogram; the number of pixels in the second grayscale histogram corresponding to the grayscale that is less than a preset first grayscale threshold and greater than a preset second grayscale threshold is set to 0 to obtain a target grayscale histogram; wherein the preset first grayscale threshold is less than the preset second grayscale threshold.

[0044] In an optional embodiment, the determination module is further configured to obtain first brightness information and first contrast information of the road image; determine target type information corresponding to the road image based on the first brightness information and the first contrast information, as well as brightness information ranges and contrast information ranges corresponding to images of different types of information, respectively; wherein the type information includes type information requiring enhancement processing and type information not requiring enhancement processing; if the target type information is type information requiring enhancement processing, perform an equalization process on the first grayscale histogram using an equalization algorithm to obtain a second grayscale histogram;

[0045] The determination module is further configured to use the first grayscale histogram as the target grayscale histogram corresponding to the first sub-image if the target type information is type information that does not require enhancement processing.

[0046] In an optional embodiment, the device further includes:

[0047] The second detection module is used to input the road image into a pre-trained neural network model for detecting spilled objects, and trigger the acquisition module if it is determined based on the neural network model that there are spilled objects in the road image.

[0048] In an optional embodiment, the second detection module is further used to binarize the road image, perform morphological processing on the binarized road image using a preset morphological operation window, and input the morphologically processed road image into a pre-trained neural network model for spilled object detection.

[0049] In an optional embodiment, the acquisition module is further used to obtain the tilt angle and rotation angle of the camera that shoots the road image; construct a rotation matrix based on the tilt angle and rotation angle; and use an affine transformation algorithm to perform rotation correction on the road image based on the rotation matrix.

[0050] In an optional embodiment, the device further includes:

[0051] The first training module is used to divide each first sample image in the first training set into multiple sample sub-images, determine the sample grayscale histogram corresponding to the sample sub-image for the multiple sample sub-images, and input the sample grayscale histogram and the first label information corresponding to the sample sub-image into the GMM model to be trained, wherein the first label information is the label information of whether the sample sub-image contains spilled objects or not; determine the sample feature vector corresponding to the sample sub-image based on each Gaussian distribution initialized in the GMM model to be trained, and determine a first prediction result of whether the sample sub-image contains spilled objects based on the sample feature vector; determine a first loss value based on the first prediction result and the first label information; and train the GMM model to be trained based on the first loss value.

[0052] In an optional embodiment, the device further includes:

[0053] The second training module is used to input the second sample image and the corresponding second label information into the neural network model to be trained for each second sample image in the second training set, wherein the second label information is label information of whether the second sample image contains spilled objects or not; based on the neural network model to be trained, determine a second prediction result of whether the second sample image contains spilled objects; determine a second loss value according to the second prediction result and the second label information; and train the neural network model to be trained according to the second loss value.

[0054] In a third aspect, the present application provides an electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0055] Memory for storing computer programs;

[0056] The processor is used to implement the method when executing the program stored in the memory.

[0057] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described is implemented.

[0058] In a fifth aspect, the present application provides a computer program product, which includes an executable program, and the executable program is executed by a processor to implement the described method. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0060] Figure 1 A schematic diagram of the first road spill detection process provided in this application;

[0061] Figure 2 Schematic diagram of the training process of the GMM model provided in this application;

[0062] Figure 3 A schematic diagram of the process of determining the target grayscale histogram corresponding to the first sub-image provided by the present application;

[0063] Figure 4 A schematic diagram of the second process of determining the target grayscale histogram corresponding to the first sub-image provided by this application;

[0064] Figure 5 A schematic diagram of the second road spill detection process provided in this application;

[0065] Figure 6 Schematic diagram of the third road spill detection process provided by this application;

[0066] Figure 7 Schematic diagram of the road spill detection scenario provided for this application;

[0067] Figure 8 This is the architecture diagram of the application's combination of CNN and GMM models for spilled object detection;

[0068] Figure 9 Schematic diagram of the tilt angle and rotation angle of the camera provided for this application;

[0069] Figure 10 Schematic diagram of expanding a dataset by image rotation provided in this application;

[0070] Figure 11 Schematic diagram of the dataset augmented by salt and pepper noise provided for this application;

[0071] Figure 12 The overall flow chart of road spill detection provided for this application;

[0072] Figure 13 This is a schematic diagram of the structure of the road spilled object detection device provided in this application;

[0073] Figure 14 This is a schematic diagram of the electronic device structure provided in this application. DETAILED DESCRIPTION

[0074] In order to make the purpose and implementation of this application clearer, the exemplary implementation of this application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only part of the embodiments of this application, not all of the embodiments.

[0075] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.

[0076] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.

[0077] The terms "comprise," "include," and "have," and any variations thereof, are intended to cover but not exclude inclusion; for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0078] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functionality associated with that element.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

[0080] For ease of explanation, the above description has been made with reference to specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations are possible. The above embodiments are selected and described to better explain the principles and practical applications, so that those skilled in the art can better utilize the embodiments and various different variations of the embodiments suitable for specific use considerations.

[0081] Figure 1 The first road spill detection process provided in this application includes the following steps:

[0082] S101: Acquire a road image to be detected, and divide the road image into a plurality of first sub-images;

[0083] S102: determining a target grayscale histogram corresponding to each of the plurality of first sub-images, inputting the target grayscale histogram into a pre-trained Gaussian mixture model (GMM), and determining whether the first sub-image contains a detection result of a scattered object based on the GMM;

[0084] S103: If the road image contains a first sub-image detected to contain scattered objects, determine that the road image contains scattered objects.

[0085] The road spilled object detection method provided in the present application is applied to electronic devices, which may be tablet computers, computers, servers and other devices.

[0086] The camera captures the road image, and the electronic device acquires the road image. The camera can be a fixed snapshot camera on the road or an onboard computer in a vehicle. After acquiring the road image to be inspected, the electronic device divides the road image into multiple first sub-images. Optionally, the road image can be evenly divided into multiple first sub-images to obtain multiple first sub-images of equal size; alternatively, the road image can be unevenly divided. The first sub-images are generally rectangular sub-images.

[0087] For multiple first sub-images, determine the target grayscale histogram corresponding to the first sub-image. Optionally, count the grayscale histograms of the first sub-images as the target grayscale histogram corresponding to the first sub-image. The horizontal axis of the grayscale histogram is the grayscale value range [0, 255], and the vertical axis is the number of pixels corresponding to each grayscale value. A pre-trained Gaussian mixture model (GMM) model is deployed in the electronic device, and the target grayscale histogram is input into the GMM model. Based on the GMM model, it is determined whether the first sub-image contains a detection result of a spilled object.

[0088] Figure 2 The training process diagram of the GMM model provided in this application includes the following steps:

[0089] S201: For each first sample image in a first training set, divide the first sample image into a plurality of sample sub-images, determine a sample grayscale histogram corresponding to the sample sub-image for each of the plurality of sample sub-images, and input the sample grayscale histogram and first label information corresponding to the sample sub-image into a GMM model to be trained, wherein the first label information is label information indicating whether the sample sub-image contains spilled objects or does not contain spilled objects;

[0090] S202: Based on each Gaussian distribution initialized in the GMM model to be trained, determine the sample feature vector corresponding to the sample sub-image, and determine a first prediction result of whether the sample sub-image contains spilled objects based on the sample feature vector; determine a first loss value based on the first prediction result and the first label information; and train the GMM model to be trained based on the first loss value.

[0091] During the iterative training process, GMM model training is determined to be complete when the determined first loss value meets the requirement, or when the number of iterations meets the requirement. The GMM model trained using the above method is capable of extracting features from the grayscale histogram and outputting a detection result for whether spilled objects are present in the image corresponding to the grayscale histogram.

[0092] If the road image contains a first sub-image whose detection result indicates that the sub-image contains scattered objects, it is determined that the road image contains scattered objects; otherwise, it is determined that the road image does not contain scattered objects.

[0093] In the present application, after obtaining the road image to be detected, the road image is first divided into multiple first sub-images, and then the target grayscale histogram corresponding to the first sub-image is determined; then, based on the Gaussian mixture model (GMM), feature extraction is performed on the target grayscale histogram, and it is determined whether the corresponding first sub-image contains spilled objects. If there is a first sub-image in the road image that contains spilled objects, it is determined that there are spilled objects in the road image. This achieves intelligent automatic detection of road spilled objects. Compared with the method of personnel watching video image detection in the related art, the efficiency of road spilled object detection is improved, and the problem of missed detection of spilled objects due to personnel negligence is avoided, thereby improving the accuracy of road spilled object detection.

[0094] Figure 3 The schematic diagram of the first process of determining the target grayscale histogram corresponding to the first sub-image provided in this application includes the following steps:

[0095] S301: Counting a first grayscale histogram corresponding to the first sub-image, and performing equalization processing on the first grayscale histogram using an equalization algorithm to obtain a second grayscale histogram;

[0096] S302: Set the number of pixel points corresponding to grayscales smaller than a preset first grayscale threshold and larger than a preset second grayscale threshold in the second grayscale histogram to 0 to obtain a target grayscale histogram; wherein the preset first grayscale threshold is smaller than the preset second grayscale threshold.

[0097] After calculating the first grayscale histogram corresponding to the first sub-image, the electronic device calls the OpenCV equalization function (equalization algorithm) to equalize the first grayscale histogram to obtain a second grayscale histogram. The second grayscale histogram is then pruned, i.e., the number of pixels in the second grayscale histogram corresponding to grayscales less than a preset first grayscale threshold and greater than a preset second grayscale threshold is set to zero, thereby obtaining a target grayscale histogram. Examples of the preset first grayscale threshold are 2, 3, etc., and examples of the preset second grayscale threshold are 253, 254, etc.

[0098] The process of image enhancement is divided into the following steps.

[0099] Step 1: Image preprocessing.

[0100] If it is a color image, it must be converted to brightness space first.

[0101] Step 2: Divide the image into several small blocks.

[0102] Divide the image into small blocks (e.g. 8x8) and perform independent enhancement on each block. Suppose the image size is M×N, divided into m×n blocks, and the size of each block is: M / m*N / n.

[0103] Step 3: Histogram equalization + contrast limiting.

[0104] Each block is given a histogram, which counts the frequency of each grayscale value within the block. The histogram dimension is 256. Histogram equalization is implemented using the OpenCV function "cv::equalizeHist." Contrast limiting is performed by setting two thresholds, min and max. Pixels with grayscale values ​​below the min or above the max are not counted.

[0105] The specific steps of equalization are to call the equalization function of opencv: assuming that hist_i is the histogram of the i-th image block, then the histogram is equalized directly by calling the function cv::equalizeHist(hist_i).

[0106] Contrast limiting is another step after histogram equalization. Contrast limiting essentially clips the histogram. A histogram is a 256-dimensional statistical graph that counts the frequency of each grayscale. Contrast limiting sets the frequency of the smallest and largest grayscales (K) to zero.

[0107] For example, after histogram equalization, a statistical number with a dimension of 256 is expressed as follows:

[0108] [h1,h2,h3,h4,…,h252,h253,h254,h255,h256].

[0109] Assuming that the contrast limit K is 2, the statistical graph after contrast limitation becomes:

[0110] [0,0,h3,h4,…,h252,h253,h254,0,0].

[0111] In the present application, the first sub-image is enhanced by combining histogram equalization with contrast limitation. Histogram equalization refers to statistically analyzing the first grayscale histogram corresponding to the first sub-image, and performing equalization processing on the first grayscale histogram using an equalization algorithm to obtain a second grayscale histogram. Contrast limitation refers to setting the number of pixels in the second grayscale histogram corresponding to grayscales less than a preset first grayscale threshold and greater than a preset second grayscale threshold to 0 to obtain a target grayscale histogram. The target grayscale histogram obtained in this way is more accurate, thereby making the detection of spilled objects based on the target grayscale histogram more accurate.

[0112] Figure 4 The second process for determining the target grayscale histogram corresponding to the first sub-image provided in this application is schematically shown, comprising the following steps:

[0113] S401: Acquire first brightness information and first contrast information of a road image to be detected;

[0114] S402: Determining object type information corresponding to the road image based on the first brightness information and the first contrast information, and the brightness information ranges and contrast information ranges corresponding to images of different types of information, wherein the type information includes type information requiring enhancement processing and type information not requiring enhancement processing;

[0115] S403: If the target type information is information requiring enhancement processing, equalize the first grayscale histogram using an equalization algorithm to obtain a second grayscale histogram; set the number of pixels in the second grayscale histogram corresponding to grayscales less than a preset first grayscale threshold and greater than a preset second grayscale threshold to zero, to obtain a target grayscale histogram; wherein the preset first grayscale threshold is less than the preset second grayscale threshold;

[0116] S404: If the target type information is information of a type that does not require enhancement processing, use the first grayscale histogram as a target grayscale histogram corresponding to the first sub-image.

[0117] Lighting significantly impacts image clarity and quality. Based on image contrast and brightness, images can be categorized into six types: daytime strong light, nighttime strong light, daytime weak light, nighttime weak light, daytime normal, and nighttime normal. Analysis revealed that for normal daytime and nighttime images, the contrast between spilled objects and the background road surface is high, so specialized enhancement processing is not required for these two types of images. However, poor lighting conditions prevent effective identification of spilled objects in the other four types of images, so a contrast-adaptive enhancement algorithm is employed.

[0118] The brightness information range and contrast information range corresponding to images of different types of information, where the type information includes type information that requires enhancement processing and type information that does not require enhancement processing. The type information that requires enhancement processing includes daytime strong light type, nighttime strong light type, daytime weak light type, and nighttime weak light type; the type information that does not require enhancement processing includes daytime normal type and nighttime normal type. Determine whether it is daytime or nighttime based on the brightness information. Set a brightness threshold. If the first brightness information of the road image is greater than the brightness threshold, it is determined to be daytime; otherwise, it is determined to be nighttime. Set two contrast thresholds, namely L1 and L2, where L1 < L2. If the first contrast information of the road image is greater than L2, it is determined to be the strong light type; if the first contrast information of the road image is less than L1, it is determined to be the weak light type; if the first contrast information of the road image is between L1 and L2, it is determined to be the normal type.

[0119] Based on the first brightness information and the first contrast information of the road image, determine the target type information corresponding to the road image. If the target type information is one of the daytime strong light type, nighttime strong light type, daytime weak light type, and nighttime weak light type, perform the step of equalizing the first gray-level histogram using an equalization algorithm to obtain a second gray-level histogram. If the target type information is one of the daytime normal type and nighttime normal type, use the first gray-level histogram as the target gray-level histogram corresponding to the first sub-image.

[0120] In this application, before using the equalization algorithm to equalize the first gray-level histogram to obtain a second gray-level histogram, first determine the target type information corresponding to the road image based on the first brightness information and the first contrast information of the road image. If the target type information is type information that does not require enhancement processing, use the first gray-level histogram as the target gray-level histogram corresponding to the first sub-image. Do not perform the subsequent step of using the equalization algorithm to equalize the first gray-level histogram to obtain a second gray-level histogram, so as to balance the efficiency and accuracy of spill detection. If the target type information is type information that requires enhancement processing, perform the step of using the equalization algorithm to equalize the first gray-level histogram to obtain a second gray-level histogram; thus ensuring the accuracy of spill detection.

[0121] Figure 5 This is the schematic diagram of the second road spill detection process provided by this application, including the following steps:

[0122] S501: Obtain the road image to be detected, and input the road image into a pre-trained neural network model for spill detection;

[0123] S502: If it is determined based on the neural network model that there is no spill in the road image, then determine that there is no spill in the road image;

[0124] S503: If it is determined based on the neural network model that there are spilled objects in the road image, dividing the road image into a plurality of first sub-images;

[0125] S504: Determine, for each of the plurality of first sub-images, a target grayscale histogram corresponding to each of the first sub-images, input the target grayscale histogram into a pre-trained Gaussian mixture model (GMM), and determine, based on the GMM, whether the first sub-image contains a detection result of a scattered object.

[0126] S505: If the first sub-image detected as containing scattered objects exists in the road image, determine that the scattered objects exist in the road image;

[0127] S506: If the road image does not contain the first sub-image detected to contain the scattered objects, determine that the road image does not contain the scattered objects.

[0128] It should be noted that the method involves obtaining a road image to be inspected and dividing it into multiple first sub-images. For each of these first sub-images, first grayscale histograms corresponding to the first sub-images are calculated and then equalized using an equalization algorithm to obtain a second grayscale histogram. The number of pixels in the second grayscale histogram corresponding to grayscales less than a preset first grayscale threshold and greater than a preset second grayscale threshold is set to zero to obtain a target grayscale histogram. This method enhances each of the first sub-images in the road image. After enhancement of the first sub-images, discontinuities may appear at the edges, requiring bilinear interpolation to fuse the results between the first sub-images for a smooth transition. The enhanced luminance channel is merged with the original color channel and converted back to RGB space, resulting in a higher-quality enhanced road image. The enhanced road image is then input into a pre-trained neural network model for spilled object detection. If the neural network model determines the presence of spilled objects in the road image, subsequent steps for spilled object detection based on a GMM model are performed. If it is determined based on the neural network model that no spilled objects exist in the road image, it can be directly determined that no spilled objects exist on the road.

[0129] The training process of the neural network model includes:

[0130] For each second sample image in the second training set, the second sample image and the corresponding second label information are input into the neural network model to be trained, wherein the second label information is label information of whether the second sample image contains spilled objects or not; based on the neural network model to be trained, a second prediction result of whether the second sample image contains spilled objects is determined; according to the second prediction result and the second label information, a second loss value is determined; and according to the second loss value, the neural network model to be trained is trained.

[0131] During the iterative training process, the neural network model training is determined to be complete when the determined second loss value meets the requirement, or when the number of iterations meets the requirement. The neural network model trained using the above method is capable of extracting features from the road system and outputting a detection result for whether spilled objects are present in the road image.

[0132] In the present application, before dividing the image into a plurality of first sub-images, the road image is first input into a pre-trained neural network model for spilled object detection. If the neural network model determines that no spilled objects are present in the road image, then the road image is determined to be free of spilled objects, and subsequent spilled object detection based on the GMM model is unnecessary. If the neural network model determines that spilled objects are present in the road image, a secondary detection of spilled objects is then performed based on the GMM model, thereby ensuring the accuracy of spilled object detection.

[0133] Before inputting the road image into a pre-trained neural network model for spilled object detection, the road image is binarized. Optionally, pixels in the road image with a grayscale value greater than a set threshold are assigned a first grayscale value, while all other pixels are assigned a second grayscale value. The first grayscale value and the second grayscale value are different. For example, the first grayscale value can be 0 and the second grayscale value can be 255; or the first grayscale value can be 255 and the second grayscale value can be 0; or other different grayscale values ​​can be present between 0 and 255.

[0134] The preset morphological operation window is, for example, a 3×3, 5×5, or other size window. Morphological processing is performed on the binarized road image using the preset morphological operation window, including erosion and dilation operations, combined with opening and closing operations, to determine the location information of the spilled objects in the road image.

[0135] Morphological operations are usually applied to images after thresholding of binary or grayscale images. This application calls the opencv function cv2.threshold. The structure element defines the "window size" and shape of the morphological operation, which is commonly 3x3 or 5x5. Erosion operation (removing small noise): The erosion operation will make small white areas (such as noise) disappear and reduce the boundary of large targets. Dilation operation (restoring the target): The dilation operation is used to "fill back" the real target area that has been eroded, and can also fill small holes. Use opening and closing operations (combined operations): Use opening operations (erosion + dilation) to remove noise. Use closing operations (dilation + erosion) to fill holes.

[0136] In this application, the road image is binarized, and then a preset morphological operation window is used to perform morphological processing on the binarized road image, thereby achieving denoising and enhancement of the road image, and then the morphologically processed road image is input into a pre-trained neural network model for spilled object detection for spilled object detection, thereby improving the accuracy of spilled object detection based on the neural network model.

[0137] Figure 6 The third road spill detection process diagram provided in this application includes the following steps:

[0138] S601: Acquire a road image to be detected, and acquire the tilt angle and rotation angle of a camera that captured the road image; construct a rotation matrix based on the tilt angle and rotation angle; perform rotation correction on the road image using an affine transformation algorithm based on the rotation matrix, and divide the rotation-corrected road image into a plurality of first sub-images;

[0139] S602: Determine, for each of the plurality of first sub-images, a target grayscale histogram corresponding to each of the first sub-images, input the target grayscale histogram into a pre-trained Gaussian mixture model (GMM), and determine, based on the GMM, whether the first sub-image contains a detection result of a scattered object.

[0140] S603: If the road image contains a first sub-image whose detection result indicates that the sub-image contains scattered objects, determine that the road image contains scattered objects.

[0141] In this application, considering the adverse effects of image distortion on scattered object detection caused by issues such as improper camera geometric parameter calibration and shooting angle, after acquiring the road image to be detected, this application performs rotation correction on the road image based on the tilt and rotation angle of the camera that captured the road image. The subsequent step of dividing the rotation-corrected road image into multiple first sub-images is then performed. This improves the accuracy of scattered object detection.

[0142] Figure 7Schematic diagram of the road spilled object detection scenario provided in this application. Figure 7 The part selected in the middle box is the detected road spillage.

[0143] This application combines the neural network model (CNN) with the HSV color space in the Gaussian mixture model (GMM) to propose a video image acquisition and preprocessing technology for detecting road spills in complex traffic environments and providing early warning of incidents. The GMM-based HSV color space detection algorithm proposed in this application uses a fast and dynamic GMM background modeling method to improve the algorithm's accuracy and robustness, achieving a detection accuracy rate of 90.42% for road spills and a false alarm rate of less than 10%.

[0144] Figure 8 This application provides an architectural diagram for detecting spilled objects by combining a CNN model and a GMM model, wherein a road image is subjected to image geometry correction, illumination enhancement, image segmentation, GMM local feature generation, and a first detection result; a road image, a CNN model, CNN global feature generation, and a second detection result; and the first and second detection results are merged.

[0145] The final result is obtained by fusing the classification results of the local features (i.e., the first detection result based on the GMM features) with the classification results of the global features (i.e., the second detection result based on the CNN features). Specifically, if the probability of the presence of scattered objects determined by the CNN output is less than a threshold, the image is directly considered to be free of scattered objects. Otherwise, the foreground image block detected by the GMM local features is considered to be scattered objects.

[0146] The process of image geometric correction for road images is as follows:

[0147] Due to issues such as improper geometric parameter calibration and shooting angles of fixed roadside cameras, the images of road spillage events captured are subject to geometric distortion or abnormal relative angles. As camera product performance indicators gradually improve, the geometric distortion of images caused by geometric parameter drift or instability has diminished, making the shooting angle issue a major disadvantage in image recognition. This application utilizes a horizontal rotation correction algorithm to correct the camera angle for horizontal and vertical motion.

[0148] The specific steps are divided into three parts, and the detailed process is as follows.

[0149] Step 1: Obtain tilt angle information.

[0150] Based on the image content (such as straight line / horizon detection), use Hough Transform or Canny edge detection combined with line fitting to find the main lines in the image (such as the horizon, wall corners, desktop edges, etc.). Calculate the angle between the direction of these lines and the horizontal line, that is, the camera roll angle. For example, assuming a straight line is detected with a slope of m, its angle is: θ = arctan (m). This is the roll angle. Similarly, calculate the rotation offset of the camera, that is, the clockwise rotation angle Figure 9 Schematic diagram of the tilt angle and rotation angle of the camera provided in this application.

[0151] Step 2: Calculate the rotation matrix.

[0152] Based on the roll and pitch values, construct the rotation matrix RRR to rotate the image inversely:

[0153]

[0154] The core of image coordinate transformation is 2D affine rotation or projection matrix, which is performed using OpenCV's affine transformation function.

[0155] Step 3: Image rotation correction.

[0156] For image rotation, the function cv2.warpAffin in OpenCV is used to obtain the rotated road image.

[0157] The process of GMM local features and detection is described as follows:

[0158] The GMM model overcomes the limitations of road spill detection and maintains excellent applicability in dynamic video scenes, enabling dynamic background modeling and image acquisition. Compared to the three-primary color (RGB) model, three-channel images in the HSV color space more closely resemble the human eye's color classification, resulting in better target classification.

[0159] First, the GMM model suitable for dynamic background is used to model the background of the image and remove moving targets such as vehicles and pedestrians; secondly, with the assistance of lane line detection, the road surface is segmented as the area to be detected and irrelevant background is removed; the road surface color is modeled using the HSV model (each color is represented by hue (H), saturation (S) and value (V). The color parameters in this model are: hue (H), saturation (S), and brightness (V). The difference between the object to be detected and the road surface is used to separate the object to be detected and the road surface; the grayscale image is processed by morphological filtering to remove noise, and the expansion and corrosion operations are used to highlight the object to be detected; the support vector machine (SVM) is used to train the target detection model to remove other objects (such as lane lines) that appear on the road surface; finally, a rectangular box is used to select and display the spilled objects on the road, the event name is marked, and the system prompts an alarm message for abnormal road events and stores the information in the database.

[0160] The detailed process of using GMM to model the background and remove irrelevant background is as follows.

[0161] Step 1: Initialize the model and obtain GMM features.

[0162] Initialize K Gaussian distributions to describe the visual features of each image block. Each Gaussian contains three parameters, initial mean μ, initial variance σ ^ 2. Initial weight w. Each image block will obtain a GMM feature described by Gaussian distribution, and the Fisher vector is actually used.

[0163] This application uses the classic GMM-based Fisher Vector encoding method. The GMM model is used to model the histogram of each block, using an initial set of 16 or 32 Gaussians. After modeling is completed, the grayscale histogram of an image block is input into the GMM model, and a vector describing the probabilistic meaning of the grayscale histogram is obtained. The obtained vector is also the feature vector used to describe the image block. This feature vector is also called a Fisher Vector.

[0164] Each training image block is labeled, including background and foreground (i.e., spilled objects). The GMM features of each block are used as input to train an SVM classifier, resulting in a GMM model that can distinguish between background and foreground. For a query image, each block is fed into the SVM classifier, and the probability of each block being the foreground is determined, sorting the blocks from highest to lowest probability. All blocks with probabilities greater than a threshold T are selected as foreground, i.e., spilled objects; the remaining blocks are considered background.

[0165] It's also important to note that when training a CNN model, the generalization ability of the CNN is related to the number of training samples. A large sample size not only improves image classification accuracy but also effectively prevents overfitting. However, in actual testing, it's difficult to obtain a sufficiently large sample size for classifier training. Therefore, artificially augmenting the training data is necessary. This involves adding additional copies of the training set by performing geometric transformations on the original images that don't change their classification. This increases the size of the training set and improves the generalization ability of the CNN model. In the collected highway video image database, since random cropping and appropriate rotation do not affect highway asset attributes, various geometric transformations, such as random scaling, image rotation, mirror flipping, and adding salt and pepper noise, are used to supplement the dataset to address the insufficient sample size. After obtaining the preliminary dataset, supervised classification methods were used, and data annotation was performed using LabelImg annotation software. The model was trained using this processed and augmented dataset, and the trained model was then used for detection. The CNN generates a single vector for the entire image. The CNN output is the probability of the presence of a spilled object in the image.

[0166] Figure 10 This is a schematic diagram of expanding a dataset by image rotation provided in this application. Figure 10 In the figure, the left image is the sample image before rotation, and the right image is the sample image after rotation. Figure 11 Schematic diagram of the dataset augmented by salt and pepper noise provided for this application, Figure 11 In the figure, the left image is a sample image without added salt and pepper noise, and the right image is a sample image with added salt and pepper noise.

[0167] Figure 12 The overall flow chart of road spilled object detection provided in this application includes: inputting video image frames, lane line detection and road surface segmentation, clustering algorithm for object classification on the road surface, morphological filtering to remove noise and enhance the image, using SVM classifier (GMM model and CNN model) for spilled object detection, and judging whether spilled objects are detected. If so, frame the area where the spilled objects are located and generate an early warning. If not, input the next frame of image and repeat the lane line detection and road surface segmentation.

[0168] Figure 13 This is a schematic diagram of the structure of the road spilled object detection device provided in this application, the device includes:

[0169] An acquisition module 131 is configured to acquire a road image to be detected and divide the road image into a plurality of first sub-images;

[0170] a determination module 132 configured to determine, for each of the plurality of first sub-images, a target grayscale histogram corresponding to each of the first sub-images, input the target grayscale histogram into a pre-trained Gaussian mixture model (GMM), and determine, based on the GMM, whether the first sub-image contains a detection result of a scattered object;

[0171] The first detection module 133 is configured to determine whether there are scattered objects in the road image if a first sub-image detected as containing scattered objects exists in the road image.

[0172] The determination module 132 is specifically used to count the first grayscale histogram corresponding to the first sub-image, and use an equalization algorithm to equalize the first grayscale histogram to obtain a second grayscale histogram; the number of pixels corresponding to the grayscale less than a preset first grayscale threshold and greater than a preset second grayscale threshold in the second grayscale histogram is set to 0 to obtain a target grayscale histogram; wherein the preset first grayscale threshold is less than the preset second grayscale threshold.

[0173] The determination module 132 is further configured to obtain first brightness information and first contrast information of the road image; determine target type information corresponding to the road image based on the first brightness information and the first contrast information, as well as the brightness information ranges and contrast information ranges corresponding to images of different types of information; wherein the type information includes type information requiring enhancement processing and type information not requiring enhancement processing; and if the target type information is type information requiring enhancement processing, perform an equalization process on the first grayscale histogram using an equalization algorithm to obtain a second grayscale histogram;

[0174] The determination module 132 is further configured to use the first grayscale histogram as a target grayscale histogram corresponding to the first sub-image if the target type information is information of a type that does not require enhancement processing.

[0175] The device further comprises:

[0176] The second detection module 134 is configured to input the road image into a pre-trained neural network model for detecting spilled objects, and trigger the acquisition module if it is determined based on the neural network model that there are spilled objects in the road image.

[0177] The second detection module 134 is further configured to binarize the road image, perform morphological processing on the binarized road image using a preset morphological operation window, and input the morphologically processed road image into a pre-trained neural network model for spilled object detection.

[0178] The acquisition module 131 is further configured to acquire the tilt angle and rotation angle of the camera that captures the road image; construct a rotation matrix based on the tilt angle and rotation angle; and perform rotation correction on the road image using an affine transformation algorithm based on the rotation matrix.

[0179] The device further comprises:

[0180] The first training module 135 is used to divide each first sample image in the first training set into multiple sample sub-images, determine a sample grayscale histogram corresponding to the sample sub-image for the multiple sample sub-images, and input the sample grayscale histogram and the first label information corresponding to the sample sub-image into the GMM model to be trained, wherein the first label information is label information of whether the sample sub-image contains spilled objects or not; determine the sample feature vector corresponding to the sample sub-image based on each Gaussian distribution initialized in the GMM model to be trained, and determine a first prediction result of whether the sample sub-image contains spilled objects based on the sample feature vector; determine a first loss value based on the first prediction result and the first label information; and train the GMM model to be trained based on the first loss value.

[0181] The device further comprises:

[0182] The second training module 136 is used to input the second sample image and the corresponding second label information into the neural network model to be trained for each second sample image in the second training set, wherein the second label information is label information of whether the second sample image contains spilled objects or not; based on the neural network model to be trained, determine a second prediction result of whether the second sample image contains spilled objects; determine a second loss value according to the second prediction result and the second label information; and train the neural network model to be trained according to the second loss value.

[0183] The present application also provides an electronic device, such as Figure 14 As shown, it includes: a processor 141, a communication interface 142, a memory 143 and a communication bus 144, wherein the processor 141, the communication interface 142, and the memory 143 communicate with each other through the communication bus 144;

[0184] The memory 143 stores a computer program, and when the program is executed by the processor 141 , the processor 141 performs any of the above method steps.

[0185] The communication bus mentioned in the electronic device mentioned above may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.

[0186] The communication interface 142 is used for communication between the electronic device and other devices.

[0187] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk memory. Alternatively, the memory may be at least one storage device located away from the processor.

[0188] The above-mentioned processor can be a general-purpose processor, including a central processing unit, a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.

[0189] The present application also provides a computer storage readable storage medium, which stores a computer program that can be executed by an electronic device. When the program runs on the electronic device, the electronic device implements any of the above method steps when executing.

[0190] The present application provides a computer program product, which includes an executable program. When the executable program is executed by a processor, the method described above is implemented.

[0191] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0192] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A method for detecting spilled objects on a road, characterized in that: The method comprises: Acquire a road image to be detected, and divide the road image into a plurality of first sub-images; For the multiple first sub-images, determine a target grayscale histogram corresponding to the first sub-image, input the target grayscale histogram into a pre-trained Gaussian mixture model (GMM), and determine whether the first sub-image contains a detection result of a scattered object based on the GMM; If the road image contains a first sub-image whose detection result indicates that the sub-image contains scattered objects, it is determined that the road image contains scattered objects.

2. The method according to claim 1, wherein Determining a target grayscale histogram corresponding to the first sub-image includes: Counting a first grayscale histogram corresponding to the first sub-image, and performing equalization processing on the first grayscale histogram using an equalization algorithm to obtain a second grayscale histogram; The number of pixels corresponding to grayscales less than a preset first grayscale threshold and greater than a preset second grayscale threshold in the second grayscale histogram is set to 0 to obtain a target grayscale histogram; wherein the preset first grayscale threshold is less than the preset second grayscale threshold.

3. The method according to claim 2, wherein Before the equalization algorithm is used to perform equalization processing on the first grayscale histogram to obtain the second grayscale histogram, the method further includes: Acquiring first brightness information and first contrast information of the road image; Determining target type information corresponding to the road image based on the first brightness information and the first contrast information, and the brightness information ranges and contrast information ranges corresponding to images of different types of information, wherein the type information includes type information requiring enhancement processing and type information not requiring enhancement processing; If the target type information is information of a type requiring enhancement processing, performing an equalization process on the first grayscale histogram using an equalization algorithm to obtain a second grayscale histogram; If the target type information is type information that does not require enhanced processing, the method further includes: The first grayscale histogram is used as a target grayscale histogram corresponding to the first sub-image.

4. The method according to claim 1, wherein Before dividing the image into a plurality of first sub-images, the method further includes: The road image is input into a pre-trained neural network model for detecting spilled objects. If it is determined based on the neural network model that there are spilled objects in the road image, a subsequent step of dividing the image into multiple first sub-images is performed.

5. The method according to claim 4, wherein Before inputting the road image into a pre-trained neural network model for spilled object detection, the method further includes: The road image is binarized, a preset morphological operation window is used to perform morphological processing on the binarized road image, and the morphologically processed road image is input into a pre-trained neural network model for spilled object detection.

6. The method according to claim 1, wherein After acquiring the road image to be detected and before dividing the road image into a plurality of first sub-images, the method further includes: Obtaining the tilt angle and rotation angle of the camera that shoots the road image; constructing a rotation matrix according to the tilt angle and rotation angle; and performing rotation correction on the road image using an affine transformation algorithm according to the rotation matrix.

7. The method according to claim 1, wherein The training process of the GMM model includes: For each first sample image in the first training set, the first sample image is divided into multiple sample sub-images. For the multiple sample sub-images, a sample grayscale histogram corresponding to the sample sub-image is determined, and the sample grayscale histogram and the first label information corresponding to the sample sub-image are input into the GMM model to be trained, wherein the first label information is label information of whether the sample sub-image contains spilled objects or not; based on each Gaussian distribution initialized in the GMM model to be trained, a sample feature vector corresponding to the sample sub-image is determined, and a first prediction result of whether the sample sub-image contains spilled objects is determined based on the sample feature vector; a first loss value is determined according to the first prediction result and the first label information; and the GMM model to be trained is trained according to the first loss value.

8. The method according to claim 4, wherein The training process of the neural network model includes: For each second sample image in the second training set, the second sample image and the corresponding second label information are input into the neural network model to be trained, wherein the second label information is label information of whether the second sample image contains spilled objects or not; based on the neural network model to be trained, a second prediction result of whether the second sample image contains spilled objects is determined; according to the second prediction result and the second label information, a second loss value is determined; and according to the second loss value, the neural network model to be trained is trained.

9. A road spilled object detection device, characterized in that: The device comprises: an acquisition module, configured to acquire a road image to be detected and divide the road image into a plurality of first sub-images; a determination module, configured to determine, for each of the plurality of first sub-images, a target grayscale histogram corresponding to each of the first sub-images, input the target grayscale histogram into a pre-trained Gaussian mixture model (GMM), and determine, based on the GMM, whether the first sub-image contains a detection result of a scattered object; The first detection module is configured to determine whether there are scattered objects in the road image if a first sub-image detected as containing scattered objects exists in the road image.

10. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 8 when executing a program stored in a memory.