Visual touch fusion-based road surface anomaly sensing method and device
Through the method of fusion of vision and tactile sense, the visual and tactile timestamp marks are used for matching and correction, and the problem of inaccurate detection of a single perception method in complex pavement environments is solved, and high-precision and reliable pavement abnormality detection is achieved.
Patent Information
- Application Number
- CN202510335807.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, a single visual or single tactile perception method is difficult to accurately detect pavement abnormalities in complex and changeable pavement environments, resulting in inaccurate detection results.
Combining the advantages of visual detection and haptic detection, visual features and tactile features of pavement abnormalities are extracted through neural networks and classification algorithms, and the timestamp marks of vision and haptics are used to match and correct them, so as to achieve decision-making fusion of visual and haptic information and output multimodal perception results.
It improves the accuracy and robustness of pavement abnormality detection, reduces the possibility of error matching caused by time errors, provides multi-dimensional pavement risk assessment, and ensures the reliability and accuracy of output results.
Smart Images

Figure CN120279524A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent driving, and particularly to a method and device for road anomaly perception based on visual-tactile fusion. Background Art
[0002] Currently, the perception and positioning of the surrounding environment of an automobile mainly rely on perception components such as visual sensors, lidar, and ultrasonic radars. Different perception means can output perception information of different modalities. Combining these perception information with deep learning algorithms and computer vision technology for identification can realize the identification of road anomalies, thereby providing comprehensive guarantee for driving safety.
[0003] In related technologies, usually only a single visual perception data or a single tactile perception data is relied on to detect road anomalies and evaluate the risk level of the road surface. This method often fails to cover many anomaly features when facing a complex and changeable road surface environment, resulting in inaccurate detection results.
[0004] Based on this, there is an urgent need for a method and device for road anomaly perception based on visual-tactile fusion to solve the above technical problems. Summary of the Invention
[0005] Embodiments of the present invention provide a method and device for road anomaly perception based on visual-tactile fusion, which can achieve accurate detection of road anomalies.
[0006] In a first aspect, embodiments of the present invention provide a method for road anomaly perception based on visual-tactile fusion, including:
[0007] Inputting the real-time visual data and real-time tactile data acquired by sensors into a trained anomaly detection model respectively, and sequentially outputting the visual detection result and tactile detection result of the road surface to be detected;
[0008] According to the first timestamp of the anomaly excitation in the tactile detection result, matching the second timestamp corresponding to the same road anomaly in the visual detection result;
[0009] Retrieving the historical image corresponding to the road anomaly in the visual data according to the second timestamp, and performing decision-level fusion on the retrieval result and the tactile detection result corresponding to the first timestamp to obtain the multi-modal perception result of the road anomaly.
[0010] In a second aspect, embodiments of the present invention further provide a device for road anomaly perception based on visual-tactile fusion, including:
[0011] A detection module, configured to input the real-time visual data and real-time tactile data acquired by sensors into a trained anomaly detection model respectively, and sequentially output the visual detection result and tactile detection result of the road surface to be detected;
[0012] A matching module, configured to match, according to a first timestamp of an abnormal excitation in the tactile detection result, a second timestamp corresponding to the same road surface abnormality in the visual detection result;
[0013] A fusion module, configured to retrieve a historical image corresponding to the road surface abnormality in the visual data according to the second timestamp, and perform decision-level fusion on the retrieved result and the tactile detection result corresponding to the first timestamp to obtain a multi-modal perception result of the road surface abnormality.
[0014] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the method described in any embodiment of this specification is implemented.
[0015] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in any embodiment of this specification.
[0016] In a fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the above-described road surface abnormality perception method based on visual-tactile fusion are implemented.
[0017] An embodiment of the present invention provides a road surface abnormality perception method and device based on visual-tactile fusion. The visual features and tactile features of road surface abnormalities in a sensing element are extracted through a neural network and a classification algorithm. Subsequently, in the PV perspective (driver's perspective), timestamp tags of vision and touch are introduced to identify each road surface abnormality that appears. Then, the abnormalities in the visual-tactile information are matched and corrected according to the time difference, and the corresponding visual-tactile information is subjected to decision-level fusion according to the correction result to realize the output of visual-tactile modal fusion information of road surface abnormalities. This method uses visual detection to give early warnings for road surface abnormalities at a long distance, and uses the information output method of visual-tactile fusion to detect road surface abnormalities at a short distance, helping the driver evaluate the road surface risk situation from more dimensions; at the same time, the possibility of incorrect matching due to time error is minimized during the process of visual-tactile information fusion, making the output result have high reliability and accuracy. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a flowchart of a road surface anomaly perception method based on visual-tactile fusion provided by an embodiment of the present invention;
[0020] Figure 2 It is a schematic diagram of the driving distance of a vehicle provided by an embodiment of the present invention;
[0021] Figure 3 It is a schematic diagram of a visualization interface when the tire does not contact a road surface anomaly at the starting moment provided by an embodiment of the present invention;
[0022] Figure 4 It is a schematic diagram of a visualization interface when contacting a road surface anomaly and retrieving a corresponding visual historical image provided by an embodiment of the present invention;
[0023] Figure 5 It is a hardware architecture diagram of an electronic device provided by an embodiment of the present invention;
[0024] Figure 6 It is a structural diagram of a road surface anomaly perception device based on visual-tactile fusion provided by an embodiment of the present invention. Detailed implementation manners
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0026] As mentioned above, the vast majority of existing manufacturers rely on a single visual sensor or a single tactile sensor to realize the perception of road surface anomaly conditions. When the single visual output faces scenarios such as poor lighting conditions, bad weather, and strong light reflection, the robustness and reliability of the visual model for anomaly detection are reduced; the single tactile sensor cannot realize the perception of all road surface anomalies or early warning.
[0027] Based on this, the concept of the present invention is to combine the advantages of visual detection and tactile detection, use the visual sensor to perform real-time detection and tracking of road surface anomalies; rely on the tactile sensor to accurately detect road surface anomalies such as manhole covers and speed bumps, and realize more detailed, accurate, and higher-robustness detection of road surface anomalies.
[0028] The following describes the specific implementation manners of the above concept.
[0029] Please refer to Figure 1 , an embodiment of the present invention provides a road surface anomaly perception method based on visual-tactile fusion, and the method includes:
[0030] Step 100: Input the real-time visual data and real-time tactile data obtained by the sensor into the trained anomaly detection model respectively, and sequentially output the visual detection result and the tactile detection result of the road surface to be detected.
[0031] Step 102: According to the first timestamp of the abnormal excitation in the tactile detection result, match the second timestamp corresponding to the same road surface anomaly in the visual detection result.
[0032] Step 104: Retrieve the historical image corresponding to the road surface anomaly in the visual data according to the second timestamp, and perform decision-level fusion on the retrieved result and the tactile detection result corresponding to the first timestamp to obtain the multi-modal perception result of the road surface anomaly.
[0033] In the embodiment of the present invention, the visual features and tactile features of the road surface anomaly in the sensing element are extracted through a neural network and a classification algorithm. Subsequently, the identity information of each road surface anomaly that appears is marked with timestamps of vision and touch under the PV perspective. Then, the anomalies in the visual and tactile information are matched and corrected according to the time difference. According to the correction result, the corresponding visual and tactile information is subjected to decision-level fusion to realize the output of the visual-tactile modal fusion information of the road surface anomaly. This method uses visual detection to give early warnings for road surface anomalies at a long distance, and uses the information output method of visual-tactile fusion to detect road surface anomalies at a short distance, helping the driver to evaluate the road surface risk situation from more dimensions; at the same time, during the process of visual-tactile information fusion, the possibility of incorrect matching due to time error is minimized, making the output result have high reliability and accuracy.
[0034] The following describes Figure 1 the execution manner of each step shown.
[0035] First, for step 100, input the real-time visual data and real-time tactile data obtained by the sensor into the trained anomaly detection model respectively, and sequentially output the visual detection result and the tactile detection result of the road surface to be detected.
[0036] In the embodiment of the present invention, both the real-time visual data and the real-time tactile data are collected by in-vehicle sensing sensors and additional modules. The sensing sensor part includes a camera, a wheel center accelerometer, a suspension accelerometer, and a tire PVDF piezoelectric sensor.
[0037] The camera is placed in the middle position of the front of the vehicle to ensure that the camera's field of view accurately reflects the driver's perspective and covers the entire road area, so as to capture the road surface image in real time and achieve fine capture of road surface information; the wheel center accelerometer is used to obtain the vertical acceleration of the wheel center; the PVDF sensor embedded in the tire is tightly attached to the inside of the tire with a flexible material, so it can be used to obtain the strain information of the inner wall of the tire and can detect different road surface anomalies; the suspension accelerometer detects the vibration generated by the abnormal contact between the vehicle body and the road surface during driving in real time and provides three-axis acceleration information.
[0038] During driving, the camera is responsible for collecting the images of the road surface ahead. When the PVDF sensor attached to the inner wall of the tire touches an abnormal road surface, it generates an electrical signal. At the same time, the embedded acceleration sensor and the suspension accelerometer sensor respond to the current road condition in real time.
[0039] The additional module includes a data acquisition module, a data transmission module, a data storage module, and a power supply module. The data acquisition module performs high-frequency sampling on the PVDF piezoelectric sensor and the acceleration sensor installed in the tire to generate a real-time one-dimensional tactile information group, and transmits the tire information to the storage unit through the data transmission unit. The power supply module supplies power to the modules that require power.
[0040] Furthermore, the collected visual and tactile data is detected using the preset system software to obtain the visual detection result and the tactile detection result of the road surface to be detected. The system software includes a visual perception system, a tactile perception system, and an in-vehicle computer system.
[0041] In terms of visual detection, the embodiment of the present invention realizes the real-time detection and tracking of common road surface abnormalities such as speed bumps, rails, square manholes, and manhole covers by training a visual detection model. Specifically, the trained visual detection model is used to detect and track the road surface abnormalities captured in the picture in real time, determine that the lower third of the entire picture is the ROI area (PV perspective), and record and save the relevant information of visual detection in real time, including: the type, timestamp, ID number, detection confidence of the road surface abnormality entering the ROI area, and the detection result picture.
[0042] In terms of tactile detection, the tactile perception system in the embodiment of the present invention will record the timestamp, electrical signal, and acceleration waveform of the tactile excitation received by the tactile sensor in the tire, and then preprocess the tactile data collected by the sensor to remove the noise points in the original data, so as to perform feature screening and extraction on the signals collected by the PVDF sensor and the accelerometer in the tire, and then detect the abnormalities existing in the tactile time series data, set the abnormality threshold, and output the determination of whether there is an abnormality and the result of the abnormality degree according to the algorithm.
[0043] The in-vehicle computer, on the one hand, receives the road surface image data captured by the camera and uses the pre-trained visual detection model to detect and track road surface abnormalities. On the other hand, it receives the vehicle tactile data group, performs data analysis feature extraction, and inputs it into the pre-written and trained tactile detection model.
[0044] It should be noted that in the above process, the object detection and tracking model (such as YOLO, Fast R-CNN, etc.) and the time series data anomaly detection and classification algorithm (such as recurrent neural network (RNN), long short-term memory network (LSTM), variational autoencoder (VAE), etc.) can be used as the anomaly detection model.
[0045] Then, for step 102, according to the first timestamp of the abnormal excitation in the tactile detection result, match the second timestamp corresponding to the same road surface anomaly in the visual detection result.
[0046] In the embodiment of the present invention, the determination process of the second timestamp includes: correcting the error of the original timestamp recorded by the tactile sensor for the abnormal excitation to obtain the first timestamp; wherein, the error includes systematic error, random error and drift error; calculating the time difference between the tactile detection result and the visual detection result corresponding to the road surface anomaly according to the distance traveled by the vehicle and the vehicle speed from the disappearance of the road surface anomaly from the picture of the visual sensor to the contact with the tire; calculating the difference between the first timestamp and the time difference to obtain the second timestamp.
[0047] Specifically, during the process of the vehicle identifying the road surface anomaly, the road surface anomaly contacted by the tire corresponds to the anomaly detected by the visual detector. When the tactile sensor generates a response at a certain moment t, the visual sensor detected visual information about the anomaly Δt time periods ago. Therefore, to achieve a comprehensive identification of the characteristics of a certain road surface anomaly, it is necessary to inversely deduce the historical image information capturing the anomaly characteristics based on the response time of the tactile sensor, and perform decision-level fusion on the two-modal information, so as to achieve accurate perception of the road surface anomaly.
[0048] For example, let the time of the i-th acquisition of the tactile sensor be That is, record the original timestamp of a certain abnormal excitation; Δt is the time difference between the time when the tactile sensor detects the abnormal information of the same road surface anomaly and the time when the visual sensor detects the abnormal information.
[0049] For the original timestamp Since the tactile sensor collects information at a high frequency, there is a certain error in the timestamp of the tactile excitation received by the road surface anomaly. After analysis, it can be obtained that there are the following three sources: systematic error ∈ sys , random error ∈ rand and drift error ∈ drift , assuming that the actual time of the i-th acquisition of the tactile sensor is T t (i), then the error source at each acquisition can be expressed as:
[0050]
[0051] The total error is:
[0052] ∈total = ∈ sys + ∈ rand (i) + ∈ drift
[0053] It can be seen that the error sources respectively include three parts:
[0054] 1) Systematic Error, the fixed error generated by the tactile collector, denoted as
[0055] 2) Random Error, the sampling frequency fluctuation caused by the noise of the tactile sensor itself or external interference, which is a random error. Assume it is Gaussian white noise with a mean of 0 and a variance of σ 2 normal distribution, denoted as Then there is
[0056] 3) Drift Error, the linear or non - linear time shift that changes with time, denoted as Denote the linear drift error as α is the drift coefficient. Denote the non - linear drift error as f represents the fitted non - linear drift function.
[0057] Denote the first timestamp of the calibrated tactile sensor as T corrected (i), then there is:
[0058]
[0059] It should be noted that the first timestamp T corrected (i) and the actual timestamp T t (i) are theoretically the same value. However, due to the randomness of the error, in the actual error correction process, it is often necessary to estimate the above - mentioned errors. That is to say, what is actually corrected for the timestamp is the estimated value of the error rather than the actual value of the error. The difference between the two further leads to a slight difference between the first timestamp obtained after correction and the actual timestamp. However, the impact of this error on the time matching result can be ignored. Therefore, the corrected timestamp can be directly used to continue the subsequent steps. This process is understandable to those skilled in the art and will not be elaborated here.
[0060] For the time difference Δt, it can be calculated from the distance d traveled by the vehicle and the vehicle speed v between the disappearance of the road surface anomaly from the visual sensor's picture and the contact with the tire:
[0061]
[0062] As Figure 2 shown, the driving distance d is the sum of the horizontal distance l between the position on the road surface corresponding to the lowermost end of the camera-captured image and the camera position, and the horizontal distance l0 between the camera and the tire.
[0063] Denote the second timestamp output by the visual sensor to be matched as T v (i), then the first time T corrected (i) of the calibrated tactile sensor and the second timestamp T v (i) of the visual sensor have the following relationship:
[0064] T v (i) = T corrected (i) - Δt
[0065] Calculate the corresponding second timestamp through the above formula, and thus perform the subsequent image retrieval process.
[0066] For step 104, retrieve the historical image corresponding to the road surface anomaly in the visual data according to the second timestamp, and perform decision-level fusion on the retrieval result and the tactile detection result corresponding to the first timestamp to obtain the multi-modal perception result of the road surface anomaly.
[0067] In the embodiment of the present invention, the retrieval process includes calculating the number of consecutive image frames saved between the starting sampling moment of the road surface anomaly and the second timestamp according to the sampling frequency of the visual sensor; selecting no more than two images from the consecutive images as the historical images corresponding to the road surface anomaly with the median frame as the standard.
[0068] Specifically, to retrieve the corresponding visual image frames and the corresponding output information according to T v (i), denote N as the number of image frames saved by the camera at the moment of T v (i) during the detection process, T v (0) as the starting acquisition moment of the camera, and f as the acquisition frequency of the camera, then there is:
[0069] N = (T v (i) - T v (0)) · f
[0070] Considering that the random error ∈ rand (i) and the drift error ∈ drift (i) are not fixed, introduce the error value range Δ∈total, and at this time the visual retrieval timestamp becomes ΔT v (i), and the corresponding number of frames becomes:
[0071] ΔN = (ΔT v (i) - T v (0)) · f
[0072] In theory, the second timestamp T v (i) is an accurate time. For example, if the starting sampling moment is 1 s and the second timestamp is 3 s, then the number of image frames N is the images between the 1st s and the 3rd s. However, due to the inevitable random errors in the actual sampling process, in order to improve the accuracy of the subsequent perception process, a range of values for the random error Δ ∈ 总 is introduced. For example, if the error is 0.1 s, then at this time ΔT v (i) represents 2.9 s - 3.1 s, and the corresponding number of image frames ΔN obtained is the images between 1 s and 2.9 s or the images between 1 s and 3.1 s. Based on this, subsequent image screening is carried out.
[0073] After determining the range of the number of image frames, select the median frame and retain 1 - 2 pieces of image data as the historical images corresponding to the road surface anomaly. For example, if there are 9 images between 1 s and 2.9 s, then select the fifth image as the corresponding image; if there are 10 images between 1 s and 3.1 s, then select the fifth and sixth images as the corresponding images.
[0074] In the embodiment of the present invention, the process of fusing the retrieved historical images with the tactile information includes the following steps: determining the first weight coefficient of the tactile information and the second weight coefficient of the visual information according to the relative importance of the tactile information and the visual information in the anomaly perception process; fusing the text information of the road surface anomaly in the visual detection result and the tactile detection result according to the first weight coefficient and the second weight coefficient to obtain a fusion result; presenting the fusion result and the visual image of the road surface anomaly in a visualization interface to obtain the multi-modal perception result.
[0075] Specifically, let E(t) be the anomaly information of the tactile sensor at the first timestamp, and V(t - Δt) be the anomaly information detected by the visual sensor at the second timestamp. By introducing the weight coefficients w1 and w2 of the tactile and visual information in the fusion result, the fusion result R(t) can be obtained:
[0076] R(t) = w1·E(t) + w2·V(t - Δt)
[0077] Among them, w1 and w2 respectively reflect the relative importance of the tactile information and the visual information in the anomaly recognition process.
[0078] Then present this result in a visualization interface to obtain the multi-modal perception result. The visualization interface includes four regions: the road surface anomaly under the PV view captured by the camera in real time, the continuous frame results output by the visual detection model, the road surface anomaly image corresponding to a certain tactile anomaly excitation received by the visual module, and the text information integrated by vision and touch.
[0079] For example,Figure 3 The small box YOLOv8Tracking shown represents the real-time tracking process for road anomalies. On the one hand, it continuously senses and displays the real-time data collected by the camera. On the other hand, it can also detect and track road anomalies. The large box Tracking Visualization is black when it has not been retrieved at the beginning; when the tactile timestamp is detected, the image and text with the detected visual timestamp will be displayed on the large window, as shown in Figure 4 the image shown.
[0080] Furthermore, after presenting the results, it is also necessary to compare the visual text information and the tactile text information in the multi-modal perception results; and determine whether the two text information point to the same road anomaly according to the detection confidence; if so, determine that the perception is correct and save the multi-modal perception results, otherwise determine that the perception is incorrect and delete the multi-modal perception results.
[0081] For example, if the anomaly presented in the visual image is a speed bump, and the content described by the tactile information looks more like a stone, then it can be considered that the perception result this time is inaccurate. Through this method, the reliability of the fusion result can be further improved.
[0082] In summary, the above road perception method combining visual detection and tactile detection first uses vision to detect road anomalies far from the vehicle, providing an early warning for vehicle driving safety and helping the driver make an early decision on whether to actively avoid; for road anomalies close to the vehicle, it uses the information output method of visual-tactile fusion to provide and supplement more tactile characteristic information for the visual detection result. This information output based on multi-sensors covers more road anomaly features, helping the driver evaluate the road risk situation from more dimensions.
[0083] Furthermore, taking the closest perception timestamps of the two detection methods after processing and calculation as the matching standard, it maximally reduces the possibility of incorrect matching due to time errors. At the same time, it analyzes the error sources during the matching process and makes relevant optimization measures to minimize the difference between the calculated matching time and the real time, so as to accurately retrieve the relevant information about road anomalies for further decision fusion. The whole process achieves the matching of visual and tactile information while avoiding time errors as much as possible, and the output result has high reliability and accuracy.
[0084] As Figure 5 、 Figure 6 shown, the embodiment of the present invention provides a road anomaly perception device based on visual-tactile fusion. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. In terms of the hardware level, as Figure 5As shown in the figure, it is a hardware architecture diagram of an electronic device where a road surface anomaly perception device based on visual-tactile fusion provided by an embodiment of the present invention. In addition to Figure 5 the shown processor, memory, network interface, and non-volatile memory, the electronic device where the device is located in the embodiment usually may also include other hardware, such as a forwarding chip responsible for processing packets, etc. Taking software implementation as an example, as Figure 6 shown, as a logically meaningful device, it is formed by the CPU of its corresponding electronic device reading the corresponding computer program in the non-volatile memory into the memory and running it. A road surface anomaly perception device based on visual-tactile fusion provided by this embodiment includes:
[0085] A detection module 600, configured to respectively input the acquired real-time visual data and real-time tactile data of the sensor into a trained anomaly detection model, and sequentially output a visual detection result and a tactile detection result of the road surface to be detected;
[0086] A matching module 602, configured to match a second timestamp corresponding to the same road surface anomaly in the visual detection result according to the first timestamp of the anomaly excitation in the tactile detection result;
[0087] A fusion module 604, configured to retrieve a historical image corresponding to the road surface anomaly in the visual data according to the second timestamp, and perform decision-level fusion on the retrieved result and the tactile detection result corresponding to the first timestamp to obtain a multi-modal perception result of the road surface anomaly.
[0088] In the embodiment of the present invention, the real-time visual data is obtained by a vehicle-mounted visual sensor, the real-time tactile data is collected by a tactile sensor arranged inside the tire, and the visual detection result includes the type, timestamp, ID number, and detection confidence of the road surface anomaly in the visual sensor target area presented on the road surface image.
[0089] In the embodiment of the present invention, when the matching module 602 executes matching the second timestamp corresponding to the same road surface anomaly in the visual detection result according to the first timestamp of the anomaly excitation in the tactile detection result, it is specifically configured to perform the following operations: perform error correction on the original timestamp recorded by the tactile sensor for anomaly excitation to obtain the first timestamp; where the error includes systematic error, random error, and drift error; calculate the time difference between the tactile detection result and the visual detection result corresponding to the road surface anomaly according to the distance traveled by the vehicle and the vehicle speed between the disappearance of the road surface anomaly from the visual sensor's picture and the contact with the tire; calculate the difference between the first timestamp and the time difference to obtain the second timestamp.
[0090] In an embodiment of the present invention, when the fusion module 604 retrieves the historical image corresponding to the road surface anomaly in the visual data according to the second timestamp, it is specifically configured to perform the following operations: calculate the number of consecutive image frames saved between the starting sampling moment of the road surface anomaly and the second timestamp according to the sampling frequency of the visual sensor; select no more than two images from the consecutive images with the median frame as the standard as the historical images corresponding to the road surface anomaly.
[0091] In an embodiment of the present invention, when the fusion module 604 performs decision-level fusion on the retrieval result and the tactile detection result corresponding to the first timestamp to obtain the multimodal perception result of the road surface anomaly, it is specifically configured to perform the following operations: determine a first weight coefficient of the tactile information and a second weight coefficient of the visual information according to the relative importance of the tactile information and the visual information in the anomaly perception process; fuse the text information of the road surface anomaly in the visual detection result and the tactile detection result according to the first weight coefficient and the second weight coefficient to obtain a fusion result; present the fusion result and the visual image of the road surface anomaly in a visualization interface to obtain the multimodal perception result.
[0092] In an embodiment of the present invention, after obtaining the multimodal perception result of the road surface anomaly, it further includes: comparing the visual text information and the tactile text information in the multimodal perception result; determining whether the two text information points to the same road surface anomaly according to the detection confidence; if so, determining that the perception is correct and saving the multimodal perception result, otherwise determining that the perception is incorrect and deleting the multimodal perception result.
[0093] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on a road surface anomaly perception device based on visual-tactile fusion. In other embodiments of the present invention, a road surface anomaly perception device based on visual-tactile fusion may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.
[0094] Regarding the information interaction, execution process, etc. between the modules in the above device, since it is based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention and will not be elaborated here.
[0095] An embodiment of the present invention further provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and when the processor executes the computer program, it implements a road surface anomaly perception method based on visual-tactile fusion in any embodiment of the present invention.
[0096] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor is caused to execute a method for pavement anomaly perception based on visual-tactile fusion according to any embodiment of the present invention.
[0097] An embodiment of the present invention further provides a computer program product. The computer program product includes a computer program. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method for pavement anomaly perception based on visual-tactile fusion described in any of the above embodiments.
[0098] For the convenience of description, when describing the above system or device, it is described by function as various modules or units respectively. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.
[0099] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0100] Finally, it should also be noted that in this article, relational terms such as first, second, third, and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such a process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0101] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A method for abnormal pavement perception based on visual-tactile fusion, characterized in that Including: Input the real-time visual data and real-time tactile data obtained by sensors into a trained anomaly detection model respectively, and sequentially output the visual detection result and tactile detection result of the road surface to be detected; According to the first timestamp of the abnormal excitation in the tactile detection result, match the second timestamp corresponding to the same road surface anomaly in the visual detection result; Retrieve the historical image corresponding to the road surface anomaly in the visual data according to the second timestamp, and perform decision-level fusion on the retrieval result and the tactile detection result corresponding to the first timestamp to obtain the multi-modal perception result of the road surface anomaly.
2. The method according to claim 1, wherein The real-time visual data is obtained by a vehicle-mounted visual sensor, and the real-time tactile data is collected by a tactile sensor arranged inside the tire. The visual detection result includes the type, timestamp, ID number, and detection confidence of the road surface anomaly in the visual sensor target area presented on the road surface image.
3. The method according to claim 2, characterized in that, The step of matching the second timestamp corresponding to the same road surface anomaly in the visual detection result according to the first timestamp of the abnormal excitation in the tactile detection result includes: Perform error correction on the original timestamp recorded by the tactile sensor for abnormal excitation to obtain the first timestamp; wherein, the error includes systematic error, random error, and drift error; Calculate the time difference between the tactile detection result and the visual detection result corresponding to the road surface anomaly according to the distance traveled by the vehicle and the vehicle speed between the disappearance of the road surface anomaly from the visual sensor's screen and the contact with the tire; Calculate the difference between the first timestamp and the time difference to obtain the second timestamp.
4. The method according to claim 1, wherein The step of retrieving the historical image corresponding to the road surface anomaly in the visual data according to the second timestamp includes: Calculate the number of consecutive image frames saved from the starting sampling moment of the road surface anomaly to the second timestamp according to the sampling frequency of the visual sensor; Select no more than two images from the consecutive images with the median frame as the standard as the historical images corresponding to the road surface anomaly.
5. The method according to claim 4, characterized in that, The step of performing decision-level fusion on the retrieval result and the tactile detection result corresponding to the first timestamp to obtain the multi-modal perception result of the road surface anomaly includes: Determine the first weight coefficient of the tactile information and the second weight coefficient of the visual information according to the relative importance of the tactile information and the visual information in the anomaly perception process; Fuse the text information of the road surface anomaly in the visual detection result and the tactile detection result according to the first weight coefficient and the second weight coefficient to obtain a fusion result; Present the fusion result and the visual image of the road surface anomaly in a visualization interface to obtain the multi-modal perception result.
6. The method according to claim 2, wherein After obtaining the multi-modal perception result of the road surface anomaly, it further includes: Compare the visual text information and tactile text information in the multi-modal perception result; Determine whether the two text informations point to the same road surface anomaly according to the detection confidence; if so, determine that the perception is correct and save the multi-modal perception result, otherwise determine that the perception is incorrect and delete the multi-modal perception result.
7. A pavement anomaly perception device based on visual-tactile fusion, characterized in that, Including: A detection module, configured to respectively input the real-time visual data and real-time tactile data acquired by a sensor into a trained anomaly detection model, and sequentially output a visual detection result and a tactile detection result of a road surface to be detected; A matching module, configured to match, according to a first timestamp of an abnormal excitation in the tactile detection result, a second timestamp corresponding to the same road surface anomaly in the visual detection result; A fusion module, configured to retrieve a historical image corresponding to the road surface anomaly in the visual data according to the second timestamp, and perform decision-level fusion on the retrieval result and the tactile detection result corresponding to the first timestamp to obtain a multi-modal perception result of the road surface anomaly.
8. An electronic device, characterized in that, It includes a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the method according to any one of claims 1-6 is implemented.
9. A computer-readable storage medium, characterized in that, A computer program is stored thereon. When the computer program is executed on a computer, the computer is made to execute the method according to any one of claims 1-6.
10. A computer program product, characterized in that, It includes a computer program. When the computer program is executed by a processor, the steps of the method according to any one of claims 1-6 are implemented.