Intelligent inspection method and apparatus for unmanned construction machinery, device and storage medium

By combining multimodal data acquisition and cloud-edge collaborative algorithms with machine learning, the inspection path of unmanned construction machinery is adaptively adjusted, solving the problem of insufficient inspection route adjustment in existing technologies and realizing flexible response and efficient inspection in complex environments.

WO2026097305A1PCT designated stage Publication Date: 2026-05-15SHANDONG JIANZHU UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SHANDONG JIANZHU UNIV
Filing Date
2024-11-07
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing unmanned construction machinery inspection methods lack adaptability to complex and ever-changing operating environments, cannot flexibly adjust inspection routes, leading to missed inspections and false inspections, and lack the ability to respond quickly to emergencies, failing to detect potential safety hazards in a timely manner.

Method used

A multimodal data acquisition network is used to acquire environmental data in real time. Cloud-edge collaborative algorithms and machine learning algorithms are used for data processing and anomaly analysis. The inspection path is adaptively adjusted, and intelligent path planning algorithms are used to prioritize the inspection of abnormal areas.

Benefits of technology

This improved the accuracy and reliability of inspections, ensured that key areas were prioritized for inspection, enhanced the targeting and efficiency of inspections, and enabled the timely detection of potential safety hazards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024130420_15052026_PF_FP_ABST
    Figure CN2024130420_15052026_PF_FP_ABST
Patent Text Reader

Abstract

An intelligent inspection method and apparatus for unmanned construction machinery, a device, and a storage medium, which are applied to the technical field of construction machinery. The method comprises: using a multi-modal data acquisition network to acquire multi-modal data in an operating environment in real time (S110); using a cloud-edge collaborative algorithm to process the multi-modal data, and performing anomaly analysis and prediction on the processed multi-modal data on the basis of a machine learning algorithm (S120); and adaptively adjusting an inspection path on the basis of the prediction result, so as to implement intelligent inspection (S130). By means of the method, potential safety hazards can be found in time, and the inspection efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Intelligent inspection methods, devices, equipment and storage media for unmanned construction machinery Technical Field

[0001] This disclosure relates to the field of construction machinery technology, and further to the field of intelligent inspection technology for construction machinery, and particularly to intelligent inspection methods, devices, equipment and storage media for unmanned construction machinery. Background Technology

[0002] Existing inspection methods for unmanned construction machinery mainly rely on preset inspection routes and simple environmental perception technologies. However, relying solely on preset inspection routes lacks adaptability to complex and ever-changing working environments and cannot be flexibly adjusted according to actual conditions. This can easily lead to missed inspections of critical areas or unnecessary inspections of non-critical areas, reducing the efficiency and focus of inspections. On the other hand, simple environmental perception technologies can usually only acquire single types of data and lack a comprehensive understanding of complex environments, resulting in a higher possibility of missed or false inspections. At the same time, existing inspection methods also lack rapid response capabilities when facing emergencies, failing to promptly detect potential safety hazards and take effective countermeasures. Summary of the Invention

[0003] This disclosure provides an intelligent inspection method, device, equipment, and storage medium for unmanned construction machinery.

[0004] According to a first aspect of this disclosure, an intelligent inspection method for unmanned construction machinery is provided; the method includes:

[0005] Utilize a multimodal data acquisition network to acquire multimodal data from the work environment in real time;

[0006] The cloud-edge collaborative algorithm is used to process multimodal data, and anomaly analysis and prediction are performed on the processed multimodal data based on machine learning algorithms.

[0007] The inspection path is adaptively adjusted based on the prediction results to achieve intelligent inspection.

[0008] Among the possible implementations of the first aspect, a multimodal data acquisition network is used to acquire multimodal data from the work environment in real time, including:

[0009] The multimodal data acquisition network includes: visual sensors, radar sensors, acoustic sensors, and environmental sensors;

[0010] By utilizing visual sensors, radar sensors, acoustic sensors, and environmental sensors, data from visual sensors, radar sensors, acoustic sensors, and environmental sensors with the same timestamp in the working environment can be acquired in real time to ensure that the data from different sensors can be aligned in real time.

[0011] In some possible implementations of the first aspect, cloud-edge collaborative algorithms are used to process multimodal data, and anomaly analysis and prediction are performed on the processed multimodal data based on machine learning algorithms, including:

[0012] By using cloud-edge collaborative algorithms, multimodal data collected from various sensors is preprocessed at the edge and sent to the cloud. The cloud then performs data fusion processing on the preprocessed data and builds anomaly analysis and prediction models based on machine learning algorithms. These models are then used to perform anomaly analysis and prediction on the processed multimodal data.

[0013] In some possible implementations of the first aspect, the anomaly analysis and prediction model is generated in the following way:

[0014] Using historical multimodal data at the same timestamp as training samples, and using the next set of historical multimodal data at the same timestamp and the corresponding state information as sample labels, a training set is generated to train and update the model until an anomaly analysis and prediction model is generated.

[0015] Among the possible implementations of the first aspect, intelligent inspection is achieved by adaptively adjusting the inspection path based on the prediction results, including:

[0016] Based on the prediction results, areas that may have abnormal situations are marked, and the inspection area is automatically prioritized.

[0017] Based on the intelligent path planning algorithm, the inspection path is adaptively adjusted according to the priority of the inspection area and the location of the inspection equipment to achieve intelligent inspection.

[0018] According to a second aspect of this disclosure, an intelligent inspection device for unmanned construction machinery is provided. The device includes:

[0019] The first processing module is used to acquire multimodal data in the working environment in real time using a multimodal data acquisition network;

[0020] The second processing module is used to process multimodal data using cloud-edge collaborative algorithms, and to perform anomaly analysis and prediction on the processed multimodal data based on machine learning algorithms.

[0021] The third processing module is used to adaptively adjust the inspection path based on the prediction results in order to achieve intelligent inspection.

[0022] According to a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods described above.

[0023] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods described above.

[0024] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described above.

[0025] According to a sixth aspect of this disclosure, an intelligent inspection vehicle for unmanned construction machinery is provided, wherein the intelligent inspection vehicle applies the method described above during operation.

[0026] In this disclosure, by acquiring multimodal data from the working environment in real time, a more comprehensive understanding of the environmental conditions is achieved, reducing the possibility of missed or false detections. By utilizing cloud-edge collaborative algorithms, the amount of data processing is distributed, reducing the system's computational burden and improving the timeliness of data processing. Anomaly analysis and prediction of multimodal data based on machine learning algorithms can promptly identify potential safety hazards, improving the accuracy and reliability of inspections. Adaptively adjusting the inspection path based on the prediction results can ensure that key areas are prioritized for inspection, improving the targeting and effectiveness of inspections.

[0027] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0028] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the scope of this disclosure. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0029] Figure 1 shows a flowchart of an intelligent inspection method for unmanned construction machinery provided in an embodiment of this disclosure;

[0030] Figure 2 shows a block diagram of an intelligent inspection device for unmanned construction machinery according to an embodiment of the present disclosure;

[0031] Figure 3 shows a block diagram of an exemplary electronic device capable of implementing embodiments of the present disclosure. Embodiments of the present invention

[0032] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0033] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0034] In response to the problems mentioned in the background art, this disclosure provides an intelligent inspection method, device, equipment, and storage medium for unmanned construction machinery.

[0035] Specifically, a multimodal data acquisition network is used to acquire multimodal data in the working environment in real time; cloud-edge collaborative algorithms are used to process the multimodal data, and anomaly analysis and prediction are performed on the processed multimodal data based on machine learning algorithms; the inspection path is adaptively adjusted according to the prediction results to achieve intelligent inspection.

[0036] This approach allows for the timely detection of potential safety hazards, improving the accuracy and reliability of inspections; ensuring that key areas receive priority inspections, thus enhancing the targeting and effectiveness of inspections; and increasing inspection efficiency.

[0037] The following detailed description, with reference to the accompanying drawings and specific embodiments, illustrates an intelligent inspection method, apparatus, equipment, and storage medium for unmanned construction machinery provided in this disclosure.

[0038] Figure 1 shows a flowchart of an intelligent inspection method for unmanned construction machinery provided in an embodiment of this disclosure; as shown in Figure 1, the intelligent inspection method 100 for unmanned construction machinery may include the following steps:

[0039] The S110 utilizes a multimodal data acquisition network to acquire multimodal data from the operating environment in real time.

[0040] The multimodal data acquisition network includes: visual sensors, radar sensors, acoustic sensors, and environmental sensors.

[0041] Specifically, visual sensors, radar sensors, acoustic sensors, and environmental sensors are used to acquire data from these sensors in real time within the working environment, ensuring real-time data alignment across different sensors. Among these:

[0042] Visual sensor data includes, but is not limited to: images or video frames containing visual information about the work environment captured by monitoring cameras in the work site, and images or video frames containing visual information about the work environment captured by vision equipment of unmanned construction machinery.

[0043] Radar sensor data includes, but is not limited to: radar measurement data that provides information on the position and dynamics of objects in the environment, such as distance, speed, and direction.

[0044] Acoustic sensor data includes, but is not limited to, audio data that reflects the acoustic characteristics of the environment, such as sound intensity and frequency.

[0045] Environmental sensor data includes, but is not limited to, environmental parameters that reflect the physical conditions of the working environment, such as temperature, humidity, and light intensity.

[0046] It should be noted that each sensor's data contains corresponding identification information, which can be a specific code, tag, or metadata field. During the data acquisition phase, each sensor's data is assigned a unique identifier to ensure accurate tracking and identification throughout the entire data processing flow. This identification information allows for quick sensor location, enabling the accurate pinpointing of the source of abnormal values ​​when the model predicts them. In this way, even in complex data fusion environments, the source of anomalies can be quickly located, providing crucial clues and evidence for subsequent inspections and troubleshooting. Furthermore, the presence of identification information makes the management and maintenance of different sensor data more convenient and efficient.

[0047] Real-time acquisition of data from visual sensors, radar sensors, acoustic sensors, and environmental sensors with the same timestamps in the working environment can be achieved using high-precision clock sources such as GPS to provide a unified reference time for all sensors. Specifically, millisecond-level clock synchronization between sensors is achieved through the PPS signal and NMEA command from the GPS receiver. The PPS signal is used to reset the millisecond information to zero, and the NMEA command is used to adjust the hour, minute, and second information. Alternatively, this can also be achieved through the following methods:

[0048] If the sensor and computing platform support the IEEE 1588 or IEEE 802.1AS protocol, submicrosecond / nanosecond level clock synchronization can be achieved through PTP (Precision Time Protocol) and gPTP (Generalized Precision Time Protocol).

[0049] Messages can also be provided through the ROS (Robot Operating System) platform. filters The packet synchronizes messages from multiple sensors; it synchronizes messages from different sensors by creating a time synchronizer to ensure that their timestamps are close.

[0050] For sensors that do not support hardware synchronization, a timestamp alignment algorithm can be used to approximate synchronization. This algorithm finds the nearest neighbor frame by searching for adjacent timestamps and corrects the timestamps based on the movement of the sensor or obstacle. This approach is suitable for scenarios where the sensor acquisition frequency is inconsistent or there is a data transmission delay.

[0051] It is understandable that real-time acquisition of data from different sensors with the same timestamp in the working environment ensures that the data obtained from visual sensors, radar sensors, acoustic sensors, and environmental sensors are aligned in time, that is, each data point contains the same timestamp. This enables the model to better understand the correlation and synchronization between data from different sensors, providing a better data foundation for model training.

[0052] According to embodiments of this disclosure, data alignment with the same timestamp ensures the consistency of data from each sensor, providing a reliable foundation for subsequent data processing and anomaly analysis. At the same time, the comprehensive acquisition of multimodal data can provide more comprehensive environmental information, helping to discover anomalies that may be missed by a single sensor.

[0053] The S120 uses a cloud-edge collaborative algorithm to process multimodal data and performs anomaly analysis and prediction on the processed multimodal data based on machine learning algorithms.

[0054] Specifically, using a cloud-edge collaborative algorithm, multimodal data collected from various sensors is preprocessed at the edge and sent to the cloud. The cloud then performs data fusion processing on the preprocessed data and builds anomaly analysis and prediction models based on machine learning algorithms. These models are then used to perform anomaly analysis and prediction on the processed multimodal data. More specifically:

[0055] First, at the edge computing nodes, the multimodal data collected in real time from various sensors is cleaned by noise reduction, interpolation, and other data cleaning processes; the cleaned data is then compressed to reduce the amount of data and improve transmission efficiency; finally, the compressed data is converted to a unified data format.

[0056] Furthermore, the pre-processed multimodal data is transmitted to the cloud data center. During the transmission process, the data can also be encrypted to ensure data security.

[0057] After receiving data from the edge, the cloud data center first performs data integrity verification to ensure that all data has been received correctly and is not damaged. Then, it uses data fusion techniques such as Kalman filtering or Bayesian networks to extract the correlation and complementarity between the various modal data, integrates the multimodal data from different sensors, and generates a comprehensive and accurate multimodal dataset.

[0058] It should be noted that, in order to quickly locate the abnormal region after the model predicts the abnormal values, the identification information in the data from each sensor will be retained during data fusion.

[0059] Based on the fused multimodal dataset, the cloud data center uses machine learning algorithms to build anomaly analysis and prediction models. Considering the strong temporal characteristics of the dataset, the model can be trained using recurrent neural networks (RNNs) or their variants, such as long short-term memory networks (LSTM) and gated recurrent units (GRUs). These can capture the temporal and long-term dependencies in the data. RNNs, through their internal recurrent structure, can maintain the memory of previous input information, thus performing well when processing time series data. LSTM and GRU, as improved versions of RNNs, further solve the gradient vanishing or gradient exploding problems that RNNs may encounter when processing long sequences, and can more effectively handle data with long time spans. The model building process can include steps such as feature selection, model training, parameter tuning, and cross-validation to ensure the accuracy and generalization ability of the model.

[0060] According to embodiments of this disclosure, by utilizing a cloud-edge collaborative algorithm, preprocessing is distributed to the edge, reducing the amount of data processing in the cloud, alleviating the system's computational burden, and improving the timeliness of data processing; at the same time, the cloud-edge collaborative processing method can provide a fast response and improve the system's real-time performance.

[0061] Furthermore, anomaly analysis and prediction models can be generated in the following ways:

[0062] Using historical multimodal data at the same timestamp as training samples, and using the next set of historical multimodal data at the same timestamp and the corresponding state information as sample labels, a training set is generated to train and update the model until an anomaly analysis and prediction model is generated.

[0063] Furthermore, the machine learning model can be continuously iterated and optimized based on feedback from practical applications and new data inputs to improve its accuracy and adaptability.

[0064] According to embodiments of this disclosure, anomaly analysis and prediction of multimodal data based on machine learning algorithms can promptly detect potential safety hazards and improve the accuracy and reliability of inspections.

[0065] S130 adaptively adjusts the inspection path based on the prediction results to achieve intelligent inspection.

[0066] Specifically, based on the prediction results, areas where anomalies may exist are marked, and inspection area priorities are automatically assigned. Based on an intelligent path planning algorithm, the inspection path is adaptively adjusted according to the inspection area priorities and the location of the inspection equipment to achieve intelligent inspection. Among these:

[0067] Based on the prediction results, the areas marked as potentially showing anomalies include:

[0068] If the model predicts that a certain sensor data in the sensor data set may be abnormal at the next moment, then based on the prediction result, the identification information of the sensor data is read, the sensor is quickly located, and the area where the sensor is located is marked as an area where abnormality may exist.

[0069] Automatically prioritizing inspection areas can be based on factors such as the severity and urgency of anomalies in the prediction results, and prioritize inspection areas, classifying areas with a high probability of anomalies and a large impact range as high-priority areas.

[0070] Intelligent path planning algorithms include, but are not limited to, A* algorithm, genetic algorithm, or other planning algorithms that can achieve the same effect.

[0071] Based on the priority of the inspection area and the location of the inspection equipment, the inspection path is adaptively adjusted using an intelligent path planning algorithm. This includes taking into account factors such as the current location of the inspection equipment, the priority of the inspection area, the length of the inspection path, and time cost to generate the optimal inspection path.

[0072] Furthermore, the generated inspection path is distributed to the inspection equipment to guide it to carry out inspections according to the planned path. During the inspection process, the status and location of the inspection equipment and any abnormalities in the inspection area can be monitored in real time. If new abnormalities occur or the inspection equipment malfunctions, the inspection path can be adjusted in a timely manner. At the same time, the inspection path is dynamically adjusted based on real-time feedback and new prediction results during the inspection process to ensure that the inspection path meets the current actual needs.

[0073] Furthermore, data and information collected during the inspection process can be used to iteratively optimize the anomaly analysis and prediction model, thereby improving the model's accuracy and adaptability.

[0074] According to embodiments of this disclosure, prioritizing and focusing on inspecting abnormal areas helps to promptly identify and address potential problems, improves inspection effectiveness, and ensures that limited inspection resources are used more effectively.

[0075] The following specific embodiment is provided for further explanation to make the above content easier to understand.

[0076] Suppose that unmanned construction machinery is operating in a 1500×1500 meter construction site. The site is equipped with a multimodal data acquisition network. The site is divided into emergency areas S1, S2, and S3, and non-emergency areas S4, S5, and S6 according to the actual operational hazards. Assume that the initial position of the inspection equipment is (30, 30).

[0077] The S110 utilizes a multimodal data acquisition network to acquire multimodal data from the operating environment in real time.

[0078] The sensor configuration in the multimodal data acquisition network is as follows:

[0079] Visual sensors: n high-definition monitoring cameras are installed around the work site, and each unmanned construction machine is also equipped with binocular vision devices for near-range environmental perception.

[0080] Radar sensor: Employing millimeter-wave radar, it can measure the distance, speed, and direction of an object.

[0081] Acoustic sensor: Using a high-sensitivity microphone, it can detect sound frequency and measure sound intensity.

[0082] Environmental sensors include temperature sensors, humidity sensors, and light intensity sensors.

[0083] Further settings for timestamp synchronization:

[0084] A GPS receiver provides a uniform reference time for all sensors. The PPS signal of the GPS receiver resets the millisecond information to zero and generates a precise pulse signal every second. At the same time, the hour, minute and second information is adjusted through NMEA instructions to ensure that the time accuracy of all sensors is at the millisecond level. For example, assuming the current time is 12:34:56.789, the PPS signal resets the millisecond portion 789 to zero, and the NMEA instructions ensure the accuracy of the hour, minute and second.

[0085] The S120 uses a cloud-edge collaborative algorithm to process multimodal data and performs anomaly analysis and prediction on the processed multimodal data based on machine learning algorithms.

[0086] First, the edge processing unit preprocesses the multimodal data, which involves data cleaning, compression, and format conversion. The specific methods for data cleaning, compression, and format conversion can be chosen appropriately based on the data characteristics. Specifically:

[0087] Data cleaning primarily involves denoising multimodal data. For visual sensor data, median filtering is used to remove salt-and-pepper noise from images. For radar sensor data, mean filtering is used to remove random noise from distance measurements. For acoustic sensors, adaptive filtering is used to remove background noise from audio, dynamically adjusting filter parameters based on the spectral characteristics of the audio signal and the statistical properties of the noise to better suppress noise. For environmental sensor data, linear interpolation is used to remove outliers and smooth the data. Taking temperature sensor data as an example, if the measured value at a certain moment deviates significantly from the measured values ​​in the preceding and following time periods, linear interpolation can be used to estimate the reasonable temperature value at that moment using the temperature values ​​at adjacent time points. Next, interpolation processing can be applied to other denoised sensor data besides environmental sensor data to make the data more continuous and smooth, reducing data discontinuities and abrupt changes. For example:

[0088] Assuming the image size is 512×512 and the median filter window size is 3×3, for each pixel, the median of the gray values ​​of its surrounding 9 pixels is taken as the new gray value of that pixel.

[0089] Assuming the radar measurement data is a set of distance values ​​[10.5, 10.3, 10.7, 10.6, 10.4], after mean filtering, the new distance value is the average value of this set of data, 10.5.

[0090] Assuming the audio signal has a sampling frequency of 44.1kHz, an adaptive filtering algorithm can effectively remove low-frequency noise and burst noise from the environment, thereby improving the quality of the audio signal.

[0091] Assuming the temperature sensor measures [25.5℃, 26.0℃, 30.0℃, 25.8℃, 26.2℃] over a period of time, where 30.0℃ is clearly an outlier, we can estimate the reasonable temperature value corresponding to this outlier as 25.9℃ by using linear interpolation based on the temperature values ​​of 26.0℃ and 25.8℃ at the two consecutive time points.

[0092] Data compression primarily employs lossless compression algorithms to compress the cleaned data; for example:

[0093] For visual sensor data, the JPEG-LS compression algorithm can be used, with a compression ratio of 2:1. If the size of the cleaned image data is 1MB, it will become 500KB after compression.

[0094] For radar sensor data, Huffman coding is used for compression. Assuming the cleaned data is a set of binary codes [10101010, 11001100, 10011001], the frequency of each symbol (i.e., different binary values) in the given binary code is counted. In this example, 10101010 appears once, 11001100 appears once, and 10011001 appears once. Each symbol and its frequency are placed as a leaf node into a priority queue, sorted by frequency from smallest to largest. The two nodes with the lowest frequencies are taken from the priority queue, and a new parent node is created with a frequency equal to the sum of the frequencies of its two child nodes. These two child nodes are used as the left and right subtrees of the parent node, and the new parent node is placed into the priority queue. This process is repeated until only one node remains in the priority queue; this node is the root node of the Huffman tree. Since each symbol in this example appears... If the frequencies are the same, two symbols can be randomly selected to construct a parent node. For example, 10101010 and 11001100 are selected to create a parent node with a frequency of 2. Then, the new parent node and 10011001 are put into a priority queue. The two nodes with the lowest frequencies are taken out again to create a higher-level parent node with a frequency of 3. Thus, the Huffman tree is constructed. Next, a Huffman code is assigned to each symbol. Starting from the root node, the code is 0 when moving to the left subtree and 1 when moving to the right subtree. When the leaf node is reached, the Huffman code of the symbol is obtained. For example, the final codes are: 10101010 is coded as 00, 11001100 is coded as 01, and 10011001 is coded as 10. The data before compression is [10101010, 11001100, 10011001], and after compression it becomes [00, 01, 10]. The data length is shortened from 24 bits to 6 bits, which is about half of the original.

[0095] For acoustic sensor data, the lossless audio compression algorithm FLAC (Free Lossless Audio Codec) can be used for compression.

[0096] For environmental sensor data, run-length encoding can be used for compression. Assuming the temperature sensor data over a period of time is:

[0097] [25℃, 25℃, 25℃, 26℃, 26℃, 26℃, 27℃];

[0098] Using run-length encoding, it can be represented as "25℃×3, 26℃×3, 27℃×1", thereby reducing data storage space.

[0099] Data format conversion can unify the data formats of different sensors into JSON format. For example, visual sensor data, originally in image file format, can be converted into a JSON object containing information such as image width, height, and pixel values. Radar sensor data, originally in binary format, can be converted into a JSON object containing information such as distance, speed, and direction. Acoustic sensor data, originally in audio file format, can be converted into a JSON object containing information such as sound intensity, frequency, and duration. Environmental sensor data, originally in a specific sensor data format, can be converted into a JSON object containing information such as temperature, humidity, and light intensity values.

[0100] The above describes the preprocessing that the edge can perform on multimodal data. After processing, the data is sent to the cloud for further processing. Meanwhile, during the process of transmitting the preprocessed data from the edge to the cloud data center, the AES-256 encryption algorithm can be used to encrypt the data.

[0101] Next, after receiving the data, the cloud data center first performs a data integrity check; for example, by calculating the hash value of the received data and comparing it with the hash value calculated by the sender. If the two hash values ​​are the same, it means that the data is intact.

[0102] Then, the Kalman filter algorithm is used to fuse the multimodal data, extract the correlation and complementarity between the various modal data, and integrate the multimodal data from different sensors to generate a comprehensive and accurate multimodal dataset.

[0103] Assume a visual sensor measures the position of an object as (x1, y1), a radar sensor measures the position of the same object as (x2, y2), an acoustic sensor detects the sound intensity I and the sound frequency characteristics F emitted by the object, and an environmental sensor measures the temperature T, humidity H, and light intensity L around the object. By using a Kalman filter algorithm, combined with the measurement errors and time correlation of each sensor, the true state of the object is estimated.

[0104] Furthermore, the state equation and observation equation of the Kalman filter algorithm can be expressed as:

[0105] State equation: X(k) = A*X(k-1) + B*U(k) + W(k);

[0106] Where X(k) represents the system state at time k, and the system state can be defined as a vector containing information such as object position, sound characteristics, and environmental parameters, for example, X(k)=[x,y,I,F,T,H,L]; A is the state transition matrix, which describes the change law of the system state from one time to the next time; B is the control input matrix, which can be set to a zero matrix if there is no external control input in the embodiments provided in this disclosure; U(k) is the control input, which is usually a zero vector when there is no external control; W(k) is the process noise, which represents the uncertainty of the system state during the evolution process.

[0107] Observation equation: Z(k) = H*X(k) + V(k);

[0108] Where Z(k) represents the observation value at time k, which can be a vector composed of measurement values ​​from various sensors, such as Z(k)=[x1,y1,x2,y2,I,F,T,H,L]. However, some values ​​may not have specific measurement values ​​at certain times, such as when the object does not make a sound, the sound intensity and frequency may be zero or invalid values; H is the observation matrix, which maps the system state to the observation value space; V(k) is the observation noise, which represents the uncertainty in the sensor measurement process.

[0109] In the Kalman filtering process, the system state at the next moment is first predicted based on the state equation. Then, the predicted value is updated based on the observation equation. Through continuous iteration, the estimated value becomes increasingly closer to the true value. For example, at the initial moment, the initial estimate of the system state and the error covariance matrix are determined based on prior knowledge or initial measurements. At each subsequent moment, a prediction is made based on the state equation to obtain the predicted system state and the predicted error covariance matrix. Then, when a new sensor measurement arrives, the Kalman gain is calculated based on the observation equation. The Kalman gain is then used to fuse the predicted and observed values ​​to obtain the updated system state estimate and error covariance matrix.

[0110] In this way, data from visual sensors, radar sensors, acoustic sensors, and environmental sensors are fused, making full use of the advantages of each sensor and making up for the limitations of a single sensor. This generates a comprehensive and accurate multimodal dataset, providing a more reliable foundation for subsequent anomaly analysis and prediction.

[0111] The following example provides a way to integrate multimodal data from different sensors to generate a multimodal dataset, using exemplary data as a basis:

[0112] Assume that at time k, the visual sensor measures the position of the object as (x1, y1) = (5, 3), the radar sensor measures the position of the same object as (x2, y2) = (4.8, 3.2), the acoustic sensor detects the sound intensity emitted by the object as I = 80 dB, the sound frequency characteristic as F = 500 Hz, and the environmental sensor measures the temperature around the object as T = 25℃, the humidity as H = 60%RH, and the light intensity as L = 5000 lux.

[0113] There is no external control input, and the state transition matrix A is the identity matrix, B is the zero matrix, U(k) is the zero vector, and the process noise W(k) is a random vector with zero mean and covariance Q.

[0114] Then we have:

[0115] State vector X(k) = [x, y, I, F, T, H, L];

[0116] The observation vector Z(k) = [x1, y1, x2, y2, I, F, T, H, L] = [5, 3, 4.8, 3.2, 80, 500, 25, 60, 5000];

[0117] The predicted state vector is X P (k) = A*X(k-1), assuming the state at the previous time step was:

[0118] X(1) = [4.2, 3.1, 78, 480, 24.5, 58, 4900]. Here, we assume the previous state is a value close to the initial state to demonstrate the calculation process. In practical applications, iterative updates are required. The predicted state vector X is then... P (2) = [4.2, 3.1, 78, 480, 24.5, 58, 4900] (because A is the identity matrix).

[0119] The prediction error covariance matrix is: P P (k)=A*P(k-1)*A T +Q;

[0120] The observation equation is: Z(k) = H*X(k) + V(k);

[0121] The Kalman gain is: K(k) = P P (k)*H T *(H*P P (k)*H T +R) -1At this point, the new sensor measurement value is Z(2)=[5,3,4.8,3.2,80,500,25,60,5000];

[0122] Then the Kalman gain K(2) = [0.4, 0.4, 0.2, 0.2, 0.1, 0.1, 0.1] (This part is simplified for ease of display, so this value is an approximation).

[0123] The updated state estimate is: X(k) = X P (k)+K(k)*(Z(k)-H*X P (k)); the predicted state vector X P (2) = [4.2, 3.1, 78, 480, 24.5, 58, 4900];

[0124] Calculations show that:

[0125] Z(k)-H*X P (k)=[5,3,4.8,3.2,80,500,25,60,5000]-[4.2,3.1,78,480,24.5,58,4900]=[0.8,-0.1,-73.2,-476.8,0.5,42,0.5,2,100];

[0126] K(k)*(Z(k-H*X) P (k))=[0.32,-0.04,-14.64,-95.36,0.05,4.2,0.05,0.2,10];

[0127] Then we have:

[0128] X(2)=X P (2)+K(k)*(Z(k)-H*X P (k))=

[0129] [4.2,3.1,78,480,24.5,58,4900]+[0.32,-0.04,-14.64,-95.36,0.05,4.2,0.05,0.2,10]=[4.52,3.06,63.36,384.64,24.55,62.2,4900.05,0.2,10].

[0130] After fusion using the Kalman filter algorithm, the following dataset was generated:

[0131] X(2)=[4.52,3.06,63.36,384.64,24.55,62.2,4900.05,0.2,10].

[0132] It should be noted that, due to the complexity and length of the detailed calculation process, the above example is a simplified version of the specific calculation process. However, those skilled in the art should be able to clearly understand this example after seeing the simplified example. At the same time, this example is intended to provide a reference for technicians, and the specific calculation details can be adjusted and optimized according to different actual situations.

[0133] Furthermore, based on the fused multimodal dataset, the cloud data center uses machine learning algorithms to construct anomaly analysis and prediction models, selecting Long Short-Term Memory (LSTM) networks as the framework for anomaly analysis and prediction models. LSTM units consist of forget gates, input gates, and output gates, which can effectively handle long-term dependencies in time series data.

[0134] Among them, the forget gate determines how much information in the hidden state of the previous moment needs to be forgotten, the input gate determines how much information in the current input needs to be saved to the cell state, and the output gate determines how much information in the hidden state of the current moment needs to be output.

[0135] The training set is generated by using historical multimodal data at the same timestamp as training samples and the state information of the next set of historical multimodal data at the same timestamp corresponding to each set of historical multimodal data at the same timestamp as sample labels.

[0136] Assuming the input sequence length is T=100, the input feature dimension at each time step is d=20, and the hidden state dimension of the LSTM unit is h=50, the calculation methods for the forget gate, input gate, and output gate are as follows:

[0137] Forgotten Gate: ;

[0138] in, It is the sigmoid function. It is the weight matrix of the forget gate. It is the bias term of the forget gate, h t-1 It is the hidden state from the previous moment, x t It is the input at the current moment.

[0139] Input Gate: ;

[0140] Among them, W i b i W represents the weight matrix and bias term of the input gate. C b C It is used to calculate the state of candidate cells. The weight matrix and bias term are given, and tanh is the hyperbolic tangent function.

[0141] Output gate: ;

[0142] Among them W o b o These are the weight matrix and bias terms of the output gate. It represents the current state of the cell.

[0143] Assume there are 2000 sets of historical data. Each set contains visual sensor images (image feature vectors extracted after preprocessing, dimension 5), radar sensor measurements (distance, speed, direction, dimension 3), acoustic sensor audio data (sound intensity, frequency, etc., dimension 4), and environmental sensor parameters (temperature, humidity, light intensity, dimension 3). Therefore, the total input feature dimension is d=15. The status information of the next set of data is normal or abnormal (represented by 0 and 1). These data are divided into a training set (1600 sets), a validation set (200 sets), and a test set (200 sets) in an 8:1:1 ratio. The LSTM model is trained using the stochastic gradient descent algorithm with a learning rate of 0.0005 and a batch size of 64. The loss function is defined as the cross-entropy loss function. ;

[0144] Where N is the number of samples, y i It's a real label. These are the labels predicted by the model. After 100 iterations of training, the model's loss function on the validation set is minimized, generating an anomaly analysis and prediction model.

[0145] Based on feedback from practical applications and new data inputs, the machine learning model can be continuously iterated and optimized. For example, new inspection data can be collected every three days and added to the training set to retrain the model. At the same time, early stopping can be used to prevent overfitting. Training is stopped when the model's loss function on the validation set does not decrease for 10 consecutive epochs.

[0146] Furthermore, when new multimodal data is input into the model, it first undergoes a preprocessing step. For example, for visual sensor image feature vectors, the value of each element is mapped to the interval [0, 1]; for radar sensor measurements, the distance, velocity, and direction are standardized so that their mean is 0 and their standard deviation is 1.

[0147] Then, the preprocessed multimodal data is input into the trained LSTM model to obtain the predicted value group output by the model. An anomaly threshold group is set. When there is data in the predicted value group that is greater than the threshold, it is judged that there is an anomaly.

[0148] S130 adaptively adjusts the inspection path based on the prediction results to achieve intelligent inspection.

[0149] The inspection area is divided into emergency area and non-emergency area, corresponding to emergency factors of 0.8 and 0.4 respectively; cases slightly exceeding the threshold are defined as low-level anomalies, cases significantly exceeding the threshold are defined as medium-level anomalies, and cases exceeding the threshold by a large margin are defined as high-level anomalies, corresponding to anomaly impact factors of 0.3, 0.6, and 0.9 respectively; priority values ​​are divided into three intervals: high, medium, and low. Priority values ​​greater than 0.6 are high-priority areas, 0.3 to 0.6 are medium-priority areas, and less than 0.3 are low-priority areas.

[0150] Define the priority calculation formula: P=S×E; where P is the priority value, S is the severity factor of the abnormal situation, and E is the urgency factor.

[0151] The model's output is iterated to obtain data exceeding the threshold. Based on the data identification information, the sensor and the area where the sensor is located are located. The area is marked as an abnormal area and its priority is calculated.

[0152] For example:

[0153] An emergency zone with a high-level anomaly has a priority value of P = 0.9 × 0.8 = 0.72, and is therefore a high-priority zone; a non-emergency zone with a medium-level anomaly has a priority value of P = 0.4 × 0.6 = 0.24, and is therefore a medium-priority zone.

[0154] Next, the A* algorithm is used for intelligent path planning. The A* algorithm selects the optimal path by calculating the sum of the estimated cost and the actual cost of each node. The estimated cost is calculated using Euclidean distance.

[0155] ;

[0156] Among them, (x g y g (x) represents the coordinates of the target node. n y n ) represents the coordinates of the current node.

[0157] The actual cost is the actual distance from the starting point to the current node. Assuming the inspection equipment moves at a speed of 2 meters per second and takes 1 second to move one step, the actual cost is the number of time steps taken.

[0158] Considering the length of the inspection path, time cost, and priority of the inspection area, assuming the weight of path length is 0.4, the weight of time cost is 0.3, and the weight of priority is 0.3, the total cost is calculated as follows: C = w1L + w2T + w3P, where C is the total cost, L is the path length, T is the time cost, P is the priority of the inspection area, and w1, w2, and w3 are the corresponding weights.

[0159] Now assume that a medium-level anomaly is expected to occur in region S1 and a low-level anomaly is expected to occur in region S5; the coordinates of the inspection point in region S1 are (120, 120); the coordinates of the inspection point in region S5 are (1100, 1100); the A* algorithm is used for intelligent path planning, and the evaluation function is set as: f(n) = g(n) + h(n), where g(n) is the actual cost from the initial node to the current node n, and h(n) is the estimated cost from the current node n to the target node.

[0160] First, calculate the priority value for each region:

[0161] The severity factor S of area S1 is 0.6, the urgency factor E is 0.8, and the priority value P1 = 0.6 × 0.8 = 0.48, which belongs to the medium priority area.

[0162] The severity factor S of area S5 is 0.3, the urgency factor E is 0.4, and the priority value P2 = 0.3 × 0.4 = 0.12, which is a low priority area.

[0163] The Euclidean distance from the initial position (30, 30) to the inspection point (120, 120) in area S1 is defined as h1:

[0164] ;

[0165] The Euclidean distance from the initial position (30,30) to the inspection point (1100,1100) in area S5 is defined as h2:

[0166] .

[0167] Assuming the actual cost is 0 at the beginning and gradually increases as the path expands, in each iteration of the A* algorithm, the node with the minimum actual cost value is selected for expansion.

[0168] Since region S1 has a higher priority, region S1 is inspected first. From the initial position (30, 30) to the inspection point (120, 120) in region S1, assuming the actual movement path distance is 110 (the assumption is simplified here for ease of demonstration; in reality, the actual cost needs to be determined based on the specific movement method and terrain factors), then g1=110, f1=g1+h1=110+127.28=237.28.

[0169] After completing the inspection of area S1, proceed from the inspection point (120, 120) in area S1 to the inspection point (1100, 1100) in area S5. At this point, the new h3: h3 = 1385.94.

[0170] Assuming the actual travel distance from the inspection point in area S1 to the inspection point in area S5 is 1000 (simplifying the assumptions), then the actual cost is:

[0171] g2=g1+1000=110+1000=1110, f2=g2+h3=1110+1385.94=2495.94.

[0172] Create two sets: an open set (to store nodes to be inspected) and a closed set (to store nodes that have already been inspected). Add the starting point (30, 30), the initial point of the inspection equipment, to the open set. For each node, record its actual cost g to the starting point, its estimated cost h to the target point, and its parent node (for backtracking). Take the node with the smallest f value from the open set as the current node. If the current node is the target point, backtrack the path and return. Add the current node to the closed set. For each neighboring node of the current node: if the neighboring node is in the closed set, skip it and calculate the actual cost g of the neighboring node.

[0173] If an adjacent node is not in the open set, add it to the open set and record its h value (the distance to the target point calculated using Euclidean distance) and its parent node as the current node. If an adjacent node is already in the open set, compare the new g value with the original g value. If the new g value is smaller, update its g value and its parent node. Repeat this step until the open set is empty. Then, starting from the target point, backtrack to the starting point based on the parent node of each node to obtain the inspection path.

[0174] Following the steps above, assuming the cost of horizontal and vertical movement is 1, and the cost of diagonal movement is... The possible path from the starting point (30, 30) to the inspection point (120, 120) in area S1 is:

[0175] (30, 30)>(31, 31)>(32, 32)>...>(119, 119)>(120, 120).

[0176] After completing the inspection of area S1, the possible path from inspection point (120, 120) in area S1 to inspection point (1100, 1100) in area S5 is as follows:

[0177] (120, 120)>(121, 121)>(122, 122)>...>(1099, 1099)>(1100, 1100).

[0178] In this way, the A* algorithm intelligently plans the inspection path from the initial position to the S1 and S5 regions sequentially.

[0179] During the inspection process, the status and location of the inspection equipment and any abnormalities in each inspection area are monitored in real time. Assuming that the inspection equipment sends location information and equipment status data every 0.5 seconds, the multimodal data acquisition network continuously collects environmental data and transmits it to the cloud for analysis in real time. If new abnormalities occur or the inspection equipment malfunctions, the inspection path can be adjusted in a timely manner.

[0180] If a new high-priority anomaly area is discovered during the inspection process, the priority of the inspection area is updated according to the new anomaly. Then, the total cost of each node is recalculated, and the path is replanned based on the new total cost, prioritizing inspection of the high-priority inspection area.

[0181] As the inspection progresses, the priority of the inspection area is dynamically adjusted according to the actual situation. If a low-priority area does not have any abnormalities for a period of time, its priority can be appropriately reduced. If the abnormality of a medium-priority area continues to worsen, it can be upgraded to a high-priority area. At the same time, the path planning is updated in real time to ensure that the inspection path always meets the current actual needs and abnormalities. Data and information during the inspection process are collected for iterative optimization of the anomaly analysis and prediction model.

[0182] According to the embodiments of this disclosure, the following technical effects are achieved:

[0183] By acquiring multimodal data from the operational environment in real time, we can gain a more comprehensive understanding of the environmental conditions and reduce the possibility of missed or false detections. By utilizing cloud-edge collaborative algorithms, we can distribute the amount of data processing, reduce the system's computational burden, and improve the timeliness of data processing. Based on machine learning algorithms, we can perform anomaly analysis and prediction on multimodal data, which can promptly identify potential safety hazards and improve the accuracy and reliability of inspections. Adaptively adjusting the inspection path according to the prediction results can ensure that key areas are prioritized for inspection, thereby improving the targeting and effectiveness of inspections.

[0184] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this disclosure.

[0185] The above is an introduction to the method embodiments. The following describes the solution described in this disclosure further through device embodiments.

[0186] Figure 2 shows a block diagram of an intelligent inspection device for unmanned construction machinery according to an embodiment of the present disclosure; as shown in Figure 2, the intelligent inspection device 200 for unmanned construction machinery may include:

[0187] The first processing module 210 is used to acquire multimodal data in the working environment in real time using a multimodal data acquisition network.

[0188] The second processing module 220 is used to process multimodal data using cloud-edge collaborative algorithms, and to perform anomaly analysis and prediction on the processed multimodal data based on machine learning algorithms.

[0189] The third processing module 230 is used to adaptively adjust the inspection path based on the prediction results in order to achieve intelligent inspection.

[0190] It is understood that each module in the intelligent inspection device 200 for unmanned construction machinery shown in Figure 2 has the function of implementing each step in the intelligent inspection method 100 for unmanned construction machinery provided in this disclosure embodiment, and can achieve its corresponding technical effect. The specific working process of the described module can be referred to the corresponding process in the foregoing method embodiment. For the sake of convenience and brevity, it will not be repeated here.

[0191] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0192] Figure 3 shows a block diagram of an exemplary electronic device capable of implementing embodiments of the present disclosure.

[0193] Electronic device 300 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 300 may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0194] Electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in ROM 302 or a computer program loaded into RAM 303 from storage unit 308. RAM 303 can also store various programs and data required for the operation of electronic device 300. The computing unit 301, ROM 302, and RAM 303 are interconnected via bus 304. I / O interface 305 is also connected to bus 304.

[0195] Multiple components in electronic device 300 are connected to I / O interface 305, including: input unit 306, such as keyboard, mouse, etc.; output unit 307, such as various types of displays, speakers, etc.; storage unit 308, such as disk, optical disk, etc.; and communication unit 309, such as network card, modem, wireless transceiver, etc. Communication unit 309 allows electronic device 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0196] The computing unit 301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by the computing unit 301, one or more steps of method 100 described above may be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to perform method 100 by any other suitable means (e.g., by means of firmware).

[0197] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0198] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0199] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0200] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including voice input, speech input, or tactile input).

[0201] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0202] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0203] According to embodiments of this disclosure, this disclosure also provides an intelligent inspection vehicle for unmanned construction machinery, wherein the intelligent inspection vehicle can apply the intelligent inspection method 100 for unmanned construction machinery as described above during operation.

[0204] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0205] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure; those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors; any modifications, equivalent substitutions and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An intelligent inspection method for unmanned construction machinery, characterized in that, include: Utilize a multimodal data acquisition network to acquire multimodal data from the work environment in real time; The multimodal data is processed using a cloud-edge collaborative algorithm, and anomaly analysis and prediction are performed on the processed multimodal data based on machine learning algorithms. The inspection path is adaptively adjusted based on the prediction results to achieve intelligent inspection.

2. The method according to claim 1, characterized in that, The method of using a multimodal data acquisition network to acquire multimodal data from the work environment in real time includes: The multimodal data acquisition network includes: visual sensors, radar sensors, acoustic sensors, and environmental sensors; By utilizing the aforementioned visual sensors, radar sensors, acoustic sensors, and environmental sensors, data from visual sensors, radar sensors, acoustic sensors, and environmental sensors with the same timestamp in the working environment can be acquired in real time to ensure that the data from different sensors can be aligned in real time.

3. The method according to claim 1, characterized in that, The process of using a cloud-edge collaborative algorithm to process the multimodal data, and then performing anomaly analysis and prediction on the processed multimodal data based on machine learning algorithms, includes: By utilizing cloud-edge collaborative algorithms, multimodal data collected from various sensors is preprocessed at the edge and sent to the cloud. The cloud then performs data fusion processing on the preprocessed data and constructs anomaly analysis and prediction models based on machine learning algorithms. These anomaly analysis and prediction models are then used to perform anomaly analysis and prediction on the processed multimodal data.

4. The method according to claim 1, characterized in that, The anomaly analysis and prediction model is generated in the following way: Using historical multimodal data at the same timestamp as training samples, and using the next set of historical multimodal data at the same timestamp and the corresponding state information as sample labels, a training set is generated to train and update the model until the anomaly analysis and prediction model is generated.

5. The method according to claim 1, characterized in that, The adaptive adjustment of the inspection path based on the prediction results to achieve intelligent inspection includes: Based on the prediction results, areas that may have abnormal situations are marked, and the inspection area is automatically prioritized. Based on the intelligent path planning algorithm, the inspection path is adaptively adjusted according to the priority of the inspection area and the location of the inspection equipment to achieve intelligent inspection.

6. An intelligent inspection device for unmanned construction machinery, characterized in that, include: The first processing module is used to acquire multimodal data in the working environment in real time using a multimodal data acquisition network; The second processing module is used to process the multimodal data using a cloud-edge collaborative algorithm, and to perform anomaly analysis and prediction on the processed multimodal data based on a machine learning algorithm. The third processing module is used to adaptively adjust the inspection path based on the prediction results in order to achieve intelligent inspection.

7. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in any one of claims 1-5.

8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.

9. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-5.

10. An intelligent inspection vehicle for unmanned construction machinery, characterized in that: The intelligent inspection vehicle uses the method described in any one of claims 1-5.