Multi-modal data compression and real-time inspection method and system
By combining dual-field-of-view imaging components and edge computing modules, panoramic coverage and accurate identification are achieved, solving the problems of insufficient panoramic monitoring and untimely detail identification in existing technologies. This reduces data transmission latency, improves fault identification rate and judgment accuracy, and meets the real-time and accuracy requirements of industrial inspection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 重庆中科汽车软件创新中心
- Filing Date
- 2026-03-26
- Publication Date
- 2026-05-15
AI Technical Summary
Existing industrial inspection technologies suffer from insufficient panoramic monitoring, untimely identification of details, and severe data transmission delays, failing to meet the requirements for real-time performance and accuracy.
By employing a dual-field-of-view imaging component combined with an edge computing module, panoramic coverage without blind spots and accurate identification are achieved. Through ROI data compression and multimodal data fusion, a comprehensive fault index is generated, control commands are generated, and flexible deployment is carried out.
It achieves a combination of panoramic coverage and precise identification, reduces data transmission latency, improves fault identification rate and judgment accuracy, and meets the real-time and accuracy requirements of industrial inspection.
Smart Images

Figure CN122053797A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial inspection technology, specifically to a multimodal data compression and real-time inspection method and system. Background Technology
[0002] In the field of industrial inspection, scenarios such as power equipment, chemical pipelines, and smart manufacturing workshops place extremely high demands on the real-time performance and data transmission efficiency of inspection systems. The operational status of industrial equipment and facilities directly affects production safety and efficiency; timely and accurate detection of potential faults is crucial for avoiding major accidents and reducing economic losses. However, existing inspection technologies have the following limitations.
[0003] Currently, most inspection systems employ single-field-of-view (SFR) inspection technology. This type of technology mainly uses independent wide-field-of-view or narrow-field-of-view cameras for inspection, such as ordinary surveillance cameras (wide-field-of-view) and telephoto inspection cameras (narrow-field-of-view). Its drawbacks are quite obvious. Wide-field-of-view cameras, due to their low detail resolution, cannot meet the requirements for accurate identification of minute faults; narrow-field-of-view cameras have limited coverage, easily creating blind spots and making panoramic monitoring difficult. Furthermore, transmitting full-volume high-definition data leads to severe transmission latency issues. For example, with a single wide-field-of-view 1080P video, the transmission latency can reach over 1.5 seconds, failing to meet the real-time requirements of industrial inspections.
[0004] Although the patent "A Mobile Smart Safety Monitoring Device and Method Based on Multimodal Perception and Edge Computing" (Publication No. CN120434354A) integrates dual cameras and an edge intelligent gateway to achieve multimodal data acquisition and edge analysis, it still has many shortcomings: First, the dual cameras lack clear switching rules, making precise collaboration impossible and resulting in either insufficient panoramic monitoring or untimely detailed inspections during patrols; second, the data compression uses a general algorithm without optimization for fault characteristic areas in the patrol scenario, easily losing key information during compression and affecting the accuracy of fault detection; third, multimodal data fusion only reaches the "data overlay" level, lacking multi-feature-level fusion rules, resulting in low reliability of fault determination and failing to fully leverage the advantages of multimodal data.
[0005] The unique nature of industrial environments results in massive amounts of data. For example, each channel of full HD video data can reach 20-50 Mbps. Directly transmitting this data would quickly cause bandwidth congestion, severely impacting data transmission efficiency and stability. To address this, existing technologies employ data compression, but most use general-purpose algorithms like H.265. These general compression algorithms are not optimized for industrial inspection scenarios. During compression, critical data related to fault characteristics, such as high-temperature points on equipment or sources of abnormal noise, may be over-compressed, leading to the loss of crucial information. Meanwhile, non-fault background data is redundantly transmitted, wasting valuable bandwidth and causing transmission delays exceeding 1 second. In industrial inspection, a 1-second delay can mean the inability to detect faults promptly, missing the optimal time for intervention and posing a significant safety hazard, failing to meet the stringent requirements of real-time inspection.
[0006] To improve the efficiency of inspection systems and reduce the load on cloud platforms, existing technologies have introduced edge computing. For example, a cloud-edge collaborative video surveillance system intelligent monitoring method and device (publication number CN119906821A) utilizes edge computing to perform preliminary data processing and analysis locally, reducing data transmission volume and improving system response speed. However, the synergy between edge computing and dual-field-of-view switching is severely lacking. Edge computing is only used for fault cause analysis and is not effectively linked with data compression and field-of-view switching. Data transmission volume is not effectively controlled, and latency issues are not fundamentally resolved. Summary of the Invention
[0007] The purpose of this invention is to propose a multimodal data compression and real-time inspection method and system, which improves the real-time performance, accuracy and comprehensiveness of industrial inspection.
[0008] To achieve the above objectives, in a first aspect, the present invention proposes a multimodal data compression and real-time inspection system, comprising: The sensing module is used to simultaneously acquire image, temperature, and voiceprint trimodal data; the sensing module includes a dual-field-of-view imaging component, an infrared temperature sensor, and a voiceprint sensor. The edge computing module is used to run the AI object detection model to detect objects in the acquired images, trigger dual field-of-view switching commands based on the detection results, and perform localized compression of ROI region data. The data fusion module is used to extract features from temperature data and voiceprint data, combine them with image fault features for weighted fusion to generate a comprehensive fault index, and classify fault levels according to the comprehensive fault index. The control command module is used to generate corresponding control commands based on the fault level.
[0009] Beneficial effects of the basic solution: The perception module of this system adopts a dual-field-of-view imaging component. The wide field of view can achieve panoramic coverage of the inspection area without blind spots, avoiding the omission of potential fault points. The edge computing module runs an AI target detection model, which can identify abnormal targets in the image in real time and trigger a narrow field of view switching command to conduct a precise and detailed inspection of abnormal areas. It takes into account both panoramic coverage and accurate identification, improves the identification rate of minor faults, reduces the fault missed rate, and is conducive to the early detection of potential equipment hazards.
[0010] The edge computing module employs a localized compression algorithm based on ROI (Region of Interest), retaining and precisely compressing only data in fault-feature areas while discarding redundant background data without faults. This achieves targeted data compression with a compression ratio of 15:1, effectively reducing data transmission volume compared to existing technologies. Simultaneously, it reduces transmission latency to less than 0.3 seconds, far superior to the transmission latency of over 1 second in existing technologies, improving the real-time transmission efficiency of inspection data and ensuring that inspection personnel can obtain fault-related data in a timely manner, thus gaining time for rapid fault handling.
[0011] The data fusion module extracts feature information from temperature and voiceprint data, combines it with image fault features to construct a weighted fusion model, generates a comprehensive fault index and classifies fault levels, and achieves triple verification of image, temperature and voiceprint data. This breaks through the limitations of existing single-modal or dual-modal inspection technologies. After testing and verification, the fault judgment accuracy reaches 94%, which is more than 20% higher than the existing comparison technology (78%). It effectively reduces misjudgment and omission, provides a scientific and accurate basis for fault judgment, and reduces the cost of manual review.
[0012] The edge computing module integrates multiple functions such as target detection, field-of-view switching triggering, data compression, and multimodal fusion. It can complete the entire process without relying on cloud commands, which not only further reduces data transmission and processing latency, but also enhances the system's autonomous operation capability and avoids the impact of cloud failures or network interruptions on inspection work. At the same time, this design is adapted to the real-time inspection needs in industrial scenarios and can be flexibly deployed in various industrial equipment inspection scenarios, with strong versatility and high practicality.
[0013] As a feasible preferred solution, the dual-field-of-view imaging assembly includes a wide-field-of-view camera and a narrow-field-of-view camera coaxially integrated on the gimbal; The wide field-of-view camera is used to capture panoramic images of the inspection area; The narrow field-of-view camera is equipped with optical zoom function, which is used to focus on the target area and capture detailed images after receiving a switching command; The edge computing module also includes a dual-field-of-view switching decision unit, which is configured to: when the target detection confidence is greater than or equal to a preset threshold, or the target priority is greater than or equal to a preset level, or the target size occupies a proportion of the wide field-of-view image greater than or equal to a preset proportion, calculate the target coordinates and drive the gimbal to turn so that the target is located in the center of the narrow field-of-view camera, and automatically restore to the wide field-of-view mode after the narrow field-of-view detailed inspection is completed.
[0014] As a feasible preferred embodiment, the edge computing module further includes an ROI data compression unit configured to execute a differentiated compression strategy: The ROI region is processed using the JPEG 2000 compression standard to preserve fault characteristics; The ROI neighborhood is processed using the H.265 compression standard; For the background region, a strategy combining feature discarding and low bit rate compression is adopted. Edge detection is used to determine whether it is a featureless region. If it is, only the contour information is retained. Otherwise, a compression ratio of 1:15 is used for processing. The processed data from each region are stitched together to generate a complete compressed image.
[0015] As a feasible and preferred embodiment, the data fusion module is specifically used for: Use a local clock to timestamp the data for each modality to establish a correlation; The average temperature, maximum temperature, and temperature deviation of the target area are extracted as temperature features. The voiceprint signal is converted into a spectrum graph, and the peak frequency, peak amplitude and spectral entropy of the spectrum are extracted as voiceprint features. The comprehensive fault index is calculated using a weighted fusion algorithm, and the formula is as follows:
[0016] in, , , The scores are for temperature, voiceprint, and image modal features, respectively. , , These are the feature weights for temperature, voiceprint, and image, respectively.
[0017] As a feasible preferred solution, the control commands generated by the control command module include gimbal turning commands, alarm prompt commands, or emergency stop prompt commands; When the fault level is minor, a PTZ marker and a local log instruction are generated. When the fault level is a general fault, generate commands for pan-tilt focusing and local audible and visual alarms. When the fault level is a severe fault, a PTZ lockout, local and remote dual alarm, and emergency shutdown prompt command are generated. The control commands are in JSON format and include fields for command type, fault level, target coordinates, timestamp, and device ID.
[0018] As a feasible and preferred solution, the transmission module uses a 5G network and the MQTT protocol for data transmission; The QoS level of the MQTT protocol is set to 2, compressed image data and control commands are transmitted separately, and the control commands adopt a priority transmission mechanism. The system is configured to meet the requirements that the control command transmission delay is less than or equal to 0.1 seconds and the compressed image data transmission delay is less than or equal to 0.2 seconds.
[0019] As a feasible and preferred solution, the AI object detection model is based on the improved YOLOv8 architecture, in which the C2f module is used in the backbone layer instead of the original C3 module, and an attention mechanism is added to the neck layer. The model has undergone structured pruning and quantization, resulting in a reduced model size and inference speed that meets real-time detection requirements. The training dataset for the model covers fault samples from power equipment, chemical pipelines, and smart manufacturing workshops.
[0020] As a feasible and preferred solution, the acquisition frequency of the sensing module is linked to the dual-field-of-view imaging component: In wide field of view mode, the acquisition frequency of the infrared temperature sensor and the acoustic sensor is the first frequency. In the narrow field-of-view detailed inspection mode, the acquisition frequency of the infrared temperature sensor and the acoustic fingerprint sensor is increased to a second frequency to ensure synchronization with the image data.
[0021] As a feasible preferred embodiment, it also includes a transmission module for transmitting the compressed fused data and the control commands.
[0022] Secondly, the present invention provides a method for multimodal data compression and real-time inspection, characterized in that it utilizes the aforementioned multimodal data compression and real-time inspection system, comprising: Panoramic monitoring and target detection: a wide field-of-view camera captures panoramic images of the inspection area, and a lightweight AI model from the edge computing module performs real-time detection. When the switching conditions are met, a narrow field-of-view switching command is triggered. Narrow field of view detailed inspection and multimodal acquisition: The gimbal drives the narrow field of view camera to turn to the target area and automatically zooms to acquire detailed images. At the same time, the temperature sensor and the acoustic sensor acquire temperature data and acoustic signals of the target area at an enhanced frequency. Data compression and multimodal fusion: The edge computing module extracts the ROI and performs differential compression on narrow field-of-view images, while the data fusion module performs weighted fusion of features from each modality, calculates the comprehensive fault index, and determines the fault level. Command generation and transmission: The control command module generates corresponding control commands based on the fault level and transmits them to the remote monitoring center. The PTZ camera then performs corresponding marking, focusing, or locking actions. In case of an anomaly, if it is a serious fault, both local and remote alarms will be triggered, along with an emergency shutdown prompt. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of a multimodal data compression and real-time inspection system.
[0024] Figure 2 This is a logical diagram of a multimodal data compression and real-time inspection method. Detailed Implementation
[0025] To make the technical solution and advantages of this application clearer, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only some embodiments of the present invention, and are only used to explain this application, not to limit it. It should be noted that the technical features or combinations of technical features described in the following embodiments should not be considered isolated; they can be combined with each other to achieve better technical effects. The same reference numerals appearing in the accompanying drawings of the following embodiments represent the same features or components, and can be applied to different embodiments.
[0026] Furthermore, unless otherwise defined, the technical or scientific terms used in this invention description shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains.
[0027] The present invention will now be described in further detail with reference to the accompanying drawings.
[0028] Reference Figure 1 This disclosure provides a multimodal data compression and real-time inspection system, including a sensing module, an edge computing module, a data fusion module, a control command module, and a transmission module.
[0029] The sensing module includes a dual-field-of-view imaging component, an infrared temperature sensor, and an acoustic sensor, enabling the synchronous acquisition of image, temperature, and acoustic data.
[0030] In this embodiment, the dual-field-of-view imaging component employs a collaborative mode of a wide-field-of-view camera and a narrow-field-of-view camera. The wide-field-of-view camera has a 120° field of view, a resolution of 1280×720, and a frame rate of 25fps, used to acquire panoramic images of the inspection area. The narrow-field-of-view camera has a 20° field of view, a resolution of 1920×1080, a frame rate of 30fps, and is equipped with 10x optical zoom, used to focus on the target area and acquire detailed images. The two cameras are coaxially integrated on the gimbal, with the optical axis of the narrow-field-of-view camera coinciding with that of the wide-field-of-view camera, ensuring that the target is located in the center of the narrow field of view after switching.
[0031] In this embodiment, the infrared temperature sensor has the following parameters: detection range -20℃~300℃, accuracy ±0.3℃, and is used to collect temperature data of the target area in real time.
[0032] In this embodiment, the parameters of the voiceprint sensor are: detection range 30dB~120dB, accuracy ±1dB, used to collect voiceprint signals of the target area in real time.
[0033] The edge computing module includes a lightweight AI object detection model, a dual-field-of-view switching decision unit, and an ROI data compression unit, enabling object detection, switching triggering, and localized data compression.
[0034] The lightweight AI target detection model adopts an improved YOLOv8 architecture, based on YOLOv8-nano, with an input image size of 640×640. The backbone uses the C2f module instead of the original C3 module, reducing the computational load by 30%. An attention mechanism (CBAM) is added to the neck layer to enhance the ability to extract fault features (such as high-temperature areas of equipment and cracks in pipelines). Referring to Table 1, which shows the current industry status, the detection accuracy of this solution is improved by 12% after testing.
[0035] Table 1
[0036] The inference speed at the edge (NVIDIA Jetson Xavier NX) reaches 60fps, meeting real-time detection requirements. The model training dataset comes from an industrial inspection fault image dataset, including three scenarios: power equipment, chemical pipelines, and intelligent manufacturing workshops, with over 100,000 labeled samples covering 15 types of faults such as "high temperature, abnormal noise, rust, cracks, and loosening." Data augmentation methods such as random cropping, rotation, and brightness adjustment are used to improve the model's generalization ability. The AdamW optimizer is used with an initial learning rate of 0.001, a batch size of 16, and 100 training epochs. An early stopping strategy (stopping if the accuracy on the validation set does not improve for 5 consecutive epochs) is employed to avoid overfitting. Structured pruning is used, removing convolutional kernels with absolute weight values <0.01 from the backbone layer, with a pruning ratio of 35%, reducing the model size from 12MB to 7.8MB. The model weights are quantized from 32-bit floating-point to 8-bit integers, resulting in a 40% improvement in inference speed and a precision loss of ≤2%. The compressed model size is 7.8MB, the inference speed is 60fps, and the object detection accuracy is 92%.
[0037] The dual-field-of-view switching decision unit triggers a narrow field-of-view switching command based on the detection results of the lightweight AI model. The switching conditions include a target detection confidence level ≥ 0.6 (such as detecting fault targets such as "damaged equipment casing" or "corroded pipes"), a target priority level ≥ 8 (in this embodiment, there are 10 priority levels, set based on the severity of the fault, such as "high temperature > abnormal noise > corrosion"), or a target size occupying ≥ 5% of the wide field-of-view image (to avoid triggering by minor interference), see Table 2.
[0038] Table 2
[0039] The switching process of the dual-field-of-view switching decision unit includes: the edge computing module calculates the target's coordinates (u, v) in the wide-field-of-view image, drives the gimbal to rotate via gimbal control commands, and positions the target at the center of the narrow-field-of-view (rotation accuracy ±0.5°), with a switching response time ≤0.1s; the narrow-field-of-view camera automatically zooms until the target is clear (based on an image sharpness evaluation function, such as a variance maximization algorithm). After the narrow-field-of-view detailed inspection is completed (acquiring 3 clear images), it automatically reverts to the wide-field-of-view panoramic mode; if no new target is detected for 5 consecutive seconds, recovery is also triggered.
[0040] The ROI data compression unit optimizes data compression strategies for inspection scenarios, preserving fault features and discarding redundant background information through ROI extraction. The ROI region uses the JPEG 2000 compression standard with a compression ratio of 1:2 to ensure clear fault details (such as crack textures and temperature anomalies), achieving a peak signal-to-noise ratio (PSNR) ≥35dB after compression. The ROI neighborhood uses the H.265 compression standard with a compression ratio of 1:8 to balance clarity and compression efficiency in transition areas. The background region employs a strategy combining feature discarding and low-bitrate compression. Edge detection determines whether the background is featureless (such as walls or floors). Featureless regions are directly compressed at a ratio of up to 1:30, retaining only contour information, while feature-rich backgrounds are compressed at a ratio of 1:15. The edge computing module first divides the ROI, then calls the corresponding compression algorithm to process each region, and finally stitches the images together to generate a complete compressed image. The total compression time is ≤0.1s. Through comparative testing, the size of a single frame of a narrow field-of-view 1080P image (2.07MB / frame) was reduced to 0.14MB after differential compression, with a compression ratio of 15:1. There was no loss in the accuracy of fault feature area identification and no loss of key information in the background area.
[0041] The data fusion module is used for feature extraction and weighted fusion of temperature and voiceprint data, and generates fault levels by combining image fault features.
[0042] In this embodiment, the acquisition frequency is linked with the dual-field-of-view camera—in wide field-of-view mode, the temperature and acoustic sensor acquisition frequency is 5Hz; in narrow field-of-view detailed inspection mode, the acquisition frequency is increased to 20Hz to ensure synchronization with image data.
[0043] The local clock of the edge computing module (with an accuracy of 1ms in this embodiment) is used to timestamp the data of each modality, ensuring that the image, temperature and voiceprint data at the same time are correlated.
[0044] The feature extraction method for temperature features is as follows: extract the average temperature T of the target region. AVE Maximum temperature T MAX Temperature deviation ΔT (ΔT=T) MAX -T, where T is the normal operating temperature of the equipment); for example, if the normal operating temperature of a transformer is T≤65℃, and T=88℃ is detected, then ΔT=23℃.
[0045] The method for extracting voiceprint features is to use Fast Fourier Transform (FFT) to convert the voiceprint signal into a spectrum and extract the peak frequency f, peak amplitude A, and spectral entropy S (reflecting the complexity of the voiceprint signal). For example, the peak frequency of a motor is 50Hz when it is running normally, but f becomes 120Hz when it is abnormal, and the spectral entropy increases significantly.
[0046] The comprehensive failure index is calculated using a weighted fusion algorithm. The formula is as follows:
[0047] in, , , The scores are for temperature, voiceprint, and image modal features, respectively. , , These are the feature weights for temperature, voiceprint, and image, respectively. In this embodiment, the temperature feature weight... =0.4 (Abnormal temperature can easily cause fires, high reliability), voiceprint feature weight =0.3, image feature weight =0.3.
[0048] Fault levels are classified according to the comprehensive fault index F: F < 3 is no fault, 3 ≤ F < 6 is a minor fault, 6 ≤ F < 9 is a general fault, and F ≥ 9 is a serious fault.
[0049] Each modal feature is scored according to the degree of anomaly, ranging from 0 to 10 points. For example, temperature ΔT≤5℃ earns 1 point, 5℃<ΔT≤15℃ earns 3 points, 15℃<ΔT≤30℃ earns 7 points, and ΔT>30℃ earns 10 points; acoustic signature spectral entropy S≤0.5 earns 1 point, and S>1.2 earns 10 points; image target confidence ≥0.9 earns 10 points, and 0.6≤image target confidence <0.9 earns 6 points.
[0050] The command control module generates control commands based on the fault level. These commands include PTZ direction transmission and alarm prompts, and the command format and transmission protocol are defined.
[0051] Specifically, for minor faults (3≤F<6): generate a "PTZ mark + local record" command, the PTZ will overlay a mark box at the target position, and no alarm will be triggered; General fault (6≤F<9): Generates "PTZ Focus + Local Audible and Visual Alarm" command, PTZ continuously focuses on the target, local buzzer alarm (frequency 1Hz), indicator light flashes yellow; Serious Fault (F≥9): Generates "PTZ Lock + Local + Remote Dual Alarm + Emergency Stop Prompt" command, PTZ locks onto the target and continues filming, on-site buzzer sounds a high-frequency alarm (frequency 5Hz), indicator light stays red, and simultaneously sends an emergency stop suggestion command to the remote monitoring center.
[0052] Command format: JSON format, including command type (mark / focus / lock), fault level, target coordinates (world coordinates X, Y, Z), timestamp, and device ID fields, for example: {"cmd_type":"lock","fault_level":"serious","target_pos":{"X":12.5,"Y":8.3,"Z":2.1},"timestamp":"20251211143025","device_id":"INS001"}.
[0053] The transmission module employs 5G and MQTT protocols to transmit compressed, fused data and control commands, ensuring low latency. The MQTT protocol's QoS level is set to 2 (ensuring messages are delivered only once), the transmission port is 1883, and the heartbeat interval is 10 seconds. Compressed image data and control commands are transmitted separately, with control commands using a "priority transmission" mechanism to ensure urgent commands occupy bandwidth first. Transmission latency is ≤0.1 seconds for control commands, ≤0.2 seconds for compressed image data, and ≤0.3 seconds for the total latency, meeting real-time inspection requirements.
[0054] Taking the inspection of power substations as an example, refer to Figure 2 A multimodal data compression and real-time inspection method includes: Step S1, panoramic monitoring and target detection: The wide field of view camera collects panoramic images of the substation, and the lightweight AI model of the edge computing module detects in real time. It finds a target in the transformer area (confidence level 0.82, judged as "suspected oil leak", priority level 8), triggering the narrow field of view switching command.
[0055] Step S2, Narrow field of view detailed inspection and multimodal acquisition: The gimbal drives the narrow field of view camera to turn to the target area, automatically zooms to a clear state, and acquires a 1080P detailed image; at the same time, the temperature sensor acquires the temperature of the target area (e.g., T=72℃, T=65℃, ΔT=7℃), and the voiceprint sensor acquires the voiceprint signal (e.g., after FFT analysis, f=80Hz, spectral entropy 1.0).
[0056] Step S3, data compression and multimodal fusion: The edge computing module extracts the ROI region (boundary box of the oil leak area) of the image and uses differential compression to compress a single frame image from 2.07MB to 0.14MB; the data fusion module calculates the scores of each modality: temperature S=3 points, voiceprint S=6 points, image S=8 points, and comprehensive fault index F=5.4 points, which is judged as a minor fault.
[0057] Step S4, Command Generation and Transmission: The control command module generates a "PTZ Mark + Local Record" command, which is transmitted to the remote monitoring center via 5G and MQTT protocols; the PTZ overlays a red marker box on the narrow field-of-view image, the system records the fault information and location coordinates, and then automatically resumes wide field-of-view panoramic monitoring.
[0058] Step S5, Anomaly Handling: If a serious fault is detected, such as switch cabinet temperature T=120℃ (ΔT=55℃, S=10 minutes), acoustic signature spectrum entropy 1.5 (S=10 minutes), image confidence 0.95 (S=10 minutes), a serious fault command is triggered: the PTZ locks onto the target, the on-site red light stays on and a high-frequency alarm is triggered, and the remote center receives an emergency shutdown prompt to achieve a rapid response.
[0059] The above content is merely an embodiment of the present invention. Commonly known structures and characteristics of the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can improve and implement this solution based on the guidance provided in this application and their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.
Claims
1. A multimodal data compression and real-time inspection system, characterized in that, include: The sensing module is used to simultaneously acquire image, temperature, and voiceprint trimodal data; the sensing module includes a dual-field-of-view imaging component, an infrared temperature sensor, and a voiceprint sensor. The edge computing module is used to run the AI object detection model to detect objects in the acquired images, trigger dual field-of-view switching commands based on the detection results, and perform localized compression of ROI region data. The data fusion module is used to extract features from temperature data and voiceprint data, combine them with image fault features for weighted fusion to generate a comprehensive fault index, and classify fault levels according to the comprehensive fault index. The control command module is used to generate corresponding control commands based on the fault level.
2. The multimodal data compression and real-time inspection system according to claim 1, characterized in that, The dual-field-of-view imaging assembly includes a wide-field-of-view camera and a narrow-field-of-view camera coaxially integrated on the gimbal; The wide field-of-view camera is used to capture panoramic images of the inspection area; The narrow field-of-view camera is equipped with optical zoom function, which is used to focus on the target area and capture detailed images after receiving a switching command; The edge computing module also includes a dual-field-of-view switching decision unit, which is configured to: when the target detection confidence is greater than or equal to a preset threshold, or the target priority is greater than or equal to a preset level, or the target size occupies a proportion of the wide field-of-view image greater than or equal to a preset proportion, calculate the target coordinates and drive the gimbal to turn so that the target is located in the center of the narrow field-of-view camera, and automatically restore to the wide field-of-view mode after the narrow field-of-view detailed inspection is completed.
3. The multimodal data compression and real-time inspection system according to claim 1, characterized in that, The edge computing module also includes an ROI data compression unit, which is configured to execute a differentiated compression strategy: The ROI region is processed using the JPEG 2000 compression standard to preserve fault characteristics; The ROI neighborhood is processed using the H.265 compression standard; For the background region, a strategy combining feature discarding and low bit rate compression is adopted. Edge detection is used to determine whether it is a featureless region. If it is, only the contour information is retained. Otherwise, a compression ratio of 1:15 is used for processing. The processed data from each region are stitched together to generate a complete compressed image.
4. The multimodal data compression and real-time inspection system according to claim 1, characterized in that, The data fusion module is specifically used for: Use a local clock to timestamp the data for each modality to establish a correlation; The average temperature, maximum temperature, and temperature deviation of the target area are extracted as temperature features. The voiceprint signal is converted into a spectrum graph, and the peak frequency, peak amplitude and spectral entropy of the spectrum are extracted as voiceprint features. The comprehensive fault index is calculated using a weighted fusion algorithm, and the formula is as follows: in, , , The scores are for temperature, voiceprint, and image modal features, respectively. , , These are the feature weights for temperature, voiceprint, and image, respectively.
5. The multimodal data compression and real-time inspection system according to claim 1, characterized in that, The control commands generated by the control command module include gimbal turning commands, alarm prompt commands, or emergency stop prompt commands. When the fault level is minor, a PTZ marker and a local log instruction are generated. When the fault level is a general fault, generate commands for pan-tilt focusing and local audible and visual alarms. When the fault level is a severe fault, a PTZ lockout, local and remote dual alarm, and emergency shutdown prompt command are generated. The control commands are in JSON format and include fields for command type, fault level, target coordinates, timestamp, and device ID.
6. The multimodal data compression and real-time inspection system according to claim 1, characterized in that, The transmission module uses a 5G network and the MQTT protocol for data transmission; The QoS level of the MQTT protocol is set to 2, compressed image data and control commands are transmitted separately, and the control commands adopt a priority transmission mechanism. The system is configured to meet the requirements that the control command transmission delay is less than or equal to 0.1 seconds and the compressed image data transmission delay is less than or equal to 0.2 seconds.
7. The multimodal data compression and real-time inspection system according to claim 1, characterized in that, The AI object detection model is based on the improved YOLOv8 architecture, in which the C2f module replaces the original C3 module in the backbone layer, and an attention mechanism is added to the neck layer. The model has undergone structured pruning and quantization, resulting in a reduced model size and inference speed that meets real-time detection requirements. The training dataset for the model covers fault samples from power equipment, chemical pipelines, and smart manufacturing workshops.
8. The multimodal data compression and real-time inspection system according to claim 1, characterized in that, The acquisition frequency of the sensing module is linked to the dual-field-of-view imaging component: In wide field of view mode, the acquisition frequency of the infrared temperature sensor and the acoustic sensor is the first frequency. In the narrow field-of-view detailed inspection mode, the acquisition frequency of the infrared temperature sensor and the acoustic fingerprint sensor is increased to a second frequency to ensure synchronization with the image data.
9. The multimodal data compression and real-time inspection system according to claim 1, characterized in that, It also includes a transmission module for transmitting the compressed fused data and the control commands.
10. A method for multimodal data compression and real-time inspection, characterized in that, The system employs a multimodal data compression and real-time inspection system as described in any one of claims 1-9, comprising: Panoramic monitoring and target detection: a wide field-of-view camera captures panoramic images of the inspection area, and a lightweight AI model from the edge computing module performs real-time detection. When the switching conditions are met, a narrow field-of-view switching command is triggered. Narrow field of view detailed inspection and multimodal acquisition: The gimbal drives the narrow field of view camera to turn to the target area and automatically zooms to acquire detailed images. At the same time, the temperature sensor and the acoustic sensor acquire temperature data and acoustic signals of the target area at an enhanced frequency. Data compression and multimodal fusion: The edge computing module extracts the ROI and performs differential compression on narrow field-of-view images, while the data fusion module performs weighted fusion of features from each modality, calculates the comprehensive fault index, and determines the fault level. Command generation and transmission: The control command module generates corresponding control commands based on the fault level and transmits them to the remote monitoring center. The PTZ camera then performs corresponding marking, focusing, or locking actions. In case of an anomaly, if it is a serious fault, both local and remote alarms will be triggered, along with an emergency shutdown prompt.