Multi-modal sensing real-time edge computing method and system for smart fixture device

The intelligent fixture system, which integrates multimodal sensors and edge computing units, solves the problems of heterogeneous multimodal sensor data and limited resources in high-end manufacturing. It achieves high-precision, real-time clamping status perception and control, improves processing accuracy and anomaly detection rate, and is suitable for high-cycle manufacturing and collaborative robot workstations.

CN122363429APending Publication Date: 2026-07-10FUJIAN BOQIN PRECISION MACHINERY TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610225274.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-25
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing intelligent fixtures in high-end manufacturing suffer from strong heterogeneity of multimodal sensor data, asynchronous timing, limited edge computing resources, inability to dynamically adjust sensor mode combinations, and lack of closed-loop collaboration mechanisms. This results in fragmented state perception and wasted resources, making it difficult to meet the requirements of high precision and real-time control.

Method used

It employs a multimodal sensor array module, an edge computing unit module, and an adaptive feedback control module, integrating force, vision, acoustic, and temperature sensors for real-time data acquisition, feature extraction, and fusion. Multimodal feature fusion is achieved through a gating cross-attention mechanism, and the clamping force and attitude are dynamically adjusted through the adaptive feedback control module to achieve end-to-end closed-loop control.

Benefits of technology

It has achieved a multi-dimensional sensing system, increasing the clamping anomaly detection rate to over 98%, improving machining accuracy by 3 times, and controlling power consumption within 15W. It is suitable for high-cycle manufacturing scenarios, meets real-time requirements, and reduces the micro-displacement of workpieces in high-vibration machining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122363429A_ABST
    Figure CN122363429A_ABST
Patent Text Reader

Abstract

This application relates to the interdisciplinary field of artificial intelligence and edge computing, and discloses a multimodal sensing real-time edge computing method and system for intelligent gripper devices. It aims to solve the problems of heterogeneous and asynchronous multimodal data, limited edge computing resources, static perception strategies, and lack of closed-loop coordination between perception and control in existing technologies. The method includes: simultaneously acquiring four types of signals: force, vision, acoustics, and temperature; performing time-series alignment and standardization processing; extracting dynamic features of each modality; fusing them through a gated cross-attention mechanism to generate a comprehensive representation of the gripping state; determining stability based on a lightweight classifier and generating feedback control signals; and driving the gripper to adjust clamping force and posture in real time. By adopting the above technical solution, this application can achieve high-precision, low-latency closed-loop control, significantly improving gripping stability and machining accuracy, while also ensuring low power consumption and strong robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of interdisciplinary technology of artificial intelligence and edge computing, and specifically relates to a multimodal sensing real-time edge computing method and system for intelligent clamping devices. Background Technology

[0002] In the fields of high-end manufacturing and flexible automation, intelligent fixtures, as key interfaces connecting workpieces and processing equipment, directly impact processing accuracy, production efficiency, and system reliability through their sensing and response capabilities. With the deepening of Industry 4.0 and intelligent manufacturing, fixtures are no longer limited to mechanical positioning but are gradually evolving into intelligent terminals integrating sensing, decision-making, and execution, placing higher demands on real-time performance, multi-source information fusion capabilities, and edge autonomous decision-making capabilities.

[0003] Among these, the synergy between multimodal sensing and edge computing has become a core technological direction for improving the performance of intelligent grippers. This direction aims to capture the physical state during the gripping process in real time by integrating multiple sensors such as force, vision, vibration, and temperature, and to complete feature extraction, state recognition, and control feedback at the edge side close to the data source, thereby reducing communication latency, alleviating the burden on the cloud, and enhancing system robustness.

[0004] Existing technologies still face significant bottlenecks in achieving multimodal sensing and edge computing for intelligent fixtures: First, multimodal sensor data is highly heterogeneous and asynchronous in timing, lacking an efficient and lightweight fusion mechanism, leading to fragmented state perception. Second, edge computing resources are limited, making it difficult to simultaneously meet the dual requirements of high-precision model inference and millisecond-level control response under low power conditions. Third, existing systems mostly employ static sensing strategies, failing to dynamically adjust the sensor modality combination and computational load according to the processing task, resulting in resource waste or sensing blind spots. Finally, the lack of a closed-loop coordination mechanism between sensor data and control commands makes it difficult to achieve real-time linkage from "perceived anomaly" to "active adjustment." These problems severely restrict the application efficiency of intelligent fixtures in high-dynamic, high-precision manufacturing scenarios, urgently requiring a real-time collaborative method and system that can achieve deep integration of multimodal sensing and edge computing. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and to provide a multimodal sensing real-time edge computing method and system for intelligent clamping devices, which can effectively solve the problems in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: On one hand, a multimodal sensing real-time edge computing system for an intelligent clamping device, comprising the following components: a multimodal sensor array module, used to simultaneously acquire four types of physical signals—force, vision, acoustic, and temperature—during the clamping process of the workpiece, and output raw multimodal sensing data streams; an edge computing unit module, deployed on the clamping body or a nearby control node, used to perform real-time preprocessing, feature extraction, and fusion inference on the raw multimodal sensing data streams to generate clamping state evaluation results; an adaptive feedback control module, used to dynamically adjust the clamping force, attitude, or clamping strategy of the clamp according to the clamping state evaluation results to maintain a stable clamping state of the workpiece during processing; and a communication interface module, used to realize low-latency data interaction between the multimodal sensor array module, the edge computing unit module, and the adaptive feedback control module, and to support state synchronization with the host computer system; On the other hand, a multimodal sensing real-time edge computing method for an intelligent clamping device includes the following steps: Step S110, simultaneously acquiring force signals, visual images, acoustic signals, and temperature signals through a multimodal sensor array module during the clamping action, forming a four-channel raw multimodal sensing data stream; Step S120, inputting the raw multimodal sensing data stream into an edge computing unit module, performing time alignment, noise suppression, and normalization on each channel's data to generate standardized multimodal sensing data; Step S130, performing dynamic feature extraction based on a sliding window on the force timing signal in the standardized multimodal sensing data to obtain a clamping force fluctuation feature vector; performing local texture and deformation feature encoding based on a lightweight convolutional neural network on the visual image to obtain a visual deformation feature vector; and performing short-time Fourier transform and time-frequency energy division on the acoustic signal. The following steps are performed: Step S140: Modeling is performed to obtain acoustic anomaly feature vectors; Trend slope and steady-state deviation analysis is performed on temperature signals to obtain thermal stability feature vectors; Step S150: Input the clamping force fluctuation feature vector, visual deformation feature vector, acoustic anomaly feature vector, and thermal stability feature vector into the multimodal feature fusion inference engine, calculate the dynamic correlation weights between each modality through a gated cross-attention mechanism, and generate a weighted fusion to generate a comprehensive clamping state representation vector; Step S160: Based on the comprehensive clamping state representation vector, determine whether the current clamping state is in a stable range through a pre-trained lightweight classifier. If it is determined to be unstable, a feedback control signal containing clamping force adjustment, attitude fine-tuning angle, and clamping strategy switching instructions is generated; Step S170: Transmit the feedback control signal to the adaptive feedback control module to drive the clamping actuator to perform real-time adjustments and complete closed-loop control.

[0007] Preferably, the multimodal sensor array module includes an embedded micro force sensor array, a micro CMOS image sensor, a MEMS microphone array, and a thermocouple temperature sensor. The embedded micro force sensor array is distributed on the contact surface of the fixture at a spacing of 450 nm and has a sampling frequency of not less than 1 kHz. The micro CMOS image sensor has a resolution of not less than 640×480 and a frame rate of not less than 30 fps. The MEMS microphone array operates in a frequency band covering 20 Hz to 20 kHz. The thermocouple temperature sensor has a response time of less than 100 ms and a measurement accuracy of ±0.5℃.

[0008] Preferably, the edge computing unit module adopts a heterogeneous computing architecture, including an ARM Cortex-A series general-purpose processor core and an NPU neural network acceleration unit. The general-purpose processor core is responsible for data preprocessing and task scheduling, while the NPU unit is dedicated to performing forward inference of convolutional neural networks and attention mechanisms. The overall inference latency is controlled within 20ms, meeting the real-time requirements.

[0009] Furthermore, the gated cross-attention mechanism in the multimodal feature fusion inference engine calculates the query vector. Key vector The attention weights are assigned between modalities using the value vector V, and the formula for calculating the attention weights is as follows:

[0010] in, Let be the dimension of the key vector. Generated from the current dominant modality features, V is generated from other auxiliary modal features, and the contribution of low-confidence modalities is dynamically suppressed by gating units to ensure the robustness of the fusion result.

[0011] In addition, the lightweight classifier adopts a depthwise separable convolutional structure, which includes three convolutional layers and one fully connected output layer. The number of parameters is controlled within 50KB. It is deployed on the NPU of the edge computing unit and has a classification accuracy of no less than 95% and a false alarm rate of less than 2%.

[0012] Preferably, after receiving the feedback control signal, the adaptive feedback control module converts the clamping force adjustment into a motor drive current command through a PID controller. At the same time, it combines the data from the six-axis attitude sensor to perform spatial attitude compensation on the end effector of the fixture, with a compensation accuracy of ±0.1 degrees, to ensure that the workpiece does not shift or vibrate during high-speed machining.

[0013] Furthermore, the communication interface module adopts a Time-Sensitive Networking (TSN) protocol stack to ensure that the transmission jitter of multimodal data streams between modules inside the fixture is less than 50μs, and supports state synchronization with the host MES system via OPC UA protocol, with a synchronization period of 100ms.

[0014] Compared with the prior art, the present invention has the following beneficial effects: By integrating four types of sensors—force, vision, acoustics, and temperature—into the clamp body, a multi-dimensional perception system covering the entire clamping process is constructed, overcoming the blind spot problem of state perception caused by traditional clamps relying on only a single force sensor, and increasing the clamping anomaly detection rate to over 98%. By adopting a heterogeneous computing architecture deployed at the edge and a lightweight multimodal fusion model, the end-to-end closed-loop latency from raw data acquisition to control command generation is less than 50ms, meeting the real-time requirements of high-cycle intelligent manufacturing scenarios. By introducing a gated cross-attention mechanism for dynamic weighted fusion of multimodal features, the interference of single-modal noise or failure on the overall judgment is effectively suppressed, and the state recognition accuracy can still be maintained at more than 85% even in scenarios where some sensors fail. By using an adaptive feedback control module to achieve joint dynamic adjustment of clamping force and attitude, the micro-displacement of workpieces during high-vibration machining processes such as milling and drilling is significantly reduced. The measured standard deviation of displacement is reduced from 12.3μm in traditional fixtures to 3.8μm, and the machining accuracy is improved by more than 3 times. The overall power consumption of the system is controlled within 15W, and it can operate stably for a long time on a mobile fixture platform without external power supply. It is suitable for emerging application scenarios such as flexible manufacturing cells and collaborative robot workstations. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the overall technical architecture of the multimodal sensing real-time edge computing method and system for intelligent clamping devices proposed in this invention. Detailed Implementation

[0016] Please refer to Figure 1 To further illustrate the technical means and effects of the present invention in order to achieve the intended purpose, the following detailed description of the specific implementation methods, structures, features and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0017] Example 1 This embodiment addresses the real-time clamping status monitoring and adaptive control requirements of intelligent fixtures in the high-precision manufacturing field when machining complex, high-value metal workpieces such as aero-engine blades. In this scenario, the workpiece material typically possesses high strength and hardness, and the machining process involves high cutting forces, high vibrations, and localized temperature rises, placing extremely high demands on clamping stability. Any minute displacement or vibration can lead to workpiece scrapping and significant losses. This embodiment constructs an intelligent fixture system integrating multimodal sensing, edge computing, and adaptive feedback control to achieve comprehensive perception, real-time evaluation, and precise control of the workpiece clamping status during machining.

[0018] First, a multimodal sensor array module is integrated into the intelligent gripper body. This module is designed to simultaneously acquire and provide multi-dimensional, high-frequency physical signals. Specifically, the array includes: Embedded micro-force sensor array: Directly distributed on the contact surface between the fixture and the workpiece with an extremely fine pitch of 450 nanometers (e.g., achieved through thin-film pressure sensors or MEMS force sensor array technology). This array does not merely provide the total clamping force, but rather forms a high-resolution map of the contact pressure distribution. Each micro-force sensor samples independently at a frequency set to at least 1 kHz, ensuring the capture of transient mechanical responses caused by cutting force fluctuations, tool runout, etc., during machining. The raw force signals are output as a matrix of pressure values ​​(e.g., in Pascals or Newtons per square millimeter), accurately reflecting the local details of the workpiece's stress state.

[0019] Miniature CMOS image sensors: Typically integrated into the jaws of a clamping device or at a fixed location near the clamping surface, these sensors have a resolution of at least 640×480 pixels and a frame rate of at least 30fps. Equipped with a macro optics system and auxiliary illumination (such as a ring LED), these sensors focus on the workpiece-clamp contact area and adjacent workpiece surfaces. They are used to capture real-time visual information such as microscopic deformations, sliding marks, sparks at the tool-workpiece contact point, smoke, or localized material peeling on the workpiece surface. The raw visual data is output as a continuous sequence of RGB image frames, providing a basis for subsequent deformation analysis.

[0020] MEMS microphone arrays are typically mounted in a compact layout at multiple key points on the fixture body, such as near the clamping surface or the tool entry point. The array operates in a frequency band from 20Hz to 20kHz, capturing the full spectrum of acoustic signals generated during machining, including cutting chatter, tool wear noise, workpiece sliding friction noise, and coolant jet noise. Multiple microphone arrays allow for sound source localization, distinguishing workpiece vibration noise from ambient background noise. The raw acoustic data is output as a multi-channel, high-sampling-rate (e.g., 48kHz) digital audio stream.

[0021] Thermocouple temperature sensors: Response time less than 100 milliseconds, measurement accuracy ±0.5°C. These sensors are strategically deployed in critical areas of the fixture-workpiece contact point (such as along the cutting heat conduction path) and inside the fixture body. They continuously monitor temperature changes at the clamping interface and temperature rise within the fixture body due to machining heat conduction and its own drive system heating, providing real-time data for assessing workpiece thermal expansion, clamping force attenuation, and the risk of localized overheating. Raw temperature data is output as a continuous time series (e.g., °C).

[0022] All the raw multimodal data streams from these sensors are transmitted internally with high efficiency and low latency through a communication interface module. This module uses a Time-Sensitive Networking (TSN) protocol stack, an Ethernet standard optimized for industrial real-time communication. TSN ensures that the transmission jitter of the multimodal data stream between the sensor array module and the edge computing unit module is strictly controlled within 50 microseconds, greatly guaranteeing the time synchronization and real-time performance of the data and avoiding lag in status judgment caused by transmission delays. Simultaneously, this communication interface module also supports status synchronization with the host computer manufacturing execution system (MES) via the OPC UA protocol, with a synchronization period set to 100 milliseconds, ensuring that the overall operating status of the fixture, fault alarms, and maintenance requirements can be promptly reported to the workshop management level.

[0023] Upon receiving the raw multimodal sensing data stream, the edge computing unit module is responsible for real-time, efficient processing and inference. This module employs a heterogeneous computing architecture, integrating an ARM Cortex-A series general-purpose processor core and an NPU neural network acceleration unit. The general-purpose processor core primarily handles data preprocessing tasks, including complex timing alignment, noise suppression, data normalization, and overall system task scheduling and resource management. The NPU neural network acceleration unit is optimized for performing high-density convolutional neural network operations and attention-based forward inference, such as image feature encoding and multimodal feature fusion. This architecture effectively offloads and accelerates computational tasks, keeping the overall inference latency from raw data input to clamping state evaluation output within 20 milliseconds, fully meeting the real-time requirements of high-cycle manufacturing scenarios. The edge computing unit is deployed within the fixture body or an adjacent control cabinet, minimizing the physical distance to sensors and actuators and reducing communication overhead.

[0024] Throughout the closed-loop control process, the intelligent fixture system executes according to the following steps: Step S110: Simultaneously collect force signals, visual images, acoustic signals and temperature signals through the multimodal sensor array module when the fixture performs the clamping action to form a four-channel raw multimodal sensing data stream.

[0025] Once the fixture completes the clamping action on the workpiece through its drive mechanism (such as a servo motor-driven jaw) and enters the processing stage, the multimodal sensor array module immediately activates its data acquisition function. Specifically: Force signal acquisition: An embedded miniature force sensor array continuously acquires normal and tangential pressure values ​​of the contact surface with the workpiece at a frequency of over 1000 times per second. These pressure values ​​are quantized into digital signals by an AD converter, forming a dynamically changing two-dimensional pressure map (e.g., a 10x10 matrix) that accurately reflects the force distribution on the workpiece in the clamping area. The instantaneous pressure value at each sensing point is part of the raw force data stream, typically measured in Newtons per square millimeter.

[0026] Visual image acquisition: A miniature CMOS image sensor captures high-resolution image frames of the clamping area and workpiece surface at a fixed rate of 30 frames per second within a preset field of view. The raw data for each frame is a 640×480 pixel RGB value matrix. Image acquisition time is rigidly synchronized with the timestamps of force, acoustic, and temperature data acquisition to ensure the temporal correlation of all data. The image data stream includes compressed or uncompressed pixel arrays.

[0027] Acoustic signal acquisition: A MEMS microphone array captures ambient acoustic waveforms in the 20Hz to 20kHz frequency band at a sampling rate of 48kHz. These raw waveform data are converted into digital sequences by a high-precision ADC, forming a multi-channel acoustic data stream. By synchronizing the data from multiple microphones in time, it can be further used for sound source localization.

[0028] Temperature signal acquisition: The thermocouple temperature sensor converts the sensed temperature (in degrees Celsius) into a voltage signal in real time with a sub-millisecond response speed, and then converts it into a digital temperature value through a high-precision ADC. These temperature values ​​are output in the form of a continuous time series, with each sampling point accompanied by a precise timestamp.

[0029] All four channels of raw sensor data streams, after being digitized and initially encapsulated, are efficiently and with low jitter transmitted to the edge computing unit module via the Time-Sensitive Networking (TSN) protocol stack, awaiting further processing. The data encapsulation format follows a predefined protocol, including sensor ID, timestamp, data type, and payload.

[0030] Step S120: Input the original multimodal sensing data stream into the edge computing unit module, and perform time alignment, noise suppression and normalization processing on the data of each channel to generate standardized multimodal sensing data.

[0031] The ARM Cortex-A general-purpose processor core of the edge computing unit module first receives the multimodal raw data stream from the communication interface module and initiates the data preprocessing process.

[0032] Timing Alignment: Due to slight differences in the physical characteristics and sampling mechanisms of different sensors, precise timing alignment is required. The system utilizes embedded hardware timestamps in all data streams to accurately synchronize data from different modalities to a unified time base. For modalities with inconsistent sampling frequencies (e.g., force 1kHz, vision 30fps), interpolation (such as linear interpolation, cubic spline interpolation) or resampling techniques are used to ensure that synchronized data points or frames for all modalities can be obtained at any given time point. For example, interpolated force, acoustic, and temperature data are padded between vision frames.

[0033] Noise Suppression: Customized noise suppression algorithms are employed based on the characteristics of each modality of data. Force data: Kalman filter or wavelet transform is applied to denoise the sensor to effectively filter out random noise, transient glitches caused by mechanical vibration and environmental electromagnetic interference, and retain the true mechanical change trend.

[0034] Visual images: Non-local means or median filters are used to remove shot noise, salt-and-pepper noise, and artifacts caused by uneven illumination, thereby enhancing image quality and feature region contrast.

[0035] Acoustic signals: Using spectral subtraction combined with adaptive filters (such as the LMS algorithm), steady-state background noise in the machining environment (such as equipment fan noise and hydraulic pump noise) is removed, highlighting abnormal sounds generated by the interaction between the workpiece and the tool.

[0036] Temperature data: Apply a moving average filter or exponential smoothing method to smooth transient temperature jumps caused by environmental fluctuations or sensor thermal noise, ensuring the stability and accuracy of temperature trends.

[0037] Normalization: To eliminate differences in the dimensions and numerical ranges of data from different modalities and to avoid excessive influence of certain modalities on model training, all denoised data undergoes normalization. For example, Min-Max normalization is used to scale the data to the [0,1] interval, or Z-score normalization is used to convert the data into a distribution with a mean of 0 and a standard deviation of 1. Normalization parameters (such as maximum, minimum, mean, and standard deviation) are statistically learned and preset based on historical large-scale normalized clamping data. After the above processing, standardized multimodal sensing data with high signal-to-noise ratio, time synchronization, and consistent dimensions is generated.

[0038] Step S130: Perform dynamic feature extraction based on a sliding window on the force time-series signal in the standardized multimodal sensing data to obtain the clamping force fluctuation feature vector; perform local texture and deformation feature encoding based on a lightweight convolutional neural network on the visual image to obtain the visual deformation feature vector; perform short-time Fourier transform and time-frequency energy distribution modeling on the acoustic signal to obtain the acoustic anomaly feature vector; perform trend slope and steady-state deviation analysis on the temperature signal to obtain the thermal stability feature vector.

[0039] This step is mainly handled by the NPU neural network acceleration unit of the edge computing unit module, which aims to extract core features that can characterize the clamping state from standardized data.

[0040] Dynamic feature extraction of force sensing time-series signals: The standardized force-pressure map sequence was processed using a fixed-size sliding window (e.g., window length of 500 ms and step size of 100 ms). Within each window, the following dynamic characteristics were calculated: Statistical characteristics: mean, standard deviation, variance, kurtosis, skewness, maximum value, minimum value, peak-to-peak value, etc., reflect the overall level and fluctuation range of the clamping force.

[0041] Time-domain characteristics: root mean square (RMS), zero-crossing rate, waveform factor, etc., characterize the energy and periodicity of the force signal.

[0042] Frequency domain characteristics: By performing Fast Fourier Transform (FFT) on the force signal within the window, the main frequency, the energy proportion of each frequency band (such as low-frequency vibration and high-frequency flutter), bandwidth, etc. are extracted to reveal potential vibration modes.

[0043] These features combine to form a high-dimensional clamping force fluctuation feature vector.

[0044] Encoding of local texture and deformation features in visual images: For the standardized visual image sequences, feature extraction is performed using a lightweight convolutional neural network (e.g., a simplified version of MobileNetV2 or EfficientNetLite) pre-trained on an NPU. The network is designed to capture local texture and subtle deformation features of the images with low computational cost. Early convolutional layers: extract low-level visual features of the image, such as edges, corners, and texture primitives, like the contours of scratches on the workpiece surface and indentations from fixtures.

[0045] Mid-depth separable convolutional layers: capture more complex local texture patterns and geometric deformations, such as tiny chipping at workpiece edges, scratches caused by sliding of clamping contact surfaces, and changes in coolant flow patterns.

[0046] Subsequent attention mechanism layer: Focuses on the region in the image most relevant to clamping instability. For example, when the clamping force fluctuates, the model will pay more attention to the image changes of the contact area between the workpiece and the fixture.

[0047] These features are encoded into a compact visual deformation feature vector, such as a 256-dimensional vector, through a global average pooling layer and a fully connected layer.

[0048] Short-time Fourier transform and time-frequency energy distribution modeling of acoustic signals: The standardized acoustic signal is framed using the Hannine Window and subjected to a Short-Time Fourier Transform (STFT) to generate a spectrogram, which is the instantaneous frequency energy distribution matrix. Based on this, the following features are further extracted: Time-frequency domain characteristics: Calculate energy, energy entropy, spectral centroid, bandwidth, etc. within a specific frequency range (e.g., tool chatter frequency range, material friction sound frequency range) to identify the occurrence of abnormal frequency bands.

[0049] Mel frequency cepstral coefficients (MFCCs): This set of features can effectively characterize the spectral envelope of acoustic signals and is highly robust for distinguishing different types of abnormal noises (such as friction noise, impact noise, and tool breakage noise).

[0050] These features together constitute an acoustic anomaly feature vector, reflecting the acoustic fingerprint under clamping conditions.

[0051] Analysis of the trend slope and steady-state deviation of the temperature signal: The standardized temperature time series was analyzed as follows: Trend slope analysis: Linear regression analysis is used to calculate the slope of temperature change within a fixed time window. A sustained negative slope may indicate increased heat dissipation due to clamping looseness, while a positive slope may indicate localized friction or accumulation of cutting heat.

[0052] Steady-state deviation analysis: Calculates the deviation between the current temperature and the preset normal clamping steady-state temperature baseline, or the deviation from the moving average temperature over a past period. Deviation values ​​exceeding a threshold indicate thermal anomalies.

[0053] These analytical results are combined into a concise thermal stability feature vector, such as a 3-dimensional vector containing the current temperature, the rate of temperature change, and the deviation from the baseline.

[0054] Step S140: Input the clamping force fluctuation feature vector, visual deformation feature vector, acoustic anomaly feature vector and thermal stability feature vector into the multimodal feature fusion inference engine, calculate the dynamic correlation weight between each modality through the gating cross-attention mechanism, and generate a weighted fusion to generate a comprehensive characterization vector of the clamping state.

[0055] The multimodal feature fusion inference engine, running on the NPU of the edge computing unit module, is responsible for deep fusion of the four modal feature vectors extracted in the previous step. Its core is a gated cross-attention mechanism.

[0056] This mechanism calculates the query vector ( ), key vector ( Attention weights are assigned between modalities using a value vector (V). The formula for calculating attention weights is as follows:

[0057] in, The dimension of the key vector serves as a scaling factor to prevent the softmax gradient from vanishing due to an excessively large inner product. In multimodal fusion, the feature vector of one modality is used as the query vector. Key vectors used to "query" other modalities The relevance of these values ​​to the current query modality is assessed. These relevances are then transformed into attention weights via a softmax function, which are applied to the value vector V of other modalities.

[0058] The specific integration process is as follows: Query generation: For example, select the clamping force fluctuation feature vector as the main modality, and transform it into a query vector through a separate fully connected layer or a small neural network. .

[0059] Key value generation: Simultaneously, the visual deformation feature vector, acoustic anomaly feature vector, and thermal stability feature vector are each generated into their corresponding key vectors through their respective fully connected layers. Sum value vector .

[0060] Cross-attention calculation:

[0061]

[0062]

[0063] Similarly, each modality can be used as a query modality to perform cross-attention calculations with other modalities, resulting in feature representations that enhance each other across multiple modalities.

[0064] Dynamic suppression via gating units: After the attention output of each modality, a gating unit (e.g., a fully connected layer with a sigmoid activation function) is introduced to dynamically generate a gating value based on the current modality's confidence or quality score (e.g., signal-to-noise ratio or effectiveness predicted by an auxiliary classifier). This gating value is multiplied by the features of the corresponding modality, thereby dynamically suppressing the contributions of low-confidence or noise-affected modalities. For example, when an acoustic sensor has low confidence due to excessive ambient noise, its weight in the final fusion is reduced to ensure the robustness of the fusion result.

[0065] Final weighted fusion: All gated modal feature vectors, along with the outputs of their cross-attention, are then integrated through a final weighted summation layer or another fully connected layer to generate a unified clamping state comprehensive representation vector containing multimodal information. This vector can comprehensively and accurately reflect the current clamping state of the workpiece during processing; for example, it may encode a complex state of "slight slippage accompanied by high-frequency vibration and local temperature rise."

[0066] Step S150: Based on the comprehensive representation vector of the clamping state, determine whether the current clamping state is in a stable range using a pre-trained lightweight classifier. If it is determined to be unstable, generate a feedback control signal that includes the clamping force adjustment amount, the attitude fine-tuning angle, and the clamping strategy switching command.

[0067] The clamped state comprehensive representation vector is then fed into a lightweight classifier deployed on the edge computing unit (NPU). This classifier employs a depthwise separable convolutional structure, containing three convolutional layers and one fully connected output layer, with the number of parameters strictly controlled to within 50KB to ensure extremely low latency inference at the edge. The classifier is designed to achieve efficient and high-accuracy real-time state discrimination, with a classification accuracy of no less than 95% and a false positive rate of less than 2%.

[0068] State determination: This classifier is pre-trained to recognize various gripping states, including but not limited to: Stable clamping: The workpiece is stable within a preset range of force, deformation, acoustics, and temperature.

[0069] Micro-sliding: The relative displacement between the workpiece and the fixture contact surface occurs at the sub-millimeter level.

[0070] Slight vibration: The workpiece vibrates beyond the threshold under the action of processing force, but has not yet become unstable.

[0071] Local deformation: Plastic or elastic deformation caused by excessive clamping or cutting force on or inside the workpiece surface.

[0072] Thermal instability: Abnormal temperature rise in the clamping area leads to thermal expansion of the workpiece or thermal creep of the clamp.

[0073] Potential for detachment: A sharp drop in clamping force poses a risk of workpiece detachment.

[0074] Feedback control signal generation: If the classifier determines that the current clamping state is unstable (i.e., not a "stable clamping" state), the system will automatically generate a corresponding feedback control signal based on the specific type and severity of the detected instability. This signal is not a simple alarm, but contains specific, quantified adjustment instructions: Clamping force adjustment: For example, indicating an increase or decrease of 20 Newtons in clamping force. The adjustment amount is determined based on the severity of the condition and historical optimization experience to avoid overcorrection.

[0075] Fine-tuning of attitude angles: For example, indicating that the end effector of the fixture rotates +0.05 degrees around the X-axis and -0.03 degrees around the Y-axis to compensate for slight tilting or rotation of the workpiece during machining.

[0076] Clamping strategy switching instructions: In some severe cases, it may be necessary to switch to a more aggressive strategy. For example, "activate preload mode", "pause machining and reclamp", "request the host computer to adjust the toolpath or cutting parameters", etc. These instructions are high-level decision signals and may trigger more complex response processes.

[0077] Step S160: The feedback control signal is transmitted to the adaptive feedback control module to drive the fixture actuator to make real-time adjustments and complete closed-loop control.

[0078] The generated feedback control signal is transmitted to the adaptive feedback control module via the TSN protocol stack of the communication interface module with extremely low latency (e.g., total transmission time less than 5 milliseconds). This module is the execution core of the intelligent fixture; it receives and parses the feedback signal, and then drives the fixture's physical actuators to make real-time adjustments.

[0079] Clamping Force Adjustment: After receiving the clamping force adjustment, the adaptive feedback control module converts the adjustment into a precise motor drive current command through its internal PID controller (proportional-integral-derivative controller). For example, if a 20-Newton increase in clamping force is required, the PID controller calculates the corresponding servo motor current increment. This servo motor, through a high-precision reducer and lead screw mechanism, drives the clamp jaws to move with micron-level precision, thereby accurately changing the clamping force on the workpiece. The PID parameters (Kp, Ki, Kd) are precisely calibrated based on the mechanical characteristics of the clamp and the elastic modulus of the workpiece material to ensure rapid response while avoiding overshoot and oscillation.

[0080] Attitude Compensation: Simultaneously, the adaptive feedback control module performs spatial attitude compensation on the fixture end effector by combining real-time attitude data acquired by the six-axis attitude sensor (typically including a three-axis accelerometer and a three-axis gyroscope) integrated on the fixture body. The attitude sensor provides real-time position and orientation information of the fixture in space. The system compares the attitude fine-tuning angle in the feedback signal with the current attitude, and through inverse kinematics calculation, precisely controls the micro-actuator (such as a piezoelectric ceramic actuator or a micro servo motor) inside the fixture end effector to fine-tune the spatial attitude of the workpiece with a compensation accuracy of ±0.1 degrees. This precise attitude compensation can effectively counteract the slight tilting or rotation caused by uneven cutting forces or the workpiece's own weight, ensuring that the workpiece always maintains an ideal geometric position during high-speed machining, thereby significantly reducing the amount of micro-displacement of the workpiece during high-vibration machining processes such as milling and drilling.

[0081] Clamping strategy switching: If the feedback control signal contains a clamping strategy switching instruction, such as "re-clamp," the adaptive feedback control module will, according to a preset program sequence, first release part of the clamping force, then reposition the workpiece based on visual or force feedback and reapply the clamping force. For the instruction "request the host computer to adjust the toolpath," the MES system is notified via the OPC UA protocol, and the MES system coordinates the CNC machine tool to adjust the machining strategy.

[0082] The entire process forms an end-to-end closed-loop control from perception to decision-making to execution, realizing real-time, proactive, and intelligent control of the fixture's workpiece clamping state, ensuring the stability and reliability of the high-precision machining process. Actual test results show that this closed-loop control system can significantly reduce the standard deviation of workpiece displacement during machining from 12.3 micrometers in traditional fixtures to 3.8 micrometers, improving machining accuracy by more than three times and increasing the detection rate of clamping anomalies to over 98%.

[0083] Example 2 This embodiment applies the invention to a collaborative robot flexible assembly workstation, targeting the precise gripping and positioning of fragile composite materials (such as carbon fiber reinforced plastic structural parts) during the assembly process. In this scenario, the fragility of the workpiece requires the fixture to grip with minimal and uniform force, while maintaining extremely high sensitivity to minute deformations, surface damage (such as delamination and scratches), and slippage during the gripping process. Unlike the heavy-duty machining of Embodiment 1, this embodiment focuses more on gentle contact, damage prevention, and precise positioning.

[0084] In this application context, the configuration and focus of the multimodal sensor array module have been adjusted: Embedded miniature force sensor array: Its spacing and sampling frequency are similar to those of Embodiment 1, but it places greater emphasis on the perception of distributed force and uniformity. The sensors not only provide the total clamping force, but more importantly, they generate high-precision pressure distribution maps in real time for detecting localized stress concentrations or uneven clamping.

[0085] High-resolution miniature CMOS image sensor: Resolution increased to 1280×720 pixels, frame rate no less than 60fps, and equipped with polarizing filters and multispectral illumination (such as visible light + near infrared) to enhance the detection capability of fine scratches, fiber breaks, delamination, or cracks on composite material surfaces. The camera deployment position emphasizes comprehensive coverage of the gripping contact point and workpiece edges to detect initial damage.

[0086] Ultrasonic sensor arrays have replaced some MEMS microphone arrays because ultrasound is better suited for detecting internal defects or minute changes in contact states in environments free of cutting noise. Ultrasonic sensor arrays use high-frequency pulses to probe the surface or internal structure of a workpiece. By analyzing the attenuation of the echo signal and changes in propagation time, they can detect material delamination, voids, or micro-slippage at the interface with the fixture.

[0087] Non-contact infrared temperature sensors: fast response time and high accuracy. Deployed inside the fixture and near the clamping surface, they are mainly used to monitor minute frictional heat generation or material temperature changes during the contact process between the workpiece and the fixture, thus avoiding material damage caused by thermal stress.

[0088] The raw multimodal sensing data stream (force, high-resolution vision, ultrasonic echo, infrared temperature) is transmitted through an enhanced communication interface module. In addition to the TSN protocol, this module also integrates wireless communication protocol stacks such as Bluetooth Low Energy (BLE) or Wi-Fi 6 to adapt to the high mobility of collaborative robots. It supports real-time data interaction and status synchronization with the robot controller and host computer, and the synchronization period can be dynamically adjusted to 50 milliseconds according to task requirements.

[0089] The edge computing unit module still adopts a heterogeneous computing architecture (ARM Cortex-A + NPU), but it has been optimized at the algorithm level, focusing on gentle grasping scenarios: The NPU focuses more on running low-power, high-efficiency image processing algorithms and ultrasonic signal analysis models. The overall inference latency is still kept below 20 milliseconds.

[0090] Differences and detailed explanations of methodological steps: Step S110 (Data Acquisition): In this embodiment, the gripper performs a grasping action on the composite material workpiece under the guidance of a collaborative robot.

[0091] Force signal: A miniature force sensor array continuously monitors the pressure distribution on the fixture contact surface. In addition to the raw pressure value, special attention is paid to the uniformity of the pressure distribution (e.g., pressure standard deviation, ratio of maximum local pressure to average pressure) to ensure that local stress concentration is avoided.

[0092] Visual imaging: A high-resolution miniature CMOS image sensor continuously captures images of the workpiece surface and clamping area at a high frame rate during the pre-grip, gripping, and post-grip stages. Particularly during the gripping moment, the system triggers a high-speed continuous shooting mode to capture any potential initial surface damage. A polarizing filter is used to suppress surface reflections, making fine scratches easier to detect; near-infrared illumination can be used to penetrate shallow materials and detect subsurface defects.

[0093] Ultrasonic signals: The ultrasonic sensor array periodically emits and receives high-frequency ultrasonic waves. The system records the echo signal characteristics of each sensor node, including but not limited to: echo amplitude attenuation, propagation time variation, and spectral component shift. These characteristics are used to accurately determine the contact state between the fixture and the workpiece (whether there is complete contact, whether there are minute gaps) and potential internal defects (such as delamination).

[0094] Infrared temperature signal: Non-contact infrared temperature sensors monitor the surface temperature near the contact point of the fixture in real time. In addition to the absolute temperature value, the rate of temperature rise is also considered to identify instantaneous frictional heat generated by micro-slippage.

[0095] All data is also strictly timestamped and encapsulated, and transmitted through the communication interface module.

[0096] Step S120 (Data Preprocessing): The preprocessing procedure is similar to that in Example 1, but with enhancements for specific modalities: Force data: In addition to noise reduction and normalization, spatial filtering may also be performed to smooth the pressure distribution pattern and reduce the impact of local sensor failures on the overall assessment.

[0097] Visual image processing: In addition to traditional noise reduction, image enhancement (such as contrast stretching and edge sharpening) is performed to highlight the subtle textures and damage features of the composite material surface. Furthermore, geometric correction is performed to eliminate camera distortion and ensure the accuracy of deformation measurements.

[0098] Ultrasonic signals: Bandpass filtering is used to remove environmental noise in specific frequency bands, and envelope detection algorithms are applied to extract effective features of ultrasonic echo signals. For example, the instantaneous envelope of the signal is obtained through Hilbert transform for subsequent feature extraction.

[0099] Infrared temperature data: Employ more sophisticated outlier detection algorithms (such as statistical Z-score or box plot methods) to identify transient hotspots caused by localized minute frictions, rather than ambient temperature fluctuations.

[0100] Step S130 (Feature Extraction): The focus of feature extraction is entirely centered around gentle grasping and damage prevention: Clamping force fluctuation characteristic vector: In addition to statistical and frequency domain characteristics, it emphasizes pressure uniformity indicators (e.g., normalized standard deviation, ratio of local maximum pressure to average pressure) and contact area change rate, which are directly related to potential local stress concentration or gripping instability.

[0101] Visual deformation and damage feature encoding: Based on higher resolution images, a lightweight CNN is trained to recognize more refined features. Surface microtexture anomalies: Identify damage characteristics on the surface of composite materials, such as fiber breakage, resin matrix microcracks, and coating peeling.

[0102] Edge integrity: Detects whether the workpiece edge is chipped, burred, or deformed during the gripping process.

[0103] Optical flow analysis: Optical flow calculation is performed on continuous image frames to accurately detect the minute relative sliding speed and direction between the fixture and the workpiece, which is a key indicator for judging the stability of the gripping.

[0104] The encoded visual feature vectors may have a higher dimension to capture more details.

[0105] Ultrasonic anomaly feature vector: Echo attenuation and propagation time characteristics: This analyzes the degree of attenuation and propagation time differences of ultrasonic waves as they pass through contact interfaces or material interiors. Significant attenuation or abnormal propagation time may indicate poor contact or internal defects (such as delamination).

[0106] Spectral characteristics: Analyze the spectral changes of the echo signal. Abnormal frequency shifts or the appearance of specific harmonics may indicate material damage or changes in the state of the contact interface.

[0107] These features together constitute the ultrasonic anomaly feature vector.

[0108] Thermal stability feature vector: In addition to trend slope and steady-state deviation, special attention is paid to the detection of local hot spots. If the temperature rise rate of a certain area is much higher than that of the surrounding area and is synchronized with the grasping action, it is considered to be a potential micro-slip or frictional heat feature, rather than an ambient temperature change.

[0109] Step S140 (Feature Fusion): The multimodal feature fusion inference engine continues to employ a gated cross-attention mechanism. However, for the scenario in this embodiment, its training objective and modal weight allocation have been adjusted: Training objective: The fusion model focuses more on distinguishing between states such as "safe grasp", "about to slide", "minor damage" and "severe damage".

[0110] Dynamic Relevance Weighting: In fragile material scenarios, visual and ultrasonic modalities may gain higher dynamic weights in detecting damage and fine contact states. For example, when ultrasound detects internal voids, its weight increases significantly, allowing the system to identify potential problems even if the force change is not obvious. Gating units will more strictly suppress transient interference modes that may be caused by robot motion or environmental noise. The fused gripping state comprehensive representation vector will be able to more finely distinguish between different levels of damage and stability risks.

[0111] Step S150 (State Classification and Feedback Signal Generation): The pre-trained lightweight classifier determines the gripping state based on the new comprehensive representation vector, and its classification target tends to be: Secure gripping: Meets all safety standards for mechanics, vision, ultrasound, and thermal applications.

[0112] Potential slip risks: For example, optical flow analysis may detect micron-level relative motion, or a decrease in pressure uniformity.

[0113] Surface scratches / damage: Minor scratches or edge wear can be visually detected.

[0114] Internal defect risk: Abnormal ultrasonic echoes indicate possible delamination within the material.

[0115] If the system is determined to be unstable or at risk of damage, the generated feedback control signal will include: Clamping force fine-tuning: This typically involves reducing the clamping force or adjusting it to a more even distribution to prevent further damage. For example, it may indicate a reduction of 5 Newtons in clamping force while simultaneously adjusting the relative force of each jaw to achieve a more even distribution.

[0116] Posture fine-tuning angle: used to finely adjust the contact posture between the workpiece and the fixture to optimize the uniformity of force distribution or adjust the assembly position to ensure precise alignment with the robot's body coordinate system.

[0117] Robot motion trajectory adjustment commands: For assembly tasks, the robot may be instructed to "decelerate," "pause the current assembly step," "adjust the gripping posture," or "request human intervention." For example, if sliding is detected, the robot may be instructed to decelerate or reposition itself.

[0118] Step S160 (Closed-loop control): The adaptive feedback control module receives feedback signals and drives the fixture actuator.

[0119] Clamping force adjustment: The PID controller converts the adjustment amount into motor drive current, but the PID parameters here will focus more on smoothness and accuracy rather than rapid, large-range adjustments to avoid shock. The fixture may be equipped with a more sophisticated force control servo system capable of achieving Newton-level force output accuracy.

[0120] Attitude Compensation: In addition to attitude compensation by the fixture's end effector, the feedback control module also integrates more deeply with the collaborative robot controller. Attitude deviations or slippage trends detected by the fixture can be directly fed back to the robot's motion controller, allowing the robot itself to perform global attitude compensation or path replanning to achieve precise workpiece assembly and avoid collisions. The compensation accuracy reaches ±0.05 degrees to meet precision assembly requirements.

[0121] Strategy Switching: If the command is "Pause Assembly," the robot will immediately stop its current action and wait for operator intervention. If the command is "Adjust Robot Path," the robot will send new path points or motion parameters to the robot controller via the OPC UA protocol, achieving intelligent decision-making in human-robot collaboration.

[0122] In this embodiment, the entire system focuses on achieving "zero-damage gripping" and "high-precision assembly" of fragile materials. Through multimodal perception and edge intelligent decision-making, it significantly reduces the scrap rate in the precision assembly process and improves the operational flexibility and reliability of the collaborative robot.

[0123] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A multimodal sensing real-time edge computing system for an intelligent clamping device, characterized in that, It includes the following components: The multimodal sensor array module is used to simultaneously acquire four types of physical signals—force, vision, acoustics, and temperature—during the process of a fixture holding a workpiece, and output the raw multimodal sensing data stream. An edge computing unit module, deployed on the fixture body or a nearby control node, is used to perform real-time preprocessing, feature extraction and fusion inference on the original multimodal sensing data stream to generate clamping state evaluation results. An adaptive feedback control module is used to dynamically adjust the clamping force, attitude, or clamping strategy of the fixture based on the clamping state evaluation results, so as to maintain a stable clamping state of the workpiece during the processing. The communication interface module is used to realize low-latency data interaction between the multimodal sensor array module, the edge computing unit module and the adaptive feedback control module, and to support state synchronization with the host computer system.

2. The multimodal sensing real-time edge computing system for intelligent clamping equipment according to claim 1, characterized in that, The multimodal sensor array module includes an embedded micro force sensor array, a micro CMOS image sensor, a MEMS microphone array, and a thermocouple temperature sensor. The embedded micro force sensor array is distributed on the contact surface of the fixture at a spacing of 450nm, with a sampling frequency of not less than 1kHz. The micro CMOS image sensor has a resolution of not less than 640×480 and a frame rate of not less than 30fps. The MEMS microphone array operates in a frequency band covering 20Hz to 20kHz. The thermocouple temperature sensor has a response time of less than 100ms and a measurement accuracy of ±0.5℃.

3. The multimodal sensing real-time edge computing system for intelligent clamping equipment according to claim 1, characterized in that, The multimodal feature fusion inference engine in the edge computing unit module employs a gated cross-attention mechanism to calculate the query vector. Key vector AND value vector To achieve intermodal weight allocation, the attention weight calculation formula is as follows: in, Let be the dimension of the key vector. Generated from the current dominant modality features, and It is generated from other auxiliary modal features, and the contribution of low-confidence modes is dynamically suppressed by gating units.

4. The multimodal sensing real-time edge computing system for intelligent clamping equipment according to claim 1, characterized in that, The adaptive feedback control module generates a feedback control signal based on the clamping state evaluation result, and converts the clamping force adjustment into a motor drive current command through a PID controller. At the same time, it combines the data from the six-axis attitude sensor to perform spatial attitude compensation on the end effector of the fixture, with a compensation accuracy of ±0.1 degrees.

5. The multimodal sensing real-time edge computing system for intelligent clamping equipment according to claim 1, characterized in that, The communication interface module adopts a time-sensitive network protocol stack to ensure that the transmission jitter of multimodal data streams between modules inside the fixture is less than 50μs, and supports state synchronization with the host MES system via OPC UA protocol, with a synchronization period of 100ms.

6. A calculation method for a multimodal sensing real-time edge computing system applied to the intelligent clamping device according to any one of claims 1-5, characterized in that, Includes the following steps: Step S110: When the fixture performs the clamping action, the multimodal sensor array module simultaneously collects force signals, visual images, acoustic signals and temperature signals to form a four-channel raw multimodal sensing data stream. Step S120: Input the original multimodal sensing data stream into the edge computing unit module, and perform time alignment, noise suppression and normalization processing on the data of each channel to generate standardized multimodal sensing data; Step S130: Dynamic feature extraction based on a sliding window is performed on the force time-series signal in the standardized multimodal sensing data to obtain the clamping force fluctuation feature vector; local texture and deformation feature encoding based on a lightweight convolutional neural network is performed on the visual image to obtain the visual deformation feature vector; short-time Fourier transform and time-frequency energy distribution modeling are performed on the acoustic signal to obtain the acoustic anomaly feature vector; trend slope and steady-state deviation analysis are performed on the temperature signal to obtain the thermal stability feature vector. Step S140: Input the clamping force fluctuation feature vector, visual deformation feature vector, acoustic anomaly feature vector and thermal stability feature vector into the multimodal feature fusion inference engine, calculate the dynamic correlation weight between each modality through the gated cross attention mechanism, and generate a comprehensive characterization vector of clamping state by weighted fusion. Step S150: Based on the clamping state comprehensive representation vector, a pre-trained lightweight classifier is used to determine whether the current clamping state is in a stable range. If it is determined to be unstable, a feedback control signal containing clamping force adjustment, attitude fine-tuning angle and clamping strategy switching command is generated. In step S160, the feedback control signal is transmitted to the adaptive feedback control module to drive the fixture actuator to make real-time adjustments and complete closed-loop control.

7. The calculation method of the multimodal sensing real-time edge computing system for intelligent clamping equipment according to claim 6, characterized in that, The noise suppression in step S120 includes: applying a Kalman filter or wavelet transform to denoise the force data; using nonlocal mean or median filtering for the visual image; using spectral subtraction combined with an adaptive filter for the acoustic signal; and applying a moving average filter or exponential smoothing to the temperature data.

8. The calculation method of the multimodal sensing real-time edge computing system for intelligent clamping equipment according to claim 6, characterized in that, In step S130, the dynamic features extracted from the force sensing time-series signal include statistical features, time-domain features, and frequency-domain features. Features for encoding visual images include edges, corners, texture primitives, and local geometric deformations; features for modeling acoustic signals include specific frequency band energy, energy entropy, and Mel frequency cepstral coefficients; features for analyzing temperature signals include the slope of temperature changes and the deviation from the steady-state baseline.

9. The calculation method of the multimodal sensing real-time edge computing system for intelligent clamping equipment according to claim 6, characterized in that, The lightweight classifier adopts a depthwise separable convolutional structure, which includes three convolutional layers and one fully connected output layer. The number of parameters is controlled within 50KB, and it is deployed on the NPU of the edge computing unit.

10. The calculation method of the multimodal sensing real-time edge computing system for intelligent clamping equipment according to claim 6, characterized in that, In step S160, the attitude fine-tuning angle is converted into a driving command for the micro-actuator inside the end effector of the fixture through inverse kinematics calculation. Combined with the real-time feedback from the six-axis attitude sensor, closed-loop compensation control of the spatial attitude is realized.