Yarn breakage posture recognition method and system based on multi-modal perception

By synchronously acquiring multimodal signals of yarn breakage events using acoustic and tension sensors, constructing a tension residual map and performing in-depth inversion, the problems of limited illumination and cross-modal inconsistency in traditional yarn attitude recognition are solved, enabling accurate identification and reliable decision-making regarding yarn breakage attitude.

CN121542818BActive Publication Date: 2026-04-24DONGHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DONGHUA UNIV
Filing Date
2026-01-20
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Traditional yarn posture recognition methods struggle to reliably extract features of slender structures under complex lighting conditions. The tension signal has a low dimension and is inconsistent with the visual signal, resulting in large errors in yarn breakage posture recognition and poor cross-modal fusion performance.

Method used

Multimodal signals of yarn breakage events are acquired synchronously by acoustic and tension sensors to construct a tension residual spectrum. A multimodal deep inversion model is then used for time alignment and feature fusion, including acoustic branch, tension branch, cross-modal fusion layer and inversion head. The broken yarn attitude parameters and confidence level are output, and a hierarchical arbitration strategy is designed.

Benefits of technology

It achieves accurate and robust identification of yarn breakage posture under complex lighting conditions, improving the intelligence level and processing reliability of spinning production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121542818B_ABST
    Figure CN121542818B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of textiles, and discloses a yarn breakage posture recognition method and system based on multi-modal perception. The method comprises the following steps: synchronously collecting multi-modal signals of a yarn breakage event through an acoustic sensor and a tension sensor, constructing a joint two-dimensional representation atlas containing an acoustic spectrum and a tension attenuation curve, and recording the tension residual atlas; inputting the tension residual atlas into a pre-trained posture inversion model to obtain breakage posture parameters and corresponding confidence; and arbitrating according to the confidence output by the inversion head. The application can effectively solve the problems of limited visual perception under complex illumination, low tension signal dimension and inconsistent cross-modal in the traditional method, and realizes accurate and robust recognition of the spatial posture of the broken yarn.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of textile technology, specifically to a method and system for yarn breakage posture recognition based on multimodal perception. Background Technology

[0002] Yarn breakage is a common anomaly during high-speed operation of spinning equipment. To achieve intelligent processing such as automatic splicing and automatic yarn clearing, it is necessary to identify the spatial posture of the broken yarn in real time, such as the orientation of the broken end, the direction of curling, the bending angle, or the attachment position. Traditional yarn posture recognition mainly relies on single visual information, which has the following main drawbacks:

[0003] First, the lighting conditions in spinning workshops are complex and variable. Components such as yarn bobbins, guide rods, and metal frames can cause strong reflections, shadows, and localized overexposure, making it difficult for vision-based detection methods to reliably extract the fine structural features of yarn. Especially when the yarn is broken, it is often in a state of low tension, localized adhesion, or entanglement, and its minute bends, curls, or twists are almost indistinguishable in the image, leading to a significant increase in posture detection errors.

[0004] Secondly, existing tension monitoring methods typically only provide low-dimensional mechanical signals, such as coarse-grained signals with significant delays output by mechanical tension components or micro-force sensors. These signals are difficult to directly map to the specific spatial orientation of the yarn break. Furthermore, inherent temporal misalignment, amplitude inconsistencies, and differences in physical semantics exist between tension signals and visual signals, rendering traditional multimodal fusion methods ineffective in practical applications.

[0005] Furthermore, existing technologies, when dealing with yarn breakage scenarios, often assume that different modalities (such as visual and tension) are synchronously consistent during the event. However, this consistency does not hold true in reality: for example, after a yarn breaks, a sudden change in the tension signal will occur first, while the visually stable broken end may be delayed due to inertial oscillation; or, when the yarn adheres to the surface of the yarn guide rod, it is almost invisible visually, but the tension still exhibits abnormal fluctuations. These cross-modal inconsistencies prevent rule-based fusion algorithms from accurately inferring the attitude.

[0006] Therefore, how to fully explore the subtle relationship between optical information and tension changes in yarn breakage scenarios in order to solve the problem of attitude misjudgment caused by cross-modal inconsistency has become a technical bottleneck that urgently needs to be overcome. Summary of the Invention

[0007] The purpose of this invention is to provide a method and system for yarn breakage posture recognition based on multimodal perception, so as to solve the problems mentioned in the background art.

[0008] This invention provides a method for yarn breakage posture recognition based on multimodal perception, comprising the following steps:

[0009] Step S1: Multimodal signals of yarn breakage events are synchronously acquired by acoustic sensors and tension sensors, and time alignment is performed based on the breakage trigger time to construct a joint two-dimensional characterization map containing acoustic spectrum and tension decay curve, denoted as tension residual map; wherein, the horizontal axis of the tension residual map is time, and the vertical axis contains acoustic frequency band energy sequence and tension decay sequence.

[0010] Step S2: Input the tension residual map into the pre-trained attitude inversion model; the attitude inversion model includes an acoustic branch, a tension branch, a cross-modal fusion layer, and an inversion head, wherein the acoustic branch is used to extract time-frequency texture features, the tension branch is used to extract tension attenuation physical features, the cross-modal fusion layer realizes inter-modal feature interaction through an attention mechanism, and the inversion head outputs decapitation attitude parameters and corresponding confidence scores, wherein the decapitation attitude parameters include orientation angle, tilt angle, and category label;

[0011] Step S3: Arbitrate based on the confidence level output by the inversion head: if the confidence level is higher than the first threshold, directly output the decapitation posture parameters and trigger the actuator action; if the confidence level is between the first threshold and the second threshold, call the vision module for verification; if the confidence level is lower than the second threshold, record the event and start the manual verification process.

[0012] The present invention also provides a yarn breakage posture recognition system based on multimodal perception, the system comprising:

[0013] Joint characterization map construction module: Multimodal signals of yarn breakage events are synchronously acquired by acoustic sensors and tension sensors, and time alignment is performed based on the breakage trigger time to construct a joint two-dimensional characterization map containing acoustic spectrum and tension decay curve, denoted as tension residual map; wherein, the horizontal axis of the tension residual map is time, and the vertical axis contains acoustic frequency band energy sequence and tension decay sequence;

[0014] Attitude inversion module: The tension residual map is input into the pre-trained attitude inversion model; the attitude inversion model includes an acoustic branch, a tension branch, a cross-modal fusion layer and an inversion head, wherein the acoustic branch is used to extract time-frequency texture features, the tension branch is used to extract tension attenuation physical features, the cross-modal fusion layer realizes inter-modal feature interaction through an attention mechanism, and the inversion head outputs decapitation attitude parameters and corresponding confidence scores, wherein the decapitation attitude parameters include orientation angle, tilt angle and category label;

[0015] Arbitration and Execution Module: Arbitrates based on the confidence level output by the inversion head: If the confidence level is higher than the first threshold, the decapitation posture parameters are directly output and the execution mechanism is triggered; if the confidence level is between the first and second thresholds, the vision module is called for verification; if the confidence level is lower than the second threshold, the event is recorded and the manual verification process is initiated.

[0016] This invention achieves high-precision synchronous alignment and fusion of acoustic and tension signals at the moment of yarn breakage, constructs a tension residual map as an attitude fingerprint, and innovatively designs a multimodal deep inversion model. This model effectively solves the problems of limited visual perception under complex lighting conditions, low tension signal dimension, and cross-modal inconsistency in traditional methods. It realizes accurate and robust identification of the spatial attitude (direction angle, tilt angle, category) of broken yarn, and realizes hierarchical intelligent decision-making and automatic action triggering based on confidence level, which can significantly improve the intelligence level and processing reliability of spinning production. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a yarn breakage posture recognition method based on multimodal perception disclosed in an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram of the attitude inversion model disclosed in an embodiment of the present invention;

[0019] Figure 3 This is another structural schematic diagram of the attitude inversion model disclosed in the embodiments of the present invention;

[0020] Figure 4 This is a schematic diagram of a yarn breakage posture recognition system based on multimodal perception disclosed in an embodiment of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figure 1 This invention provides a method for yarn breakage posture recognition based on multimodal perception, comprising the following steps:

[0023] Step S1: Multimodal signals of yarn breakage events are synchronously acquired by acoustic sensors and tension sensors, and time alignment is performed based on the breakage trigger time to construct a joint two-dimensional characterization map containing acoustic spectrum and tension decay curve, denoted as tension residual map; wherein, the horizontal axis of the tension residual map is time, and the vertical axis contains acoustic frequency band energy sequence and tension decay sequence.

[0024] Yarn breakage is a common anomaly in high-speed spinning equipment. To achieve intelligent processing such as automatic splicing and yarn clearing, the primary task is to acquire effective data reflecting the physical state at the moment of breakage. Traditional single-vision or tension monitoring methods, in complex industrial environments, struggle to accurately perceive the subtle, transient spatial posture of the yarn after breakage due to issues such as light interference, low-dimensional signals, and delays. Therefore, this invention simultaneously acquires signals from both acoustic and tension modes.

[0025] Specifically, acoustic sensors (such as industrial microphones) and tension sensors deployed on the spinning equipment are used to simultaneously monitor and listen to the yarn's operating status. When the yarn breaks, the instantaneous fiber breakage generates a unique sound wave signal, while the tension drops sharply from its operating value and exhibits a specific attenuation pattern. These two signals are closely correlated in the time domain but have a slight delay.

[0026] First, the accurate fracture trigger moment is detected and determined, which can be based on the point of abrupt change in acoustic signal energy most sensitive to the fracture event. Then, using this fracture trigger moment as the zero point, the continuously acquired buffered data of the acoustic and tension signals are truncated and aligned to eliminate the timing deviations caused by physical propagation and sensor response.

[0027] Based on the aligned synchronization data segment, short-time Fourier transform and other time-frequency analyses are performed on the acoustic signal to obtain an acoustic time-spectrum reflecting the changes in its frequency components over time. Attenuation curve fitting is performed on the tension signal to extract its attenuation time constant, peak value, initial slope, and other physical parameter sequences. Finally, the acoustic time-spectrum (representing multiple acoustic frequency band energy sequences on the vertical axis) and the tension attenuation sequence are aligned and spliced ​​along a unified time axis (horizontal axis) to form a joint two-dimensional representation map integrating acoustic and tension information, namely, the tension residual map. It can be understood that this tension residual map serves as the physical information carrier for subsequent high-precision attitude inversion.

[0028] Step S2: Input the tension residual map into the pre-trained attitude inversion model; the attitude inversion model includes an acoustic branch, a tension branch, a cross-modal fusion layer, and an inversion head, wherein the acoustic branch is used to extract time-frequency texture features, the tension branch is used to extract tension attenuation physical features, the cross-modal fusion layer realizes inter-modal feature interaction through an attention mechanism, and the inversion head outputs decapitation attitude parameters and corresponding confidence scores, wherein the decapitation attitude parameters include orientation angle, tilt angle, and category label;

[0029] The constructed tension residual map contains fingerprint information of the decapitation posture, which needs to be used for deep feature mining and inversion using a deep learning model. For this purpose, this step inputs the tension residual map into a pre-trained posture inversion model, which employs a multi-branch neural network architecture. (See [link to relevant documentation]). Figure 2 It mainly includes the following four functional modules:

[0030] Acoustic Branch: This branch is mainly responsible for processing the acoustic time-frequency part of the tension residual spectrum. Specifically, it is used to automatically extract deep time-frequency texture features from the time-frequency image that can characterize events such as breakage transients, yarn swing, or collisions with machine parts.

[0031] Tension Branch: This branch is mainly responsible for processing the tension decay sequence in the tension residual spectrum. Specifically, it is used to capture the physical characteristics of tension decay, such as the dynamic decay and oscillation mode of the tension signal after breakage. These characteristics are directly related to the physical state of the yarn, such as the breakage direction and the shrinkage force.

[0032] Cross-modal fusion layer: Given the potential temporal and semantic inconsistencies (i.e., cross-modal inconsistencies) between acoustic and tension signals in decapitation events, simple feature splicing yields limited results. Therefore, this invention introduces a cross-attention mechanism. This layer uses the output of the tension branch as the key and value, and the output of the acoustic branch as the query. This allows the physical semantics of the tension features to actively guide and focus attention on relevant time-frequency components in the acoustic features, thereby achieving refined intermodal feature interaction, effectively fusing complementary information, and suppressing conflicts.

[0033] Inversion Head: Used to receive high-level feature representations after cross-modal fusion, and to perform nonlinear transformation and mapping through a series of fully connected layers. Its output includes two parts: (1) head-breaking posture parameters, specifically including continuous-valued spatial orientation angles and tilt angles (used for precise control of the robotic arm), and discrete-valued head-breaking orientation category labels (such as up-slant, down-slant, side-slant, etc.); (2) the confidence level corresponding to each predicted head-breaking posture parameter, which can be generated based on techniques such as heteroscedastic regression, and is used to quantify the self-evaluation of the posture inversion model on the reliability of the current prediction results.

[0034] Step S3: Arbitrate based on the confidence level output by the inversion head: if the confidence level is higher than the first threshold, directly output the decapitation posture parameters and trigger the actuator action; if the confidence level is between the first threshold and the second threshold, call the vision module for verification; if the confidence level is lower than the second threshold, record the event and start the manual verification process.

[0035] Considering that attitude inversion models may encounter extreme situations in complex and ever-changing production environments where training data is not fully covered, directly trusting all prediction results may be risky. Therefore, this invention designs and adopts a confidence-based hierarchical arbitration strategy to achieve an optimal balance between reliability, efficiency, and cost.

[0036] Specifically, the decapitation posture parameters and their confidence levels output by the inversion head are acquired in real time. Two confidence thresholds are preset (first threshold > second threshold), forming three decision intervals:

[0037] If the output confidence level is higher than the first threshold, it indicates that the attitude inversion model is very confident in the current recognition result. At this time, the attitude result is directly adopted, and subsequent actuators (such as joint robotic arms, alarms, or shutdown devices) are immediately triggered to perform corresponding actions, achieving rapid automatic processing and ensuring production efficiency.

[0038] If the confidence level falls between the first and second thresholds, it indicates that the attitude inversion model's judgment has some uncertainty. To avoid erroneous operations, the action is not executed directly; instead, the vision module is invoked for verification. The vision module (such as an auxiliary industrial camera) captures an image of the current decapitation scene and uses the image information to perform secondary verification or correction on the attitude inversion model's prediction results, forming a final decision before triggering the action. This significantly improves the system's robustness without significantly affecting efficiency.

[0039] If the confidence level is below the second threshold, it indicates that the current situation may be very rare or complex, and neither the pose inversion model nor visual assistance can reliably determine it. Therefore, the complete multimodal data and prediction results of the event are recorded and marked as requiring manual review. Furthermore, these samples can be used for subsequent expert analysis and incremental model learning, driving the continuous evolution of the system.

[0040] This invention achieves high-precision synchronous alignment and fusion of acoustic and tension signals at the moment of yarn breakage, constructs a tension residual map as an attitude fingerprint, and innovatively designs a multimodal deep inversion model. This model effectively solves the problems of limited visual perception under complex lighting conditions, low tension signal dimension, and cross-modal inconsistency in traditional methods. It realizes accurate and robust identification of the spatial attitude (direction angle, tilt angle, category) of broken yarn, and realizes hierarchical intelligent decision-making and automatic action triggering based on confidence level, which can significantly improve the intelligence level and processing reliability of spinning production.

[0041] As an example, before step S1, the method further includes: real-time monitoring of the yarn running status, and when the change in acoustic signal energy or tension signal exceeds a preset threshold, determining it as a breakage event and determining the breakage trigger time.

[0042] During the high-speed, continuous operation of spinning equipment, yarn breakage can occur at any time. This invention utilizes acoustic and tension sensors deployed on the production line to continuously and in real-time acquire and monitor the yarn's operating status. In practice, the audio stream acquired by the acoustic sensors is continuously analyzed, and its short-time energy is calculated; simultaneously, the signal value output by the tension sensors is continuously read.

[0043] Meanwhile, to accurately capture the moment of breakage, acoustic energy thresholds and tension change thresholds are set for the two signal characteristics mentioned above. The acoustic energy threshold is set based on statistical analysis of the sound characteristics of a large number of historical breakage events. At the moment of breakage, yarn fibers generate a concentrated (typically 2-10kHz) instantaneous acoustic pulse with a sudden increase in energy. When the short-term energy of the acoustic signal calculated in real time exceeds this threshold, it indicates that a breakage may have occurred. The tension change threshold is set based on the normal spinning tension fluctuation range. When yarn breaks, its tension drops sharply from the working value, with the slope or absolute change far exceeding normal fluctuations. When the change in tension signal (such as the difference between adjacent sampling points or the decrease within a certain time window) exceeds this threshold, it also indicates that a breakage may have occurred.

[0044] It should be noted that under certain extreme operating conditions, such as when ambient noise can significantly mask the breakage sound, a sudden change in the tension signal can serve as a reliable trigger. Conversely, when the yarn breaks due to extreme slack, resulting in minimal tension change, the acoustic signal becomes a more sensitive indicator. This invention employs dual-path triggering logic, using either acoustic or tension signals, ensuring that breakage events can be captured with high sensitivity and reliability under various complex production environments, thereby minimizing the missed detection rate.

[0045] If the signal from any of the aforementioned sensors exceeds its corresponding preset threshold, a breakage event is determined to have occurred, and this exceeding moment is precisely recorded as the global breakage trigger moment t0. It can be understood that this breakage trigger moment t0 is the sole time reference for the entire subsequent multimodal data processing flow; all signal interception, alignment, and map construction operations in step S1 are centered around this moment.

[0046] As an example, time alignment is performed based on the fracture triggering moment to construct a joint two-dimensional characterization map containing the acoustic spectrum and tension decay curve, including:

[0047] Step S11: Taking the fracture triggering time as the center, extract acoustic signal segments and tension signal segments of preset time lengths from the continuous acquisition buffers of the acoustic sensor and tension sensor, respectively.

[0048] In this step, after determining the precise fracture trigger time t0 in step S1, the two independent data streams are immediately operated on using this time as the absolute center. The continuous acquisition buffers (e.g., circular buffers) of the acoustic and tension sensors are continuously maintained, storing the latest historical data.

[0049] From the acoustic buffer, extract audio sampling data of a preset length (e.g., a total of 100 milliseconds, i.e., t0 ± 50 ms) centered at t0. Understandably, the length of this time window is pre-designed to fully capture the sound transients generated at the moment of break (typically lasting from a few milliseconds to tens of milliseconds) and their possible brief reverberations or reflections, while avoiding the inclusion of too much irrelevant environmental background noise.

[0050] From the tension buffer, extract tension or vibration sampling data of a preset length (e.g., a total of 400 milliseconds, i.e., t0 ± 200 ms) centered at t0. Understandably, the longer time window is to fully capture the complete dynamic process of tension from its working value before fracture, through a sudden drop, to decay to a steady state (or the generation of low-frequency oscillations).

[0051] It should be noted that the preset lengths of the two signal segments are different because sound events are short, while the mechanical response process is relatively slower and longer. This step, through this targeted truncation, can achieve preliminary data regularization while preserving the complete event information of each segment, which is beneficial for subsequent fine alignment and fusion.

[0052] Step S12: Perform a short-time Fourier transform on the acoustic signal segment to generate an acoustic time-spectrum; and perform attenuation curve fitting on the tension signal segment to extract its attenuation feature sequence as a function of time.

[0053] In this step, a Short-Time Fourier Transform (STFT) is applied to the extracted acoustic signal segment. STFT is a classic time-frequency analysis method. Its principle is to divide a long signal into many short, overlapping time segments (windows), and then perform a Fourier transform on each segment to obtain the pattern of signal frequency components changing over time. Through this processing, the original amplitude-time one-dimensional waveform is converted into a frequency-time-energy two-dimensional time-spectrum diagram, i.e., the acoustic time-spectrum diagram mentioned above. In this spectrum, the acoustic signature characteristics, such as the broadband or specific frequency band (e.g., 2-10kHz) energy bursts caused by yarn breakage and possible harmonic structures, are clearly presented.

[0054] Simultaneously, attenuation curve fitting and analysis were performed on the extracted tension signal segments. After yarn breakage, the tension release process can be approximated using an exponential decay model (e.g., ...). Where A is the attenuation amplitude, representing the dynamic change in tension from its peak at the moment of fracture to its steady state; τ is the attenuation time constant; C is the steady-state offset or baseline value, representing the steady-state value after the tension attenuation process ends; and t is the relative time, relative to the fracture trigger moment. The original tension-time curve is described by the origin (where is the time elapsed since the start of the decay process). This step uses algorithms such as nonlinear least squares to fit the original tension-time curve to the aforementioned exponential decay model.

[0055] By performing a fitting operation, not only can a smooth fitting curve be generated, but also a sequence of key physical parameters can be extracted, mainly including: the decay time constant. The parameters reflect the rate of tension dissipation and are related to the yarn material and the constraint of the break point; the peak value and initial slope characterize the impact intensity at the moment of breakage; the residual or oscillation mode, the difference between the fitted curve and the original data, characterizes complex behaviors such as yarn swaying and entanglement. These physical parameters are organized into a decay characteristic sequence in chronological order. This sequence is highly abstract and has a clear physical meaning, directly related to the changes in the mechanical state of the yarn.

[0056] Step S13: Align the acoustic time-spectrum map with the attenuation feature sequence on the time axis and stitch them together along the feature dimension to generate the joint two-dimensional characterization map.

[0057] The acoustic time spectrum (horizontal axis is time) and tension decay feature sequence (index also represents time) generated in step S12 are both truncated with t0 as the reference in step S11. However, due to the small offset that may be introduced by the sensor sampling rate and initial processing, they need to be finally aligned to ensure that the horizontal coordinates (time axis) of the two feature maps completely coincide at point t0 and to align the time scale within the entire analysis window.

[0058] The aligned acoustic time-spectrum and tension decay feature sequence are considered as two two-dimensional matrices, where the acoustic time-spectrum is [number of frequency channels × time step], and the decay feature sequence is [number of feature parameters × time step]. They are then concatenated vertically along the feature dimension (i.e., the vertical axis). Specifically, the energy values ​​of each frequency channel in the acoustic time-spectrum constitute the upper half of the joint two-dimensional characterization map, and the feature parameter values ​​of each feature parameter in the tension decay feature sequence constitute the lower half. The horizontal axis (time axis) remains consistent and continuous after concatenation.

[0059] The final generated matrix is ​​the tension residual map, a two-dimensional array with multiple rows (acoustic frequency bands, tension features) and multiple columns (time points). It can be understood that this tension residual map retains the richness of the original multimodal information, including both time-frequency texture and physical parameters, and the acoustic response and mechanical state at any given time point correspond vertically within the same column of the matrix. Furthermore, the aforementioned tension residual map unifies complex multi-sensor, heterogeneous signals into a regular, image-like structure that can be directly processed by models such as convolutional networks.

[0060] As an example, the acoustic branch uses a network structure containing two-dimensional convolutional blocks and temporal convolutional modules to perform deep feature extraction on the acoustic time-frequency part of the tension residual spectrum to obtain the time-frequency texture features;

[0061] The tension branch extracts temporal features from the tension decay sequence portion of the tension residual map through a network structure containing one-dimensional convolutional or gated recurrent units to obtain the physical features of tension decay.

[0062] The cross-modal fusion layer employs a cross-attention mechanism, using the tension attenuation physical features as keys and values, and the time-frequency texture features as queries, to perform feature reweighting and fusion, thereby obtaining the intermodal feature interactions.

[0063] The inversion head maps the intermodal feature interactions into continuous orientation angles, tilt angles, and class labels through a fully connected layer, thus obtaining the decapitation posture parameters; and, based on heteroscedastic regression, it synchronously outputs the confidence scores of each predicted decapitation posture parameter.

[0064] In this embodiment, the structure and function of each component of the attitude inversion model are described in detail below:

[0065] The input to the acoustic branch is the upper half of the tension residual spectrum, i.e., the acoustic time-spectrum. To extract effective features from this two-dimensional spatiotemporal signal, a composite neural network structure is employed for this acoustic branch, as follows:

[0066] Two-dimensional convolutional blocks are composed of multiple stacked two-dimensional convolutional layers, activation function layers, and pooling layers. The two-dimensional convolutional kernels slide on the frequency-time plane of the time-spectrum, which can efficiently capture the texture patterns, edges, and energy accumulation areas of local regions (such as specific frequency bands within a specific time period). These patterns correspond to different acoustic features at the moment of fracture, such as sharp impulses, broadband noise, or specific resonances.

[0067] Temporal Convolution Module: Following the 2D convolution block, this module captures the long-term dependencies of features over time. It can employ 1D temporal convolution or a temporal convolutional network (TCN) to model feature sequences along the time axis, thereby understanding the complete dynamic process of acoustic events from occurrence to evolution to dissipation.

[0068] Through the cascaded processing of the above structure, the acoustic branch can abstract deep, discriminative time-frequency texture features from the original time-frequency spectrum. These features are key clues for inferring the yarn breakage mode and initial motion direction.

[0069] The input to the tension branch is the lower half of the tension residual spectrum, namely the tension decay characteristic sequence. This tension decay characteristic sequence is essentially composed of multiple physical parameters (such as the decay constant). The trajectory of changes in peak value, slope, etc. over time is typical of time series data. Therefore, this branch adopts a network structure that is adept at processing sequence data, as follows:

[0070] One-dimensional convolutional networks: By sliding one-dimensional convolutional kernels across a time series, it is possible to extract the combined relationships and short-term patterns between multiple feature parameters within a local time window, such as the steepness of the decay initiation phase or the period of oscillation.

[0071] Gated Recurrent Unit (GRU): As a variant of Recurrent Neural Network (RNN), GRU can effectively capture long-term dependencies in time series through its internal gating mechanism (update gate and reset gate). It is well-suited for modeling the entire continuous process of tension decaying from peak to stability and for memorizing important states in the process.

[0072] Understandably, regardless of whether a one-dimensional convolutional network or a GRU is used, the purpose of tension branch is to extract tension decay physical features that can characterize the essence of fracture mechanics from the decay feature sequence. These features directly reflect the dynamic behavior and stress state of the yarn after it breaks.

[0073] For cross-modal fusion layers, considering that time-frequency texture features and tension attenuation physical features originate from different sensors, and may experience information mismatch (i.e., cross-modal inconsistency) due to environmental interference or the characteristics of the event itself, simple feature stitching often yields poor results. Therefore, this invention introduces a cross-attention mechanism to achieve refined fusion.

[0074] Specifically, the cross-modal fusion layer uses the tension attenuation physical features output by the tension branch as keys and values, and the time-frequency texture features output by the acoustic branch as queries. During attention calculation, each acoustic feature query queries all tension feature keys to calculate attention weights. These weights reflect the correlation between the acoustic feature and each tension feature.

[0075] Then, these weights are used to perform a weighted summation of the tension feature values, resulting in a context vector that incorporates tension information. This vector is then combined with the original acoustic query features. Furthermore, the explicit physical semantics contained in the tension features (such as decay rate and impact intensity) guide and recalibrate the acoustic features, emphasizing the time-frequency components that correspond to the current mechanical state and suppressing inconsistencies that may be caused by noise or irrelevant reflections. In this way, intermodal feature interaction can be achieved, generating a unified feature representation that includes both acoustic details and is constrained by physical laws.

[0076] The unified high-level features obtained after cross-modal fusion layer processing are fed into the inversion head. The inversion head consists of several fully connected layers and is responsible for performing the final regression and classification tasks. It maps the fused features into continuous numerical outputs (i.e., regression values ​​of orientation angle and tilt angle) for precise control of the actuators; at the same time, it maps them into discrete classification outputs (i.e., orientation category labels, such as "upward", "downward", "lateral") for rapid qualitative judgment.

[0077] Furthermore, to assess the reliability of each prediction, the attitude inversion model integrates heteroscedastic regression. Unlike ordinary regression, which only predicts the mean, heteroscedastic regression predicts both the mean and variance of the output value (reflecting uncertainty). In this attitude inversion model, in addition to outputting the predicted values ​​of the attitude parameters, the inversion head also simultaneously outputs a confidence score characterizing the variance of each predicted value. A low confidence score means that the model considers the prediction to have high uncertainty under the current input.

[0078] This implementation method accurately extracts the time-frequency texture features of the fracture acoustic pattern through the acoustic branch and effectively analyzes the physical temporal features of tension attenuation through the tension branch. It also innovatively adopts a cross-attention fusion mechanism that guides the acoustic features with the physical features of tension, which can effectively overcome the cross-modal inconsistency problem of traditional multimodal methods in yarn breakage scenarios. This enables high-precision and interpretable inversion of the breakage posture (angle, tilt angle, category) and generates quantitative confidence scores simultaneously, which can provide accurate support for subsequent reliable hierarchical decision-making.

[0079] As an example, please refer to Figure 3 The attitude inversion model further includes a viewpoint-aware transformation module, which generates cross-viewpoint attention parameters related to intermodal feature interactions through multi-level feature abstraction and attention distillation mechanisms; correspondingly, the cross-modal fusion layer obtains the intermodal feature interactions based on the cross-viewpoint attention parameters and the cross-attention mechanism.

[0080] As an example, the viewpoint-aware transformation module generates cross-viewpoint attention parameters related to intermodal feature interactions through multi-level feature abstraction and attention distillation mechanisms, including:

[0081] The time-frequency texture features and the tension attenuation physical features are projected onto a high-dimensional shared semantic space to obtain aligned acoustic semantic features and tension semantic features.

[0082] Cross-attention calculation is performed on the acoustic semantic features and tension semantic features to obtain an initial attention map; multi-scale feature extraction is performed on the initial attention map through a convolutional neural network to generate attention distillation features with spatial awareness.

[0083] The attention distillation features are input into a parameter prediction network, which includes an adaptive gating mechanism to dynamically adjust the information flow based on the correlation between features, and finally outputs the cross-view attention parameters.

[0084] In this embodiment, the present invention further introduces a viewpoint-aware transformation module into the attitude inversion model. This module aims to dynamically generate cross-viewpoint attention parameters to guide cross-modal fusion, ensuring that the effective association between modes (acoustic and tension) is not a fixed template but varies with the specific spatial attitude, i.e., the viewpoint, of the severed head. This viewpoint-aware transformation module automatically senses the viewpoint corresponding to the current input through a multi-level processing flow and generates appropriate fusion guidance parameters.

[0085] First, the viewpoint-aware transformation module receives time-frequency texture features from the acoustic branch and tension attenuation physical features from the tension branch. Although both have been deeply extracted, their feature spaces still exhibit heterogeneity. To establish a comparable basis, this invention utilizes a projection network (such as a fully connected layer) to map these two types of features to a high-dimensional shared semantic space. In this space, the acoustic semantic features not only contain the texture information of the sound but are also endowed with abstract semantics related to the physical state; the tension semantic features not only contain mechanical attenuation information but also incorporate the abstract semantics of the possible sound-generating background. Understandably, this mapping aligns features from different physical sources at the same semantic level.

[0086] Next, the viewpoint perception transformation module performs cross-attention calculation on the aligned acoustic semantic features and tension semantic features. This calculation uses one side as the query and the other as the key-value pair to generate an initial attention map that reflects the strength of the association between the two sides among all elements in the semantic space.

[0087] The initial attention map itself contains rich modal relationship patterns, but may contain redundancy or noise. To address this, a lightweight convolutional neural network (CNN) is introduced to extract features from the initial attention map at multiple scales. Convolutional kernels of different scales can capture various relational structures, ranging from fine-grained local relationships to global overall patterns, thereby generating a more discriminative and robust attention distillation feature. Understandably, this feature has distilled out the core spatial patterns of intermodal interactions.

[0088] Finally, the distilled attention features are input into the parameter prediction network. The key to this network lies in its built-in adaptive gating system, which uses weights calculated based on the features themselves to control the information flow through different channels or paths. This adaptive gating system dynamically adjusts the information flow paths within the network based on the strength of correlation reflected within the input attention-distilled features. Specifically, for highly correlated feature combinations, the gating mechanism opens important paths to reinforce learning; for weakly correlated or irrelevant parts, it suppresses information flow to prevent interference.

[0089] After nonlinear transformation and focusing by this adaptive gating network, a set of structured, dimensionally accurate cross-view attention parameters is finally output. It can be understood that this set of cross-view attention parameters is a quantitative answer to the question of how to best utilize tension features to guide and modulate acoustic features under a specific decapitation posture.

[0090] The cross-view attention parameters generated above are fed to the cross-modal fusion layer in real time. As a result, the cross-modal fusion layer no longer performs only the standard, fixed-parameter cross-attention calculation, but incorporates the cross-view attention parameters as dynamic modulation factors into the calculation process.

[0091] The perspective-aware transformation module introduced in this invention aligns and distills attention by mapping acoustic features and tension features to a high-dimensional shared semantic space, and generates dynamic cross-perspective attention parameters using an adaptive gating mechanism. This enables the cross-modal fusion layer to adaptively adjust the fusion strategy according to the specific yarn breakage posture scenario (perspective), effectively solving the problem that traditional fixed fusion modes have difficulty handling cross-modal inconsistencies in complex and variable yarn breakage scenarios. As a result, it significantly improves the accuracy of posture inversion, the model's scene adaptability, and robustness to edge cases.

[0092] As an example, the cross-modal fusion layer obtains the intermodal feature interactions based on the cross-view attention parameters and the cross-attention mechanism, including:

[0093] The cross-perspective attention parameter is decomposed into query correction factor, key correction factor, and value correction factor;

[0094] The query correction factor is subjected to a Hadamard product with the query vector in the cross-attention mechanism to generate an enhanced query vector; the key correction factor and value correction factor are respectively tensor-concatenated with the key vector and value vector in the cross-attention mechanism to form a multi-dimensional interactive representation.

[0095] Based on the enhanced query vector and the multi-dimensional interaction representation, multi-level attention computation is performed to obtain a fusion feature representation with cross-perspective perception capability.

[0096] This implementation utilizes a lightweight linear transformation layer to handle cross-view attention parameters. Decompose: , , .in For learnable weight matrix, Given the bias vector, the query correction factor is obtained. , bond correction factor Sum value correction factor All dimensions .

[0097] Let the query vector in the original cross-attention mechanism be... (Derived from acoustic features), the key vector is (Derived from tension features), the value vector is (and (Same origin). The fusion calculation is performed as follows:

[0098] (1) Generate enhanced query vector: For the query vector Apply a correction factor to each row (i.e., each query position). : ,in This represents the Hadamard product (element-by-element multiplication). Broadcast to Same shape This yields the enhanced query vector. This operation allows each feature dimension of the query to be adaptively scaled according to the current perspective.

[0099] (2) Constructing a multi-dimensional interactive representation: Incorporating the key correction factor With value correction factor Compared with the original key vector respectively Sum value vector Concatenate along the feature dimension:

[0100]

[0101]

[0102] Among them, the correction factor and Broadcast to shape Then the components are concatenated. This yields the extended key representation. Sum value representation Its feature dimensions are expanded to This integrates original physical characteristics with perspective guidance information.

[0103] (3) Perform multi-level attention calculation: Use the following two-level cascaded attention modules:

[0104] The first layer of attention is calculated as follows:

[0105] , where Softmax is the activation function of the final output layer.

[0106] This layer uses enhanced queries. With original bond Calculate the weights and apply them to the original values. This achieves initial information filtering guided by perspective. The first layer output is then processed through residual connections and layer normalization (LayerNorm). .

[0107] Second-level attention calculation: As a query, with the expanded key Sum Perform the calculation:

[0108] .

[0109] Understandably, this layer utilizes information that expands upon the viewpoint. and To enable deeper interaction.

[0110] (4) The output of the second layer attention is transformed and mapped nonlinearly through a feedforward neural network (FFN): ,in , x represents the feature vector or feature matrix input to the FFN layer, b1 is the bias vector of the first linear transformation, and b2 is the bias vector of the second linear transformation.

[0111] Finally, residual connectivity and layer normalization are applied again to obtain a fused feature representation with cross-perspective perception capabilities. : .

[0112] Through the specific tensor operations and network forward process described above, the cross-modal fusion layer realizes the injection of high-level perspective perception parameters layer by layer and element by element into the core computation of the attention mechanism, enabling the feature interaction process to achieve dynamic and adaptive adjustment closely related to the pose scene, thereby helping to improve the modeling accuracy of the pose inversion model for complex multimodal relationships.

[0113] Please see Figure 4 This invention also provides a yarn breakage posture recognition system 200 based on multimodal perception, the system comprising:

[0114] Joint characterization map construction module 201: Simultaneously acquires multimodal signals of yarn breakage events through acoustic sensors and tension sensors, and performs time alignment based on the breakage trigger time to construct a joint two-dimensional characterization map containing acoustic spectrum and tension decay curve, denoted as tension residual map; wherein, the horizontal axis of the tension residual map is time, and the vertical axis contains acoustic frequency band energy sequence and tension decay sequence;

[0115] Attitude inversion module 202: Inputs the tension residual map into the pre-trained attitude inversion model; the attitude inversion model includes an acoustic branch, a tension branch, a cross-modal fusion layer and an inversion head, wherein the acoustic branch is used to extract time-frequency texture features, the tension branch is used to extract tension attenuation physical features, the cross-modal fusion layer realizes inter-modal feature interaction through an attention mechanism, and the inversion head outputs decapitation attitude parameters and corresponding confidence scores, wherein the decapitation attitude parameters include orientation angle, tilt angle and category label;

[0116] Arbitration and Execution Module 203: Arbitrates based on the confidence level output by the inversion head: If the confidence level is higher than the first threshold, the decapitation posture parameters are directly output and the execution mechanism is triggered; if the confidence level is between the first threshold and the second threshold, the vision module is called for verification; if the confidence level is lower than the second threshold, the event is recorded and the manual verification process is initiated.

[0117] As an example, the joint characterization map construction module 201 is also used for:

[0118] The yarn running status is monitored in real time. When the change in acoustic signal energy or tension signal exceeds a preset threshold, it is determined as a breakage event and the breakage trigger time is determined.

[0119] As an example, the joint characterization map construction module 201 is specifically used for:

[0120] Centered on the fracture triggering moment, acoustic signal segments and tension signal segments of preset time lengths are extracted from the continuous acquisition buffers of the acoustic sensor and tension sensor, respectively.

[0121] The acoustic signal segment is subjected to a short-time Fourier transform to generate an acoustic time-spectrum; and the tension signal segment is fitted with an attenuation curve to extract its attenuation feature sequence over time.

[0122] The acoustic time-spectrum map and the attenuation feature sequence are aligned with the time axis and spliced ​​along the feature dimension to generate the joint two-dimensional characterization map.

[0123] As an example, the acoustic branch uses a network structure containing two-dimensional convolutional blocks and temporal convolutional modules to perform deep feature extraction on the acoustic time-frequency part of the tension residual spectrum to obtain the time-frequency texture features;

[0124] The tension branch extracts temporal features from the tension decay sequence portion of the tension residual map through a network structure containing one-dimensional convolutional or gated recurrent units to obtain the physical features of tension decay.

[0125] The cross-modal fusion layer employs a cross-attention mechanism, using the tension attenuation physical features as keys and values, and the time-frequency texture features as queries, to perform feature reweighting and fusion, thereby obtaining the intermodal feature interactions.

[0126] The inversion head maps the intermodal feature interactions into continuous orientation angles, tilt angles, and class labels through a fully connected layer, thus obtaining the decapitation posture parameters; and, based on heteroscedastic regression, it synchronously outputs the confidence scores of each predicted decapitation posture parameter.

[0127] As an example, the attitude inversion model also includes a viewpoint-aware transformation module, which generates cross-viewpoint attention parameters related to intermodal feature interactions through multi-level feature abstraction and attention distillation mechanisms; correspondingly, the cross-modal fusion layer obtains the intermodal feature interactions based on the cross-viewpoint attention parameters and the cross-attention mechanism.

[0128] As an example, the viewpoint-aware transformation module generates cross-viewpoint attention parameters related to intermodal feature interactions through multi-level feature abstraction and attention distillation mechanisms, including:

[0129] The time-frequency texture features and the tension attenuation physical features are projected onto a high-dimensional shared semantic space to obtain aligned acoustic semantic features and tension semantic features.

[0130] Cross-attention calculation is performed on the acoustic semantic features and tension semantic features to obtain an initial attention map; multi-scale feature extraction is performed on the initial attention map through a convolutional neural network to generate attention distillation features with spatial awareness.

[0131] The attention distillation features are input into a parameter prediction network, which includes an adaptive gating mechanism to dynamically adjust the information flow based on the correlation between features, and finally outputs the cross-view attention parameters.

[0132] As an example, the cross-modal fusion layer obtains the intermodal feature interactions based on the cross-view attention parameters and the cross-attention mechanism, including:

[0133] The cross-perspective attention parameter is decomposed into query correction factor, key correction factor, and value correction factor;

[0134] The query correction factor is subjected to a Hadamard product with the query vector in the cross-attention mechanism to generate an enhanced query vector; the key correction factor and value correction factor are respectively tensor-concatenated with the key vector and value vector in the cross-attention mechanism to form a multi-dimensional interactive representation.

[0135] Based on the enhanced query vector and the multi-dimensional interaction representation, multi-level attention computation is performed to obtain a fusion feature representation with cross-perspective perception capability.

[0136] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for yarn breakage posture recognition based on multimodal perception, characterized in that, The methods and steps include the following: Step S1: Multimodal signals of yarn breakage events are synchronously acquired by acoustic sensors and tension sensors, and time alignment is performed based on the breakage trigger time to construct a joint two-dimensional characterization map containing acoustic spectrum and tension decay curve, denoted as tension residual map; wherein, the horizontal axis of the tension residual map is time, and the vertical axis contains acoustic frequency band energy sequence and tension decay sequence. Step S2: Input the tension residual map into the pre-trained attitude inversion model; the attitude inversion model includes an acoustic branch, a tension branch, a cross-modal fusion layer, and an inversion head, wherein the acoustic branch is used to extract time-frequency texture features, the tension branch is used to extract tension attenuation physical features, the cross-modal fusion layer realizes inter-modal feature interaction through an attention mechanism, and the inversion head outputs decapitation attitude parameters and corresponding confidence scores, wherein the decapitation attitude parameters include orientation angle, tilt angle, and category label; Step S3: Arbitrate based on the confidence level output by the inversion head: if the confidence level is higher than the first threshold, directly output the decapitation posture parameters and trigger the actuator action; if the confidence level is between the first threshold and the second threshold, call the vision module for verification; if the confidence level is lower than the second threshold, record the event and start the manual verification process. The attitude inversion model also includes a viewpoint-aware transformation module, which generates cross-viewpoint attention parameters related to intermodal feature interactions through multi-level feature abstraction and attention distillation mechanisms; correspondingly, the cross-modal fusion layer obtains the intermodal feature interactions based on the cross-viewpoint attention parameters and the cross-attention mechanism. The viewpoint perception transformation module generates cross-viewpoint attention parameters related to intermodal feature interactions through multi-level feature abstraction and attention distillation mechanisms, including: The time-frequency texture features and the tension attenuation physical features are projected onto a high-dimensional shared semantic space to obtain aligned acoustic semantic features and tension semantic features. Cross-attention calculation is performed on the acoustic semantic features and tension semantic features to obtain an initial attention map; multi-scale feature extraction is performed on the initial attention map through a convolutional neural network to generate attention distillation features with spatial awareness. The attention distillation features are input into a parameter prediction network, which includes an adaptive gating mechanism to dynamically adjust the information flow based on the correlation between features, and finally outputs the cross-view attention parameters.

2. The yarn breakage posture recognition method based on multimodal perception according to claim 1, characterized in that: Before step S1, the method further includes: real-time monitoring of the yarn running status, and when the change in acoustic signal energy or tension signal exceeds a preset threshold, determining it as a breakage event and determining the breakage trigger time.

3. The yarn breakage posture recognition method based on multimodal perception according to claim 1, characterized in that: Time alignment was performed based on the fracture triggering moment to construct a joint two-dimensional characterization map containing acoustic spectrum and tension decay curve, including: Step S11: Taking the fracture triggering time as the center, extract acoustic signal segments and tension signal segments of preset time lengths from the continuous acquisition buffers of the acoustic sensor and tension sensor, respectively. Step S12: Perform a short-time Fourier transform on the acoustic signal segment to generate an acoustic time-spectrum; and perform attenuation curve fitting on the tension signal segment to extract its attenuation feature sequence as a function of time. Step S13: Align the acoustic time-spectrum map with the attenuation feature sequence on the time axis and stitch them together along the feature dimension to generate the joint two-dimensional characterization map.

4. The yarn breakage posture recognition method based on multimodal perception according to claim 3, characterized in that: The acoustic branch uses a network structure containing two-dimensional convolutional blocks and temporal convolutional modules to perform deep feature extraction on the acoustic time-frequency part of the tension residual spectrum to obtain the time-frequency texture features. The tension branch extracts temporal features from the tension decay sequence portion of the tension residual map through a network structure containing one-dimensional convolutional or gated recurrent units to obtain the physical features of tension decay. The cross-modal fusion layer employs a cross-attention mechanism, using the tension attenuation physical features as keys and values, and the time-frequency texture features as queries, to perform feature reweighting and fusion, thereby obtaining the intermodal feature interactions. The inversion head maps the intermodal feature interactions into continuous orientation angles, tilt angles, and class labels through a fully connected layer, thus obtaining the decapitation posture parameters; and, based on heteroscedastic regression, it synchronously outputs the confidence scores of each predicted decapitation posture parameter.

5. The yarn breakage posture recognition method based on multimodal perception according to claim 1, characterized in that: The cross-modal fusion layer obtains the intermodal feature interactions based on the cross-view attention parameters and the cross-attention mechanism, including: The cross-perspective attention parameter is decomposed into query correction factor, key correction factor, and value correction factor; The query correction factor is subjected to a Hadamard product with the query vector in the cross-attention mechanism to generate an enhanced query vector; the key correction factor and value correction factor are respectively tensor-concatenated with the key vector and value vector in the cross-attention mechanism to form a multi-dimensional interactive representation. Based on the enhanced query vector and the multi-dimensional interaction representation, multi-level attention computation is performed to obtain a fusion feature representation with cross-perspective perception capability.

6. A yarn breakage posture recognition system based on multimodal perception, characterized in that: The system includes: Joint characterization map construction module: Multimodal signals of yarn breakage events are synchronously acquired by acoustic sensors and tension sensors, and time alignment is performed based on the breakage trigger time to construct a joint two-dimensional characterization map containing acoustic spectrum and tension decay curve, denoted as tension residual map; wherein, the horizontal axis of the tension residual map is time, and the vertical axis contains acoustic frequency band energy sequence and tension decay sequence; Attitude inversion module: The tension residual map is input into the pre-trained attitude inversion model; the attitude inversion model includes an acoustic branch, a tension branch, a cross-modal fusion layer and an inversion head, wherein the acoustic branch is used to extract time-frequency texture features, the tension branch is used to extract tension attenuation physical features, the cross-modal fusion layer realizes inter-modal feature interaction through an attention mechanism, and the inversion head outputs decapitation attitude parameters and corresponding confidence scores, wherein the decapitation attitude parameters include orientation angle, tilt angle and category label; Arbitration and Execution Module: Arbitrates based on the confidence level output by the inversion head: if the confidence level is higher than the first threshold, the decapitation posture parameters are directly output and the execution mechanism is triggered; if the confidence level is between the first and second thresholds, the vision module is called for verification; if the confidence level is lower than the second threshold, the event is recorded and the manual verification process is initiated. The attitude inversion model also includes a viewpoint-aware transformation module, which generates cross-viewpoint attention parameters related to intermodal feature interactions through multi-level feature abstraction and attention distillation mechanisms; correspondingly, the cross-modal fusion layer obtains the intermodal feature interactions based on the cross-viewpoint attention parameters and the cross-attention mechanism. The viewpoint perception transformation module generates cross-viewpoint attention parameters related to intermodal feature interactions through multi-level feature abstraction and attention distillation mechanisms, including: The time-frequency texture features and the tension attenuation physical features are projected onto a high-dimensional shared semantic space to obtain aligned acoustic semantic features and tension semantic features. Cross-attention calculation is performed on the acoustic semantic features and tension semantic features to obtain an initial attention map; multi-scale feature extraction is performed on the initial attention map through a convolutional neural network to generate attention distillation features with spatial awareness. The attention distillation features are input into a parameter prediction network, which includes an adaptive gating mechanism to dynamically adjust the information flow based on the correlation between features, and finally outputs the cross-view attention parameters.

7. A yarn breakage posture recognition system based on multimodal perception according to claim 6, characterized in that: The joint characterization map construction module is also used for: The yarn running status is monitored in real time. When the change in acoustic signal energy or tension signal exceeds a preset threshold, it is determined as a breakage event and the breakage trigger time is determined.

8. A yarn breakage posture recognition system based on multimodal perception according to claim 6, characterized in that: The joint characterization map construction module is specifically used for: Centered on the fracture triggering moment, acoustic signal segments and tension signal segments of preset time lengths are extracted from the continuous acquisition buffers of the acoustic sensor and tension sensor, respectively. The acoustic signal segment is subjected to a short-time Fourier transform to generate an acoustic time-spectrum; and the tension signal segment is fitted with an attenuation curve to extract its attenuation feature sequence over time. The acoustic time-spectrum map and the attenuation feature sequence are aligned with the time axis and spliced ​​along the feature dimension to generate the joint two-dimensional characterization map.

Citation Information

Patent Citations

  • Neural network-based textile industry broken yarn identification method and system

    CN121234141A

  • Equipment vibration measurement device and method based on machine vision and laser radar

    CN121346956A