Deep well casting molten aluminum liquid leakage sound online monitoring method and system

By improving the dual-stream lightweight aluminum leakage sound recognition network, and combining the natural logarithmic linear filter bank and the adaptive Linformer cross-modal attention mechanism, the lag problem of molten aluminum leakage monitoring in deep well casting is solved, and accurate identification and rapid early warning are achieved in complex environments.

CN121783459APending Publication Date: 2026-04-03SOUTH CHINA UNIV OF TECH +3
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Monitoring methods for molten aluminum leakage during deep well casting are subject to lag. Existing acoustic monitoring methods are difficult to accurately identify leakage events in high humidity and high noise environments, and existing technologies lack accuracy and real-time performance under complex noise conditions.

Method used

An improved dual-stream lightweight aluminum leakage sound recognition network is adopted, which combines a natural logarithmic linear filter bank and an adaptive Linformer cross-modal attention mechanism to construct a local sub-band enhanced convolutional branch and a global temporal dependency modeling branch, thereby realizing real-time recognition and adaptive monitoring of leakage sound.

Benefits of technology

It can identify leakage events in high humidity and high noise environments in advance, has dynamic adaptive capabilities, improves the accuracy and reliability of monitoring, and realizes rapid early warning and proactive prevention and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121783459A_ABST
    Figure CN121783459A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of aluminum processing safety monitoring and industrial acoustic detection, and relates to a deep well casting molten aluminum liquid leakage online monitoring method and system based on underwater sound recognition. The method comprises the following steps: in an offline stage, collecting environmental sound and leakage sound samples by using a hydrophone, extracting natural logarithmic linear (Ln-linear) spectrum features, and obtaining a deployable model through double-flow lightweight convolution structure training and quantification; in the online stage, mixed sound signals are collected in real time, a model is input to output leakage confidence, a threshold value is adaptively updated in combination with historical data, a grading alarm mechanism is constructed by adopting continuous confidence accumulation, sound-light alarm is triggered or control signals are output to a PLC / DCS, and second-level early warning and emergency disposal are achieved. The system has the advantages of high real-time performance, high anti-interference capability, linkage control and the like, and is suitable for safety protection of a deep well casting scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of deep well casting, specifically relating to an online monitoring method and system for the sound of molten aluminum leakage in deep well casting. Background Technology

[0002] In deep-well casting, molten aluminum solidifies into ingots by direct contact with circulating cooling water within a water-sealed zone. If a leak occurs during this stage, the hot metal and cold water will undergo a violent, transient vaporization reaction, releasing high-temperature, high-pressure steam that may trigger a chain reaction of deflagrations. This poses a serious threat to personnel and equipment. Such explosions are extremely energetic and spread rapidly, often forming within seconds, posing a significant threat to the safety of production personnel, equipment, and the entire plant area. Similar accidents have occurred repeatedly both domestically and internationally in recent years, indicating a significant lag in existing monitoring and early warning systems.

[0003] Currently, on-site monitoring of leak risks mainly relies on infrared thermal imaging / video recognition or contact temperature sensors. However, the internal environment of deep wells is generally characterized by high humidity steam, white fog condensation, and strong reflection from metal walls, making infrared images prone to obstruction or false triggering. Some companies have attempted to monitor leak sounds using the sound pressure threshold method, but deep well workshops simultaneously contain strong noise sources such as fans, hydraulic pumps, and chain metal impacts. A single threshold method is difficult to reliably distinguish abnormal events under complex noise conditions, resulting in insufficient engineering reliability.

[0004] Compared to images and temperature signals, the initial stage of a leak is often accompanied by short-duration, high-frequency anomalous acoustic events such as sharp hissing sounds and pulse bursts. These sounds appear earlier than temperature rise and visible liquid flow diffusion, exhibiting stronger predictability and perceptibility. Therefore, acoustic monitoring is considered a crucial pathway for achieving second-level early warning of molten aluminum leaks in deep-well casting. In recent years, sound source demixing and acoustic emission (AE) monitoring technologies have been validated in scenarios such as pipeline leaks and valve failures, enabling the extraction of transient anomalous signals under complex noise conditions. However, the sounds of molten aluminum leaks in deep-well casting environments are characterized by short duration, easy energy attenuation, complex frequency band distribution, and easy submersion by background noise. Existing methods still have shortcomings in terms of recognition accuracy, real-time performance, and model adaptability. Summary of the Invention

[0005] The main objective of this invention is to overcome the shortcomings and deficiencies of the prior art and provide an online sound monitoring method and system for molten aluminum leakage in deep well casting. This method and system can identify leakage events in advance under high humidity, high noise and complex working conditions, has dynamic adaptive capabilities and can be linked with existing production control systems, so as to achieve rapid early warning and proactive prevention and control of leakage accidents.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] In a first aspect, the present invention provides an online monitoring method for the sound of leakage of molten aluminum liquid in deep well casting, comprising the following steps:

[0008] Offline modeling stage: Hydrophones are deployed in the deep well water-sealed area to collect sound samples. The sound samples are preprocessed and labeled to construct a training dataset. The spectrum obtained by performing a short-time Fourier transform on the training dataset is used to extract the spectral features of the sound through a natural logarithmic linear filter bank. An improved dual-stream lightweight aluminum leakage sound recognition network is then trained. The improved dual-stream lightweight aluminum leakage sound recognition network has a multi-band convolutional branch at the front end to reduce the feature dimensionality and enhance the multi-scale information expression. The improved dual-stream lightweight aluminum leakage sound recognition network inputs the time-frequency features in parallel to the local branch and the global branch. The local branch uses two-dimensional convolution combined with a multi-scale frequency attention mechanism to enhance the local time-frequency feature correlation modeling ability. The global branch achieves long sequence dependency modeling through dimensionality reduction projection and a lightweight Transformer coding layer. After the local branch and the global branch are interactively fused through an adaptive Linformer cross-modal attention mechanism, the features are integrated using a low-rank bottleneck structure, and the multi-label recognition results are output after global pooling and a classification head.

[0009] Online monitoring phase: Real-time acquisition of sound signals from the deep well water-sealed area, conversion into natural logarithmic linear spectral characteristics, input into the dual-flow lightweight model for inference, and output of leakage confidence; dynamic updating of the discrimination threshold based on historical data, and graded alarms using a continuous confidence time accumulation strategy; when the accumulation time reaches the first threshold, an audible and visual alarm is triggered, and when it reaches the second threshold, a control signal is output to link with the upper-level system.

[0010] As a preferred technical solution, in the offline modeling stage, the spectral features of the sound are extracted using a natural logarithmic linear filter bank, specifically as follows:

[0011] The preprocessed audio signal is subjected to a short-time Fourier transform to obtain a spectrum. A linear filter bank divides the frequency axis into equal intervals to preserve high-frequency details. The formula for the center frequency is:

[0012] ;

[0013] Applying natural logarithmic compression to the power spectrum energy, the formula is as follows:

[0014] ;

[0015] in For the first Output energy of a linear filter:

[0016]

[0017] This is the complex spectrum after the short-time Fourier transform. For the first The frequency response of a linear filter.

[0018] As a preferred technical solution, the multi-band convolution branch adopts a sub-band multi-convolution feature preprocessing module, which divides the acoustic vector into sub-bands and sets adjacent frequency band overlap. Each sub-band is processed independently by a lightweight blueprint convolutional network and then fused and reduced in dimension by 1×1 convolution.

[0019] In the global branch, given the input sequence The attention calculation process is as follows:

[0020]

[0021]

[0022] in Indicates a single-head dimension. , , , To make the projection matrix learnable, the number of channels C in this lightweight design is limited to a low dimension, which significantly reduces the amount of computation while ensuring attention modeling capability.

[0023]

[0024] Then through the low-rank projection matrix , Compress the key and value into an approximate representation of length ℓ:

[0025]

[0026] The cross-branch interaction phase employs an adaptive Linformer cross-channel attention mechanism, where global features and local features are respectively... First, the input is mapped to a lower-dimensional space through dimensionality reduction. :

[0027] Attention is calculated as follows:

[0028]

[0029] After mapping the output back to the original channel dimension via projection, it is added to the global feature residual. This preserves global semantic consistency, introduces local discriminative information, and avoids the limitations of conventional Transformers on long sequences. Complexity.

[0030] As a preferred technical solution, the branch fusion stage adopts a low-rank bottleneck structure, and the splicing feature is set as follows: First, the dimension is reduced and mapped to the bottleneck dimension r through convolution:

[0031] ;

[0032] Then, we move back to the output dimension:

[0033] ;

[0034] in This reduces the number of parameters and computational cost of the fusion layer from [previous data]. Significantly reduced to This enables efficient cross-modal feature fusion.

[0035] As a preferred technical solution, the offline modeling stage also includes the following:

[0036] Calculate the mean of the background confidence sequence on the validation set. with standard deviation and initialize the threshold: ;

[0037] The trained model is then quantized and compressed, and optimized using pruning or knowledge distillation to obtain a lightweight model that can be deployed in real time on edge devices.

[0038] As a preferred technical solution, the continuous confidence time accumulation strategy includes integral accumulation and moving average accumulation;

[0039] The leak event log includes timestamps, confidence curves, cumulative duration, and audio path, and the recorded information is written to the log database.

[0040] As a preferred technical solution, the method also includes: an online update phase: periodically transmitting suspected leaked data back to the offline modeling module, updating model parameters through incremental learning, and achieving adaptive optimization of the system;

[0041] During the online update and adaptive optimization phase of the model, suspected leaked data sent back to the offline modeling module needs to be manually confirmed before data augmentation and retraining are performed. By introducing new samples and updating parameters, combined with quantization compression, the latest deployable model is obtained, thereby achieving continuous optimization of the system.

[0042] Secondly, the present invention provides a real-time monitoring system for leakage of molten aluminum liquid in deep well casting for implementing the method, characterized in that it includes a multi-channel sound acquisition device, an edge processing device, a communication interface device, an alarm and control output device, a human-computer interaction and storage device, and a cloud server.

[0043] The multi-channel sound acquisition device consists of multiple hydrophones and their preamplifier and analog-to-digital conversion circuits. The analog-to-digital conversion circuit is located on the hydrophone side or in the edge processing device and supports USB, RS485, I²S or PoE interfaces to connect to the edge processing device.

[0044] The edge processing device includes an embedded processor or industrial computer, has an audio input interface connected to a multi-channel sound acquisition device, adopts a desktop, wall-mounted or DIN rail mounting structure, and is equipped with a backup power supply or UPS module.

[0045] The communication interface device includes any one or more combinations of Ethernet, RS485, WiFi, and 4G;

[0046] The alarm and control output device includes an audible and visual alarm and a 4-20mA, Modbus or OPC UA output interface, is powered by AC 220V or DC 24V and has a relay dry contact output interface.

[0047] The human-computer interaction and storage device includes a local display screen, a button or touch screen interface, and a data storage module;

[0048] The cloud server is used for centralized management of monitoring site data and remote equipment maintenance, and supports remote parameter configuration and system upgrades.

[0049] As a preferred technical solution, the hydrophone is a piezoelectric ceramic hydrophone, and its placement position is 1.2 to 1.5 m below the surface of the deep well water.

[0050] As a preferred technical solution, the edge processing device is equipped with the aforementioned online sound monitoring method for leaking molten aluminum liquid in deep well casting, which is used to process the collected sound signals in real time and output leakage discrimination indicators.

[0051] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0052] (1) The present invention proposes an Ln-linear natural logarithmic linear filter spectrum to replace the conventional Mel spectrum or linear spectrum, which significantly improves the amplitude contrast of small leakage sound in the high frequency region and effectively enhances the distinguishability of fine-grained texture features.

[0053] (2) The present invention constructs a dual-stream structure with local sub-band enhanced convolutional branch and global temporal dependency modeling branch, which simultaneously captures different variation patterns of short bursts and continuous bubbles, and has a higher leakage category separation degree than traditional single-branch CNN.

[0054] (3) The present invention uses the dynamic adaptive threshold of the fusion index and the continuous confidence time accumulation strategy to jointly judge, which can suppress false triggering caused by instantaneous impact noise and quickly form a high confidence response under continuous leakage, effectively improving the accuracy and reliability of alarm. Attached Figure Description

[0055] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0056] Figure 1 This is a flowchart of the online monitoring method for leakage of molten aluminum liquid in deep well casting according to an embodiment of the present invention;

[0057] Figure 2 A comparison of the spectra of aluminum leakage sound using different feature extraction methods in embodiments of the present invention;

[0058] Figure 3 This is a network structure diagram of the aluminum leakage sound recognition model in an embodiment of the present invention;

[0059] Figure 4 This is a flowchart illustrating the dynamic threshold discrimination and hierarchical alarm logic of an embodiment of the present invention.

[0060] Figure 5 This is a schematic diagram of the online monitoring system for leakage of molten aluminum liquid in deep well casting, according to an embodiment of the present invention. Detailed Implementation

[0061] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.

[0062] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.

[0063] This embodiment provides an online monitoring method for the sound of molten aluminum leakage in deep well casting. Underwater microphones are deployed in the water-sealed area of ​​the casting well to collect multi-source mixed sound signals in real time, including equipment operation sounds, liquid flow noise, and the sound of the interaction between molten aluminum and cooling water. The collected signals are preprocessed and converted into a spectrum based on a natural logarithmic linear filter bank (Ln-linear). Simultaneously, features such as MFCC, logarithmic Mel spectrum, or linear spectrum energy can be combined to achieve a comprehensive expression of multi-dimensional acoustic information. A dual-stream lightweight recognition model is input, consisting of a local two-dimensional convolutional branch and a global temporal dependency modeling branch. The former highlights the fine-grained features of brief bursts and high-frequency sharp signals, while the latter captures the cross-frame temporal dependencies during the continuous evolution of the leakage. The two branches output a comprehensive confidence index at the fusion layer. To improve the model's adaptability under different water-sealed depths, equipment operating states, and diurnal noise variations, this invention introduces a dynamic adaptive threshold update algorithm, allowing the alarm criteria to automatically adjust with environmental changes. Considering that leakage sounds often exhibit a short-intermittent-continuous evolutionary characteristic, this invention further incorporates a continuous confidence time accumulation mechanism to avoid false alarms caused by instantaneous high confidence levels. During the alarm determination phase, this invention constructs a tiered response mechanism: when the accumulated time reaches the first alarm threshold, an audible and visual alarm is triggered to alert on-duty personnel; when the duration further reaches the second alarm threshold, emergency control measures such as stopping irrigation, limiting flow, or increasing water seal pressure are implemented in conjunction with a higher-level system such as a PLC or DCS via a communication interface.

[0064] like Figure 1 As shown in the figure, this embodiment provides an online monitoring method for the sound of molten aluminum leakage in deep well casting, including an offline modeling stage, an online monitoring stage, and an online updating stage, as detailed below:

[0065] S1. In the offline modeling stage, environmental sound and leakage sound samples are collected using hydrophones, cleaned, labeled, and enhanced. Features such as Ln-linear spectrum and Mel-frequency cepstral coefficients are extracted, and a deployable model is obtained through training and quantization using a two-stream lightweight convolutional structure. The specific content is as follows:

[0066] S11. Hydrophone Deployment and Data Acquisition: Deploy at least one piezoelectric ceramic hydrophone below the surface of the deep well, at a distance of 1.2–1.5 m below the water surface. The reason for using a hydrophone is that it can directly acquire the leakage sound signal in the water, avoiding energy attenuation at the water-air interface and air noise interference, significantly improving the signal-to-noise ratio and response time.

[0067] S12. Data preprocessing and dataset construction: Collect various audio samples, including normal operating sounds, equipment operation sounds, manual water discharge sounds, and simulated leak sounds. Clean and filter the raw data, and generate "leakage" and "background" labels through manual judgment or clustering rules to form a training dataset.

[0068] S13. Data Augmentation: To enhance the model's sensitivity to weak leakage events, perform data augmentation operations such as cropping, splicing, noise addition, and time shifting on the data.

[0069] S14. Feature extraction: The feature extraction method is used to extract features from the spectrum obtained by performing a short-time Fourier transform on the preprocessed audio signal and then applying a natural logarithmic linear filter bank (Ln-linear).

[0070] Furthermore, step S14 specifically includes:

[0071] S141, Pre-weighting;

[0072] Pre-emphasis is the process of enhancing the high-frequency components of a signal using a high-pass filter, aiming to compensate for high-frequency losses caused by the sound source and transmission path. Pre-emphasis is typically implemented using a filter, which has the following form:

[0073]

[0074] in, The original signal, The signal after pre-emphasis. This is the pre-emphasis factor (typically between 0.95 and 1), which determines the degree to which the filter enhances the signal. Pre-emphasis helps improve the recognizability of high-frequency information by highlighting high-frequency components.

[0075] S142, Normalization;

[0076] The purpose of normalization is to adjust the amplitude of the audio signal to a uniform range, preventing excessively high or low amplitude amplitudes from causing model training instability or accuracy degradation during subsequent processing. Common normalization methods include normalizing the maximum value of the signal, as shown in the following formula:

[0077]

[0078] in, It is the original signal. It is the normalized signal, maximum value It is the largest absolute value in the signal. After normalization, the amplitude of the signal is compressed to between 0 and 1, which helps with subsequent processing and analysis.

[0079] S143, framing, windowing;

[0080] Audio signals are continuous. To perform spectral analysis, they need to be segmented into multiple short frames. Each frame contains a fixed number of sampling points to capture frequency changes within a short timeframe. Typically, there is some overlap between adjacent frames. The framing and windowing process is as follows:

[0081] First, the audio signal is divided according to frame length Perform frame segmentation. Assume the sampling frequency is... The duration of each frame is Then the number of sampling points per frame is:

[0082]

[0083] Then, a window function is applied to the signal in each frame to reduce edge effects. To reduce spectral leakage, a Hamming window is used, whose window function is defined as:

[0084]

[0085] in, It is the first Hanming window One coefficient, It is the length of the window function. The purpose of the Hamming window is to smoothly weight the signal in each frame, so that the amplitude of the signal gradually decreases at the edges of the frame, thereby effectively reducing spectral leakage.

[0086] The signal after windowing is:

[0087]

[0088] in, It is the original signal. This is the windowed signal. In this way, the signal amplitude at the edges of the frame is smoothly reduced, avoiding the impact of edge effects on the spectral analysis results.

[0089] S144, Short-Time Fourier Transform;

[0090] The Short-Time Fourier Transform (STFT) is an analytical method that transforms a time-domain signal into a time-frequency domain representation. By segmenting the signal into different time periods and applying a window function, the STFT can capture the spectral characteristics of the signal as it changes over time, thereby obtaining local frequency information. Applying the STFT to an audio signal can be expressed as:

[0091]

[0092] in, This indicates the sound signal after windowing processing. This represents the number of points in the Fourier transform. This transform yields the distribution characteristics of the sound signal in the time-frequency domain, providing a foundation for subsequent sound recognition and enhancement.

[0093] S145, Feature extraction methods;

[0094] Existing feature extraction methods often use Mel-scale-based methods, which transform the time-domain signal into Mel-scale frequencies through a Mel filter bank after a short-time Fourier transform. While this aligns with the laws of sound perception, it suffers from insufficient resolution in the high-frequency range, easily leading to the loss of high-frequency pulse and harmonic details in the sound of molten aluminum leakage. Figure 2 Therefore, in this embodiment, after the original time-domain signal is transformed to the time-frequency domain via a short-time Fourier transform, a feature extraction method using a natural logarithmic linear filter bank (Ln-linear) is selected, as follows:

[0095] A linear filter bank is divided into equal intervals along the frequency axis, with a center frequency of:

[0096]

[0097] That is, firstly, a linearly divided filter bank is used to preserve high-frequency details, and then natural logarithmic compression is applied to the power spectrum energy:

[0098]

[0099] in For the first Output energy of a linear filter:

[0100]

[0101] The complex spectrum is the result of the short-time Fourier transform (STFT). For the first The frequency response of a linear filter.

[0102] The use of the Ln-linear feature extraction method preserves high-frequency details and avoids the sparse distribution of Mel at high frequencies, thus better capturing pulse and harmonic features. While logarithm compression can significantly compress the energy dynamic range, excessive compression can weaken the details of high-energy pulses and harmonics. Using the natural logarithm (ln) provides a moderate compression, balancing the feature distribution while better preserving the high-frequency details of aluminum leakage sound, such as... Figure 2 As shown.

[0103] S15. An improved dual-stream lightweight aluminum leakage sound recognition network was developed. Based on the traditional end-to-end architecture, which uses convolutional networks to extract local time-frequency features of aluminum leakage sound and Transformers to model global contextual relationships, the network structure was optimized to meet the real-time requirements of on-site detection and embedded deployment. Specifically:

[0104] S151. A sub-band multi-frequency-band convolutional structure is introduced in the feature preprocessing part of the network front-end to achieve fine-grained frequency domain modeling of the original features. First, the acoustic vectors extracted from the features are divided into sub-frequency bands, and appropriate overlap is set between adjacent frequency bands to maintain spectral continuity. Each sub-frequency band is processed independently by a lightweight blueprint convolutional (BSConv) network. BSConv combines the advantages of depthwise separable convolution and dilated convolution, which can significantly reduce the number of parameters and computational complexity while maintaining the receptive field.

[0105] Features extracted through frequency band branching are then fused and dimensionality reduced using 1×1 convolutions, effectively compressing feature dimensions and network size while preserving spectral identification capabilities. Compared to traditional single-band convolutional structures, this multi-band parallel mechanism processes local channel and local spectral information only in each sub-band, reducing the total number of parameters to approximately 1 / N of the original (where N is the number of sub-bands), achieving efficient model lightweighting while ensuring spectral integrity.

[0106] S152, An improved dual-stream lightweight aluminum leakage sound recognition network is used to efficiently realize aluminum leakage sound recognition.

[0107] Traditional convolutional-transformer two-stream networks, by combining local feature extraction and global context modeling, exhibit excellent time-frequency feature representation capabilities in acoustic signal analysis. Convolutional neural networks (CNNs), relying on strong inductive bias, can efficiently capture local details such as transient impacts and spectral patterns; while the Transformer, based on a self-attention mechanism, models long-range dependencies across time and frequency bands, compensating for the limited receptive field of convolution. The synergy between the two achieves a balance between local discriminativity and global consistency, thus maintaining high robustness and generalization ability in complex and non-stationary acoustic environments.

[0108] In this embodiment, as Figure 3 As shown, specifically, the feature preprocessing part is input in parallel to the local and global branches. The local branch adopts a structure combining two-dimensional convolution and multi-scale attention, and adaptively emphasizes frequency bands with significant energy transitions through pooling and weighting mechanisms at different receptive field scales. This design can enhance the response to local high-frequency transients and narrowband harmonics in the aluminum leakage spectrum, enabling the model to identify weak but stable harmonic peaks and spectral envelope abrupt changes. After channel fusion, the multi-scale convolution features are mapped by a nonlinear activation function to obtain more discriminative local spectral texture features.

[0109] Meanwhile, the global branch performs dimensionality reduction projection on the preprocessed output, mapping the two-dimensional spectrogram into a one-dimensional temporal feature vector. It introduces a one-dimensional temporal context module and position-encoded convolutions to enhance temporal dependencies, and then captures global dependencies within a long time window through a lightweight Transformer coding layer. This branch effectively characterizes the regular changes in the aluminum leakage sound signal over time, including transient sounds during the aluminum leakage occurrence phase, energy decay during the termination phase, and periodic impact sounds, thus achieving effective integration of cross-frame feature information while maintaining temporal continuity.

[0110] The outputs of the two branches are interactively fused via an adaptive Linformer cross-attention mechanism, achieving dynamic complementarity between local time-frequency features and global temporal context. This mechanism preserves key cross-scale information under low-rank approximation, thereby enhancing the model's ability to focus on aluminum leakage sound features under complex background noise. The fused features are further integrated through a low-rank bottleneck structure, compressing redundant dimensions and improving feature compactness and robustness. Finally, multi-label recognition output for aluminum leakage sound is achieved through global pooling and a lightweight multi-layer classification head, which can distinguish different types of aluminum leakage sound patterns and their occurrence intensity.

[0111] Overall, this network structure is customized to address the high-frequency transients, narrowband harmonics, and spectral envelope abrupt changes of aluminum leakage sound through collaborative modeling of local and global branches. By combining Linformer attention and low-rank bottleneck fusion mechanisms, it significantly reduces the number of parameters and computational complexity while maintaining the ability to efficiently capture long-term and fine-grained spectral information, thus balancing real-time performance with embedded deployment requirements and exhibiting excellent aluminum leakage sound recognition performance.

[0112] The global branch employs a lightweight Transformer encoding layer based on a multi-head attention mechanism. Given an input sequence... The attention calculation process is as follows:

[0113]

[0114]

[0115] in Indicates a single-head dimension. , , , This is a learnable projection matrix. In this lightweight design, the number of channels C is limited to a low dimension (e.g., 32), which significantly reduces computational cost while maintaining attention modeling capabilities.

[0116]

[0117] Then through the low-rank projection matrix , Compress the key and value into an approximate representation of length ℓ:

[0118]

[0119] The cross-branch interaction phase employs an adaptive Linformer cross-channel attention mechanism. Let the global features and local features be respectively... First, the input is mapped to a lower-dimensional space through dimensionality reduction. :

[0120] Attention is calculated as follows:

[0121]

[0122] After mapping the output back to the original channel dimension via projection, it is added to the global feature residual. This preserves global semantic consistency, introduces local discriminative information, and avoids the limitations of conventional Transformers on long sequences. Complexity.

[0123] The branch fusion phase employs a low-rank bottleneck structure. Let the splicing feature be... First, the dimension is reduced and mapped to the bottleneck dimension r through convolution:

[0124]

[0125] Then, we move back to the output dimension:

[0126]

[0127] in This reduces the number of parameters and computational cost of the fusion layer from [previous data]. Significantly reduced to This enables efficient cross-modal feature fusion.

[0128] In the overall process, the input aluminum leakage audio signal is processed by spectral feature extraction and normalization, and then input in parallel to the local and global branches for feature encoding. After the two-stream features are interacted through a cross-attention mechanism, the output results are fused through low-rank bottleneck and global pooling to generate a compact representation. Finally, after maintaining feature consistency through identity mapping, the model output results are judged by an adaptive threshold strategy to achieve the recognition and classification of aluminum leakage sounds. This architecture takes into account the collaborative modeling of local and global acoustic information, and combines a lightweight attention mechanism and a low-rank fusion structure to significantly reduce the number of parameters and computational complexity while ensuring the real-time performance of the model in embedded environments.

[0129] S16, Threshold initialization and model lightweighting, specifically:

[0130] Calculate the mean of the background confidence sequence on the validation set. with standard deviation and initialize the threshold: ;

[0131] The trained model is then subjected to INT8 or float16 quantization compression, combined with pruning or knowledge distillation optimization, to obtain a lightweight model that can be deployed in real time on edge devices.

[0132] S2. Online Monitoring Phase: Sound signals from the deep well water-sealed area are acquired in real time, converted into Ln-linear spectral features, and input into the dual-flow lightweight model for inference, outputting leakage confidence. The discrimination threshold is dynamically updated based on historical data, and a continuous confidence time accumulation strategy is used for tiered alarms. When the accumulated time reaches the first threshold, an audible and visual alarm is triggered; when it reaches the second threshold, a control signal is output to link with the upper-level system. Details are as follows:

[0133] S21. Real-time data acquisition and preprocessing: Continuously acquire mixed raw sound signals from the deep well water-sealed area, upload them to the edge processing device and convert them into Ln-linear spectral feature vectors;

[0134] S22, Model Inference and Discriminant Index Generation: Input the spectrogram into the dual-stream lightweight model: Local branches extract local enhancement features, and global branches achieve temporal feature modeling through dimensionality reduction and lightweight attention encoding; The features of the two branches are adaptively fused to form a discriminant index C(n), which is used for subsequent confidence calculation and aluminum leakage sound recognition.

[0135] S23. Adaptive threshold update, as detailed below:

[0136] During the online determination process, the threshold T(n) can be updated according to changes in background noise using either a statistical or sliding decay method:

[0137] Statistical: ;

[0138] Sliding decay type: ;

[0139] in This is the attenuation coefficient. This invention is not limited to a specific form and can adaptively switch according to operating conditions. For example... Figure 4 As shown, the threshold T(n) is used to perform frame-level comparison of the discrimination index C(n). When C(n) exceeds T(n), the instantaneous discrimination result A(n) is triggered. This result serves as the input reference for subsequent confidence time accumulation and alarm judgment logic.

[0140] S24. Continuous confidence time accumulation: To suppress false alarms caused by instantaneous noise, this invention uses cumulative confidence or moving average methods to calculate the duration of leakage.

[0141] Integral type: ;

[0142] Moving average type: ;

[0143] When the cumulative confidence level Exceeding the first threshold (Approximately 300ms) triggers an audible and visual alarm; if the second threshold is reached... (Approximately 3 seconds) At this time, a linkage signal is output to the PLC / DCS via Modbus or 4-20mA interface to execute the measures of stopping pouring or limiting flow.

[0144] The reason for using the time accumulation method is that the sound of aluminum liquid leakage has the evolutionary characteristics of "intermittent-continuous". If only instantaneous peak values ​​are used for judgment, it is easy to mistakenly identify mechanical impact or single bubble sound as leakage event. Time integration can effectively distinguish between occasional noise and real leakage process.

[0145] S25. Leakage event recording and uploading: The system automatically records the timestamp, confidence curve, cumulative duration and audio path of the leakage event and writes it to the log database for subsequent analysis and tracing.

[0146] S3, Online Update and Model Iteration Phase;

[0147] Once an alarm event is confirmed by on-site personnel, its corresponding audio segment and cumulative confidence curve can be marked and archived, and periodically sent back to the offline modeling module as incremental samples for retraining or fine-tuning model parameters, thereby improving the stability of the system under different water levels, equipment conditions, or nighttime noise conditions.

[0148] To avoid uncontrollable drift caused by online learning, this embodiment adopts the method of "adding to the training library after manual confirmation" for updates, rather than fully automatic learning, in order to ensure system controllability and engineering safety.

[0149] like Figure 5 As shown, this embodiment provides an online sound monitoring system for leakage of molten aluminum liquid in deep well casting, including:

[0150] A multi-channel sound acquisition device consists of multiple hydrophones and their preamplifier and analog-to-digital converter circuits. The analog-to-digital converter circuits can be located on the hydrophone side or in the edge processing device.

[0151] An edge processing device, including an embedded processor or an industrial computer, is provided with an audio input interface for connection to the multi-channel sound acquisition device;

[0152] Communication interface device, including Ethernet, RS485, WiFi, 4G or any one or more combinations thereof;

[0153] Alarm and control output devices, including audible and visual alarms and 4-20mA, Modbus or OPC UA output interfaces;

[0154] Human-computer interaction and storage device, including local display screen, button or touch screen interface and data storage module;

[0155] Optional remote servers or cloud platforms can be used for centralized management of data from monitoring sites or for remote equipment maintenance.

[0156] The multi-channel sound acquisition device supports connection to the edge processing device via USB, RS485, I²S, or PoE interfaces. The edge processing device can be desktop, wall-mounted, or DIN rail mounted and can be configured with a backup power supply or UPS module to ensure continuous operation in the event of a power outage. The alarm and control output device is powered by AC 220V or DC 24V and has a relay dry contact output interface. The remote server or cloud platform supports remote parameter configuration and system upgrades to achieve centralized management and adaptive optimization of the monitoring system.

[0157] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0158] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0159] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A method for online monitoring of sound leakage from molten aluminum in deep well casting, characterized in that, Includes the following steps: Offline modeling phase: Hydrophones are deployed in the deep well water-sealed area to collect sound samples, and the sound samples are preprocessed and labeled to build a training dataset; The spectrum obtained by performing a short-time Fourier transform on the training dataset is used to extract the spectral features of the sound through a natural logarithmic linear filter bank. An improved dual-stream lightweight aluminum leakage sound recognition network is then trained. The improved dual-stream lightweight aluminum leakage sound recognition network has a multi-band convolutional branch at the front end to reduce the feature dimensionality and enhance the multi-scale information representation. The improved dual-stream lightweight aluminum leakage sound recognition network inputs the time-frequency features in parallel to the local branch and the global branch. The local branch uses two-dimensional convolution combined with a multi-scale frequency attention mechanism to enhance the local time-frequency feature correlation modeling ability. The global branch achieves long sequence dependency modeling through dimensionality reduction projection and a lightweight Transformer coding layer. After the local branch and the global branch are interactively fused through an adaptive Linformer cross-modal attention mechanism, the features are integrated using a low-rank bottleneck structure, and the multi-label recognition results are output after global pooling and a classification head. Online monitoring phase: Real-time acquisition of sound signals from the deep well water-sealed area, conversion into natural logarithmic linear spectral characteristics, input into the dual-flow lightweight model for inference, and output of leakage confidence; dynamic updating of the discrimination threshold based on historical data, and graded alarms using a continuous confidence time accumulation strategy; when the accumulation time reaches the first threshold, an audible and visual alarm is triggered, and when it reaches the second threshold, a control signal is output to link with the upper-level system.

2. The method for online monitoring of sound leakage from molten aluminum in deep well casting according to claim 1, characterized in that, In the offline modeling stage, the spectral features of the sound are extracted using a natural logarithmic linear filter bank, specifically as follows: The preprocessed audio signal is subjected to a short-time Fourier transform to obtain a spectrum. A linear filter bank divides the frequency axis into equal intervals to preserve high-frequency details. The formula for the center frequency is: ; Applying natural logarithmic compression to the power spectrum energy, the formula is as follows: ; in For the first Output energy of a linear filter: This is the complex spectrum after the short-time Fourier transform. For the first The frequency response of a linear filter.

3. The method for online monitoring of sound leakage from molten aluminum in deep well casting according to claim 1, characterized in that, The multi-band convolution branch adopts a sub-band multi-convolution feature preprocessing module, which divides the acoustic vector into sub-bands and sets adjacent band overlap. Each sub-band is processed independently by a lightweight blueprint convolutional network and then fused and reduced in dimension by 1×1 convolution. In the global branch, given the input sequence The attention calculation process is as follows: in Indicates a single-head dimension. ,、 ,、 ,、 To make the projection matrix learnable, the number of channels C in this lightweight design is limited to a low dimension, which significantly reduces the amount of computation while ensuring attention modeling capability. Then through the low-rank projection matrix , Compress the key and value into an approximate representation of length ℓ: The cross-branch interaction phase employs an adaptive Linformer cross-channel attention mechanism, where global features and local features are respectively... First, the input is mapped to a lower-dimensional space through dimensionality reduction. : Attention is calculated as follows: After mapping the output back to the original channel dimension via projection, it is added to the global feature residual. This preserves global semantic consistency, introduces local discriminative information, and avoids the limitations of conventional Transformers on long sequences. Complexity.

4. The method for online monitoring of sound leakage from molten aluminum in deep well casting according to claim 3, characterized in that, The branch fusion stage adopts a low-rank bottleneck structure, and the splicing feature is set as follows. First, the dimension is reduced and mapped to the bottleneck dimension r through convolution: ; Then, we move back to the output dimension: ; in This reduces the number of parameters and computational cost of the fusion layer from [previous data]. Significantly reduced to This enables efficient cross-modal feature fusion.

5. The method for online monitoring of sound leakage from molten aluminum in deep well casting according to claim 1, characterized in that, The offline modeling phase also includes the following: Calculate the mean of the background confidence sequence on the validation set. with standard deviation and initialize the threshold: ; The trained model is then quantized and compressed, and optimized using pruning or knowledge distillation to obtain a lightweight model that can be deployed in real time on edge devices.

6. The method for online monitoring of sound leakage from molten aluminum in deep well casting according to claim 1, characterized in that, The continuous confidence time accumulation strategy includes integral accumulation and moving average accumulation; The leak event log includes timestamps, confidence curves, cumulative duration, and audio path, and the recorded information is written to the log database.

7. The method for online monitoring of sound leakage from molten aluminum in deep well casting according to claim 1, characterized in that, The method also includes: an online update phase: periodically transmitting suspected leaked data back to the offline modeling module, updating model parameters through incremental learning, and achieving adaptive optimization of the system; During the online update and adaptive optimization phase of the model, suspected leaked data sent back to the offline modeling module needs to be manually confirmed before data augmentation and retraining are performed. By introducing new samples and updating parameters, combined with quantization compression, the latest deployable model is obtained, thereby achieving continuous optimization of the system.

8. A real-time monitoring system for leakage of molten aluminum liquid in deep well casting for implementing the method of any one of claims 1-7, characterized in that, It includes multi-channel sound acquisition devices, edge processing equipment, communication interface devices, alarm and control output devices, human-computer interaction and storage devices, and cloud servers; The multi-channel sound acquisition device consists of multiple hydrophones and their preamplifier and analog-to-digital conversion circuits. The analog-to-digital conversion circuit is located on the hydrophone side or in the edge processing device and supports USB, RS485, I²S or PoE interfaces to connect to the edge processing device. The edge processing device includes an embedded processor or industrial computer, has an audio input interface connected to a multi-channel sound acquisition device, adopts a desktop, wall-mounted or DIN rail mounting structure, and is equipped with a backup power supply or UPS module. The communication interface device includes any one or more combinations of Ethernet, RS485, WiFi, and 4G; The alarm and control output device includes an audible and visual alarm and a 4-20mA, Modbus or OPC UA output interface, powered by AC 220V or DC 24V and equipped with a relay dry contact output interface. The human-computer interaction and storage device includes a local display screen, a button or touch screen interface, and a data storage module; The cloud server is used for centralized management of monitoring site data and remote equipment maintenance, and supports remote parameter configuration and system upgrades.

9. The real-time monitoring system for leakage of molten aluminum liquid in deep well casting according to claim 8, characterized in that, The hydrophone is a piezoelectric ceramic hydrophone, and it is installed at a depth of 1.2 to 1.5 meters below the surface of the deep well.

10. The real-time monitoring system for leakage of molten aluminum liquid in deep well casting according to claim 8, characterized in that, The edge processing device is equipped with the online sound monitoring method for leakage of molten aluminum liquid in deep well casting as described in claim 1, which is used to process the collected sound signals in real time and output leakage discrimination indicators.