Audio data management system and method based on sound console
Through the audio data management system based on the mixer, an intelligent correlation model is built using the LSTM and Transformer model, which solves the problem that tuners have difficulty in processing audio data in real time, realizes real-time abnormality detection and processing of audio data, and improves the collaborative efficiency and version control capabilities of audio production.
Patent Information
- Application Number
- CN202510870075.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-26
AI Technical Summary
In large-scale live performances or complex audio production environments, it is difficult for tuners to grasp the status changes of all sound sources in real time, resulting in the failure to deal with sound quality defects or equipment failures in a timely manner. The existing technology lacks the ability to deeply correlate multi-source heterogeneous data such as audio data, operation sequences and equipment status, and lacks refined version management and collaborative processing mechanisms in cloud tuning scenarios where multi-person collaborates.
The audio data management system based on the mixer is adopted. Through time domain, frequency domain and time frequency analysis, combined with long and short-term memory network LSTM and Transformer models, it captures the time dynamics and feature interactions of audio data and operation sequences, builds an intelligent correlation model, analyzes audio data in real time, executes processing strategies, and implements collaborative version management in the cloud.
Real-time abnormality detection and processing of audio data is realized, the collaboration efficiency and version control capabilities of audio production are improved, and the sound quality stability and normal operation of the equipment is ensured.
Smart Images

Figure CN120544607A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tuning project management, in particular to an audio data management system and method based on a mixing console. Background Art
[0002] In large-scale live performances or complex audio production environments, real-time management and mixing of audio data are crucial to sound quality and live experience.
[0003] Traditional audio management relies heavily on the engineer's experience and real-time monitoring, supplemented by basic hardware monitoring. However, this model has significant shortcomings. In complex audio scenarios, engineers struggle to fully and comprehensively understand the changing status of all sound sources in real time. Information overload or delayed responses can lead to missed critical adjustments, resulting in poor sound quality or unresolved equipment failures. Existing technologies often lack the ability to deeply correlate and analyze heterogeneous data from multiple sources, including audio data, operation sequences, and device status. This makes it difficult to automatically identify and distinguish different types of anomalies, including instrument failures, microphone issues, audience noise, or device switching. This often requires engineers to rely on their experience, which is inefficient, subjective, and prone to errors in handling anomalies. While remote operation is possible in collaborative cloud-based audio mixing scenarios, the lack of sophisticated version management, collaborative processing mechanisms, and intuitive version comparison and evaluation tools hinders efficient collaboration between engineers. Historical operation tracing and version comparison analysis are also challenging, hindering the management and quality control of complex projects. Summary of the Invention
[0004] The object of the present invention is to provide an audio data management system and method based on a mixing console to solve the problems raised in the prior art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] In a first aspect, the present invention provides an audio data management method based on a mixing console, comprising:
[0007] Collect audio data, operation sequence data, equipment status and physical parameter data, and spectrum energy distribution data, and pre-process the collected data;
[0008] Through time domain, frequency domain, and time-frequency analysis, we extract the time domain, frequency domain, and time-frequency features of audio data, as well as the timing and state features of operation sequence data. We use the long short-term memory (LSTM) network to capture the temporal dynamics of each feature, and the self-attention mechanism of the Transformer to capture the long-range dependencies and interactions between features. Combined with a multimodal fusion strategy, we organically combine the processed features to construct an association model for audio data and operation sequence data.
[0009] The model analyzes the collected data to determine the on-site sound conditions and existing problems. When an instrument's audio signal, frequency band energy, or device parameters are abnormal, it determines that the instrument has failed or has left the stage normally. When the level, frequency band energy, or timestamp offset of the human voice activity are abnormal, it is determined to be a microphone anomaly. When the noise level, noise spectrum, or time matching in the audience area is abnormal, it is determined to be audience noise. When the audio signal or spectrum characteristics change, or when the device operation log and timestamp matching are abnormal, it is determined to be a device switch. The model takes corresponding measures according to the different judgments.
[0010] In a cloud-based mixing environment, the mixer's operations are recorded as independent steps and versions, allowing engineers to collaborate on the same audio, create and manage their own mixing versions, and support version comparison, merging, evaluation, and visual analysis of parameter differences.
[0011] In conjunction with the first aspect, in a first implementation of the first aspect of the present application, collecting audio data, operation sequence data, device status and physical parameter data, and spectrum energy distribution data, and preprocessing the collected data includes:
[0012] Audio data from instruments, vocals, and audience channels is collected through a high-precision multi-channel audio interface. Operation sequence data is collected through the console's internal logic interface. Device status and physical parameter data are collected through embedded sensors. Real-time spectrum energy distribution data is collected through a spectrum analyzer, recording energy changes in each frequency band and spectrum dynamic characteristics.
[0013] Synchronize the collected data by adding timestamps during data collection to ensure that the data stream is aligned based on a unified time base; determine the target format that needs to be unified and use a programming language to convert all data into a unified format; normalize or standardize numerical data.
[0014] In combination with the first aspect, in the second implementation of the first aspect of the present application, the time domain, frequency domain, and time-frequency features of the audio data, as well as the timing and state features of the operation sequence data, are extracted through time domain, frequency domain, and time-frequency analysis; the long short-term memory network LSTM is used to capture the respective temporal dynamics, and the long-distance dependencies and interactions between the features are captured through the self-attention mechanism of the Transformer; the processed features are organically combined with a multimodal fusion strategy to construct an association model of the audio data and the operation sequence data, including:
[0015] The pre-processed audio data is subjected to time domain analysis, frequency domain analysis, and time-frequency analysis. Time domain analysis divides the signal into short time frames, calculates the root mean square energy and zero-crossing rate of each frame, and reveals the basic intensity characteristics of the signal and its dynamic characteristics that change over time. Frequency domain analysis uses Fourier transform to convert the signal from the time domain to the frequency domain, identifying the various frequency components that make up the signal and the corresponding energy distribution. Time-frequency analysis divides the signal into time-varying windows and performs frequency domain analysis within each window to generate a time-frequency spectrum that displays the time-varying characteristics of the signal's frequency content. Key features from these analysis results are extracted to form a sequence of time domain feature vectors, frequency domain feature vectors, and time-frequency feature vectors, which constitute an audio feature vector sequence.
[0016] Extract timing features and state features from the preprocessed operation sequence data; process each event or state point in the operation sequence one by one in chronological order, review past operations based on the current time point, and extract timing features; query the value or state of each relevant component at the current time point, extract state features, and combine all the timing features and state features extracted for the current time point into a fixed-length numerical vector in a predefined order and format. Concatenate the feature vectors corresponding to each time point in the sequence in chronological order to obtain a two-dimensional sequence of timing feature vectors;
[0017] The extracted audio feature vector sequence and the time series feature vector sequence of the operation sequence are input into their respective LSTM networks. The LSTM learns and outputs a representation of the dynamic change pattern of each sequence in the time dimension. The dynamic representation of the audio features and the dynamic representation of the operation features output by the LSTM are sent as input sequences to the Transformer model. The Transformer's self-attention mechanism is used to calculate the feature correlation weights within and between sequences. The multi-head attention mechanism and multi-layer perceptron are used to capture the complex interactions and long-range dependencies between features.
[0018] A multimodal fusion layer is designed to receive the temporal dynamic representation output by the LSTM and the interaction and dependency representation output by the Transformer. An attention mechanism is used to organically combine feature representations from different models that capture different information, generating a joint feature representation that simultaneously reflects audio content, operational behavior, and interactions. This joint feature representation is then input into the final classifier to construct an association model.
[0019] Use labeled data to train the association model, deploy the trained association model, receive real-time collected and pre-processed audio data and operation sequence data, dynamically update the internal state, and output the association analysis results in real time, specifically the audio data, operation sequence data, and the patterns of their association.
[0020] In combination with the first aspect, in a third implementation of the first aspect of the present application, when the audio signal, frequency band energy, or device parameter of the musical instrument is abnormal, determining whether the musical instrument has failed or left the scene normally includes:
[0021] The audio signal and operation sequence of the musical instrument are monitored in real time and analyzed in combination with the association model. When the energy value of a specific frequency band is continuously lower than the preset silence threshold within the preset fault silence time, and there are no valid performance operation instructions during this period, the association model predicts that this state is inconsistent with the current music context and determines that it is an instrument fault. When the energy value or harmonic structure of a specific frequency band of the signal exceeds the preset fault noise threshold within the preset fault noise time, and the amplitude of the change exceeds the preset fault allowable fluctuation range, and the association model fails to identify patterns related to playing techniques or environmental factors, it is determined that it is an instrument fault. In order to rule out poor contact of the device connection cable, the device status data related to the instrument channel is checked, and the spectrum characteristics of the abnormal period are analyzed. When the spectrum characteristics are different from those of the instrument's own fault, it is determined to be a line problem rather than an instrument fault.
[0022] When it is detected that the instrument is pressed or receives a mute command, and after the command is issued, the signal energy value drops below the preset silence threshold within the preset departure decay time, and the association model confirms that the operation is consistent with the historical normal pattern, it is judged as a normal departure; when the signal energy value is lower than the preset silence threshold within the preset departure structure time, and the silence period matches the current music structure within the preset time tolerance threshold, and the association model does not detect any abnormal pattern, it is judged as a normal departure; when the signal energy value linearly decreases within the preset gradient time, its rate is lower than the preset departure permission threshold, and finally falls below the preset silence threshold, and the association model confirms that this process is consistent with the historical normal operation pattern, it is judged as a normal departure; to rule out the situation where the performer stops playing but the instrument is still in the on state and is just temporarily not in use, the device status data and operation sequence data are checked to see if there is any power operation or effector switching, and the spectrum is analyzed. If there is a harmonic structure unique to the instrument, it is judged that the instrument was turned on and temporarily not in use when the performance stopped, rather than a normal departure.
[0023] In combination with the first aspect, in a fourth implementation of the first aspect of the present application, when the level, frequency band energy, or timestamp offset of the human voice activity is abnormal, determining that the microphone is abnormal includes:
[0024] The collected vocal audio data is monitored in real time and analyzed in combination with the association model; the vocal area is identified using voice activity detection technology, and the vocal level, specific frequency band energy and timestamp sequence of the vocal activity in the area are extracted; the real-time extracted vocal level, frequency band energy and timestamp offset of the vocal activity are compared with the preset threshold; when the vocal level continuously exceeds the preset normal level range and the duration exceeds the preset sound abnormality duration threshold, the association model fails to identify the pattern related to the singer's deliberate volume adjustment and abnormal distance from the microphone, and it is judged as volume abnormality; when the energy level of non-human voice background noise continuously exceeds the preset noise threshold and the duration exceeds the preset noise duration threshold, the association model fails to identify the pattern related to the singer's breathing, environmental noise or singing skills, and it is judged as abnormal noise; when the timestamp sequence of the vocal activity has a long blank period and the duration exceeds the preset silence duration threshold, the association model fails to identify the pattern related to the singer's normal pause, improvisation or normal silence, and it is judged as a signal interruption.
[0025] In combination with the first aspect, in a fifth implementation of the first aspect of the present application, when the noise level, noise spectrum, or time matching in the audience area is abnormal, determining that it is auditorium noise includes:
[0026] The audio signal from the microphone array in the audience area is monitored in real time and analyzed in combination with the association model; the real-time noise level, noise spectrum characteristics and time pattern of noise activity of the signal are extracted; the real-time extracted noise level, noise spectrum characteristics and time pattern of noise activity are compared with the preset threshold value; when the noise level continuously exceeds the preset maximum allowable noise level threshold, the noise spectrum characteristics continuously deviate from the preset normal range threshold, or the noise activity time pattern does not match the normal interaction pattern threshold of the current performance stage, and the duration of the abnormal state exceeds the preset minimum duration threshold for judging the audience noise, the association model determines that the noise pattern does not match the normal interaction pattern of the current performance stage and is judged to be audience noise; in order to exclude crosstalk from musical instruments in other areas, it is necessary to compare the signals of other channels. When similar sounds are detected in other channels at the same time and the spectral characteristics of the original instruments are retained, it is judged to be crosstalk from musical instruments in other areas rather than audience noise.
[0027] In combination with the first aspect, in a sixth implementation of the first aspect of the present application, when an audio signal or spectrum feature changes or an abnormality occurs in the match between the device operation log and the timestamp, determining that the device is switching includes:
[0028] The audio signals, spectrum feature changes, and operation logs and timestamps of related equipment of each channel of the mixing console are monitored in real time, and analyzed in combination with the association model; the detected audio signal mutation amplitude, the rate and amplitude of spectrum feature changes, and the matching time difference between the timestamp of the device operation log and the audio signal mutation timestamp are compared with the preset threshold; when the audio signal mutation amplitude exceeds the preset device switching signal mutation threshold, the spectrum feature change rate or amplitude exceeds the preset device switching spectrum change threshold, and the matching time difference between the timestamp of the device operation log and the audio signal mutation timestamp is within the preset device switching time matching tolerance threshold, or the device operation log records the switching instruction, and the duration of the abnormal state exceeds the preset device switching judgment minimum duration threshold, the association model confirms that the detected state meets the characteristic pattern of device switching and is judged as device switching.
[0029] In combination with the first aspect, in a seventh implementation of the first aspect of the present application, performing corresponding processing according to different judged situations includes:
[0030] When an instrument failure is determined, the time point of the failure is marked on the timeline of the audio data, the audio channel corresponding to the failed instrument is marked, the volume of the channel is automatically reduced to silence, the technician is notified, and the details of the failure are recorded in the log; when it is determined to be a normal departure, the time point of the departure is marked on the timeline of the audio data, the audio channel corresponding to the departing instrument or singer is marked, the volume of the channel is controlled to smoothly reduce according to the preset fade-out curve until it is silent, and the departure information is recorded in the log;
[0031] When it is determined that the microphone is abnormal, it is handled according to the specific situation; when the volume is abnormal, the automatic gain control algorithm is used to dynamically and smoothly adjust the gain of the microphone channel to restore it to the preset normal level range based on the singer's singing habits, historical data of song performances, the progress of the current song and the preset strategy; when it is abnormal noise, the time domain audio signal collected by the microphone is converted to the frequency domain through short-time Fourier transform to obtain the amplitude spectrum and phase spectrum of each frame signal, and the noise audio is short-time Fourier transformed to calculate the average amplitude spectrum as an estimate of the noise amplitude spectrum; using spectral subtraction, the estimated noise amplitude spectrum is subtracted from the signal amplitude spectrum, and the obtained pure voice amplitude spectrum is combined with the original signal phase spectrum to form a processed spectrum, which is then inversely short-time Fourier transformed and converted back to the time domain audio signal to obtain the noise-reduced audio; when the signal is interrupted, an attempt is made to automatically reconnect the microphone signal; if the reconnection is successful, the volume and equalization settings of the channel are restored according to the historical data and song progress, and if the reconnection fails, the channel is muted, the technician is notified to check and replace the equipment immediately, and the details of the signal interruption are recorded in the log;
[0032] When an abnormal auditorium noise is detected, the specific time range in which the abnormal noise occurred is marked on the audio data timeline. All auditorium audio input channels within this time period are marked as abnormal noise. The auditorium volume balance is automatically adjusted based on the noise type, and the proportion of the auditorium volume in the main mix is reduced. The type, duration, and treatment measures of the abnormal noise are recorded in the log.
[0033] When it is determined that a device switch occurs, the time point at which the switch occurs is marked on the timeline of the audio data, the audio channels involved in the switch are marked, the switch type is confirmed based on the device operation log, and the parameters of the audio channels corresponding to the new device are automatically adjusted to ensure that the audio signal is not discontinuous or sudden during the switching process. At the same time, detailed switching operations are recorded in the log.
[0034] In conjunction with the first aspect, in an eighth implementation of the first aspect of the present application, in the cloud-based audio mixing environment, the tuner's operations are recorded as independent operation steps and versions, enabling engineers to collaboratively process the same audio, create and manage their own mixing versions, and support version comparison, merging, evaluation, and visual analysis of parameter differences, including:
[0035] The cloud server records and stores the tuner's operation steps and timestamps, and automatically generates a mixing version; allows multiple engineers to collaborate and create independent version branches; supports version comparison, displaying the audio waveforms, spectrum differences and corresponding operation differences of different mixing versions; supports version merging, allowing engineers to selectively merge operations and parameters from different version branches, automatically handle conflicts and provide conflict resolution suggestions; supports listening evaluation of the generated mixing version, records evaluation results and feedback; provides a graphical display of parameter differences, and displays parameter changes between different versions in the form of charts or timelines, making it easier for engineers to intuitively analyze the differences between versions.
[0036] In a second aspect, the present invention provides an audio data management system based on a mixing console, comprising:
[0037] Data acquisition and preprocessing module: includes a data acquisition unit and a data preprocessing unit; the data acquisition unit collects audio data, operation sequence data, device status and physical parameter data, and spectrum energy distribution data; the data preprocessing unit synchronizes the collected data, unifies the data format, and normalizes or standardizes the numerical data;
[0038] Feature extraction and association model building module: includes a feature extraction unit, an association model building unit, and a model training unit. The feature extraction unit performs multi-dimensional feature extraction on pre-processed audio data and operation sequence data to form a feature vector sequence and a time series feature vector sequence. The association model building unit inputs the feature sequences of audio and operation into the LSTM to learn temporal dynamic patterns, and then sends the LSTM output to the Transformer to capture the complex interactions and long-distance dependencies between features. A multimodal fusion layer is designed to generate a joint feature representation and build an association model. The model training unit trains the association model, receives real-time collected and pre-processed data, dynamically updates the status, and outputs analysis results in real time, including audio data, operation sequence data, and interrelated patterns.
[0039] Anomaly Detection and Judgment Module: This module includes an anomaly detection unit and a situation judgment unit. The anomaly detection unit uses a trained association model to monitor audio data, operation sequence data, device status and physical parameter data, and spectral energy distribution data in real time to identify abnormal data. The situation judgment unit then determines the situation of abnormal data detected by the anomaly detection unit, classifying it as instrument failure, normal exit, microphone anomaly, audience noise, and device switching.
[0040] Exception handling module: includes an exception handling unit and a log recording unit; wherein the exception handling unit executes the corresponding processing strategy according to the judgment result of the situation judgment unit; the log recording unit records detailed information;
[0041] Cloud collaboration and version management module: includes version recording and branch management unit, version comparison and merging unit and evaluation feedback and visualization unit; among them, the version recording and branch management unit records the tuner's operation steps and timestamps on the cloud server, generates a mixing version, supports multi-engineer collaboration, and creates independent version branches for parallel work; the version comparison and merging unit displays the differences between different mixing versions, supports version merging, allows engineers to selectively merge operations of other branches, automatically handles conflicts and provides conflict resolution suggestions; the evaluation feedback and visualization unit supports the evaluation of generated mixing versions, records evaluation results and feedback, and provides a graphical display of parameter differences.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] 1. This invention integrates multimodal data, extracts features using time domain, frequency domain and time-frequency analysis, and adopts LSTM and Transformer deep learning models to capture the respective temporal dynamics and complex interactions and long-range dependencies between features, and constructs an intelligent association model between audio and operation sequences.
[0044] 2. Based on the constructed correlation model, the present invention analyzes the collected data in real time, realizes the detection and judgment of anomalies, and executes corresponding processing strategies according to the judgment results.
[0045] 3. This invention introduces cloud-based collaborative version management to improve the collaborative efficiency and version control capabilities of audio production. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A schematic diagram of the steps of a method for managing audio data based on a mixing console according to the present invention;
[0047] Figure 2 The present invention is a system structure diagram of an audio data management system based on a mixing console. DETAILED DESCRIPTION
[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0049] Example: Figure 1-Figure 2 As shown, the present invention provides a technical solution.
[0050] like Figure 1 As shown in a schematic diagram of the steps of a method for managing audio data based on a mixing console, the present invention provides a method for managing audio data based on a mixing console, comprising:
[0051] Step S100: collecting audio data, operation sequence data, device status and physical parameter data, and spectrum energy distribution data, and preprocessing the collected data;
[0052] Specifically, the system collects audio data from instruments, vocals, and audience channels through a high-precision multi-channel audio interface; collects operation sequence data through the internal logic interface of the mixing console; collects device status and physical parameter data through embedded sensors; and collects real-time spectrum energy distribution data through a spectrum analyzer, recording energy changes in each frequency band and spectrum dynamic characteristics.
[0053] Synchronize the collected data by adding timestamps during data collection to ensure that the data stream is aligned based on a unified time base; determine the target format that needs to be unified and use a programming language to convert all data into a unified format; normalize or standardize numerical data.
[0054] In a specific embodiment, multi-track audio data is collected through a 24-channel Neumann U87 microphone array with a sampling rate of 48kHz and a bit depth of 24 bits, and audio streams for strings, brass, vocals, and audience feedback are recorded separately. The peak level of the string channel reaches -6dBFS, and the average background noise of the audience channel is -50dBFS. Operation sequence data is collected through the Yamaha CL series mixing console API interface, recording the change curve of the main fader smoothly transitioning from -12dB to -3dB, the real-time change data of the reverberation decay time adjusting from 2.5s to 4.0s, and the timestamp of the drum solo scene activation. The device status is collected through embedded sensors, and the temperature of the channel 1 microphone is monitored to be 26.4°C, the line current is 0.8A, and the connection status is displayed as locked. The spectrum analyzer collects spectrum data every 100ms, recording dynamic changes of -20dB energy in the 1kHz band and -35dB energy in the 10kHz band.
[0055] In the preprocessing stage, Python's PyAudioAnalysis library is used to synchronize the time of the collected 24 tracks of audio. The timestamp error is calibrated to within ±1ms using the NTP protocol. All data is converted into a unified structure based on the JSON format, and Min-Max normalization is performed on the numerical data to map the values to the [0, 1] interval.
[0056] Step S200: Through time domain, frequency domain, and time-frequency analysis, the time domain, frequency domain, and time-frequency features of the audio data, as well as the timing and state features of the operation sequence data, are extracted. The long short-term memory network (LSTM) is used to capture the temporal dynamics of each feature, and the self-attention mechanism of the Transformer is used to capture the long-range dependencies and interactions between features. Combined with a multimodal fusion strategy, the processed features are organically combined to construct an association model of the audio data and the operation sequence data.
[0057] Specifically, the pre-processed audio data is subjected to time domain analysis, frequency domain analysis, and time-frequency analysis; the time domain analysis divides the signal into short time frames, calculates the root mean square energy and zero-crossing rate of each frame, and reveals the basic intensity characteristics of the signal and its dynamic characteristics that change over time; the frequency domain analysis uses Fourier transform to convert the signal from the time domain to the frequency domain, and identifies the various frequency components that constitute the signal and the corresponding energy distribution; the time-frequency analysis divides the signal into time-varying windows and performs frequency domain analysis in each window to generate a time-frequency spectrum to show the time-varying characteristics of the signal's frequency content; the key features of these analysis results are extracted to form a sequence of time domain feature vectors, frequency domain feature vectors, and time-frequency feature vectors to form an audio feature vector sequence;
[0058] Extract timing features and state features from the preprocessed operation sequence data; process each event or state point in the operation sequence one by one in chronological order, review past operations based on the current time point, and extract timing features; query the value or state of each relevant component at the current time point, extract state features, and combine all the timing features and state features extracted for the current time point into a fixed-length numerical vector in a predefined order and format. Concatenate the feature vectors corresponding to each time point in the sequence in chronological order to obtain a two-dimensional sequence of timing feature vectors;
[0059] The extracted audio feature vector sequence and the temporal feature vector sequence of the operation sequence are input into their respective configured LSTM networks. The LSTM learns and outputs a representation of the dynamic change pattern of each sequence in the time dimension. The dynamic representation of the audio features and the dynamic representation of the operation features output by the LSTM are sent as input sequences to the Transformer model. The Transformer's self-attention mechanism is used to calculate the feature correlation weights within and between sequences. The dynamic representations of the audio and operation sequences are projected into different feature spaces through linear transformation. Position encoding is added to introduce sequence order information. For each position in the sequence, the correlation score between it and all positions in the sequence and all positions in the other sequence is calculated. This is measured using scaled dot product attention and obtained by interacting with the query, key, and value matrices. These scores are normalized by softmax and become feature correlation weights, which represent the importance of each position to the current position. The multi-head attention mechanism further parallelizes this process, capturing different types of correlations through multiple independent attention heads. The outputs of each head are spliced and linearly projected to obtain a feature representation that integrates the complex dependencies within and between sequences.
[0060] A multimodal fusion layer is designed to receive the temporal dynamic representation output by the LSTM and the interaction and dependency representation output by the Transformer. An attention mechanism is used to organically combine feature representations from different models that capture different information, generating a joint feature representation that simultaneously reflects audio content, operational behavior, and interactions. This joint feature representation is then input into the final classifier to construct an association model.
[0061] Collect multimodal data segments containing audio data and corresponding operation sequences, ensure that the two are precisely synchronized in time, assign clear predefined labels to each data segment, describe the audio content, operation behavior and its associated patterns; ensure that the audio features and operation features of each time step strictly correspond to the labels through data alignment, and divide the data into training set, validation set and test set, use the training set and validation set to train the association model, optimize the model parameters through forward propagation, loss calculation, backpropagation and parameter update, monitor the performance and adjust the hyperparameters on the validation set at the same time, after training, use the test set to evaluate the final model, select the model version with the best performance and strong generalization ability to save, deploy the trained association model, receive the real-time collected and preprocessed audio data and operation sequence data, dynamically update the internal state, and output the association analysis results in real time, specifically the audio data, operation sequence data and the pattern of association between them.
[0062] In a specific embodiment, feature extraction is performed on the preprocessed audio data, the 48kHz audio signal of the string channel is divided into short time frames of 20ms, the root mean square energy and zero crossing rate of each frame are calculated, and a time domain feature vector sequence is formed; a 1024-point fast Fourier transform is used to perform frequency domain analysis, and the main frequency components at 440Hz, energy -10dB and 880Hz, energy -15dB are identified, and a frequency domain feature vector sequence is generated; a time-frequency analysis is performed by short-time Fourier transform to generate a 128×128 time-spectrum diagram, and it is captured that the 440Hz frequency component appears clearly in frames 200-250. The energy is significantly enhanced from -15dB to -5dB, forming a time-frequency feature vector sequence, and finally forming an audio feature vector sequence containing 3000 time domain feature vectors, frequency domain feature vectors and time-frequency feature vectors; the timing features are extracted from the operation sequence data, and the average frequency of the fader position adjustment operation is calculated to be 0.5 times / second, and the average interval between two operations is 2.1 seconds. The state features are extracted, and the current fader position is -3.5dB, the knob angle is 72 degrees, and the activated effector is the reverberation unit, forming an operation sequence feature vector sequence containing 1500 timing feature vectors and state feature vectors.
[0063] The audio feature vector sequence is input into an LSTM network with 128 hidden units, and the operation sequence feature vector sequence is input into an LSTM network with 64 hidden units. The LSTM outputs 3000 and 1500 dynamic change pattern representations, respectively. The audio dynamic representation and operation dynamic representation output by the LSTM are used as input sequences and sent to the Transformer model. Eight attention heads are set, and the feature correlation weights are calculated using the self-attention mechanism. A multimodal fusion layer is designed, and a gating mechanism is used to fuse the temporal dynamic representation output by the LSTM and the interaction representation output by the Transformer to generate a joint feature representation containing 4096 dimensions. The joint feature representation is input into the final three-layer fully connected classifier to construct an association model.
[0064] The association model is trained using a labeled dataset containing 1,000 samples. After training, the model is deployed to receive and analyze the current audio data and operation sequence data in real time, dynamically update the internal state, and output the association analysis results in real time.
[0065] Step S300: The model analyzes the collected data to determine the on-site sound conditions and existing problems. When an instrument's audio signal, frequency band energy, or device parameters are abnormal, it is determined that the instrument has failed or has left the stage normally. When the level, frequency band energy, or timestamp offset of the human voice activity are abnormal, it is determined that there is a microphone anomaly. When the noise level, noise spectrum, or time matching in the audience area is abnormal, it is determined to be audience noise. When the audio signal or frequency spectrum characteristics change, or when there is an abnormality in the matching between the device operation log and the timestamp, it is determined that there is a device switch.
[0066] Specifically, the audio signal and operation sequence of the musical instrument are monitored in real time and analyzed in combination with the association model; when the energy value of a specific frequency band is continuously lower than the preset silence threshold within the preset fault silence time, and there is no valid performance operation instruction during the period, the association model predicts that this state is inconsistent with the current music context and it is judged as a musical instrument fault; when the energy value or harmonic structure of a specific frequency band of the signal exceeds the preset fault noise threshold within the preset fault noise time, and the amplitude of the change exceeds the preset fault allowable fluctuation range, the association model fails to identify the pattern related to the performance technique or environmental factors and it is judged as a musical instrument fault; in order to rule out the situation where the equipment connection cable has poor contact, the equipment status data related to the instrument channel is checked, and the spectrum characteristics of the abnormal period are analyzed. When it is different from the spectrum characteristics of the instrument's own fault, it is judged to be a line problem rather than an instrument fault;
[0067] When it is detected that the instrument is pressed or receives a mute command, and after the command is issued, the signal energy value drops below the preset silence threshold within the preset departure decay time, and the association model confirms that the operation is consistent with the historical normal pattern, it is judged as a normal departure; when the signal energy value is lower than the preset silence threshold within the preset departure structure time, and the silence period matches the current music structure within the preset time tolerance threshold, and the association model does not detect any abnormal pattern, it is judged as a normal departure; when the signal energy value linearly decreases within the preset gradient time, its rate is lower than the preset departure permission threshold, and finally falls below the preset silence threshold, and the association model confirms that this process is consistent with the historical normal operation pattern, it is judged as a normal departure; to rule out the situation where the performer stops playing but the instrument is still in the on state and is just temporarily not in use, the device status data and operation sequence data are checked to see if there is any power operation or effector switching, and the spectrum is analyzed. If there is a harmonic structure unique to the instrument, it is judged that the instrument was turned on and temporarily not in use when the performance stopped, rather than a normal departure.
[0068] The collected vocal audio data is monitored in real time and analyzed in combination with the association model; the vocal area is identified using voice activity detection technology, and the vocal level, specific frequency band energy and timestamp sequence of the vocal activity in the area are extracted; the real-time extracted vocal level, frequency band energy and timestamp offset of the vocal activity are compared with the preset threshold; when the vocal level continuously exceeds the preset normal level range and the duration exceeds the preset sound abnormality duration threshold, the association model fails to identify the pattern related to the singer's deliberate volume adjustment and abnormal distance from the microphone, and it is judged as volume abnormality; when the energy level of non-human voice background noise continuously exceeds the preset noise threshold and the duration exceeds the preset noise duration threshold, the association model fails to identify the pattern related to the singer's breathing, environmental noise or singing skills, and it is judged as abnormal noise; when the timestamp sequence of the vocal activity has a long blank period and the duration exceeds the preset silence duration threshold, the association model fails to identify the pattern related to the singer's normal pause, improvisation or normal silence, and it is judged as a signal interruption.
[0069] The audio signal from the microphone array in the audience area is monitored in real time and analyzed in combination with the association model; the real-time noise level, noise spectrum characteristics and time pattern of noise activity of the signal are extracted; the real-time extracted noise level, noise spectrum characteristics and time pattern of noise activity are compared with the preset threshold value; when the noise level continuously exceeds the preset maximum allowable noise level threshold, the noise spectrum characteristics continuously deviate from the preset normal range threshold, or the noise activity time pattern does not match the normal interaction pattern threshold of the current performance stage, and the duration of the abnormal state exceeds the preset minimum duration threshold for judging the audience noise, the association model determines that the noise pattern does not match the normal interaction pattern of the current performance stage and is judged to be audience noise; in order to exclude crosstalk from musical instruments in other areas, it is necessary to compare the signals of other channels. When similar sounds are detected in other channels at the same time and the spectral characteristics of the original instruments are retained, it is judged to be crosstalk from musical instruments in other areas rather than audience noise.
[0070] The audio signals, spectrum feature changes, and operation logs and timestamps of related equipment of each channel of the mixing console are monitored in real time, and analyzed in combination with the association model; the detected audio signal mutation amplitude, the rate and amplitude of spectrum feature changes, and the matching time difference between the timestamp of the device operation log and the audio signal mutation timestamp are compared with the preset threshold; when the audio signal mutation amplitude exceeds the preset device switching signal mutation threshold, the spectrum feature change rate or amplitude exceeds the preset device switching spectrum change threshold, and the matching time difference between the timestamp of the device operation log and the audio signal mutation timestamp is within the preset device switching time matching tolerance threshold, or the device operation log records the switching instruction, and the duration of the abnormal state exceeds the preset device switching judgment minimum duration threshold, the association model confirms that the detected state meets the characteristic pattern of device switching and is judged as device switching.
[0071] In one specific embodiment, the model analyzes the collected data in real time to determine the on-site sound conditions. It monitors that the energy value of the 1000Hz frequency band of the audio signal of the violin channel is continuously lower than the -60dBFS silence threshold during the 10-second fault silence period from 15:23:00 to 15:23:10. During this period, there are no violin playing operation instructions in the operation sequence, and the association model predicts that this silent state is inconsistent with the current music segment that should be played continuously, and it is determined to be a violin fault.
[0072] The vocal channel was monitored in real time, and voice activity detection technology was used to identify the vocal area. The vocal activity timestamp sequence from 15:27:10 to 15:27:30 was extracted. The vocal level in this area was continuously between -15dBFS and -10dBFS, exceeding the preset normal level range of -20dBFS to -12dBFS. The duration exceeded the sound abnormality duration threshold of 3 seconds, and the association model failed to identify patterns related to the singer's deliberate adjustment, and it was judged to be a volume abnormality.
[0073] Real-time monitoring of the microphone array in the audience area revealed that the noise level from 15:30:00 to 15:30:45 continuously exceeded the maximum allowable noise level threshold of -25dBFS. The noise spectrum characteristics in the frequency band above 5000Hz continuously deviated from the normal range. The noise activity time pattern did not match the normal interaction pattern of the current performance stage. The abnormal state lasted for more than 10 seconds, which is the minimum duration threshold for determining audience noise. The association model determined that the noise pattern did not match the normal interaction pattern and was therefore judged to be audience noise.
[0074] Step S400: performing corresponding processing according to different judged situations;
[0075] Specifically, when it is determined that an instrument malfunctions, the time point at which the malfunction occurred is marked on the timeline of the audio data, the audio channel corresponding to the malfunctioning instrument is marked, the volume of the channel is automatically lowered to mute, the technician is notified, and the details of the malfunction are recorded in the log; when it is determined that the departure is normal, the time point at which the departure occurred is marked on the timeline of the audio data, the audio channel corresponding to the departing instrument or singer is marked, the volume of the channel is controlled to smoothly decrease according to a preset fade-out curve until it is muted, and the departure information is recorded in the log;
[0076] When it is determined that the microphone is abnormal, it is handled according to the specific situation; when the volume is abnormal, the automatic gain control algorithm is used to dynamically and smoothly adjust the gain of the microphone channel to restore it to the preset normal level range based on the singer's singing habits, historical data of song performances, the progress of the current song and the preset strategy; when it is abnormal noise, the time domain audio signal collected by the microphone is converted to the frequency domain through short-time Fourier transform to obtain the amplitude spectrum and phase spectrum of each frame signal, and the noise audio is short-time Fourier transformed to calculate the average amplitude spectrum as an estimate of the noise amplitude spectrum; using spectral subtraction, the estimated noise amplitude spectrum is subtracted from the signal amplitude spectrum, and the obtained pure voice amplitude spectrum is combined with the original signal phase spectrum to form a processed spectrum, which is then inversely short-time Fourier transformed and converted back to the time domain audio signal to obtain the noise-reduced audio; when the signal is interrupted, an attempt is made to automatically reconnect the microphone signal; if the reconnection is successful, the volume and equalization settings of the channel are restored according to the historical data and song progress, and if the reconnection fails, the channel is muted, the technician is notified to check and replace the equipment immediately, and the details of the signal interruption are recorded in the log;
[0077] When an abnormal auditorium noise is detected, the specific time range in which the abnormal noise occurred is marked on the audio data timeline. All auditorium audio input channels within this time period are marked as abnormal noise. The auditorium volume balance is automatically adjusted based on the noise type, and the proportion of the auditorium volume in the main mix is reduced. The type, duration, and treatment measures of the abnormal noise are recorded in the log.
[0078] When it is determined that a device switch occurs, the time point at which the switch occurs is marked on the timeline of the audio data, the audio channels involved in the switch are marked, the switch type is confirmed based on the device operation log, and the parameters of the audio channels corresponding to the new device are automatically adjusted to ensure that the audio signal is not discontinuous or sudden during the switching process. At the same time, detailed switching operations are recorded in the log.
[0079] In a specific embodiment, it is determined that an instrument failure occurs in the violin channel at 15:23:00, and the time point is immediately marked on the audio data timeline, the violin channel is marked as a failure, its volume is automatically lowered to silent, and a notification "Channel 1 violin failure, please check" is sent to the background, and "15:23:00 Channel 1 violin failure, silent processing" is recorded in the log.
[0080] Abnormal noise was detected in the vocal channel between 15:27:10 and 15:27:30. The collected audio signal was subjected to STFT transformation, and the average amplitude spectrum of the noise audio was calculated. Spectral subtraction was used to subtract the noise amplitude spectrum from the original signal amplitude spectrum to obtain the pure speech amplitude spectrum. This spectrum was combined with the original phase spectrum and an inverse STFT was performed to obtain the noise-reduced audio. The log entry reads "Voice noise reduction processed for channel 5 from 15:27:10 to 15:27:30."
[0081] A device switch is determined to have occurred at 15:31:20. This time point is marked on the timeline, and the main output channel is marked as a device switch. The operation log confirms that this is a scene switch, and the volume, pan, and effect parameters of each channel in the new scene are automatically adjusted to ensure a smooth switching process. The log records "Scene switch at 15:31:20, channel parameters automatically adjusted."
[0082] Step S500: In the cloud-based tuning environment, the tuner's operations are recorded as independent operation steps and versions, allowing engineers to collaboratively process the same audio, create and manage their own mixing versions, and support version comparison, merging, evaluation, and visual analysis of parameter differences.
[0083] Specifically, the cloud server records and stores the tuner's operation steps and timestamps, and automatically generates a mixing version; allows multiple engineers to collaborate and create independent version branches; supports version comparison, displaying the audio waveforms, spectrum differences and corresponding operation differences of different mixing versions; supports version merging, allowing engineers to selectively merge operations and parameters from different version branches, automatically handle conflicts and provide conflict resolution suggestions; supports listening evaluation of the generated mixing version, records evaluation results and feedback; provides a graphical display of parameter differences, and displays parameter changes between different versions in the form of charts or timelines, making it easier for engineers to intuitively analyze the differences between versions.
[0084] In one specific embodiment, while working on Song X in a cloud-based tuning environment, engineer A records his or her steps: adjusting the master output fader from -3dB to -1dB at 15:40:05 and adding a compressor to the bass on Channel 2 at 15:40:20, with a threshold of -18dB and a ratio of 4:1. Version v1.0 is automatically generated, storing the operations and timestamps. Engineer B creates a separate branch, v1.1, from v1.0. At 15:45:10, he adjusts the EQ for the vocal on Channel 1, adjusting the 1kHz gain from 0dB to +3dB. The version comparison function reveals differences between v1.0 and v1.1 in the EQ parameters for Channel 1 and the effects on Channel 2, visualizing subtle changes in the waveform and spectrum. Engineer A decides to merge the branches, retaining B's EQ adjustments and discarding B's compressor addition. The system automatically resolves the conflicting master output fader adjustments. After the merge, version v1.2 is generated. Engineer A used the listening evaluation feature to give v1.2 a score of 8 / 10, noting, "Vocals are clearer, but the low end is slightly lacking." A visualization of the parameter differences shows that between v1.0 and v1.2, the 1kHz parameter on Channel 1 changed from 0dB to +3dB at 15:45:10, and the compressor parameter on Channel 2 was added at 15:40:20.
[0085] like Figure 2As shown in the system structure diagram of an audio data management system based on a mixing console, the present invention provides an audio data management system based on a mixing console, comprising:
[0086] Data acquisition and preprocessing module: includes a data acquisition unit and a data preprocessing unit; the data acquisition unit collects audio data, operation sequence data, device status and physical parameter data, and spectrum energy distribution data; the data preprocessing unit synchronizes the collected data, unifies the data format, and normalizes or standardizes the numerical data;
[0087] Feature extraction and association model building module: includes a feature extraction unit, an association model building unit, and a model training unit. The feature extraction unit performs multi-dimensional feature extraction on pre-processed audio data and operation sequence data to form a feature vector sequence and a time series feature vector sequence. The association model building unit inputs the feature sequences of audio and operation into the LSTM to learn temporal dynamic patterns, and then sends the LSTM output to the Transformer to capture the complex interactions and long-distance dependencies between features. A multimodal fusion layer is designed to generate a joint feature representation and build an association model. The model training unit trains the association model, receives real-time collected and pre-processed data, dynamically updates the status, and outputs analysis results in real time, including audio data, operation sequence data, and interrelated patterns.
[0088] Anomaly Detection and Judgment Module: This module includes an anomaly detection unit and a situation judgment unit. The anomaly detection unit uses a trained association model to monitor audio data, operation sequence data, device status and physical parameter data, and spectral energy distribution data in real time to identify abnormal data. The situation judgment unit then determines the situation of abnormal data detected by the anomaly detection unit, classifying it as instrument failure, normal exit, microphone anomaly, audience noise, and device switching.
[0089] Exception handling module: includes an exception handling unit and a log recording unit; wherein the exception handling unit executes the corresponding processing strategy according to the judgment result of the situation judgment unit; the log recording unit records detailed information;
[0090] Cloud collaboration and version management module: includes version recording and branch management unit, version comparison and merging unit and evaluation feedback and visualization unit; among them, the version recording and branch management unit records the tuner's operation steps and timestamps on the cloud server, generates a mixing version, supports multi-engineer collaboration, and creates independent version branches for parallel work; the version comparison and merging unit displays the differences between different mixing versions, supports version merging, allows engineers to selectively merge operations of other branches, automatically handles conflicts and provides conflict resolution suggestions; the evaluation feedback and visualization unit supports the evaluation of generated mixing versions, records evaluation results and feedback, and provides a graphical display of parameter differences.
[0091] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
Claims
1. A method for managing audio data based on a mixing console, characterized in that: include: Collect audio data, operation sequence data, equipment status and physical parameter data, and spectrum energy distribution data, and pre-process the collected data; Through time domain, frequency domain, and time-frequency analysis, we extract the time domain, frequency domain, and time-frequency features of audio data, as well as the timing and state features of operation sequence data. We use the long short-term memory (LSTM) network to capture the temporal dynamics of each feature, and the self-attention mechanism of the Transformer to capture the long-range dependencies and interactions between features. Combined with a multimodal fusion strategy, we organically combine the processed features to construct an association model for audio data and operation sequence data. The model analyzes the collected data to determine the on-site sound conditions and existing problems. When an instrument's audio signal, frequency band energy, or device parameters are abnormal, it determines that the instrument has failed or has left the stage normally. When the level, frequency band energy, or timestamp offset of the human voice activity are abnormal, it is determined to be a microphone anomaly. When the noise level, noise spectrum, or time matching in the audience area is abnormal, it is determined to be audience noise. When the audio signal or spectrum characteristics change, or when the device operation log and timestamp matching are abnormal, it is determined to be a device switch. The model takes corresponding measures according to the different judged situations. In a cloud-based mixing environment, the mixer's operations are recorded as independent steps and versions, allowing engineers to collaborate on the same audio, create and manage their own mixing versions, and support version comparison, merging, evaluation, and visual analysis of parameter differences.
2. The audio data management method based on a mixing console according to claim 1, characterized in that: The collecting of audio data, operation sequence data, device status and physical parameter data, and spectrum energy distribution data, and preprocessing of the collected data include: Audio data from instruments, vocals, and audience channels is collected through a high-precision multi-channel audio interface. Operation sequence data is collected through the console's internal logic interface. Device status and physical parameter data are collected through embedded sensors. Real-time spectrum energy distribution data is collected through a spectrum analyzer, recording energy changes in each frequency band and spectrum dynamic characteristics. Synchronize the collected data by adding timestamps during data collection to ensure that the data stream is aligned based on a unified time base; determine the target format that needs to be unified and use a programming language to convert all data into a unified format; normalize or standardize numerical data.
3. The audio data management method based on a mixing console according to claim 1, characterized in that: The method extracts the time domain, frequency domain and time-frequency features of audio data, as well as the timing and state features of sequence data through time domain, frequency domain and time-frequency analysis; uses the long short-term memory network (LSTM) to capture the temporal dynamics of each feature, and uses the self-attention mechanism of the Transformer to capture the long-range dependencies and interactions between features; Combined with the multimodal fusion strategy, the processed features are organically combined to build an association model between audio data and operation sequence data, including: The pre-processed audio data is subjected to time domain analysis, frequency domain analysis, and time-frequency analysis. Time domain analysis divides the signal into short time frames, calculates the root mean square energy and zero-crossing rate of each frame, and reveals the basic intensity characteristics of the signal and its dynamic characteristics that change over time. Frequency domain analysis uses Fourier transform to convert the signal from the time domain to the frequency domain, identifying the various frequency components that make up the signal and the corresponding energy distribution. Time-frequency analysis divides the signal into time-varying windows and performs frequency domain analysis within each window to generate a time-frequency spectrum that displays the time-varying characteristics of the signal's frequency content. Key features from these analysis results are extracted to form a sequence of time domain feature vectors, frequency domain feature vectors, and time-frequency feature vectors, which constitute an audio feature vector sequence. Extract timing features and state features from the preprocessed operation sequence data; process each event or state point in the operation sequence one by one in chronological order, review past operations based on the current time point, and extract timing features; query the value or state of each relevant component at the current time point, extract state features, and combine all the timing features and state features extracted for the current time point into a fixed-length numerical vector in a predefined order and format. Concatenate the feature vectors corresponding to each time point in the sequence in chronological order to obtain a two-dimensional sequence of timing feature vectors; The obtained audio feature vector sequence and the temporal feature vector sequence of the operation sequence are respectively input into their respective configured LSTM networks. The LSTM learns and outputs a representation of the dynamic change pattern of each sequence in the time dimension. The dynamic representation of the audio features and the dynamic representation of the operation features output by the LSTM are sent as input sequences to the Transformer model. The Transformer's self-attention mechanism is used to calculate the feature correlation weights within and between sequences. The multi-head attention mechanism and multi-layer perceptron are used to capture the complex interactive relationships and long-range dependencies between features. A multimodal fusion layer is designed to receive the temporal dynamic representation output by the LSTM and the interaction and dependency representation output by the Transformer. An attention mechanism is used to organically combine feature representations from different models that capture different information, generating a joint feature representation that simultaneously reflects audio content, operational behavior, and interactions. This joint feature representation is then input into the final classifier to construct an association model. Use labeled data to train the association model, deploy the trained association model, receive real-time collected and pre-processed audio data and operation sequence data, dynamically update the internal state, and output the association analysis results in real time, specifically the audio data, operation sequence data, and the patterns of their association.
4. The audio data management method based on a mixing console according to claim 1, characterized in that: When the audio signal, frequency band energy or device parameters of the musical instrument are abnormal, determining whether the musical instrument has failed or left the scene normally includes: The audio signal and operation sequence of the musical instrument are monitored in real time and analyzed in combination with the association model. When the energy value of a specific frequency band is continuously lower than the preset silence threshold within the preset fault silence time, and there are no valid performance operation instructions during this period, the association model predicts that this state is inconsistent with the current music context and determines that it is an instrument fault. When the energy value or harmonic structure of a specific frequency band of the signal exceeds the preset fault noise threshold within the preset fault noise time, and the amplitude of the change exceeds the preset fault allowable fluctuation range, and the association model fails to identify patterns related to playing techniques or environmental factors, it is determined that it is an instrument fault. In order to rule out poor contact of the device connection cable, the device status data related to the instrument channel is checked, and the spectrum characteristics of the abnormal period are analyzed. When the spectrum characteristics are different from those of the instrument's own fault, it is determined to be a line problem rather than an instrument fault. When it is detected that the instrument is pressed or receives a mute command, and after the command is issued, the signal energy value drops below the preset silence threshold within the preset departure decay time, and the association model confirms that the operation is consistent with the historical normal pattern, it is judged as a normal departure; when the signal energy value is lower than the preset silence threshold within the preset departure structure time, and the silence period matches the current music structure within the preset time tolerance threshold, and the association model does not detect any abnormal pattern, it is judged as a normal departure; when the signal energy value linearly decreases within the preset gradient time, its rate is lower than the preset departure permission threshold, and finally falls below the preset silence threshold, and the association model confirms that this process is consistent with the historical normal operation pattern, it is judged as a normal departure; to rule out the situation where the performer stops playing but the instrument is still in the on state and is just temporarily not in use, the device status data and operation sequence data are checked to see if there is any power operation or effector switching, and the spectrum is analyzed. If there is a harmonic structure unique to the instrument, it is judged that the instrument was turned on and temporarily not in use when the performance stopped, rather than a normal departure.
5. The audio data management method based on a mixing console according to claim 1, characterized in that: When the level, frequency band energy, or timestamp offset of the human voice activity is abnormal, it is determined that the microphone is abnormal, including: The collected vocal audio data is monitored in real time and analyzed in combination with the association model; the vocal area is identified using voice activity detection technology, and the vocal level, specific frequency band energy and timestamp sequence of the vocal activity in the area are extracted; the real-time extracted vocal level, frequency band energy and timestamp offset of the vocal activity are compared with the preset threshold; when the vocal level continuously exceeds the preset normal level range and the duration exceeds the preset sound abnormality duration threshold, the association model fails to identify the pattern related to the singer's deliberate volume adjustment and abnormal distance from the microphone, and it is judged as volume abnormality; when the energy level of non-human voice background noise continuously exceeds the preset noise threshold and the duration exceeds the preset noise duration threshold, the association model fails to identify the pattern related to the singer's breathing, environmental noise or singing skills, and it is judged as abnormal noise; when the timestamp sequence of the vocal activity has a long blank period and the duration exceeds the preset silence duration threshold, the association model fails to identify the pattern related to the singer's normal pause, improvisation or normal silence, and it is judged as a signal interruption.
6. The audio data management method based on a mixing console according to claim 1, characterized in that: When the noise level, noise spectrum or time matching in the audience area is abnormal, it is judged as audience noise, including: The audio signal from the microphone array in the audience area is monitored in real time and analyzed in combination with the association model; the real-time noise level, noise spectrum characteristics and time pattern of noise activity of the signal are extracted; the real-time extracted noise level, noise spectrum characteristics and time pattern of noise activity are compared with the preset threshold value; when the noise level continuously exceeds the preset maximum allowable noise level threshold, the noise spectrum characteristics continuously deviate from the preset normal range threshold, or the noise activity time pattern does not match the normal interaction pattern threshold of the current performance stage, and the duration of the abnormal state exceeds the preset minimum duration threshold for judging the audience noise, the association model determines that the noise pattern does not match the normal interaction pattern of the current performance stage and is judged to be audience noise; in order to exclude crosstalk from musical instruments in other areas, it is necessary to compare the signals of other channels. When similar sounds are detected in other channels at the same time and the spectral characteristics of the original instruments are retained, it is judged to be crosstalk from musical instruments in other areas rather than audience noise.
7. The audio data management method based on a mixing console according to claim 1, characterized in that: When the audio signal or spectrum characteristics change or the device operation log matches the timestamp abnormally, it is determined to be a device switch, including: The audio signals, spectrum feature changes, and operation logs and timestamps of related equipment of each channel of the mixing console are monitored in real time, and analyzed in combination with the association model; the detected audio signal mutation amplitude, the rate and amplitude of spectrum feature changes, and the matching time difference between the timestamp of the device operation log and the audio signal mutation timestamp are compared with the preset threshold; when the audio signal mutation amplitude exceeds the preset device switching signal mutation threshold, the spectrum feature change rate or amplitude exceeds the preset device switching spectrum change threshold, and the matching time difference between the timestamp of the device operation log and the audio signal mutation timestamp is within the preset device switching time matching tolerance threshold, or the device operation log records the switching instruction, and the duration of the abnormal state exceeds the preset device switching judgment minimum duration threshold, the association model confirms that the detected state meets the characteristic pattern of device switching and is judged as device switching.
8. The audio data management method based on a mixing console according to claim 1, characterized in that: The corresponding processing according to different judged situations includes: When an instrument failure is determined, the time point of the failure is marked on the timeline of the audio data, the audio channel corresponding to the failed instrument is marked, the volume of the channel is automatically reduced to silence, the technician is notified, and the details of the failure are recorded in the log; when it is determined to be a normal departure, the time point of the departure is marked on the timeline of the audio data, the audio channel corresponding to the departing instrument or singer is marked, the volume of the channel is controlled to smoothly reduce according to the preset fade-out curve until it is silent, and the departure information is recorded in the log; When it is determined that the microphone is abnormal, it is handled according to the specific situation; when the volume is abnormal, the automatic gain control algorithm is used to dynamically and smoothly adjust the gain of the microphone channel to restore it to the preset normal level range based on the singer's singing habits, historical data of song performances, the progress of the current song and the preset strategy; when it is abnormal noise, the time domain audio signal collected by the microphone is converted to the frequency domain through short-time Fourier transform to obtain the amplitude spectrum and phase spectrum of each frame signal, and the noise audio is short-time Fourier transformed to calculate the average amplitude spectrum as an estimate of the noise amplitude spectrum; using spectral subtraction, the estimated noise amplitude spectrum is subtracted from the signal amplitude spectrum, and the obtained pure voice amplitude spectrum is combined with the original signal phase spectrum to form a processed spectrum, which is then inversely short-time Fourier transformed and converted back to the time domain audio signal to obtain the noise-reduced audio; when the signal is interrupted, an attempt is made to automatically reconnect the microphone signal; if the reconnection is successful, the volume and equalization settings of the channel are restored according to the historical data and song progress, and if the reconnection fails, the channel is muted, the technician is notified to check and replace the equipment immediately, and the details of the signal interruption are recorded in the log; When an abnormal auditorium noise is detected, the specific time range in which the abnormal noise occurred is marked on the audio data timeline. All auditorium audio input channels within this time period are marked as abnormal noise. The auditorium volume balance is automatically adjusted based on the noise type, and the proportion of the auditorium volume in the main mix is reduced. The type, duration, and treatment measures of the abnormal noise are recorded in the log. When it is determined that a device switch occurs, the time point at which the switch occurs is marked on the timeline of the audio data, the audio channels involved in the switch are marked, the switch type is confirmed based on the device operation log, and the parameters of the audio channels corresponding to the new device are automatically adjusted to ensure that the audio signal is not discontinuous or sudden during the switching process. At the same time, detailed switching operations are recorded in the log.
9. The audio data management method based on a mixing console according to claim 1, characterized in that: In the cloud-based audio mixing environment, the mixer's operations are recorded as independent steps and versions, allowing engineers to collaborate on the same audio, create and manage their own mixing versions, and support version comparison, merging, evaluation, and visual analysis of parameter differences, including: The cloud server records and stores the tuner's operation steps and timestamps, and automatically generates a mixing version; allows multiple engineers to collaborate and create independent version branches; supports version comparison, displaying the audio waveforms, spectrum differences and corresponding operation differences of different mixing versions; supports version merging, allowing engineers to selectively merge operations and parameters from different version branches, automatically handle conflicts and provide conflict resolution suggestions; supports listening evaluation of the generated mixing version, records evaluation results and feedback; provides a graphical display of parameter differences, and displays parameter changes between different versions in the form of charts or timelines, making it easier for engineers to intuitively analyze the differences between versions.
10. An audio data management system based on a mixing console, using the audio data management method based on a mixing console according to any one of claims 1 to 9, characterized in that: include: Data acquisition and preprocessing module: includes a data acquisition unit and a data preprocessing unit; the data acquisition unit collects audio data, operation sequence data, device status and physical parameter data, and spectrum energy distribution data; the data preprocessing unit synchronizes the collected data, unifies the data format, and normalizes or standardizes the numerical data; Feature extraction and association model building module: includes a feature extraction unit, an association model building unit, and a model training unit. The feature extraction unit performs multi-dimensional feature extraction on pre-processed audio data and operation sequence data to form a feature vector sequence and a time series feature vector sequence. The association model building unit inputs the feature sequences of audio and operation into the LSTM to learn temporal dynamic patterns, and then sends the LSTM output to the Transformer to capture the complex interactions and long-distance dependencies between features. A multimodal fusion layer is designed to generate a joint feature representation and build an association model. The model training unit trains the association model, receives real-time collected and pre-processed data, dynamically updates the status, and outputs analysis results in real time, including audio data, operation sequence data, and interrelated patterns. Anomaly Detection and Judgment Module: This module includes an anomaly detection unit and a situation judgment unit. The anomaly detection unit uses a trained association model to monitor audio data, operation sequence data, device status and physical parameter data, and spectral energy distribution data in real time to identify abnormal data. The situation judgment unit then determines the situation of abnormal data detected by the anomaly detection unit, classifying it as instrument failure, normal exit, microphone anomaly, audience noise, and device switching. Exception handling module: includes an exception handling unit and a log recording unit; wherein the exception handling unit executes the corresponding processing strategy according to the judgment result of the situation judgment unit; the log recording unit records detailed information; Cloud collaboration and version management module: includes version recording and branch management unit, version comparison and merging unit and evaluation feedback and visualization unit; among them, the version recording and branch management unit records the tuner's operation steps and timestamps on the cloud server, generates a mixing version, supports multi-engineer collaboration, and creates independent version branches for parallel work; the version comparison and merging unit displays the differences between different mixing versions, supports version merging, allows engineers to selectively merge operations of other branches, automatically handles conflicts and provides conflict resolution suggestions; the evaluation feedback and visualization unit supports the evaluation of generated mixing versions, records evaluation results and feedback, and provides a graphical display of parameter differences.
Citation Information
Patent Citations
Audio expression method and system for device status information
CN108489519A
Digital sound console based on card insertion technology
CN114337887A
Digital tuning equipment early warning supervision system and method based on Internet of Things
CN116016119A
Method for monitoring abnormal operation of sound console based on historical data
CN116155426A
Sound console operation control system based on Internet of Things control
CN117596282A