Method and system for repairing audio data loss in AI recording scene
By using timestamps and verification codes to verify audio data in AI recording scenarios, combined with deep analysis and device status matching recovery algorithms, real-time repair and optimization of audio data is achieved, and the playback interruption and sound quality degradation caused by audio data loss is solved, improving the user experience.
Patent Information
- Application Number
- CN202510430060.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
During the audio playback and recording of Android devices, audio data is easily lost or damaged, resulting in playback interruption and sound quality degradation, and lack of real-time repair solutions, affecting the user experience.
It provides a method for repairing audio data loss in AI recording scenarios. By obtaining audio data, checking based on timestamps and verification codes, marking lost data fragments, conducting in-depth analysis to determine the data location and quantity, matching the device status selection recovery algorithm, performing data recovery and optimizing audio data through spectrum analysis.
It realizes real-time detection and repair of audio data loss in AI recording scenarios, improves the stability and user experience of audio applications, and overcomes the defects of the existing technology that cannot repair lost audio data in real time.
Smart Images

Figure CN119943090A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech processing, and in particular to a method and system for repairing audio data loss in an AI recording scenario. Background Art
[0002] With the widespread use of Android devices in multimedia applications such as audio playback and recording, the problem of audio data loss has gradually become a major pain point in user experience. During the transmission, storage or processing of audio data, data loss or corruption may occur, resulting in playback interruptions, reduced sound quality and other problems.
[0003] Currently, there is a lack of an efficient and intelligent solution to repair lost audio data in real time, which affects the stability of audio applications and user experience. Summary of the invention
[0004] The main purpose of the present invention is to provide a method and system for repairing audio data loss in an AI recording scenario, aiming to overcome the defect that the lost audio data cannot be repaired in real time.
[0005] To achieve the above object, the present invention provides a method for repairing audio data loss in an AI recording scenario, comprising the following steps: Obtain audio data from AI devices in recording scenarios; Based on the integrity of the timestamp and the check code, the audio data is checked to mark the audio segment with lost data; Perform in-depth analysis on the audio clips marked with missing data to determine the specific location of the missing data in the audio clips and the amount of missing data; Obtain the device status of the AI device, match the corresponding audio data recovery algorithm based on the device status, and perform data recovery on the audio segment where the lost data is located to obtain preliminary recovered audio data; The initially restored audio data is subjected to quality optimization processing based on spectrum analysis to obtain target audio data that meets playback requirements.
[0006] Further, based on the integrity of the timestamp and the check code, the audio data is checked, including: Timestamping the audio data at fixed time intervals, dividing the audio data into a plurality of data blocks, and generating a checksum for each data block; By comparing the continuity of adjacent timestamps and the correctness of the check code, it is determined whether the audio data is complete. If the timestamp is interrupted or the check code is wrong, the corresponding audio segment is marked as an audio segment with lost data.
[0007] Furthermore, the marked audio segment with lost data is deeply analyzed to determine the specific location of the lost data in the audio segment and the amount of lost data, including: Using machine learning algorithms, the frequency, amplitude, and phase characteristics of audio clips are extracted, and the characteristic differences between normal audio and audio clips with lost data are identified, thereby determining the specific location and amount of lost data.
[0008] Further, the audio segment where the lost data is located is restored to obtain initially restored audio data, including: The audio data within a preset range before and after the lost data is input into the audio data recovery algorithm, and the audio data recovery algorithm fills or reconstructs the lost part to obtain preliminary recovered audio data.
[0009] Furthermore, the audio data of the AI device in the recording scenario is obtained, including: Start the recording program of the AI device, put the recording operation and noise reduction operation into the same independent thread for processing, and extract the recording operation into an independent process; transfer data through AIDL to obtain the audio data obtained by recording; The audio data obtained by recording is stored based on the buffer pool, and another thread is opened to reduce the noise of the audio data in the buffer pool to obtain the audio data in the recording scene.
[0010] Furthermore, the corresponding audio data recovery algorithm is matched based on the device state, including: Inputting the device status into a deep neural network model, the device status including network connection quality, remaining storage space, current CPU load rate, and memory idle rate; In the hidden layer of the deep neural network model, a structure of alternating multiple convolutional layers and pooling layers is set; the convolutional layer uses convolution kernels of different sizes to extract features of the input device state and capture the correlation between device states at different scales; the pooling layer uses the maximum pooling operation to reduce the dimension of the features extracted by the convolutional layer to obtain key features.
[0011] After multiple convolution and pooling operations, the data is input into the fully connected layer for processing, the processed feature data is mapped to the index space of multiple preset audio data recovery algorithms, and the probability value of each recovery algorithm being selected is output; The recovery algorithm with the largest probability value is selected as the corresponding audio data recovery algorithm.
[0012] Further, after obtaining the target audio data that meets the playback requirements, the method further comprises: Comparing difference information in multiple dimensions between the audio data in the recording scenario and the target audio data that meets the playback requirements; Constructing a data matrix based on the difference information; obtaining device information of the AI device and constructing it into a character matrix; extracting matrix features of multiple dimensions of the data matrix and the character matrix; Performing mutation processing on the character matrix based on the data matrix to obtain a mutation matrix; A management authority code is generated based on the variation matrix and matrix features of multiple dimensions to perform authority management on the audio data that meets the playback requirements.
[0013] Furthermore, the character matrix is subjected to mutation processing based on the data matrix to obtain a mutation matrix, including: Performing quantum state encoding processing on the data matrix and the character matrix, mapping the elements in the data matrix and the character matrix to different quantum states of quantum bits, respectively, to obtain a quantum encoding matrix pair; Using quantum superposition and entanglement characteristics, a quantum entanglement operation is performed on the quantum coding matrix pair to generate an entangled quantum matrix; The entangled quantum matrix is customized to be transformed through a quantum gate operation sequence to obtain a variant quantum matrix; the quantum gate operation sequence is generated based on an algorithmic rule that changes dynamically over time; The variation quantum matrix is subjected to quantum measurement decoding processing to obtain a variation matrix.
[0014] Furthermore, generating a management authority code based on the variation matrix and matrix features of multiple dimensions includes: Combining the variation matrix and matrix features of multiple dimensions into a high-dimensional vector, and converting the high-dimensional vector into a polynomial function using a polynomial fitting algorithm; Based on the polynomial function, a corresponding two-dimensional curve is drawn on a two-dimensional plane; Constructing a three-dimensional surface graph in a three-dimensional space according to the coefficients and degrees of the polynomial function; Extracting features from the polynomial function, the two-dimensional curve, and the three-dimensional surface graph, and combining the extracted features into a combined feature vector; The combined feature vector is encrypted based on a chaotic encryption algorithm to obtain a ciphertext; the ciphertext is binary-encoded to obtain the management authority code.
[0015] The present invention also provides a system for repairing audio data loss in an AI recording scenario, comprising: The acquisition module is used to obtain the audio data of the AI device in the recording scenario; A verification module, used to verify the audio data based on the integrity of the timestamp and the check code, and mark the audio segment with lost data; An analysis module is used to perform in-depth analysis on the audio segment with the marked missing data to determine the specific location of the missing data in the audio segment and the amount of the missing data; A recovery module is used to obtain the device status of the AI device, match the corresponding audio data recovery algorithm based on the device status, and recover the audio segment where the lost data is located to obtain preliminary recovered audio data; The optimization module is used to perform quality optimization processing on the initially restored audio data based on spectrum analysis to obtain audio data that meets the playback requirements.
[0016] The present invention also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0017] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned methods are implemented.
[0018] The method and system for repairing audio data loss in an AI recording scenario provided by the present invention include: obtaining audio data of an AI device in a recording scenario; verifying the audio data based on the integrity of a timestamp and a checksum, and marking the audio segment of the lost data; performing in-depth analysis on the marked audio segment of the lost data to determine the specific location of the lost data in the audio segment and the amount of lost data; obtaining the device status of the AI device, matching the corresponding audio data recovery algorithm based on the device status, and performing data recovery on the audio segment where the lost data is located to obtain initially recovered audio data; performing quality optimization processing based on spectrum analysis on the initially recovered audio data to obtain target audio data that meets playback requirements. In the present invention, in a recording scenario, it is detected in real time whether there is lost data, and the lost data is repaired in real time, overcoming the defect that the lost audio data cannot be repaired in real time at present. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic diagram of the steps of a method for repairing audio data loss in an AI recording scenario in one embodiment of the present invention; Figure 2 It is a structural block diagram of a system for repairing audio data loss in an AI recording scenario in one embodiment of the present invention; Figure 3 It is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.
[0020] The implementation, functional features and advantages of the present invention will be further described in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0022] Reference Figure 1 In one embodiment of the present invention, a method for repairing audio data loss in an AI recording scenario is provided, comprising the following steps: Step S1, obtaining audio data of the AI device in a recording scenario; Step S2, verifying the audio data based on the integrity of the timestamp and the check code, and marking the audio segment with lost data; Step S3, performing in-depth analysis on the marked audio segment with lost data to determine the specific location of the lost data in the audio segment and the amount of lost data; Step S4, obtaining the device state of the AI device, matching the corresponding audio data recovery algorithm based on the device state, and performing data recovery on the audio segment where the lost data is located to obtain initially recovered audio data; Step S5, performing quality optimization processing based on spectrum analysis on the initially restored audio data to obtain target audio data that meets the playback requirements.
[0023] In this embodiment, as described in step S1 above, obtaining the audio data of the AI device in the recording scenario is to obtain the original data for subsequent processing, which is the basis of the entire repair process. Specifically, with the help of the built-in recording function module of the AI device, for example, using recording tools such as AlsaRecord, by setting appropriate parameters (such as sampling rate, number of channels, audio encoding format, etc.) to perform recording operations, the sound signal is converted into digital audio data. Different AI devices may use different recording interfaces or software frameworks, but the core is to collect sound in the environment and convert it into a processable data format.
[0024] As described in step S2 above, based on the integrity of the timestamp and the check code, the audio data is checked, the audio segments of the lost data are marked, the lost parts of the audio data are identified, and the problem areas are located for subsequent recovery work.
[0025] During the recording process, a timestamp is added to each audio data block to record the order and time information of data generation; at the same time, the checksum of each data block is calculated, such as CRC (cyclic redundancy check) code, MD5, etc. During the check, the timestamps of adjacent data blocks are compared to see if they are continuous. If there is an abnormal time interval, it is preliminarily determined that data is lost; then the checksum of the data block is recalculated and compared with the original checksum. If they are inconsistent, it is determined that there is a problem with the data block, and the audio segment containing the data block is marked as an audio segment with lost data. This double check mechanism can improve the accuracy of detection and avoid misjudgment.
[0026] As described in step S3 above, the marked audio segment of lost data is deeply analyzed to determine the specific location and amount of lost data in the audio segment, accurately locate the detailed information of the lost data, and provide a basis for selecting an appropriate recovery algorithm.
[0027] Use audio signal processing technology and machine learning algorithms for analysis. Starting from the audio signal features, extract features such as frequency, amplitude, and phase, and compare the feature differences between normal audio and lost data audio. For simple loss situations, the time domain continuity of the audio signal can be analyzed to locate the loss; for complex loss situations, use machine learning models (such as convolutional neural networks, CNN) for training and identification. The model learns a large number of normal and lost data audio samples, establishes a mapping relationship between features and loss locations and the amount of lost data, and thus accurately analyzes the marked audio segments.
[0028] As described in step S4 above, the device status of the AI device is obtained, the corresponding audio data recovery algorithm is matched based on the device status, and the audio segment where the lost data is located is recovered to obtain preliminarily recovered audio data. Specifically, obtain information such as the network status (such as network connection stability, bandwidth), storage status (remaining storage space, read and write speed), and computing resource status (CPU usage, memory occupancy) of the AI device. Establish a correspondence between the device status and the recovery algorithm. For example, when the network is good, the storage space is sufficient, and the computing resources are abundant, select a high-precision recovery algorithm based on a complex model, such as an audio reconstruction algorithm based on deep learning; when device resources are tight, select a simple and efficient algorithm, such as a linear interpolation algorithm. After the algorithm is selected, the audio segment where the lost data is located and related information are input into the algorithm, and the preliminarily recovered audio data is obtained after calculation.
[0029] As described in step S5 above, the initially recovered audio data is subjected to quality optimization processing based on spectrum analysis to obtain target audio data that meets the playback requirements. Using spectrum analysis technology, the audio data is converted from the time domain to the frequency domain, and its frequency components and energy distribution are analyzed. By comparing the spectrum characteristics of normal audio, the abnormal frequency components or energy deviations in the restored audio are found. The abnormal parts are adjusted and optimized through signal processing methods such as filtering and equalization. For example, a bandpass filter is used to remove high-frequency noise in the restored audio, and an equalizer is used to adjust the volume balance of different frequency bands, so that the optimized audio is closer to normal audio in terms of spectrum characteristics, meeting the auditory quality requirements during playback.
[0030] In one embodiment, the audio data is verified based on the integrity of the timestamp and the check code, including: Timestamping the audio data at fixed time intervals, dividing the audio data into a plurality of data blocks, and generating a checksum for each data block; By comparing the continuity of adjacent timestamps and the correctness of the check code, it is determined whether the audio data is complete. If the timestamp is interrupted or the check code is wrong, the corresponding audio segment is marked as an audio segment with lost data.
[0031] In this embodiment, during the collection or transmission of audio data, a timestamp is added to the audio data at a pre-set fixed time interval (such as every 10 milliseconds). A timestamp is an identifier that records the time sequence of data generation or processing, which can be accurate to milliseconds or even higher. By timestamping the audio data, not only can the audio data be given a time sequence, but the continuity of the data can also be checked based on the time sequence during subsequent verification. For example, if the start time of the audio data collection is 0 seconds and the fixed time interval is 10 milliseconds, the timestamp of the first data block may be marked as 0:00:00.010, the second data block as 0:00:00.020, and so on. In this way, the audio data has a clear sequence identification in the time dimension.
[0032] Divide the audio data into multiple data blocks, each of which contains a certain amount of audio samples. The size of the data block is usually determined based on actual application requirements and technical implementation, such as audio sampling rate, encoding format, and system performance. Generally speaking, the common data block size may be between a few hundred bytes and a few thousand bytes. The purpose of dividing the data blocks is to more conveniently manage and verify the audio data. Each data block can be used as an independent verification unit to improve the accuracy and efficiency of the verification.
[0033] A checksum is generated for each data block. The checksum is a piece of code calculated by a specific algorithm based on the data content in the data block. Common checksum algorithms include cyclic redundancy check (CRC), hash functions (such as MD5, SHA-1, etc.). Taking the CRC algorithm as an example, it performs a specific polynomial division operation on the binary data in the data block to obtain a remainder as a checksum. The role of the checksum is that when an error occurs during data transmission or storage, the recalculated checksum is likely to be inconsistent with the original checksum, so that data errors can be detected. Since the checksum is generated based on the content of the data block, any change in any bit of data in the data block will cause the checksum to change, so the checksum can effectively detect the integrity of the data.
[0034] When verifying audio data, check the timestamps of adjacent data blocks in sequence. Under normal circumstances, the difference between the timestamps of adjacent data blocks should be equal to the preset fixed time interval. For example, if the fixed time interval is 10 milliseconds, then the difference between the timestamps of two adjacent data blocks should always be 10 milliseconds. If the difference between adjacent timestamps is found to be greater than 10 milliseconds, it means that data may be lost during this period, that is, the timestamp has been interrupted. Timestamp interruption is one of the important clues to determine whether audio data may be lost, and it can help quickly locate the audio data location range where there may be problems.
[0035] For each data block, after receiving or reading the data, the checksum of the data block is recalculated and compared with the checksum originally generated and stored. If the recalculated checksum is exactly the same as the original checksum, it means that there is no data error in the transmission or storage of the data block; on the contrary, if the two are inconsistent, it means that the data in the data block has changed, which may be due to transmission errors, storage failures, etc. The correctness comparison of the checksum can accurately detect data errors within the data block, which is an important basis for judging the integrity of audio data.
[0036] When the timestamp is interrupted or the checksum is wrong, the audio segment containing these problem data blocks will be marked as an audio segment with lost data. The marking method can be to set a specific identification bit in the data structure, or to create an index table that records the location and range of the lost data. For example, a field can be added to the metadata of the audio data to mark whether a certain audio segment has data loss; or the start time, end time, data block number and other information of the audio segment with lost data can be recorded in a special error log. The purpose of marking the audio segment with lost data is to provide a clear goal and scope for subsequent data recovery work, so as to facilitate more targeted processing of lost data.
[0037] In one embodiment, performing in-depth analysis on the marked audio segment with lost data to determine the specific location of the lost data in the audio segment and the amount of lost data includes: Using machine learning algorithms, the frequency, amplitude, and phase characteristics of audio clips are extracted, and the characteristic differences between normal audio and audio clips with lost data are identified, thereby determining the specific location and amount of lost data.
[0038] In this embodiment, firstly, a complex signal such as audio is converted into quantifiable and analyzable feature data, which provides a basis for subsequent identification of the difference between normal and lost data audio segments.
[0039] Audio signals are composed of sine waves of different frequencies, and different sounds have unique frequency distributions. A common method for extracting frequency features is to use Fast Fourier Transform (FFT) to convert audio signals from the time domain to the frequency domain. Through FFT, the energy distribution of audio at different frequencies can be obtained. For example, vowels and consonants in speech signals have obvious differences in frequency. The high-frequency part may contain more consonant information, while the low-frequency part is related to vowels. In actual operation, the audio clip is divided into multiple small time windows, and the audio signal in each window is FFTed to obtain the frequency feature vector of the window.
[0040] Amplitude reflects the strength of an audio signal. In audio processing, amplitude features can be directly obtained from time domain signals. Common amplitude features include root mean square (RMS), peak value, etc. The root mean square value measures the average energy of an audio signal, while the peak value indicates the maximum strength of the signal at a certain moment. For example, in music audio, the amplitude of the climax is usually larger than that of the prelude. By calculating the amplitude features at different time points in an audio clip, a curve showing the amplitude changing over time can be obtained.
[0041] Phase describes the relative position of different frequency components of an audio signal. Although phase features are not as widely concerned as frequency and amplitude features in many audio processing tasks, in some cases, they are very important for accurately restoring audio signals. Methods such as Hilbert transform can be used to extract the phase information of audio signals. Phase features play an important role in processing multi-channel audio or in scenarios where accurate audio reconstruction is required.
[0042] By comparing the features of normal audio and audio clips with lost data, we can find the key feature differences that can distinguish the two, and provide a basis for determining the location and amount of lost data. First, we need to collect a large number of normal audio clips and audio clips with known lost data as training data. This data should be representative, covering different types of audio (such as speech, music, ambient sound, etc.) and different degrees of data loss.
[0043] Select appropriate machine learning algorithms, such as support vector machines (SVM), decision trees, neural networks, etc., to train the training data. Taking neural networks as an example, the extracted frequency, amplitude, and phase features are used as input, and the audio clips are marked as normal or missing data as output. By continuously adjusting the weights and biases of the neural network, the model can learn the feature differences between normal audio and audio clips with missing data. The trained model can classify new audio clips to determine whether they are audio clips with missing data. At the same time, by analyzing the decision-making process of the model, the features that have a greater impact on the classification results, that is, the key feature differences between normal audio and audio clips with missing data, can be found. For example, it may be found that the energy of the audio clip with missing data is significantly lower than that of normal audio at certain specific frequencies, or the amplitude changes are more irregular.
[0044] Based on the identified feature differences, the specific location of the lost data in the audio clip is accurately located, and the amount of lost data is calculated to provide accurate information for subsequent data recovery. By analyzing the audio clip frame by frame or window by window, the feature differences are combined to determine whether each small part is lost data. For example, if the frequency characteristics of a window are significantly different from normal audio, and the difference conforms to the feature pattern of the audio clip with lost data, it is considered that there is lost data in the window. By recording the start and end positions of these abnormal windows, the specific location of the lost data in the audio clip can be determined.
[0045] After determining the location of the lost data, the amount of lost data can be calculated based on the audio sampling rate, bit rate and other parameters. For example, if the audio sampling rate is 44100 Hz, that is, 44100 samples are collected per second, each sample is 16 bits, and the time range of the lost data is determined to be 1 second, the amount of lost data can be calculated to be 44100 × 16 bits. In actual calculations, factors such as the encoding format of the data may need to be considered to more accurately calculate the amount of lost data.
[0046] In one embodiment, data recovery is performed on the audio segment where the lost data is located to obtain initially recovered audio data, including: The audio data within a preset range before and after the lost data is input into the audio data recovery algorithm, and the audio data recovery algorithm fills or reconstructs the lost part to obtain preliminary recovered audio data.
[0047] In this embodiment, in the audio segment where the lost data has been marked, the starting point and the end point of the lost data are used as a reference, and the audio data before and after are selected according to the preset range parameters. The size of this preset range is set based on experience and understanding of audio characteristics, and different audio types and loss situations may require different ranges. For example, for voice audio, the audio data within 100 milliseconds before and after the lost data is set as a reference; and for music audio, due to the continuity of its melody and rhythm, the range will be expanded to 200-500 milliseconds before and after. Choosing the right range is critical. If the range is too small, it may not provide enough information for the algorithm to accurately recover the lost data. If the range is too large, it may introduce too much irrelevant information, increase the algorithm calculation burden and may affect the recovery effect.
[0048] With the help of a specially designed audio data recovery algorithm, the lost part of the data is processed based on the input surrounding audio data to restore the lost audio content. The audio data within the preset range before and after the lost data determined in the first step is accurately input into the audio data recovery algorithm in the format and data organization required by the algorithm. Different algorithms have different requirements for input data. Some algorithms may require audio data to be input in a specific sampling rate, number of channels, quantization accuracy and other formats, so the selected audio data may need to be converted to the corresponding format and preprocessed before input. For example, the algorithm requires the input of mono audio data with a sampling rate of 16kHz, while the original selected audio data is stereo with a sampling rate of 44.1kHz, so it needs to be converted to a format that meets the algorithm requirements before input.
[0049] Through algorithmic calculation and processing, based on the input surrounding audio data, attempts are made to fill in the gaps in the lost data, or to reconstruct the lost audio data so that it can be restored to its state before the loss to a certain extent.
[0050] Filling is a relatively simple and direct recovery method, which is suitable for situations where the amount of lost data is small or the audio features are relatively stable. A common filling method is based on interpolation algorithms, such as linear interpolation. Assuming that the audio sample values at two time points before and after the lost data are known, the linear interpolation algorithm will calculate the value of the middle missing sample according to the linear relationship between these two values. If the frequency and amplitude changes of the audio data are relatively stable, linear interpolation can better fill the gaps in the lost data and maintain the basic continuity of the audio. There is also a method based on spline interpolation, which can fit the lost data by constructing a smooth curve. Compared with linear interpolation, it can generate a smoother audio transition effect and reduce the sudden changes or unnaturalness that may occur in the audio after filling.
[0051] When the amount of lost data is large or the audio features are complex, reconstruction is required. Reconstruction algorithms based on machine learning have been widely used in recent years, such as audio generation models based on deep learning. Taking the generative adversarial network (GAN) as an example, it consists of two parts: the generator and the discriminator. The generator is responsible for trying to generate the audio data of the lost part based on the input surrounding audio data, and the discriminator determines the similarity between the generated audio data and the real audio data. Through continuous adversarial training between the generator and the discriminator, the generator gradually learns to generate more realistic audio data to reconstruct the lost part. There are also algorithms based on recurrent neural networks (RNN) and their variants (such as LSTM, GRU), which can process the time series characteristics of audio data and reconstruct the lost part by learning the front-to-back dependencies of the audio. They are especially suitable for audio with obvious time series characteristics, such as speech and music with strong melody.
[0052] After being processed by the audio data recovery algorithm, a preliminary version of the audio data that restores the lost part is obtained. Although there may still be some imperfections, it provides a basis for further optimization and processing. After the algorithm completes the filling or reconstruction of the lost part, the restored audio data of the lost part is integrated with the non-lost part of the original audio. When integrating, it is necessary to ensure the continuity of the audio data on the timeline and the smooth transition of the audio features to avoid obvious breakpoints or sudden changes in audio quality. For example, when splicing the restored part and the non-lost part of the audio, it is necessary to perform appropriate fusion processing on the audio signal at the junction, such as using a fade-in and fade-out method to make the transition of the audio at the junction more natural.
[0053] In one embodiment, obtaining audio data of an AI device in a recording scenario includes: Start the recording program of the AI device, put the recording operation and noise reduction operation into the same independent thread for processing, and extract the recording operation into an independent process; transfer data through AIDL to obtain the audio data obtained by recording; The audio data obtained by recording is stored based on the buffer pool, and another thread is opened to reduce the noise of the audio data in the buffer pool to obtain the audio data in the recording scene.
[0054] In this embodiment, first, the predefined recording interface or related function in the AI device system is called to trigger the start of the recording program. This process usually involves interaction with the underlying hardware driver of the device, so that the recording device (such as a microphone) starts to collect sound signals and converts them into electrical signals. For example, in an AI device based on the Android system, the MediaRecorder class or AudioRecord class provided by the system can be used to start the recording function, and the basic properties of the recording can be determined by setting relevant parameters (such as audio source, sampling rate, number of channels, encoding format, etc.).
[0055] Under the thread management mechanism of the operating system, create a new independent thread. In this thread, write the code logic for the recording operation and the noise reduction operation at the same time. The recording operation continuously collects audio data according to the set parameters, and the noise reduction operation processes the collected audio data in real time. For example, use the noise reduction algorithm in some open source audio processing libraries (such as FFmpeg) to perform real-time filtering and denoising operations on the recorded data. By placing these two operations in the same thread, they can share resources within the thread, reduce the overhead of data transmission and synchronization, and improve overall processing efficiency.
[0056] Then, according to the process management mechanism of the operating system, a new process is created to perform the recording operation. The way to create a process is different in different operating systems. In the Linux system, the fork() function can be used to create a child process to perform the recording task; in the Windows system, the CreateProcess() function can be used to implement a similar function. Placing the recording operation in an independent process gives it independent memory space and system resources, isolating it from other processes, and reducing the risk of recording failure due to the crash or abnormality of other processes.
[0057] AIDL (Android Interface Definition Language) is an inter-process communication mechanism unique to the Android system. First, define an AIDL interface file and declare the data types and methods that need to be passed. For example, define an interface for transmitting audio data, which includes methods for sending audio data. Implement this AIDL interface in the recording process and pass the audio data obtained by the recording through the interface method. In the receiving process, bind this AIDL service to obtain the audio data. AIDL will automatically handle the serialization and deserialization of data to ensure that data is transmitted accurately between different processes.
[0058] Then, create a buffer pool data structure, which can be a queue, linked list or other suitable storage structure. After the recording process obtains the audio data, it will be stored in the buffer pool in sequence. The size of the buffer pool can be adjusted according to actual needs and device resources. For example, if the device memory is sufficient and the recording time is long, a larger buffer pool can be set to store more audio data; if the device resources are limited, a smaller buffer pool can be set to ensure the rational use of resources. The buffer pool can also use a circular cache method. When the buffer pool is full, the new data will overwrite the earliest stored data to ensure the real-time nature of the data.
[0059] Finally, a new thread is created to perform the noise reduction task. This thread continuously reads audio data from the cache pool and processes the audio data using noise reduction algorithms (such as Fourier transform-based spectrum analysis noise reduction algorithms, machine learning-based noise reduction algorithms, etc.). After the noise reduction process is completed, the obtained audio data is high-quality audio data that meets the requirements of the recording scene. By opening a separate thread for noise reduction, complex noise reduction calculations in the threads of recording and other operations are avoided, thread blockage is prevented, and the overall performance of the system is improved.
[0060] In one embodiment, the audio data recovery algorithm corresponding to the device state matching includes: Inputting the device status into a deep neural network model, the device status including network connection quality, remaining storage space, current CPU load rate, and memory idle rate; In the hidden layer of the deep neural network model, a structure of alternating multiple convolutional layers and pooling layers is set; the convolutional layer uses convolution kernels of different sizes to extract features of the input device state and capture the correlation between device states at different scales; the pooling layer uses the maximum pooling operation to reduce the dimension of the features extracted by the convolutional layer to obtain key features.
[0061] After multiple convolution and pooling operations, the data is input into the fully connected layer for processing, the processed feature data is mapped to the index space of multiple preset audio data recovery algorithms, and the probability value of each recovery algorithm being selected is output; The recovery algorithm with the largest probability value is selected as the corresponding audio data recovery algorithm.
[0062] In this embodiment, first, the device status data such as network connection quality, remaining storage space, CPU current load rate, and memory idle rate are sorted and preprocessed. For example, the network connection quality may be expressed as quantitative indicators such as signal strength and bandwidth; the remaining storage space and memory idle rate can be directly presented in the form of bytes or percentage of total capacity; the CPU current load rate is a value indicating how busy the CPU is. These data may have different dimensions and value ranges, so they need to be normalized and mapped to a unified interval (such as [0, 1]) for model processing. The processed device status data is used as the input of the model and is passed to the deep neural network model in the input format specified by the model (such as vector form).
[0063] The convolution layer uses convolution kernels of different sizes, which is one of the key innovations of this technology. Convolution kernels of different sizes can capture the correlation between device status data at different scales. Smaller convolution kernels are suitable for capturing local, detailed features, such as the relationship between small fluctuations in network connection quality in a short period of time and instantaneous changes in CPU load; larger convolution kernels are better at obtaining global, macro features, such as the long-term correlation between remaining storage space and memory idle rate when the overall network environment is stable. Through the convolution operation, the convolution kernel is slid on the input data, and a weighted sum is performed on each position to generate a feature map. These feature maps contain the device status feature information extracted by convolution, enriching the expression of the data.
[0064] The pooling layer uses the maximum pooling operation, and its main function is to reduce the dimension of the features extracted by the convolution layer. In the maximum pooling process, the feature map is divided into small areas (such as 2x2 pixel blocks), and then the maximum value in each area is selected as the representative value of the area. This can greatly reduce the amount of data and reduce the computational complexity of the model without losing too much key information. At the same time, maximum pooling also has a certain anti-interference ability, which can enhance the robustness of the model to small changes in input data. For example, when there is some slight noise or fluctuation in the device status data, maximum pooling can ignore these unimportant details and retain the most representative features.
[0065] After multiple convolution and pooling operations, the features of the data have been highly extracted and compressed. The fully connected layer processes these features comprehensively. Each of its neurons is connected to all neurons in the previous layer, which can fully learn the complex nonlinear relationship between features. The fully connected layer performs weighted summation of the input features through the weight matrix and bias vector, and performs nonlinear transformation through activation functions (such as ReLU function) to further integrate and optimize the features.
[0066] The feature data processed by the fully connected layer is mapped to the index space of multiple preset audio data recovery algorithms. This mapping process is achieved through a series of weight matrix calculations. The weights learned by the model will establish a connection between the feature data and different recovery algorithms. Finally, the model outputs the probability value of each recovery algorithm being selected. These probability values reflect the applicability and rationality of each recovery algorithm under the current device state. For example, if the model outputs a high probability value for a complex recovery algorithm based on deep learning, it means that the algorithm is more suitable for audio data recovery when the current device resources are sufficient (such as good network connection quality, low CPU load, and high memory idle rate).
[0067] Finally, among the probability values of each recovery algorithm output by the model, the algorithm with the largest probability value is found through comparison operation. This algorithm is considered to be the most suitable audio data recovery algorithm under the current device state. For example, if the model outputs the probability values of three recovery algorithms based on linear interpolation, machine learning, and deep learning, which are 0.2, 0.3, and 0.5 respectively, then the recovery algorithm based on deep learning is selected as the corresponding audio data recovery algorithm. This probability-based selection method makes full use of the deep neural network model's ability to learn the complex relationship between device status and recovery algorithms, realizes the intelligent selection of recovery algorithms, and improves the efficiency and quality of audio data recovery.
[0068] In one embodiment, after obtaining the target audio data that meets the playback requirements, the method includes: Comparing difference information in multiple dimensions between the audio data in the recording scenario and the target audio data that meets the playback requirements; Constructing a data matrix based on the difference information; obtaining device information of the AI device and constructing it into a character matrix; extracting matrix features of multiple dimensions of the data matrix and the character matrix; Performing mutation processing on the character matrix based on the data matrix to obtain a mutation matrix; A management authority code is generated based on the variation matrix and matrix features of multiple dimensions to perform authority management on the audio data that meets the playback requirements.
[0069] In this embodiment, first, comparison is made from multiple dimensions, such as the time domain dimension of the audio, comparing the amplitude, duration, waveform shape, etc. of the audio; the frequency domain dimension, comparing the frequency components, frequency distribution, energy spectrum, etc.; and the comparison can also be made from the perceptual dimensions such as the timbre and sound quality of the audio. The difference values in these dimensions are calculated using a digital signal processing algorithm, for example, by calculating the difference in the amplitude of two audio signals at the same time point, the difference in the energy of the same frequency interval in the spectrum diagram, etc., a series of quantized difference data is obtained.
[0070] The size of the matrix is determined according to the dimension and quantity of the difference information. For example, if M difference data points are obtained from N dimensions, a data matrix with N rows and M columns can be constructed. The difference data of each dimension is filled into the corresponding rows and columns of the matrix in turn, so that each element in the matrix represents the audio difference information under a specific dimension. In this way, the audio differences can be systematically stored and analyzed through the data matrix.
[0071] Then, the relevant information of the AI device is encoded and converted into a matrix form that is easy to process. The device information is one of the important factors in generating the permission code, and is combined with the audio difference information to increase the uniqueness and security of the permission code. Collect various types of information about AI devices, such as device model, device ID, hardware configuration parameters (such as CPU model, memory capacity), software version, etc. Encode this information, for example, using ASCII code or other custom encoding methods to convert character information into numerical values. Then, determine the size of the character matrix according to the type and quantity of device information, and fill the encoded numerical values into the matrix according to certain rules to form a character matrix.
[0072] Then, key features are extracted from the two matrices. These features can represent the core information of the matrices and are used for subsequent mutation processing and permission code generation, reducing the amount of data while retaining key information, improving processing efficiency and the accuracy of permission codes. For data matrices and character matrices, a variety of feature extraction methods can be used. For example, principal component analysis (PCA) can be used to reduce the dimension of the matrix and obtain the main components as features; or the eigenvalues and eigenvectors of the matrix can be calculated and used as the features of the matrix.
[0073] Furthermore, the character matrix is mutated through the data matrix so that the information of the two matrices can be integrated with each other, which further increases the complexity and uniqueness of the permission code generation and improves the security of permission management.
[0074] Specifically, a mutation rule is defined, for example, operating the elements at corresponding positions in the character matrix according to the values of the elements in the data matrix. It can be an arithmetic operation, such as addition and multiplication, or a complex encryption operation. For example, the value of an element in the data matrix is used as a key to perform encryption transformation on the elements at corresponding positions in the character matrix. All elements of the character matrix are processed according to this rule to obtain a mutated matrix. The mutated matrix integrates audio difference information and device information and is highly unique.
[0075] The variation matrix and matrix features are used to generate a unique management permission code, which is used to manage the permissions of audio data that meets the playback requirements, ensuring that only authorized devices or users can access and use the audio data. The variation matrix and matrix features of multiple dimensions are merged, for example, the matrix features are input as parameters into a hash function, and the variation matrix is used as the input data of the hash function. The hash function calculates the input data and generates a hash value of a fixed length. This hash value is the management permission code. Other encryption algorithms, such as symmetric encryption algorithms or asymmetric encryption algorithms, can also be used to generate permission codes in combination with variation matrices and matrix features. The generated permission code can be stored locally on the device or on the server to verify the user or device's access rights to audio data.
[0076] In one embodiment, performing mutation processing on the character matrix based on the data matrix to obtain a mutation matrix includes: Performing quantum state encoding processing on the data matrix and the character matrix, mapping the elements in the data matrix and the character matrix to different quantum states of quantum bits, respectively, to obtain a quantum encoding matrix pair; Using quantum superposition and entanglement characteristics, a quantum entanglement operation is performed on the quantum coding matrix pair to generate an entangled quantum matrix; The entangled quantum matrix is customized to be transformed through a quantum gate operation sequence to obtain a variant quantum matrix; the quantum gate operation sequence is generated based on an algorithmic rule that changes dynamically over time; The variation quantum matrix is subjected to quantum measurement decoding processing to obtain a variation matrix.
[0077] In this embodiment, the matrix data is converted into a quantum state representation so that subsequent processing can be performed using the characteristics of quantum computing. Quantum bits have characteristics such as superposition and entanglement, which can bring advantages to data processing that classical computing does not have. For each element in the data matrix and the character matrix, it is mapped to different quantum states of the quantum bit according to the encoding rules. For example, the element can be mapped to different superposition state coefficients according to its numerical value. In this way, each element in the data matrix and the character matrix is encoded as a quantum state of a quantum bit, thereby obtaining two quantum coding matrices to form a quantum coding matrix pair.
[0078] Through quantum entanglement, a strong correlation is established between two quantum coding matrices, so that the state change of one matrix will instantly affect the state of the other matrix, thereby increasing the complexity and correlation of the data.
[0079] Quantum entanglement refers to a special correlation between multiple quantum bits, which makes their quantum states cannot be described independently. Quantum entanglement gates (such as CNOT gates) are used to operate on the corresponding quantum bits in the quantum coding matrix pair. For example, for the quantum bits in corresponding positions in the data matrix and the character matrix, the CNOT gate can be used to entangle them. The CNOT gate takes a control bit and a target bit as input. If the control bit is in the "1" state, the state of the target bit is flipped; if the control bit is in the "0" state, the state of the target bit remains unchanged. By performing similar entanglement operations on each pair of corresponding quantum bits in the quantum coding matrix pair, the quantum bits in the two matrices are entangled with each other to form an overall entangled quantum matrix.
[0080] Then, complex transformations are performed on the entangled quantum matrix to further increase the randomness and uniqueness of the data, generate a quantum matrix with high variability, and provide higher security for the management authority code generated subsequently. Quantum gates are the basic logic units used to operate quantum bits in quantum computing. Common quantum gates include Pauli gates (such as X gates, Y gates, and Z gates), Hadamard gates, and phase gates.
[0081] Design an algorithmic rule that changes dynamically over time to generate a sequence of quantum gate operations. For example, the quantum gates to be used and the order in which they are operated can be determined based on the current timestamp, the system's random number generator, or other dynamic factors. This sequence of quantum gate operations is applied to the entangled quantum matrix in turn, and each quantum bit in the matrix is operated on. Each operation changes the quantum state of the quantum bit. After a series of quantum gate operations, the entangled quantum matrix is transformed into a variant quantum matrix, the state of which is highly uncertain and unique.
[0082] Then, the variant quantum matrix of the quantum state is converted back to the classical matrix form for subsequent use in generating management authority codes. Quantum measurement is the process of collapsing a quantum state into a classical state. Each quantum bit in the variant quantum matrix is measured, and the measurement result causes the quantum bit to collapse to a "0" or "1" state.
[0083] Reverse decoding is performed based on the measurement results and previous encoding rules to convert the measurement results of the quantum bits into classical values. For example, if the matrix element values were previously mapped to superposition state coefficients, then now the corresponding matrix element values are inferred based on the measured "0" or "1" state combined with the previous mapping relationship. The measurement and decoding results of all quantum bits are recombined into a matrix form to obtain the final variant matrix. This variant matrix combines the information of the data matrix and the character matrix, and has undergone complex transformations of quantum processing. It has high security and uniqueness and can be used to generate management authority codes.
[0084] In one embodiment, generating a management authority code based on the variation matrix and matrix features of multiple dimensions includes: Combining the variation matrix and matrix features of multiple dimensions into a high-dimensional vector, and converting the high-dimensional vector into a polynomial function using a polynomial fitting algorithm; Based on the polynomial function, a corresponding two-dimensional curve is drawn on a two-dimensional plane; Constructing a three-dimensional surface graph in a three-dimensional space according to the coefficients and degrees of the polynomial function; Extracting features from the polynomial function, the two-dimensional curve, and the three-dimensional surface graph, and combining the extracted features into a combined feature vector; The combined feature vector is encrypted based on a chaotic encryption algorithm to obtain a ciphertext; the ciphertext is binary-encoded to obtain the management authority code.
[0085] In this embodiment, the variation matrix is first stretched into a long one-dimensional vector in a specific order (such as row by row or column by column), and the matrix features of multiple dimensions are also organized into vector form. Then, these vectors are connected in sequence to form a high-dimensional vector containing all the information. Then, the polynomial fitting algorithm is used to process this high-dimensional vector. The algorithm will find a suitable polynomial so that this polynomial conforms to the data characteristics represented by the high-dimensional vector as much as possible, thereby converting the high-dimensional vector into a polynomial function.
[0086] The abstract polynomial function is presented in a visual way, and the characteristics of the data are mined and displayed from a two-dimensional perspective, providing more information dimensions for the subsequent generation of management permission codes. In a two-dimensional plane coordinate system, a series of points that are evenly distributed or have a specific pattern are selected as the values of the horizontal coordinates. Substitute these horizontal coordinate values into the previously obtained polynomial function to calculate the corresponding vertical coordinate values. Then, the points corresponding to these horizontal and vertical coordinates are plotted on the plane, and finally these points are connected in sequence to obtain the corresponding curve of this polynomial function on the two-dimensional plane.
[0087] Further utilizing the information contained in the polynomial function, by constructing a three-dimensional graph, increasing the dimension and complexity of the data, more unique features can be mined, and the security and uniqueness of the management authority code can be improved. Specifically, a three-dimensional space coordinate system is constructed, and the degree of the polynomial function is used to determine the value of one axis in the coordinate system, and then the different coefficients of the polynomial function correspond to the values of the other two axes. According to different combinations of the coefficients and degrees of the polynomial function, a series of points are determined in the three-dimensional space. These points are connected to construct a three-dimensional surface graph. The shape, curvature and other features of this graph contain information about the polynomial function, and also indirectly contain information about the variation matrix and matrix features.
[0088] Extract key information that can represent its characteristics from different forms of data presentation (polynomial function, two-dimensional curve, three-dimensional surface graphics), and then integrate this information to form a comprehensive feature vector, so as to more comprehensively reflect the characteristics of the original data and provide richer materials for generating management authority codes. For polynomial functions, extract some attributes that can reflect its characteristics, such as coefficients, highest order, etc.; for two-dimensional curves, extract features such as the turning point of the curve, the degree of fluctuation of the curve, and the intersection of the curve and the coordinate axis; for three-dimensional surface graphics, extract the degree of curvature of the surface, the concave and convexity of the surface, and the stretching in different directions. Arrange the features extracted from these three data forms and arrange them into a vector in a certain order. This vector is the combined feature vector.
[0089] The combined feature vector is encrypted and converted into a difficult-to-crack ciphertext form, and then the final management permission code is generated through binary encoding to ensure the security of access rights to the audio data. The original information can only be obtained through a specific decryption method, preventing the permission code from being easily stolen or forged.
[0090] The chaotic encryption algorithm is used to encrypt the combined feature vector. The chaotic encryption algorithm uses the characteristics of the chaotic system to perform complex transformations on the combined feature vector, turning it into a seemingly random ciphertext form. Then, the generated ciphertext is binary-encoded and converted into a binary sequence consisting of 0 and 1. This binary sequence is the final management permission code. When verifying the permission of audio data, only those with the correct decryption method and the corresponding permission code can pass the verification and access the audio data.
[0091] In one embodiment, generating a management authority code based on the variation matrix and matrix features of multiple dimensions includes: Combining the characters in the variation matrix in sequence to obtain a variation character sequence; Simulate the matrix characteristics of each dimension into a curve respectively; Connecting the ends of the curves to obtain a combined curve, and superimposing the combined curve on the variant character sequence, and calculating the spatial relationship between each character in the variant character sequence and the combined curve, and taking the character whose spatial relationship meets the preset conditions as the target character; The target characters are combined to obtain the management authority code.
[0092] In this embodiment, first, the characters in the variation matrix are integrated in order to form a character sequence, in preparation for subsequent association and analysis with the curve. The character sequence contains the information of the variation matrix and is one of the basic data for generating the management authority code. To clarify the order of arrangement of the characters in the variation matrix, the characters in the matrix can usually be taken out one by one in the order of row priority or column priority of the matrix. For example, if the variation matrix is a two-dimensional matrix, starting from the first character of the first row, all the characters of the row are taken out in sequence in a row-first manner, and then the characters of the second row are taken, and so on, until the characters of the entire matrix are taken out. The taken characters are connected in sequence to obtain a variation character sequence.
[0093] Furthermore, the abstract matrix features are transformed into intuitive curve forms to facilitate subsequent spatial relationship analysis and feature extraction. The matrix features of different dimensions show their changing rules and characteristics in the form of curves, providing more information dimensions for generating management authority codes. For each dimension of the matrix features, according to its data characteristics and changing rules, a suitable method is selected to simulate the curve. For example, if the matrix features are a series of values, these values can be used as the ordinates, and the sequence number or time can be used as the abscissa to draw a continuous curve through interpolation and other methods. Assuming that the matrix features of a certain dimension are a set of values [2, 4, 6, 8], these values can be respectively corresponded to the points with abscissas of 1, 2, 3, and 4, and then these points can be connected into a curve using linear interpolation or other more complex interpolation methods. Different curve models, such as sine curves, polynomial curves, etc., can also be used according to the type of matrix features to better fit and represent the changes in features.
[0094] Then, combine the curves of multiple dimensions into an overall curve and superimpose it with the variant character sequence. By analyzing the spatial relationship between the characters and the curves, select the characters that meet the specific conditions. These characters will be used as the key elements for generating the management permission code, making the generation of the permission code have unique rules and complexity. Connect the curves of each dimension obtained in the previous step in a certain order to form a continuous combined curve. For example, there are two curves A and B. Connect the end point of curve A with the starting point of curve B to obtain a combined curve.
[0095] Superimpose the combined curve with the variant character sequence. The superposition here can be understood as displaying the character sequence and the curve simultaneously in a unified space or coordinate system. For example, the character sequence can be placed on a plane, and then the combined curve can be drawn on the same plane, so that the character and the curve have a relative positional relationship in space. Define a method for calculating the spatial relationship between the character and the combined curve, such as calculating the distance from the character to the curve, whether the character is above or below the curve, etc. Preset some conditions, such as the distance from the character to the curve is less than a certain threshold, or the character is in a specific area of the curve, etc. For each character in the variant character sequence, calculate its spatial relationship with the combined curve. When the spatial relationship of a character meets the preset conditions, the character is marked as the target character.
[0096] Finally, the selected target characters are combined to form the final management permission code, which is generated based on the comprehensive information of the variation matrix and matrix features. It has uniqueness and complexity and is used to manage the permissions of audio data that meets the playback requirements. According to the order in which the target characters are selected in the variation character sequence, these target characters are connected in sequence to obtain a new character sequence, which is the management permission code. The above management permission code can be used to verify the access rights of users or devices to audio data. Only with the correct management permission code can the corresponding operations be performed on the audio data.
[0097] Reference Figure 2 In another embodiment of the present invention, a system for repairing audio data loss in an AI recording scenario is provided, comprising: The acquisition module is used to obtain the audio data of the AI device in the recording scenario; A verification module, used to verify the audio data based on the integrity of the timestamp and the check code, and mark the audio segment with lost data; An analysis module is used to perform in-depth analysis on the audio segment with the marked missing data to determine the specific location of the missing data in the audio segment and the amount of the missing data; A recovery module is used to obtain the device status of the AI device, match the corresponding audio data recovery algorithm based on the device status, and recover the audio segment where the lost data is located to obtain preliminary recovered audio data; The optimization module is used to perform quality optimization processing on the initially restored audio data based on spectrum analysis to obtain audio data that meets the playback requirements.
[0098] In this embodiment, for the specific implementation of each module in the above system embodiment, please refer to the above method embodiment, which will not be repeated here.
[0099] Reference Figure 3In an embodiment of the present invention, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. Among them, the processor designed by the computer is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0100] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0101] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0102] In summary, the method and system for repairing audio data loss in an AI recording scenario provided in an embodiment of the present invention include: obtaining audio data of an AI device in a recording scenario; verifying the audio data based on the integrity of a timestamp and a checksum, and marking the audio segment of the lost data; performing an in-depth analysis on the marked audio segment of the lost data to determine the specific location of the lost data in the audio segment and the amount of lost data; obtaining the device status of the AI device, matching the corresponding audio data recovery algorithm based on the device status, and performing data recovery on the audio segment where the lost data is located to obtain initially recovered audio data; performing quality optimization processing based on spectrum analysis on the initially recovered audio data to obtain target audio data that meets the playback requirements. In the present invention, in a recording scenario, it is detected in real time whether there is lost data, and the lost data is repaired in real time, overcoming the defect that the lost audio data cannot be repaired in real time at present.
[0103] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided by the present invention and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM.
[0104] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0105] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for repairing audio data loss in an AI recording scenario, characterized in that: The following steps are involved: Obtain audio data from AI devices in recording scenarios; Based on the integrity of the timestamp and the check code, the audio data is checked to mark the audio segment with lost data; Perform in-depth analysis on the audio clips marked with missing data to determine the specific location of the missing data in the audio clips and the amount of missing data; Obtain the device status of the AI device, match the corresponding audio data recovery algorithm based on the device status, and perform data recovery on the audio segment where the lost data is located to obtain preliminary recovered audio data; The initially restored audio data is subjected to quality optimization processing based on spectrum analysis to obtain target audio data that meets playback requirements.
2. The method for repairing audio data loss in an AI recording scenario according to claim 1, characterized in that: Based on the integrity of the timestamp and the check code, the audio data is checked, including: Timestamping the audio data at fixed time intervals, dividing the audio data into a plurality of data blocks, and generating a checksum for each data block; By comparing the continuity of adjacent timestamps and the correctness of the check code, it is determined whether the audio data is complete. If the timestamp is interrupted or the check code is wrong, the corresponding audio segment is marked as an audio segment with lost data.
3. The method for repairing audio data loss in an AI recording scenario according to claim 1, characterized in that: Perform in-depth analysis on the audio clips marked with missing data to determine the specific location and amount of missing data in the audio clips, including: Using machine learning algorithms, the frequency, amplitude, and phase characteristics of audio clips are extracted, and the characteristic differences between normal audio and audio clips with lost data are identified, thereby determining the specific location and amount of lost data.
4. The method for repairing audio data loss in an AI recording scenario according to claim 1, characterized in that: Perform data recovery on the audio segment where the lost data is located to obtain preliminary recovered audio data, including: The audio data within a preset range before and after the lost data is input into the audio data recovery algorithm, and the audio data recovery algorithm fills or reconstructs the lost part to obtain preliminary recovered audio data.
5. The method for repairing audio data loss in an AI recording scenario according to claim 1, characterized in that: Get the audio data of the AI device in the recording scenario, including: Start the recording program of the AI device, put the recording operation and noise reduction operation into the same independent thread for processing, and extract the recording operation into an independent process; transfer data through AIDL to obtain the audio data obtained by recording; The audio data obtained by recording is stored based on the buffer pool, and another thread is opened to reduce the noise of the audio data in the buffer pool to obtain the audio data in the recording scene.
6. The method for repairing audio data loss in an AI recording scenario according to claim 1, characterized in that: The audio data recovery algorithm corresponding to the device state matching includes: Inputting the device status into a deep neural network model, the device status including network connection quality, remaining storage space, current CPU load rate, and memory idle rate; In the hidden layer of the deep neural network model, a structure of alternating multiple convolutional layers and pooling layers is set; the convolutional layer uses convolution kernels of different sizes to extract features of the input device states and capture the correlation between device states at different scales; the pooling layer uses a maximum pooling operation to reduce the dimension of the features extracted by the convolutional layer to obtain key features; After multiple convolution and pooling operations, the data is input into the fully connected layer for processing, the processed feature data is mapped to the index space of multiple preset audio data recovery algorithms, and the probability value of each recovery algorithm being selected is output; The recovery algorithm with the largest probability value is selected as the corresponding audio data recovery algorithm.
7. The method for repairing audio data loss in an AI recording scenario according to claim 1, characterized in that: After obtaining the target audio data that meets the playback requirements, the method includes: Comparing difference information in multiple dimensions between the audio data in the recording scenario and the target audio data that meets the playback requirements; Constructing a data matrix based on the difference information; obtaining device information of the AI device and constructing it into a character matrix; extracting matrix features of multiple dimensions of the data matrix and the character matrix; Performing mutation processing on the character matrix based on the data matrix to obtain a mutation matrix; A management authority code is generated based on the variation matrix and matrix features of multiple dimensions to perform authority management on the audio data that meets the playback requirements.
8. The method for repairing audio data loss in an AI recording scenario according to claim 7, characterized in that: Performing mutation processing on the character matrix based on the data matrix to obtain a mutation matrix includes: Performing quantum state encoding processing on the data matrix and the character matrix, mapping the elements in the data matrix and the character matrix to different quantum states of quantum bits, respectively, to obtain a quantum encoding matrix pair; Using quantum superposition and entanglement characteristics, a quantum entanglement operation is performed on the quantum coding matrix pair to generate an entangled quantum matrix; The entangled quantum matrix is customized to be transformed through a quantum gate operation sequence to obtain a variant quantum matrix; the quantum gate operation sequence is generated based on an algorithmic rule that changes dynamically over time; The variation quantum matrix is subjected to quantum measurement decoding processing to obtain a variation matrix.
9. The method for repairing audio data loss in an AI recording scenario according to claim 7, characterized in that: Generate a management authority code based on the variation matrix and matrix features of multiple dimensions, including: Combining the variation matrix and matrix features of multiple dimensions into a high-dimensional vector, and converting the high-dimensional vector into a polynomial function using a polynomial fitting algorithm; Based on the polynomial function, a corresponding two-dimensional curve is drawn on a two-dimensional plane; Constructing a three-dimensional surface graph in a three-dimensional space according to the coefficients and degrees of the polynomial function; Extracting features from the polynomial function, the two-dimensional curve, and the three-dimensional surface graph, and combining the extracted features into a combined feature vector; The combined feature vector is encrypted based on a chaotic encryption algorithm to obtain a ciphertext; the ciphertext is binary-encoded to obtain the management authority code.
10. A system for repairing audio data loss in an AI recording scenario, characterized in that: include: The acquisition module is used to obtain the audio data of the AI device in the recording scenario; A verification module, used to verify the audio data based on the integrity of the timestamp and the check code, and mark the audio segment with lost data; An analysis module is used to perform in-depth analysis on the audio segment with the marked missing data to determine the specific location of the missing data in the audio segment and the amount of the missing data; A recovery module is used to obtain the device status of the AI device, match the corresponding audio data recovery algorithm based on the device status, and recover the audio segment where the lost data is located to obtain preliminary recovered audio data; The optimization module is used to perform quality optimization processing on the initially restored audio data based on spectrum analysis to obtain audio data that meets the playback requirements.
Citation Information
Patent Citations
Color image storage thumbnail encryption algorithm based on genetic algorithm
CN112116672A
Audio packet loss recovery method and device, electronic equipment and storage medium
CN118155650A
Quantum image encryption method and device, computer equipment and computer readable storage medium
CN118869894A
Video data transmission method
CN119255057A
Audio frame insertion method, device, equipment, medium and program product
CN119299739A
Cited By
Recording marking method and device, recording event display method and device, and equipment
CN120786136A