An audio processing method, apparatus, device, and storage medium

By synthesizing audio watermark data with audio data and collecting, analyzing and comparing audio watermark data during playback, the problem of low audio fault detection efficiency in the prior art is solved, and more efficient and accurate fault detection is achieved.

CN117854548BActive Publication Date: 2025-05-27SHUXING TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410184938.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-19
Publication Date
2025-05-27
Estimated Expiration
2044-02-19

AI Technical Summary

Technical Problem

In the prior art, the detection efficiency of audio failures is low, mainly relying on manual intervention, and it is prone to misjudgment and misjudgment.

Method used

Audio stream data is obtained by receiving audio data from the sending device and preset audio watermark data, and the audio stream data is collected during playback, and the retrieval audio watermark data is compared with the preset audio watermark data to obtain the fault information of the target module.

Benefits of technology

It improves the efficiency of audio fault detection, reduces the dependence of manual intervention, and enhances the accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117854548B_ABST
    Figure CN117854548B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses an audio processing method, device, equipment and storage medium. The method includes: receiving audio data from a sending device, and synthesizing the audio data and preset audio watermark data to obtain audio stream data; playing the audio stream data through a sound playback device; during the process of playing the audio stream data, collecting the played audio stream data through a sound collection module in the receiving device to obtain target audio stream data; the target audio stream data includes collected audio data and collected audio watermark data; parsing the audio stream data to obtain the preset audio watermark data, and parsing the target audio stream data to obtain the collected audio watermark data; comparing the collected audio watermark data with the preset audio watermark data, and obtaining fault information of the target module according to the comparison result, where the fault information is used to indicate whether there is a fault in the receiving device. Using the embodiment of the present application can improve the efficiency of audio fault detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer application technologies, and particularly to an audio processing method, apparatus, device, and storage medium. Background Art

[0002] With the continuous development of audio technologies, the processing and transmission of audio data have become an indispensable part of people's lives. However, due to the complexity and diversity of audio data, audio failures occur from time to time. For example, problems such as audio data distortion, noise, and interruption will affect the quality and reliability of audio, bringing inconvenience and trouble to users. The diagnosis and repair of audio failures mainly rely on manual intervention, which is time-consuming and laborious in this way, and it is also prone to misjudgment and missed judgment. Based on this, how to improve the efficiency of audio failure detection is a technical problem that urgently needs to be solved at present. Summary of the Invention

[0003] The technical problem to be solved by the embodiments of this application is to provide an audio processing method, apparatus, device, and storage medium that can improve the efficiency of audio failure detection.

[0004] In a first aspect, the embodiments of this application provide an audio processing method, which includes:

[0005] Receiving audio data from a sending device, and synthesizing the audio data and preset audio watermark data to obtain audio stream data;

[0006] Playing the audio stream data through a sound playback device;

[0007] During the process of playing the audio stream data, collecting the played audio stream data through a sound collection module in a receiving device to obtain target audio stream data; the target audio stream data includes back-captured audio data and back-captured audio watermark data;

[0008] Parsing the audio stream data to obtain preset audio watermark data, and parsing the target audio stream data to obtain back-captured audio watermark data;

[0009] Comparing the back-captured audio watermark data with the preset audio watermark data, and obtaining fault information of a target module according to the comparison result, where the fault information is used to indicate whether there is a fault in the receiving device.

[0010] In a second aspect, the embodiments of this application provide an audio processing apparatus, which includes:

[0011] A receiving unit, configured to receive audio data from a sending device, and synthesize the audio data and preset audio watermark data to obtain audio stream data;

[0012] A playback unit for playing audio stream data through a sound playback device;

[0013] An acquisition unit for acquiring the played audio stream data through a sound acquisition module in the receiving device during the process of playing the audio stream data, to obtain target audio stream data; the target audio stream data includes back-acquired audio data and back-acquired audio watermark data;

[0014] A processing unit for parsing the audio stream data to obtain preset audio watermark data, and parsing the target audio stream data to obtain back-acquired audio watermark data;

[0015] The processing unit is further configured to compare the back-acquired audio watermark data with the preset audio watermark data, and obtain fault information of the target module according to the comparison result, and the fault information is used to indicate whether there is a fault in the receiving device.

[0016] In one embodiment, the processing unit parses the audio stream data to obtain preset audio watermark data, and parses the target audio stream data to obtain back-acquired audio watermark data, including:

[0017] Invoking a trained fault recognition model to parse the audio stream data to obtain preset audio watermark data, and parsing the target audio stream data to obtain back-acquired audio watermark data;

[0018] Comparing the back-acquired audio watermark data with the preset audio watermark data, and obtaining fault information of the target module according to the comparison result, including:

[0019] Invoking a trained fault recognition model to compare the back-acquired audio watermark data with the preset audio watermark data, and obtaining fault information of the target module according to the comparison result;

[0020] Wherein, the trained fault recognition model is trained with the training objective that the fault information of the training audio stream data is the same as the fault information of the training audio stream data input into the fault simulation model; the fault information of the training audio stream data is obtained by parsing the training audio stream data and the audio synthesis data through the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data, and then comparing the audio watermark data of the training audio stream data with the preset audio watermark data. The training audio stream data is obtained by inputting the fault information of the training audio stream data into the fault simulation model, so that the fault simulation model simulates the audio synthesis data, and the audio synthesis data is obtained by synthesizing the training audio data and the preset audio watermark data, and the training audio data has no fault.

[0021] In one embodiment, the processing unit is further configured to include:

[0022] Synthesize the training audio data and the preset audio watermark data to obtain audio synthesis data;

[0023] Input the fault information of the training audio stream data into the fault simulation model, so that the fault simulation model simulates the audio synthesis data to obtain the training audio stream data;

[0024] Parse the training audio stream data and the audio synthesis data through the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data;

[0025] Compare the audio watermark data of the training audio stream data with the preset audio watermark data to obtain the fault information of the training audio stream data;

[0026] Train the fault simulation model with the same fault information of the training audio stream data as the training target input into the fault simulation model to obtain the trained fault recognition model.

[0027] In one embodiment, the processing unit synthesizes the training audio data and the preset audio watermark data to obtain audio synthesis data, including:

[0028] Synthesize the training audio data and the preset audio watermark data through the encoding module in the fault recognition model to obtain audio synthesis data;

[0029] Parse the training audio stream data and the audio synthesis data through the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data, including:

[0030] Parse the training audio stream data and the audio synthesis data through the decoding module in the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data;

[0031] Compare the audio watermark data of the training audio stream data with the preset audio watermark data to obtain the fault information of the training audio stream data, including:

[0032] Compare the audio watermark data of the training audio stream data with the preset audio watermark data through the decoding module in the fault recognition model to obtain the fault information of the training audio stream data;

[0033] Train the fault simulation model with the same fault information of the training audio stream data as the training target input into the fault simulation model to obtain the trained fault recognition model, including:

[0034] Taking the fault information of the training audio stream data to be the same as the fault information of the training audio stream data input into the fault simulation model as the training objective, the encoding module and the decoding module are trained to obtain the trained fault recognition model.

[0035] In one embodiment, the processing unit synthesizes the audio data and the preset audio watermark data to obtain the audio stream data, including:

[0036] The encoding module in the trained fault recognition model synthesizes the audio data and the preset audio watermark data to obtain the audio stream data;

[0037] Parsing the audio stream data to obtain the preset audio watermark data, and parsing the target audio stream data to obtain the retrieved audio watermark data, including:

[0038] The decoding module in the trained fault recognition model parses the audio stream data to obtain the preset audio watermark data, and parses the target audio stream data to obtain the retrieved audio watermark data;

[0039] Comparing the retrieved audio watermark data with the preset audio watermark data, and obtaining the fault information of the target module according to the comparison result, including:

[0040] The decoding module in the trained fault recognition model compares the retrieved audio watermark data with the preset audio watermark data, and obtains the fault information of the target module according to the comparison result.

[0041] In one embodiment, the number of sound playback devices is multiple;

[0042] The processing unit synthesizes the audio data and the preset audio watermark data to obtain the audio stream data, including:

[0043] Synthesizing the audio data, the preset audio watermark data, and the object data of the sound playback device that outputs the audio data to obtain the audio stream data;

[0044] During the process of playing the audio stream data, the sound acquisition module acquires the played audio stream data to obtain the target audio stream data; the target audio stream data includes the retrieved audio data and the retrieved audio watermark data; including:

[0045] During the process of multiple sound playback devices playing the audio stream data, the sound acquisition module acquires at least one played audio stream data to obtain at least one target audio stream data; any target audio stream data includes one retrieved audio data, one retrieved audio watermark data, and the object data corresponding to any target audio stream data; the object data is used to identify the sound playback device that plays the audio stream data corresponding to any target audio stream data;

[0046] Compare the collected audio watermark data with the preset audio watermark data, and obtain the fault information of the target module according to the comparison result, including:

[0047] For any target audio stream data, compare the collected audio watermark data corresponding to the any target audio stream data with the preset audio watermark data to obtain a comparison result;

[0048] If the comparison result indicates that the collected audio watermark data corresponding to any target audio stream data does not match the preset audio watermark data, generate the fault information of the sound playback device identified by the object data corresponding to the any target audio stream data, and the fault information is used to indicate that the sound playback device identified by the object data corresponding to the any target audio stream data has a fault.

[0049] In one embodiment, the processing unit synthesizes the audio data, the preset audio watermark data, and the object data of the sound playback device that outputs the audio data to obtain audio stream data, including:

[0050] For any sound playback device, obtain the private key of the any sound playback device;

[0051] Based on the private key of the any sound playback device, encrypt the object data corresponding to the any sound playback device to obtain the encrypted object data;

[0052] Synthesize the audio data, the preset audio watermark data, and the encrypted object data to obtain audio stream data; wherein, the audio stream data is played by any sound playback device;

[0053] Parse the audio stream data to obtain the preset audio watermark data, and parse the target audio stream data to obtain the collected audio watermark data, including:

[0054] After the audio stream data from any sound playback device is collected by the sound collection module, parse the audio stream data to obtain the collected audio data, the collected audio watermark data, and the encrypted object data;

[0055] Decrypt the encrypted object data with the public key of any sound playback device to obtain the object data.

[0056] In one embodiment, the processing unit parses the audio stream data to obtain the preset audio watermark data, and parses the target audio stream data to obtain the collected audio watermark data, including:

[0057] Perform delay estimation on the audio stream data and the target audio stream data to obtain the delay of the target audio stream data relative to the audio stream data;

[0058] Adjust the time axis of the target audio stream data based on the delay of the target audio stream data relative to the audio stream data, align the target audio stream data and the audio stream data, and obtain the aligned target audio stream data and audio stream data;

[0059] Parse the aligned target audio stream data and audio stream data to obtain the recollected audio watermark data and the preset audio watermark data.

[0060] In a third aspect, an embodiment of the present invention provides a computer device, which includes a memory, a communication interface, and a processor. Among them, the memory, the communication interface, and the processor are interconnected; the memory stores a computer program, and the processor calls the computer program stored in the memory to implement the method in the first aspect above.

[0061] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method in the first aspect above is implemented.

[0062] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program, and the computer program is stored in a computer storage medium; a processor of a computer device reads the computer program from the computer storage medium, and the processor executes the computer program, so that the computer device executes the method in the first aspect above.

[0063] In the embodiment of the present application, audio data from a sending device is received, and the audio data and the preset audio watermark data are synthesized to obtain audio stream data; the audio stream data is played through a sound playback device; during the process of playing the audio stream data, the played audio stream data is collected through a sound collection module in the receiving device to obtain target audio stream data; the target audio stream data includes recollected audio data and recollected audio watermark data; the preset audio watermark data is obtained by parsing the audio stream data, and the recollected audio watermark data is obtained by parsing the target audio stream data; the recollected audio watermark data and the preset audio watermark data are compared, and the fault information of the target module is obtained according to the comparison result, and the fault information is used to indicate whether there is a fault in the receiving device. By adding audio watermark data to the audio data and parsing the audio watermark data when diagnosing audio faults, fault information can be obtained, improving the efficiency of audio fault detection. Description of the Drawings

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the following will describe the drawings required to be used in the embodiments of the present application or the background art.

[0065] Figure 1 It is a schematic structural diagram of a model training system provided by an embodiment of the present application;

[0066] Figure 2 It is a schematic structural diagram of an audio processing system provided by an embodiment of the present application;

[0067] Figure 3 It is a schematic flowchart of an audio processing method provided by an embodiment of the present application;

[0068] Figure 4 It is a schematic structural diagram of another audio processing system provided by an embodiment of the present application;

[0069] Figure 5 It is a schematic flowchart of another audio processing method provided by an embodiment of the present application;

[0070] Figure 6 It is a schematic structural diagram of an audio processing device provided by an embodiment of the present application;

[0071] Figure 7 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0072] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0073] In the specific implementation manners of the present application, data related to users, such as users' audio information, etc., are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with local laws, regulations, and standards.

[0074] Please refer to Figure 1 , Figure 1 It is a schematic structural diagram of a model training system provided by an embodiment of the present application.

[0075] Exemplarily, the training audio data and the preset audio watermark data are synthesized by a fault recognition model to obtain audio synthesis data. The audio watermark data is a technology that embeds specific information into an audio signal. The purpose is to add some hidden identifiers or information to the audio data without affecting the listening experience. In this solution, the preset audio watermark data is used to detect audio data faults. Optionally, the fault recognition model can be a Recurrent Neural Networks (RNN). RNN is a neural network model suitable for processing sequential data. When processing audio signals, it can be used to model and process sequential data, thereby realizing the embedding and parsing of audio watermark data. In the scenario of this solution, the audio data can be regarded as a time series, and RNN can model it and learn the features therein, including the information of the embedded audio watermark data. When embedding the audio watermark data, RNN can encode the audio watermark data into the audio in a certain way by fusing the audio watermark data information into the representation of the audio signal, thereby realizing the embedding of the audio watermark data. When parsing the audio watermark data, RNN can be trained to recognize and extract the audio watermark data information embedded in the audio signal, thereby realizing the parsing of the audio watermark data.

[0076] Optionally, the fault recognition model can be a Generative Adversarial Networks (GAN). Generator network: It accepts random noise or other inputs and generates audio data with audio watermark data. The goal of the generator is to make the generated audio as similar as possible to the real audio while containing the information of the embedded audio watermark data. Discriminator network: The discriminator network is trained to distinguish between real audio and the audio generated by the generator. The goal of the discriminator is to accurately identify the generated audio and make it difficult to distinguish the generated audio from the real audio. Adversarial training: In adversarial training, the generator and the discriminator compete with each other to promote the improvement of each other's performance. The generator tries to deceive the discriminator and generate more and more realistic audio, while the discriminator tries to improve its ability to distinguish between generated audio and real audio. Parser network: Using the audio data generated by the generator, a parser network can be trained to learn how to parse the audio watermark data information embedded in the audio. The goal of the parser network is to extract the audio watermark data as accurately as possible and make it robust to various generated audio data. Optimization process: By adjusting the weights of the generator, discriminator, and parser, the entire system can be optimized to realize the generation of audio data with audio watermark data and the accurate parsing of the embedded audio watermark data. This method combining generative adversarial networks and parser networks can improve the performance of embedding and parsing audio watermark data during the training process, while making the generated audio data more realistic and maintaining the stability of the audio watermark data.

[0077] Further, audio data, preset audio watermark data, and object data of a sound playback device for outputting audio data can be synthesized to obtain audio stream data, where the object data can refer to the combination of a device identifier and a user identifier as the object data.

[0078] In one implementation, training audio data and preset audio watermark data are synthesized to obtain audio synthesis data, including:

[0079] The training audio data and the preset audio watermark data are synthesized by an encoding module in the fault identification model to obtain audio synthesis data.

[0080] In this embodiment, the fault identification model includes an encoding module. The training audio data and the preset audio watermark data are synthesized by the encoding module of the fault identification model to obtain audio synthesis data.

[0081] Input the fault information of the training audio stream data into the fault simulation model so that the fault simulation model simulates the audio synthesis data to obtain the training audio stream data.

[0082] Among them, a specified audio fault type can be input so that the fault simulation model performs fault simulation according to the input fault type. Fault information of a specified module can also be input, for example, a playback module, a sound collection module, and a processing module, etc. The processing module can refer to various processing modules before playing the audio stream data, such as a sampling rate conversion module, a sound effect processing module, etc.

[0083] The fault identification model analyzes the training audio stream data and the audio synthesis data to obtain the audio watermark data of the training audio stream data and the preset audio watermark data.

[0084] In one implementation, the fault identification model analyzes the training audio stream data and the audio synthesis data to obtain the audio watermark data of the training audio stream data and the preset audio watermark data, including:

[0085] The fault identification model analyzes the training audio stream data and the audio synthesis data through a decoding module in the fault identification model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data.

[0086] The audio watermark data of the training audio stream data and the preset audio watermark data are compared to obtain the fault information of the training audio stream data.

[0087] In this embodiment, according to the difference between the audio watermark data of the training audio stream data and the preset audio watermark data, the fault information is obtained, where the fault information can include whether there is a fault and the category of the fault.

[0088] Further, based on the object data, determine the device and user identifier where the fault information appears.

[0089] In one implementation, compare the audio watermark data of the training audio stream data with the preset audio watermark data to obtain the fault information of the training audio stream data, including:

[0090] Compare the audio watermark data of the training audio stream data with the preset audio watermark data through the decoding module in the fault recognition model to obtain the fault information of the training audio stream data.

[0091] In this embodiment, the decoding module in the fault recognition model obtains the fault information based on the difference between the audio watermark data of the training audio stream data and the preset audio watermark data, where the fault information may include whether there is a fault and the category of the fault. Train the fault simulation model with the fault information of the training audio stream data being the same as the fault information of the training audio stream data input to the fault simulation model as the training target to obtain the trained fault recognition model.

[0092] In one implementation, train the fault simulation model with the fault information of the training audio stream data being the same as the fault information of the training audio stream data input to the fault simulation model as the training target to obtain the trained fault recognition model, including:

[0093] Train the encoding module and the decoding module with the fault information of the training audio stream data being the same as the fault information of the training audio stream data input to the fault simulation model as the training target to obtain the trained fault recognition model.

[0094] In this embodiment, the fault simulation model simulates the input audio fault and uses the fault information as label data. With the fault information of the training audio stream data being the same as the fault information of the training audio stream data input to the fault simulation model as the training target, that is, the fault information of the training audio stream data being the same as the label data as the training target, train the encoding module and the decoding module to obtain the trained fault recognition model.

[0095] In the embodiments of the present application, the encoding module and the decoding module adapted to the audio diagnosis scenario are obtained through an end-to-end network training mode. Therefore, the preset audio watermark data embedded in the audio data is suitable for diagnosing various audio fault problems.

[0096] Please refer to Figure 2 , Figure 2It is a schematic diagram of the architecture of an audio processing system provided by an embodiment of the present application. Exemplarily, a receiving device receives audio data from a transmitting device, and synthesizes the audio data and preset audio watermark data to obtain audio stream data; the audio stream data is played through a sound playback device; during the process of playing the audio stream data, the played audio stream data is collected by a sound collection module in the receiving device to obtain target audio stream data; the target audio stream data includes back-captured audio data and back-captured audio watermark data; the preset audio watermark data is obtained by parsing the audio stream data, and the back-captured audio watermark data is obtained by parsing the target audio stream data; the back-captured audio watermark data and the preset audio watermark data are compared, and fault information of the target module is obtained according to the comparison result, and the fault information is used to indicate whether there is a fault in the receiving device.

[0097] Please refer to Figure 3 , Figure 3 It is a schematic flowchart of an audio processing method provided by an embodiment of the present application. As Figure 3 shown, the audio processing method includes but is not limited to steps S301-S305, where:

[0098] S301, receive audio data from a transmitting device, and synthesize the audio data and preset audio watermark data to obtain audio stream data.

[0099] In this embodiment, when receiving the audio data sent by the transmitting device, the preset audio watermark data can be added to the audio data to check whether there is a fault in the playback device. For example, in a live video call scenario, when the host's device is the transmitting device and the live viewer's device is the receiving device, adding the preset audio watermark data to the audio data sent by the host can detect whether there is a fault in the receiving device.

[0100] In one implementation, synthesizing the audio data and the preset audio watermark data to obtain audio stream data includes:

[0101] Synthesize the audio data and the preset audio watermark data through an encoding module in the trained fault recognition model to obtain audio stream data.

[0102] In this embodiment, the audio data and the preset audio watermark data are synthesized through an encoding module in the trained fault recognition model to obtain audio stream data.

[0103] S302, play the audio stream data through a sound playback device.

[0104] In this embodiment, the sound playback device can be the receiving device or a separate device. For example, it can be the speaker of the receiving device or a sound device.

[0105] S303. During the process of playing the audio stream data, the audio stream data being played is collected by the sound collection module in the receiving device to obtain target audio stream data. The target audio stream data includes the recollected audio data and the recollected audio watermark data.

[0106] In this embodiment, during the process of playing the audio stream data, the audio stream data being played is collected by the sound collection module in the receiving device to obtain target audio stream data. The sound collection module may refer to the microphone of the receiving device. The audio stream data is recollected through the microphone to obtain target audio stream data, and the target audio stream data includes the recollected audio data and the recollected audio watermark data.

[0107] S304. Parse the audio stream data to obtain the preset audio watermark data, and parse the target audio stream data to obtain the recollected audio watermark data.

[0108] In this embodiment, if there is an audio fault, the preset audio watermark data and the recollected audio watermark data are different; if there is no audio fault, the preset audio watermark data and the recollected audio watermark data are the same.

[0109] In one implementation, parsing the audio stream data to obtain the preset audio watermark data and parsing the target audio stream data to obtain the recollected audio watermark data includes:

[0110] Perform delay estimation on the audio stream data and the target audio stream data to obtain the delay of the target audio stream data relative to the audio stream data.

[0111] Based on the delay of the target audio stream data relative to the audio stream data, adjust the time axis of the target audio stream data to align the target audio stream data with the audio stream data, and obtain the aligned target audio stream data and audio stream data.

[0112] Parse the aligned target audio stream data and audio stream data to obtain the recollected audio watermark data and the preset audio watermark data.

[0113] In this embodiment, first determine the delay of the target audio stream data relative to the audio stream data. For example, through the cross-correlation method, cross-correlation calculates the cross-correlation function between two audio data to find the time delay value between them. The peak of the cross-correlation function corresponds to the best alignment position, thus providing an estimate of the time delay. Using the obtained time delay value, alignment can be achieved by adjusting the time axis of the target audio stream data forward or backward by the time delay value. The purpose of parsing after aligning the target audio stream data and the audio stream data is to ensure that the target audio stream data and the audio stream data are consistent in time, so as to accurately compare their features and conduct comparative analysis.

[0114] In one implementation, parsing the audio stream data to obtain preset audio watermark data, and parsing the target audio stream data to obtain the retrieved audio watermark data, including:

[0115] Invoking the trained fault recognition model to parse the audio stream data to obtain preset audio watermark data, and parsing the target audio stream data to obtain the retrieved audio watermark data.

[0116] In one implementation, parsing the audio stream data to obtain preset audio watermark data, and parsing the target audio stream data to obtain the retrieved audio watermark data, including:

[0117] Parsing the audio stream data through the decoding module in the trained fault recognition model to obtain preset audio watermark data, and parsing the target audio stream data to obtain the retrieved audio watermark data.

[0118] In this embodiment, through an end-to-end training method, the encoding module and the decoding module of the fault recognition model are trained, and the decoding module of the trained fault recognition model can parse the audio watermark data in the audio data.

[0119] S305. Comparing the retrieved audio watermark data with the preset audio watermark data, and obtaining the fault information of the target module according to the comparison result, where the fault information is used to indicate whether there is a fault in the receiving device.

[0120] In this embodiment, the retrieved audio watermark data is compared with the preset audio watermark data, and the fault information of the target module is obtained according to the comparison result, where the target module may refer to the playback module, the sound collection module, and the processing module of the playback device, etc., and the processing module may refer to various processing modules before playing the audio stream data, such as the sample rate conversion module, the sound effect processing module, etc.

[0121] In one implementation, comparing the retrieved audio watermark data with the preset audio watermark data, and obtaining the fault information of the target module according to the comparison result, including:

[0122] Invoking the trained fault recognition model to compare the retrieved audio watermark data with the preset audio watermark data, and obtaining the fault information of the target module according to the comparison result;

[0123] Among them, the trained fault recognition model is trained with the goal that the fault information of the training audio stream data is the same as that of the training audio stream data input into the fault simulation model. The fault information of the training audio stream data is obtained by parsing the training audio stream data and the audio synthesis data through the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data, and then comparing the audio watermark data of the training audio stream data with the preset audio watermark data. The training audio stream data is obtained by inputting the fault information of the training audio stream data into the fault simulation model, so that the fault simulation model simulates the audio synthesis data. The audio synthesis data is synthesized from the training audio data and the preset audio watermark data, and there is no fault in the training audio data.

[0124] In this embodiment, the trained fault recognition model compares the retrieved audio watermark data with the preset audio watermark data, and obtains the fault information of the target module according to the comparison result. The target module may refer to the playback module, the sound collection module, and the processing module of the playback device, etc. The processing module may refer to various processing modules before playing the audio stream data, such as the sample rate conversion module, the sound effect processing module, etc.

[0125] In one implementation, comparing the retrieved audio watermark data with the preset audio watermark data and obtaining the fault information of the target module according to the comparison result includes:

[0126] Comparing the retrieved audio watermark data with the preset audio watermark data through the decoding module in the trained fault recognition model, and obtaining the fault information of the target module according to the comparison result.

[0127] In this embodiment, the decoding module in the trained fault recognition model compares the retrieved audio watermark data with the preset audio watermark data, and obtains the fault information of the target module according to the comparison result. When training the fault recognition model, training the encoding module and the decoding module in an end-to-end manner can enable the decoding module to parse and obtain the audio watermark data, and obtain the fault information according to the obtained audio watermark data. The fault information may include whether there is a fault, the fault of the target module, and the type of the fault. For example, there is a problem in the target module. The target module may refer to the playback module, the sound collection module, and the processing module of the playback device, etc. The processing module may refer to various processing modules before playing the audio stream data, such as the sample rate conversion module, the sound effect processing module, etc. The type of the fault may include problems such as no sound, stuttering, and noise.

[0128] Furthermore, when the fault information is determined, the fault can be processed through the fault processing module. For example, for the problem of no sound caused by the audio device being preempted, the device priority usage right can be obtained with a higher priority, etc.

[0129] In an embodiment of the present application, audio data from a sending device is received, and the audio data and preset audio watermark data are synthesized to obtain audio stream data; the audio stream data is played through a sound playback device; during the process of playing the audio stream data, the played audio stream data is collected by a sound collection module in the receiving device to obtain target audio stream data; the target audio stream data includes back-captured audio data and back-captured audio watermark data; the preset audio watermark data is obtained by parsing the audio stream data, and the back-captured audio watermark data is obtained by parsing the target audio stream data; the back-captured audio watermark data and the preset audio watermark data are compared, and fault information of the target module is obtained according to the comparison result, and the fault information is used to indicate whether there is a fault in the receiving device. By adding the audio watermark data to the audio data and parsing the audio watermark data when diagnosing audio faults, fault information can be obtained, and the audio fault detection efficiency can be improved.

[0130] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of another audio processing system provided by an embodiment of the present application. Exemplarily, a receiving device receives audio data from a sending device, and synthesizes the audio data, preset audio watermark data, and object data of a sound playback device that outputs the audio data to obtain audio stream data; the audio stream data is played through a plurality of sound playback devices; during the process of playing the audio stream data by the plurality of sound playback devices, at least one audio stream data played is collected by a sound collection module to obtain at least one target audio stream data; any target audio stream data includes one piece of back-captured audio data, one piece of back-captured audio watermark data, and object data corresponding to any target audio stream data; the object data is used to identify the sound playback device that plays the audio stream data corresponding to any target audio stream data; the preset audio watermark data is obtained by parsing the audio stream data, and the back-captured audio watermark data is obtained by parsing any target audio stream data; for any target audio stream data, the back-captured audio watermark data corresponding to any target audio stream data and the preset audio watermark data are compared to obtain a comparison result; if the comparison result indicates that the back-captured audio watermark data corresponding to any target audio stream data does not match the preset audio watermark data, fault information of the sound playback device identified by the object data corresponding to any target audio stream data is generated, and the fault information is used to indicate that there is a fault in the sound playback device identified by the object data corresponding to any target audio stream data.

[0131] Based on the above description, please refer to Figure 5 , Figure 5 which is a schematic flowchart of another audio processing method provided by an embodiment of the present application. As shown in Figure 5 the audio processing method includes but is not limited to steps S501 - S505, where:

[0132] S501, Synthesize the audio data, the preset audio watermark data, and the object data of the sound playback device that outputs the audio data to obtain the audio stream data.

[0133] In this embodiment, there are multiple sound playback devices. When the receiving device receives the audio data, it synthesizes the audio data, the preset audio watermark data, and the object data of the sound playback device that outputs the audio data to obtain the audio stream data, where the object data may refer to the combination of the sound playback device identifier and the user identifier using the sound playback device as the object data.

[0134] In one implementation, synthesizing the audio data, the preset audio watermark data, and the object data of the sound playback device that outputs the audio data to obtain the audio stream data includes:

[0135] For any sound playback device, obtain the private key of any sound playback device;

[0136] Based on the private key of any sound playback device, encrypt the object data corresponding to any sound playback device to obtain the encrypted object data;

[0137] Synthesize the audio data, the preset audio watermark data, and the encrypted object data to obtain the audio stream data; wherein, the audio stream data is played through any sound playback device.

[0138] In this embodiment, by obtaining the private key of any sound playback device, using the private key to encrypt the object data corresponding to any sound playback device to obtain the encrypted object data, and synthesizing the audio data, the preset audio watermark data, and the encrypted object data to obtain the audio stream data. Through the method of encrypting the object data, information security can be guaranteed.

[0139] Optionally, technologies such as digital signature and hash function can be used to encrypt the object data to generate the ciphertext of the object data.

[0140] S502, Play the audio stream data through multiple sound playback devices.

[0141] In this embodiment, there are multiple sound playback devices, and any sound playback device is uniquely identified by the object data.

[0142] S503, During the process of playing the audio stream data through multiple sound playback devices, collect at least one of the played audio stream data through the sound collection module to obtain at least one target audio stream data.

[0143] In this embodiment, any target audio stream data includes a captured audio data, a captured audio watermark data, and object data corresponding to any target audio stream data; the object data is used to identify the sound playback device that plays the audio stream data corresponding to any target audio stream data. During the process of playing the audio stream data, at least one target audio stream data is obtained by collecting, through the sound collection module in the receiving device, the played audio stream data, where the sound collection module may refer to the microphone of the receiving device, and at least one audio stream data is captured through the microphone to obtain at least one target audio stream data.

[0144] S504. Parse the audio stream data to obtain the preset audio watermark data, and parse any target audio stream data to obtain the captured audio watermark data.

[0145] In this embodiment, any target audio stream data is parsed to obtain the captured audio data, the captured audio watermark data, and the object data of the sound playback device.

[0146] In one implementation, parsing the audio stream data to obtain the preset audio watermark data and parsing the target audio stream data to obtain the captured audio watermark data includes:

[0147] After the audio stream data from any sound playback device is collected through the sound collection module, parse the audio stream data to obtain the captured audio data, the captured audio watermark data, and the encrypted object data;

[0148] Decrypt the encrypted object data with the public key of any sound playback device to obtain the object data.

[0149] In this embodiment, the parsed object data is the encrypted object data, and it is necessary to decrypt the object data with the public key corresponding to the private key of any sound playback device to obtain the decrypted object data.

[0150] S505. For any target audio stream data, compare the captured audio watermark data corresponding to any target audio stream data with the preset audio watermark data to obtain a comparison result.

[0151] In this embodiment, if the comparison result indicates that the recollected audio watermark data corresponding to any target audio stream data does not match the preset audio watermark data, fault information of the sound playback device identified by the object data corresponding to any target audio stream data is generated. The fault information is used to indicate that there is a fault in the sound playback device identified by the object data corresponding to any target audio stream data, and the fault information is also used to indicate the target module with a fault. The target module may refer to the playback module, sound collection module, and processing module of the sound playback device. The processing module may refer to various processing modules before playing the audio stream data, such as the sampling rate conversion module and sound effect processing module.

[0152] Further, when the fault information is determined, the fault can be processed by a fault processing module. For example, for the problem of no sound caused by the preemption of the audio device, the device priority usage right can be obtained with a higher priority.

[0153] In the embodiment of the present application, audio data, preset audio watermark data, and object data of the sound playback device for outputting the audio data are synthesized to obtain audio stream data; the audio stream data is played through multiple sound playback devices; during the process of playing the audio stream data by multiple sound playback devices, at least one played audio stream data is collected through a sound collection module to obtain at least one target audio stream data; the preset audio watermark data is obtained by parsing the audio stream data, and the recollected audio watermark data is obtained by parsing any target audio stream data; for any target audio stream data, the recollected audio watermark data corresponding to any target audio stream data is compared with the preset audio watermark data to obtain a comparison result. By adding audio watermark data and object data to the audio data and parsing the audio watermark data and object data when diagnosing audio faults, the fault information and the sound playback device corresponding to the object data can be obtained, improving the audio fault detection efficiency.

[0154] The embodiment of the present application also provides a computer storage medium, in which program instructions are stored. When the program instructions are executed, they are used to implement the corresponding methods described in the above embodiments.

[0155] See again Figure 6 , Figure 6 is a schematic structural diagram of an audio processing device provided by an embodiment of the present application.

[0156] In one implementation manner of the audio processing device of the embodiment of the present application, the audio processing device includes the following structure:

[0157] A receiving unit 601, configured to receive audio data from a sending device, and synthesize the audio data and preset audio watermark data to obtain audio stream data;

[0158] A playback unit 602 for playing audio stream data through a sound playback device;

[0159] A collection unit 603 for collecting the played audio stream data through a sound collection module in the receiving device during the process of playing the audio stream data to obtain target audio stream data; the target audio stream data includes collected audio data and collected audio watermark data;

[0160] A processing unit 604 for parsing the audio stream data to obtain preset audio watermark data and parsing the target audio stream data to obtain collected audio watermark data;

[0161] The processing unit 604 is further configured to compare the collected audio watermark data with the preset audio watermark data and obtain fault information of the target module according to the comparison result, and the fault information is used to indicate whether there is a fault in the receiving device.

[0162] In one embodiment, the processing unit 604 parses the audio stream data to obtain preset audio watermark data and parses the target audio stream data to obtain collected audio watermark data, including:

[0163] Invoking a trained fault recognition model to parse the audio stream data to obtain preset audio watermark data and parsing the target audio stream data to obtain collected audio watermark data;

[0164] Comparing the collected audio watermark data with the preset audio watermark data and obtaining fault information of the target module according to the comparison result, including:

[0165] Invoking a trained fault recognition model to compare the collected audio watermark data with the preset audio watermark data and obtaining fault information of the target module according to the comparison result;

[0166] Wherein, the trained fault recognition model is trained with the training objective that the fault information of the training audio stream data is the same as the fault information of the training audio stream data input into the fault simulation model; the fault information of the training audio stream data is obtained by comparing the audio watermark data of the training audio stream data and the preset audio watermark data after parsing the training audio stream data and the audio synthesis data through the fault recognition model, the training audio stream data is obtained by inputting the fault information of the training audio stream data into the fault simulation model to enable the fault simulation model to simulate the audio synthesis data, the audio synthesis data is obtained by synthesizing the training audio data and the preset audio watermark data, and the training audio data has no fault.

[0167] In one embodiment, the processing unit 604 is further configured to include:

[0168] Synthesize the training audio data and the preset audio watermark data to obtain audio synthesis data;

[0169] Input the fault information of the training audio stream data into the fault simulation model, so that the fault simulation model simulates the audio synthesis data to obtain the training audio stream data;

[0170] Parse the training audio stream data and the audio synthesis data through the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data;

[0171] Compare the audio watermark data of the training audio stream data with the preset audio watermark data to obtain the fault information of the training audio stream data;

[0172] Use the fact that the fault information of the training audio stream data is the same as the fault information of the training audio stream data input into the fault simulation model as the training target to train the fault simulation model to obtain the trained fault recognition model.

[0173] In one embodiment, the processing unit 604 synthesizes the training audio data and the preset audio watermark data to obtain audio synthesis data, including:

[0174] Synthesize the training audio data and the preset audio watermark data through the encoding module in the fault recognition model to obtain audio synthesis data;

[0175] Parse the training audio stream data and the audio synthesis data through the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data, including:

[0176] Parse the training audio stream data and the audio synthesis data through the decoding module in the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data;

[0177] Compare the audio watermark data of the training audio stream data with the preset audio watermark data to obtain the fault information of the training audio stream data, including:

[0178] Compare the audio watermark data of the training audio stream data with the preset audio watermark data through the decoding module in the fault recognition model to obtain the fault information of the training audio stream data;

[0179] Use the fact that the fault information of the training audio stream data is the same as the fault information of the training audio stream data input into the fault simulation model as the training target to train the fault simulation model to obtain the trained fault recognition model, including:

[0180] Taking the fault information of the training audio stream data to be the same as the fault information of the training audio stream data input into the fault simulation model as the training objective, the encoding module and the decoding module are trained to obtain the trained fault recognition model.

[0181] In one embodiment, the processing unit 604 synthesizes the audio data and the preset audio watermark data to obtain the audio stream data, including:

[0182] The encoding module in the trained fault recognition model synthesizes the audio data and the preset audio watermark data to obtain the audio stream data;

[0183] Parsing the audio stream data to obtain the preset audio watermark data, and parsing the target audio stream data to obtain the retrieved audio watermark data, including:

[0184] The decoding module in the trained fault recognition model parses the audio stream data to obtain the preset audio watermark data, and parses the target audio stream data to obtain the retrieved audio watermark data;

[0185] Comparing the retrieved audio watermark data with the preset audio watermark data, and obtaining the fault information of the target module according to the comparison result, including:

[0186] The decoding module in the trained fault recognition model compares the retrieved audio watermark data with the preset audio watermark data, and obtains the fault information of the target module according to the comparison result.

[0187] In one embodiment, the number of sound playback devices is multiple;

[0188] The processing unit 604 synthesizes the audio data and the preset audio watermark data to obtain the audio stream data, including:

[0189] Synthesizing the audio data, the preset audio watermark data, and the object data of the sound playback device that outputs the audio data to obtain the audio stream data;

[0190] During the process of playing the audio stream data, the sound acquisition module acquires the played audio stream data to obtain the target audio stream data; the target audio stream data includes the retrieved audio data and the retrieved audio watermark data; including:

[0191] During the process of playing the audio stream data by multiple sound playback devices, the sound acquisition module acquires at least one played audio stream data to obtain at least one target audio stream data; any target audio stream data includes one retrieved audio data, one retrieved audio watermark data, and the object data corresponding to any target audio stream data; the object data is used to identify the sound playback device that plays the audio stream data corresponding to any target audio stream data.

[0192] Compare the collected audio watermark data with the preset audio watermark data, and obtain the fault information of the target module according to the comparison result, including:

[0193] For any target audio stream data, compare the collected audio watermark data corresponding to the any target audio stream data with the preset audio watermark data to obtain a comparison result;

[0194] If the comparison result indicates that the collected audio watermark data corresponding to any target audio stream data does not match the preset audio watermark data, generate the fault information of the sound playback device identified by the object data corresponding to the any target audio stream data, and the fault information is used to indicate that the sound playback device identified by the object data corresponding to the any target audio stream data has a fault.

[0195] In one embodiment, the processing unit 604 synthesizes the audio data, the preset audio watermark data, and the object data of the sound playback device that outputs the audio data to obtain the audio stream data, including:

[0196] For any sound playback device, obtain the private key of the any sound playback device;

[0197] Based on the private key of the any sound playback device, encrypt the object data corresponding to the any sound playback device to obtain the encrypted object data;

[0198] Synthesize the audio data, the preset audio watermark data, and the encrypted object data to obtain the audio stream data; wherein, the audio stream data is played by any sound playback device;

[0199] Analyze the audio stream data to obtain the preset audio watermark data, and analyze the target audio stream data to obtain the collected audio watermark data, including:

[0200] After the audio stream data from any sound playback device is collected by the sound collection module, analyze the audio stream data to obtain the collected audio data, the collected audio watermark data, and the encrypted object data;

[0201] Decrypt the encrypted object data with the public key of any sound playback device to obtain the object data.

[0202] In one embodiment, the processing unit 604 analyzes the audio stream data to obtain the preset audio watermark data, and analyzes the target audio stream data to obtain the collected audio watermark data, including:

[0203] Estimate the delay of the target audio stream data relative to the audio stream data by delaying the audio stream data and the target audio stream data to obtain the delay of the target audio stream data relative to the audio stream data;

[0204] Adjust the time axis of the target audio stream data based on the delay of the target audio stream data relative to the audio stream data, align the target audio stream data and the audio stream data, and obtain the aligned target audio stream data and audio stream data;

[0205] Parse the aligned target audio stream data and audio stream data to obtain the retrieved audio watermark data and the preset audio watermark data.

[0206] In the embodiment of the present application, the receiving unit 601 receives audio data from a sending device, and synthesizes the audio data and the preset audio watermark data to obtain audio stream data; the playing unit 602 plays the audio stream data through a sound playing device; during the process of playing the audio stream data, the collecting unit 603 collects the played audio stream data through a sound collecting module in the receiving device to obtain target audio stream data; the target audio stream data includes retrieved audio data and retrieved audio watermark data; the processing unit 604 parses the audio stream data to obtain the preset audio watermark data, and parses the target audio stream data to obtain the retrieved audio watermark data; compare the retrieved audio watermark data with the preset audio watermark data, and obtain the fault information of the target module according to the comparison result, and the fault information is used to indicate whether there is a fault in the receiving device. By adding the audio watermark data to the audio data and parsing the audio watermark data when diagnosing audio faults, the fault information can be obtained, and the audio fault detection efficiency can be improved.

[0207] See again Figure 7 , Figure 7 FIG. is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device of the embodiment of the present application includes structures such as a power supply module, and includes a processor 701, a memory 702, and a communication interface 703. Data can be exchanged between the processor 701, the memory 702, and the communication interface 703, and the corresponding audio processing method is implemented by the processor 701.

[0208] The memory 702 may include volatile memory, such as random-access memory (RAM); the memory 702 may also include non-volatile memory, such as flash memory, solid-state drive (SSD), etc.; the memory 702 may further include a combination of the above types of memories.

[0209] The processor 701 may be a central processing unit (CPU). The processor 701 may also be a combination of a CPU and a GPU. In a computer device, multiple CPUs and GPUs may be included as needed for corresponding audio processing. In one embodiment, the memory 702 is used to store program instructions. The processor 701 may call the program instructions to implement various methods involved in the above embodiments of the present application.

[0210] In the first possible implementation manner, the processor 701 of the computer device calls the program instructions stored in the memory 702 to receive audio data from a sending device, and synthesize the audio data and preset audio watermark data to obtain audio stream data; play the audio stream data through a sound playback device; during the process of playing the audio stream data, collect the played audio stream data through a sound collection module in the receiving device to obtain target audio stream data; the target audio stream data includes back-captured audio data and back-captured audio watermark data; parse the audio stream data to obtain the preset audio watermark data, and parse the target audio stream data to obtain the back-captured audio watermark data; compare the back-captured audio watermark data with the preset audio watermark data, and obtain the fault information of the target module according to the comparison result, and the fault information is used to indicate whether there is a fault in the receiving device.

[0211] In one embodiment, when the processor 701 parses the audio stream data to obtain the preset audio watermark data and parses the target audio stream data to obtain the back-captured audio watermark data, the following operations may be performed:

[0212] Call the trained fault recognition model to parse the audio stream data to obtain the preset audio watermark data, and parse the target audio stream data to obtain the back-captured audio watermark data;

[0213] Compare the back-captured audio watermark data with the preset audio watermark data, and obtain the fault information of the target module according to the comparison result, including:

[0214] Call the trained fault recognition model to compare the back-captured audio watermark data with the preset audio watermark data, and obtain the fault information of the target module according to the comparison result;

[0215] Among them, the trained fault recognition model is obtained by training the fault recognition model with the training objective that the fault information of the trained audio stream data is the same as the fault information of the training audio stream data input into the fault simulation model; the fault information of the training audio stream data is obtained by parsing the training audio stream data and the audio synthesis data through the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data, and then comparing the audio watermark data of the training audio stream data with the preset audio watermark data. The training audio stream data is obtained by inputting the fault information of the training audio stream data into the fault simulation model to enable the fault simulation model to simulate the audio synthesis data. The audio synthesis data is obtained by synthesizing the training audio data and the preset audio watermark data, and there is no fault in the training audio data.

[0216] In one embodiment, the processor 701 may further perform the following operations:

[0217] Synthesize the training audio data and the preset audio watermark data to obtain audio synthesis data;

[0218] Input the fault information of the training audio stream data into the fault simulation model to enable the fault simulation model to simulate the audio synthesis data and obtain the training audio stream data;

[0219] Parse the training audio stream data and the audio synthesis data through the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data;

[0220] Compare the audio watermark data of the training audio stream data with the preset audio watermark data to obtain the fault information of the training audio stream data;

[0221] Train the fault simulation model with the training objective that the fault information of the training audio stream data is the same as the fault information of the training audio stream data input into the fault simulation model to obtain the trained fault recognition model.

[0222] In one embodiment, when the processor 701 synthesizes the training audio data and the preset audio watermark data to obtain audio synthesis data, the following operations may be performed:

[0223] Synthesize the training audio data and the preset audio watermark data through the encoding module in the fault recognition model to obtain audio synthesis data;

[0224] Parsing the training audio stream data and the audio synthesis data through the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data includes:

[0225] The decoding module in the fault identification model parses the training audio stream data and the audio synthesis data in the fault identification model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data;

[0226] Compare the audio watermark data of the training audio stream data with the preset audio watermark data to obtain the fault information of the training audio stream data, including:

[0227] The decoding module in the fault identification model compares the audio watermark data of the training audio stream data with the preset audio watermark data to obtain the fault information of the training audio stream data;

[0228] Use the fact that the fault information of the training audio stream data is the same as the fault information of the training audio stream data input into the fault simulation model as the training target to train the fault simulation model to obtain the trained fault identification model, including:

[0229] Use the fact that the fault information of the training audio stream data is the same as the fault information of the training audio stream data input into the fault simulation model as the training target to train the encoding module and the decoding module to obtain the trained fault identification model.

[0230] In one embodiment, the processor 701 synthesizes the audio data and the preset audio watermark data to obtain the audio stream data, and the following operations can be performed:

[0231] The encoding module in the trained fault identification model synthesizes the audio data and the preset audio watermark data to obtain the audio stream data;

[0232] Parse the audio stream data to obtain the preset audio watermark data, and parse the target audio stream data to obtain the retrieved audio watermark data, including:

[0233] The decoding module in the trained fault identification model parses the audio stream data to obtain the preset audio watermark data, and parses the target audio stream data to obtain the retrieved audio watermark data;

[0234] Compare the retrieved audio watermark data with the preset audio watermark data, and obtain the fault information of the target module according to the comparison result, including:

[0235] The decoding module in the trained fault identification model compares the retrieved audio watermark data with the preset audio watermark data, and obtains the fault information of the target module according to the comparison result.

[0236] In one embodiment, the number of sound playback devices is multiple;

[0237] The processor 701 synthesizes the audio data and the preset audio watermark data to obtain audio stream data, and can perform the following operations:

[0238] Synthesize the audio data, the preset audio watermark data, and the object data of the sound playback device that outputs the audio data to obtain audio stream data;

[0239] During the process of playing the audio stream data, collect the played audio stream data through the sound collection module to obtain the target audio stream data; the target audio stream data includes the collected audio data and the collected audio watermark data; including:

[0240] During the process of playing the audio stream data by multiple sound playback devices, collect at least one played audio stream data through the sound collection module to obtain at least one target audio stream data; any target audio stream data includes one collected audio data, one collected audio watermark data, and the object data corresponding to any target audio stream data; the object data is used to identify the sound playback device that plays the audio stream data corresponding to any target audio stream data.

[0241] Compare the collected audio watermark data with the preset audio watermark data, and obtain the fault information of the target module according to the comparison result, including:

[0242] For any target audio stream data, compare the collected audio watermark data corresponding to any target audio stream data with the preset audio watermark data to obtain a comparison result;

[0243] If the comparison result indicates that the collected audio watermark data corresponding to any target audio stream data does not match the preset audio watermark data, generate the fault information of the sound playback device identified by the object data corresponding to any target audio stream data, and the fault information is used to indicate that the sound playback device identified by the object data corresponding to any target audio stream data is faulty.

[0244] In one embodiment, the processor 701 synthesizes the audio data, the preset audio watermark data, and the object data of the sound playback device that outputs the audio data to obtain audio stream data, and can perform the following operations:

[0245] For any sound playback device, obtain the private key of any sound playback device;

[0246] Based on the private key of any sound playback device, encrypt the object data corresponding to any sound playback device to obtain the encrypted object data;

[0247] Synthesize the audio data, the preset audio watermark data, and the encrypted object data to obtain audio stream data; wherein, the audio stream data is played by any sound playback device.

[0248] Parse the audio stream data to obtain preset audio watermark data, and parse the target audio stream data to obtain the retrieved audio watermark data, including:

[0249] After collecting the audio stream data from any sound playback device through the sound collection module, parse the audio stream data to obtain the retrieved audio data, the retrieved audio watermark data, and the encrypted object data;

[0250] Decrypt the encrypted object data with the public key of any sound playback device to obtain the object data.

[0251] In one embodiment, the processor 701 can perform the following operations to parse the audio stream data to obtain the preset audio watermark data and parse the target audio stream data to obtain the retrieved audio watermark data:

[0252] Estimate the delay between the audio stream data and the target audio stream data to obtain the delay of the target audio stream data relative to the audio stream data;

[0253] Based on the delay of the target audio stream data relative to the audio stream data, adjust the time axis of the target audio stream data to align the target audio stream data with the audio stream data, obtaining the aligned target audio stream data and audio stream data;

[0254] Parse the aligned target audio stream data and audio stream data to obtain the retrieved audio watermark data and the preset audio watermark data.

[0255] In the embodiment of the present application, the processor 701 receives the audio data from the sending device, synthesizes the audio data and the preset audio watermark data to obtain the audio stream data; plays the audio stream data through the sound playback device; during the process of playing the audio stream data, collects the played audio stream data through the sound collection module in the receiving device to obtain the target audio stream data; the target audio stream data includes the retrieved audio data and the retrieved audio watermark data; parse the audio stream data to obtain the preset audio watermark data, and parse the target audio stream data to obtain the retrieved audio watermark data; compare the retrieved audio watermark data with the preset audio watermark data, and obtain the fault information of the target module according to the comparison result, and the fault information is used to indicate whether there is a fault in the receiving device. By adding the audio watermark data to the audio data and parsing the audio watermark data when diagnosing audio faults, the fault information can be obtained, improving the efficiency of audio fault detection.

[0256] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by relevant hardware instructed by a computer program. This program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage media include: various media such as ROM, random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0257] The above-disclosed are only some embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of the above embodiments, and equivalent changes made according to the claims of the present application still fall within the scope covered by the present invention.

Claims

1. An audio processing method, characterized in that: The method is applied to a receiving device, and the method comprises: Receiving audio data from a sending device, and synthesizing the audio data and preset audio watermark data to obtain audio stream data; Play the audio stream data through a sound playing device; In the process of playing the audio stream data, the sound collection module in the receiving device collects the played audio stream data to obtain target audio stream data; the target audio stream data includes the collected audio data and the collected audio watermark data; Calling the trained fault recognition model to parse the audio stream data to obtain the preset audio watermark data, and parsing the target audio stream data to obtain the recovered audio watermark data; Calling the trained fault recognition model to compare the retrieved audio watermark data with the preset audio watermark data, and obtaining fault information of the target module according to the comparison result, wherein the fault information is used to indicate whether the receiving device has a fault; Among them, the trained fault recognition model is obtained by training the fault recognition model with the fault information of the training audio stream data being the same as the fault information of the training audio stream data input into the fault simulation model as the training target; the fault information of the training audio stream data is obtained by parsing the training audio stream data and the audio synthesis data by the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data, and then comparing the audio watermark data of the training audio stream data with the preset audio watermark data; the training audio stream data is obtained by inputting the fault information of the training audio stream data into the fault simulation model so that the fault simulation model simulates the audio synthesis data; the audio synthesis data is obtained by synthesizing the training audio data and the preset audio watermark data, and there is no fault in the training audio data.

2. The method according to claim 1, characterized in that The method further comprises: synthesizing the training audio data and the preset audio watermark data to obtain the audio synthesis data; Inputting fault information of the training audio stream data into the fault simulation model so that the fault simulation model simulates the audio synthesis data to obtain the training audio stream data; Parsing the training audio stream data and the audio synthesis data by using the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data; Comparing the audio watermark data of the training audio stream data with the preset audio watermark data to obtain fault information of the training audio stream data; The fault simulation model is trained with the fault information of the training audio stream data being the same as the fault information of the training audio stream data input into the fault simulation model as a training target to obtain the trained fault recognition model.

3. The method according to claim 2, characterized in that The synthesizing the training audio data and the preset audio watermark data to obtain the audio synthesis data includes: The training audio data and the preset audio watermark data are synthesized by the encoding module in the fault recognition model to obtain the audio synthesis data; The step of parsing the training audio stream data and the audio synthesis data by the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data includes: The fault recognition model parses the training audio stream data and the audio synthesis data through a decoding module in the fault recognition model to obtain the audio watermark data of the training audio stream data and the preset audio watermark data; The comparing the audio watermark data of the training audio stream data with the preset audio watermark data to obtain the fault information of the training audio stream data includes: Comparing the audio watermark data of the training audio stream data with the preset audio watermark data through a decoding module in the fault identification model to obtain fault information of the training audio stream data; The fault simulation model is trained with the fault information of the training audio stream data being the same as the fault information of the training audio stream data input to the fault simulation model as a training target to obtain the trained fault recognition model, including: The encoding module and the decoding module are trained with the fault information of the training audio stream data being the same as the fault information of the training audio stream data input into the fault simulation model as a training target to obtain the trained fault recognition model.

4. The method according to claim 3, characterized in that The step of synthesizing the audio data and preset audio watermark data to obtain audio stream data includes: The audio data and the preset audio watermark data are synthesized by the encoding module in the trained fault recognition model to obtain audio stream data; The step of parsing the audio stream data to obtain the preset audio watermark data, and parsing the target audio stream data to obtain the recovered audio watermark data, comprises: The audio stream data is parsed by a decoding module in the trained fault recognition model to obtain the preset audio watermark data, and the target audio stream data is parsed to obtain the recovered audio watermark data; The step of comparing the retrieved audio watermark data with the preset audio watermark data and obtaining the fault information of the target module according to the comparison result includes: The retrieved audio watermark data and the preset audio watermark data are compared by a decoding module in the trained fault recognition model, and the fault information of the target module is obtained according to the comparison result.

5. The method according to claim 1, characterized in that The number of the sound playing devices is multiple; The step of synthesizing the audio data and preset audio watermark data to obtain audio stream data includes: The audio data, the preset audio watermark data, and the object data of the sound playing device that outputs the audio data are synthesized to obtain audio stream data; In the process of playing the audio stream data, the audio stream data played is collected by the sound collection module to obtain the target audio stream data; the target audio stream data includes the collected audio data and the collected audio watermark data; including: In the process of playing the audio stream data by multiple sound playing devices, at least one audio stream data played is collected by the sound collecting module to obtain at least one target audio stream data; any target audio stream data includes a back-collected audio data, a back-collected audio watermark data, and object data corresponding to any target audio stream data; the object data is used to identify the sound playing device that plays the audio stream data corresponding to any target audio stream data; The step of comparing the retrieved audio watermark data with the preset audio watermark data and obtaining the fault information of the target module according to the comparison result includes: For any target audio stream data, the retrieved audio watermark data corresponding to the any target audio stream data is compared with the preset audio watermark data to obtain a comparison result; If the comparison result indicates that the retrieved audio watermark data corresponding to any target audio stream data does not match the preset audio watermark data, fault information of the sound playback device identified by the object data corresponding to any target audio stream data is generated, and the fault information is used to indicate that there is a fault in the sound playback device identified by the object data corresponding to any target audio stream data.

6. The method according to claim 5, characterized in that The step of synthesizing the audio data, the preset audio watermark data, and the object data of the sound playback device that outputs the audio data to obtain the audio stream data includes: For any sound playing device, obtaining a private key of the any sound playing device; Encrypting object data corresponding to any sound playing device based on a private key of any sound playing device to obtain encrypted object data; The audio data, the preset audio watermark data, and the encrypted object data are synthesized to obtain audio stream data; wherein the audio stream data is played through any of the sound playing devices; The step of parsing the audio stream data to obtain the preset audio watermark data, and parsing the target audio stream data to obtain the recovered audio watermark data, comprises: After the audio stream data from any sound playing device is collected by the sound collection module, the audio stream data is parsed to obtain the collected audio data, the collected audio watermark data and the encrypted object data; The encrypted object data is decrypted by using the public key of any sound playing device to obtain the object data.

7. The method according to claim 1, characterized in that The method further comprises: Performing delay estimation on the audio stream data and the target audio stream data to obtain a delay of the target audio stream data relative to the audio stream data; Based on the delay of the target audio stream data relative to the audio stream data, adjusting the time axis of the target audio stream data, aligning the target audio stream data with the audio stream data, and obtaining aligned target audio stream data and audio stream data; The aligned target audio stream data and audio stream data are parsed to obtain the recovered audio watermark data and the preset audio watermark data.

8. An audio processing device, characterized in that: include: A receiving unit, configured to receive audio data from a sending device, and synthesize the audio data and preset audio watermark data to obtain audio stream data; A playing unit, used for playing the audio stream data through a sound playing device; A collection unit, used for collecting the played audio stream data through the sound collection module in the receiving device during the process of playing the audio stream data, to obtain target audio stream data; the target audio stream data includes the collected audio data and the collected audio watermark data; A processing unit, configured to call the trained fault recognition model to parse the audio stream data to obtain the preset audio watermark data, and parse the target audio stream data to obtain the recovered audio watermark data; The processing unit is further used to call the trained fault recognition model to compare the retrieved audio watermark data with the preset audio watermark data, and obtain fault information of the target module according to the comparison result, wherein the fault information is used to indicate whether the receiving device has a fault; The trained fault recognition model is obtained by training the fault recognition model with the fault information of the training audio stream data being the same as the fault information of the training audio stream data input into the fault simulation model as the training target; The fault information of the training audio stream data is obtained by parsing the training audio stream data and the audio synthesis data through the fault identification model, obtaining the audio watermark data of the training audio stream data and the preset audio watermark data, and then comparing the audio watermark data of the training audio stream data with the preset audio watermark data. The training audio stream data is obtained by inputting the fault information of the training audio stream data into the fault simulation model so that the fault simulation model simulates the audio synthesis data. The audio synthesis data is synthesized by synthesizing the training audio data and the preset audio watermark data, and there is no fault in the training audio data.

9. A computer device, characterized in that: The computer device includes a memory, a communication interface and a processor, wherein the memory, the communication interface and the processor are interconnected; the memory stores a computer program, and the processor calls the computer program stored in the memory to implement the method described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product, characterized in that The computer program product comprises a computer program, which is stored in a computer storage medium; a processor of a computer device reads the computer program from the computer storage medium, and the processor executes the computer program, so that the computer device executes the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Anomaly checking method and device for audio playing equipment and storage medium

    CN112669883A

  • Methods, systems, apparatus, and articles of manufacture to determine performance of audience measurement meters

    US20230209287A1