Audio processing methods, devices and related equipment

CN119380731BActive Publication Date: 2026-05-26VIVO MOBILE COMM CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2024-10-24
Publication Date
2026-05-26

Smart Images

  • Figure CN119380731B_ABST
    Figure CN119380731B_ABST
Patent Text Reader

Abstract

This application discloses an audio processing method, apparatus, and related devices, belonging to the field of audio processing technology. The audio processing method provided in this application includes: acquiring N corresponding first spectral features and N watermark features, where the first spectral features are the spectral features of audio segments in the audio, and the watermark features are features of watermark information, with N being a positive integer; inputting the N first spectral features and N watermark features into a watermark processing model, and embedding each watermark feature into the corresponding first spectral feature through the generator of the watermark processing model to obtain N second spectral features; and generating watermarked audio based on the N second spectral features and the audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of audio processing technology, specifically relating to an audio processing method, apparatus and related equipment. Background Technology

[0002] Currently, electronic devices can utilize voice replication technology to determine a speaker's timbre, rhythm, and style from a small amount of genuine audio data. They can then use this timbre, rhythm, and style to synthesize audio that matches the original speaker's timbre, rhythm, and style. Typically, to prevent the misuse of voice replication technology, electronic devices can perform a watermark embedding operation after obtaining the replicated audio. This adds an audio watermark to the replicated audio, allowing other electronic devices to determine that the replicated audio is not genuine when they transmit the watermarked audio to other devices.

[0003] However, during the transmission of the replicated audio with the added audio watermark information, signal processing such as filtering and resampling may be performed on the replicated audio, which may cause the audio watermark information to be interfered with or destroyed. Therefore, the reliability of the electronic device in performing the watermark embedding operation is low. Summary of the Invention

[0004] The purpose of this application is to provide an audio processing method, apparatus, and related equipment that can improve the reliability of electronic devices performing watermark embedding operations.

[0005] In a first aspect, embodiments of this application provide an audio processing method, which includes: obtaining N first spectral features and N watermark features that correspond one-to-one, wherein the first spectral features are spectral features of audio segments in the audio, and the watermark features are features of watermark information, and N is a positive integer; inputting the N first spectral features and N watermark features into a watermark processing model, and embedding each watermark feature into the corresponding first spectral feature through the generator of the watermark processing model to obtain N second spectral features; and generating watermarked audio based on the N second spectral features and the audio.

[0006] Secondly, embodiments of this application provide an audio processing apparatus, comprising an acquisition module and a processing module. The acquisition module is configured to acquire N corresponding first spectral features and N watermark features, where the first spectral features are the spectral features of audio segments in the audio, and the watermark features are features of watermark information, where N is a positive integer. The processing module is configured to input the N first spectral features and N watermark features acquired by the acquisition module into a watermark processing model, embed each watermark feature into the corresponding first spectral feature through the generator of the watermark processing model to obtain N second spectral features; and generate watermarked audio based on the N second spectral features and the audio.

[0007] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the steps of the method described in the first aspect.

[0010] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the method described in the first aspect.

[0011] In this embodiment, the electronic device can acquire N first spectral features and N watermark features, each first spectral feature being the spectral feature of an audio segment in the audio, and each watermark feature being a feature of the watermark information. The electronic device inputs these N first spectral features and N watermark features into a watermark processing model. The generator of the watermark processing model embeds each watermark feature into the corresponding first spectral feature to obtain N second spectral features. Thus, the electronic device can generate watermarked audio based on these N second spectral features and the audio. Since the electronic device can acquire N first spectral features and N watermark features of N audio segments in the audio, and embed each watermark feature into the corresponding first spectral feature through the generator of the trained watermark processing model to obtain N second spectral features, instead of directly adding N watermark information to N audio segments, on the one hand, because the first spectral features and watermark features are not easily interfered with or destroyed by signal processing, even if signal processing is performed on the obtained watermarked audio, the watermark features will not be interfered with or destroyed. On the other hand, since the watermarking model is a trained model, after inputting N first spectral features and N watermark features into the watermarking model, the generator of the watermarking model can accurately embed each watermark feature into the corresponding first spectral feature, thereby improving the reliability of the watermarked audio. Attached Figure Description

[0012] Figure 1 This is one of the flowcharts illustrating the audio processing method provided in the embodiments of this application;

[0013] Figure 2 This is a second schematic flowchart of the audio processing method provided in the embodiments of this application;

[0014] Figure 3 This is a schematic diagram illustrating the amount of audio data in the audio processing method provided in the embodiments of this application;

[0015] Figure 4 This is the third flowchart illustrating the audio processing method provided in the embodiments of this application;

[0016] Figure 5 This is the fourth flowchart illustrating the audio processing method provided in the embodiments of this application;

[0017] Figure 6 This is a schematic diagram of a reversible block in the audio processing method provided in the embodiments of this application;

[0018] Figure 7 This is the fifth flowchart illustrating the audio processing method provided in the embodiments of this application;

[0019] Figure 8This is the sixth flowchart illustrating the audio processing method provided in the embodiments of this application;

[0020] Figure 9 This is the seventh flowchart illustrating the audio processing method provided in the embodiments of this application;

[0021] Figure 10 This is the eighth flowchart illustrating the audio processing method provided in the embodiments of this application;

[0022] Figure 11 This is a schematic diagram illustrating the amount of audio data in the audio processing method provided in the embodiments of this application;

[0023] Figure 12 This is the ninth flowchart illustrating the audio processing method provided in the embodiments of this application;

[0024] Figure 13 This is a schematic diagram of a reversible block in the audio processing method provided in the embodiments of this application;

[0025] Figure 14 This is the tenth flowchart illustrating the audio processing method provided in the embodiments of this application;

[0026] Figure 15 This is eleventh of the flowcharts illustrating the audio processing method provided in the embodiments of this application;

[0027] Figure 16 This is a schematic diagram of the structure of the audio processing device provided in the embodiments of this application;

[0028] Figure 17 This is one of the hardware structure diagrams of the electronic device provided in the embodiments of this application;

[0029] Figure 18 This is the second schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0030] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0031] The following will explain the technical terms used in the embodiments of this application.

[0032] 1. Sound replication technology

[0033] Voice replication technology is a rapidly developing text-to-speech (TTS) technology in recent years. Electronic devices can collect a speaker's voice and extract its speech features. When the speaker's voice needs to be replicated, the electronic device can directly synthesize audio based on the original audio and these speech features to obtain replicated audio. The content of the replicated audio is identical to the original text or audio, and the timbre, rhythm, and style of the replicated audio are identical to the original speaker's timbre, rhythm, and style.

[0034] 2. Audio watermark information

[0035] Audio watermarking is a special type of information hidden in an audio signal (such as copyright information or other metadata). Its purpose is to provide the copyright, source, or other relevant information of the audio without affecting the listening experience.

[0036] Audio watermarking is a technique that uses redundant data and randomness in digital audio to embed audio watermark information into the audio signal. This audio watermark information is imperceptible to the human ear and is not easy to remove.

[0037] 3. Generative Adversarial Networks

[0038] Generative Adversarial Networks (GANs) mainly consist of two parts: a generator and a discriminator. Their basic principle is learning through the adversarial interaction between these two components. The generator's task is to generate artificial samples that closely approximate the real data distribution, while the discriminator aims to distinguish between real and generated data as accurately as possible.

[0039] 4. Sliding window

[0040] A sliding window is a fixed-size or dynamically adjustable window that can slide across a data stream or dataset to process the data.

[0041] 5. Data Augmentation

[0042] Data augmentation is a commonly used technique in machine learning and deep learning. It generates new training samples by performing various transformations and processing on existing training data, thereby increasing the diversity and quantity of the dataset and improving the generalization ability and robustness of the model.

[0043] 6. Other terms

[0044] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0045] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."

[0046] The audio processing method, apparatus, and related equipment provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0047] The audio processing method provided in this application can be executed by an audio processing device, a terminal, or a functional module or entity within a terminal. This application uses an electronic device executing the audio processing method as an example to illustrate the audio processing method provided in this application.

[0048] Figure 1 A flowchart illustrating an audio processing method provided in an embodiment of this application is shown. Figure 1 As shown, an audio processing method provided in this application embodiment may include the following steps 101 to 103.

[0049] Step 101: The electronic device acquires N corresponding first spectral features and N watermark features.

[0050] In this embodiment of the application, each of the above N first spectral features is a spectral feature of an audio segment in the audio, and N is a positive integer.

[0051] In some embodiments of this application, when the electronic device obtains the above-mentioned audio through an audio processing application, the electronic device can directly acquire N first spectral features and N watermark features.

[0052] In some embodiments of this application, the aforementioned audio can be replicated audio, that is, replicated audio generated using sound replication technology. The replicated audio can be speech audio or music audio, etc. Of course, the audio can also be other types of audio, and this application does not limit this.

[0053] In some embodiments of this application, the audio segment in the above-mentioned audio can be understood as the audio segment in the audio to which a watermark is to be added.

[0054] In some embodiments of this application, the N audio segments may at least partially overlap, or the N audio segments may not overlap. The durations corresponding to the N audio segments may be at least partially the same, or the durations corresponding to the N audio segments may all be different.

[0055] In one possible implementation of this application, the aforementioned N audio segments can be predetermined audio segments in the audio. For example, N predetermined time periods can be pre-stored in the electronic device, so that the electronic device can determine the audio segment located in the first predetermined time period as the first audio segment, and the audio segment located in the second predetermined time period as the second audio segment, and so on, to determine N audio segments.

[0056] In another possible implementation of this application, the aforementioned N audio segments can be audio segments determined according to a preset duration and a preset time interval. In some examples, combined with Figure 1 ,like Figure 2 As shown, prior to step 101 above, the audio processing method provided in this application embodiment may further include step 201 below.

[0057] Step 201: The electronic device determines an audio segment from the audio at preset time intervals.

[0058] In this embodiment of the application, the duration of an audio segment is matched with a preset duration.

[0059] In some embodiments of this application, the preset duration can be any value within a first range, where the lower limit of the first range can be 1 second and the upper limit can be 1.5 seconds. The preset time interval can be any value within a second range, where the lower limit of the second range can be 0.5 seconds and the upper limit can be 1.5 seconds.

[0060] In some embodiments of this application, the electronic device may first determine an audio start position from the audio, and then determine an audio end position as the position after a preset duration from the audio start position, and determine the audio segment between the audio start position and the audio end position as an audio segment. Then, the electronic device may first determine another audio start position as the position after a preset time interval from the audio end position, and then determine another audio end position as the position after a preset duration from the other audio start position, and determine the audio segment between the other audio start position and the other audio end position as another audio segment. This process is repeated until N audio segments are determined.

[0061] For example, Figure 3 A schematic diagram of the audio data stream is shown. (For example...) Figure 3 As shown, the electronic device can determine an audio segment whose duration matches the preset duration at preset time intervals, thus identifying audio segment 1 and audio segment 2. Audio segment 1 is located in dashed box 10, and audio segment 2 is located in dashed box 11. Assuming the preset time interval is 1 second and the preset duration is 1.2 seconds, then the duration corresponding to audio segment 1 is 1 second, the duration corresponding to audio segment 2 is 1 second, and the interval between the end position of audio segment 1 and the start position of audio segment 2 is 1.2 seconds.

[0062] Thus, since electronic devices can determine an audio segment whose duration matches a preset time interval from the audio, i.e., the audio segment with added watermark information can be determined in a specific way, other electronic devices can also accurately determine the audio segment with added watermark information in the same way, without the need for the electronic device to additionally indicate the location of the audio segment with added watermark information, thereby simplifying the complexity of parsing watermark information.

[0063] In some embodiments of this application, the aforementioned first spectral feature may include at least one of the following: Mel spectral feature, spectral slope feature, spectral centroid feature, spectral smoothness feature, spectral taper feature, etc. Of course, the first spectral feature may also include other features, and this application embodiment does not limit this.

[0064] Among them, the aforementioned MEL spectral features are two-dimensional spectrograms that contain both time-domain and frequency-domain information of the audio.

[0065] In some examples, where the first spectral feature includes a Mel spectral feature, and the electronic device determines N audio segments, the electronic device can first preprocess the N audio segments, then perform a short-time Fourier transform operation on the processed N audio segments to obtain N short-time amplitude spectra, and finally input these N short-time amplitude spectra into the Mel mel filter bank to obtain N Mel spectral features. The feature dimension can be [T×D], where T and D are both positive integers. The preprocessing described above can include at least one of the following: sampling rate conversion, pre-emphasis processing, frame segmentation processing, and windowing processing.

[0066] For example, assuming that the duration of an audio segment is 1 second, the dimension of the mel spectrum feature corresponding to that audio segment is [100×128].

[0067] In this embodiment of the application, each of the above N watermark features is a feature of the watermark information.

[0068] In some embodiments of this application, the aforementioned N watermark features can be features of the same watermark information. It is understood that the N watermark features can be identical.

[0069] In some embodiments of this application, the watermark information can be a watermark code, which can be a binary code.

[0070] In some embodiments of this application, the electronic device can first obtain a first identifier, determine watermark content information based on the first identifier, and then determine watermark information based on the watermark content information. The first identifier includes at least one of the following: the device ID of the electronic device, the user ID logged into the aforementioned audio processing application, and the username logged into the aforementioned audio processing application. The electronic device stores at least one first correspondence, each first correspondence being a correspondence between an audio watermark code and an identifier. Thus, the electronic device can determine a matching identifier from at least one identifier based on the first identifier, and identify the audio watermark code corresponding to that identifier as the watermark content information.

[0071] It should be noted that the number of bits in the aforementioned audio watermark encoding can be greater than or equal to a preset number of bits, thus allowing the number of bits in the watermark content information to also be greater than or equal to the preset number of bits, thereby enabling the watermark content information to cover a wider range of identifiers. The preset number of bits can be any integer from 32 to 40. Of course, the preset number of bits can also be other numbers, and this embodiment does not limit this.

[0072] In some examples, electronic devices can directly identify the watermark content information as watermark information.

[0073] In other examples, the electronic device can first acquire a tag and watermark content information, and then determine the watermark information based on the tag and watermark content information. This involves combining... Figure 1 ,like Figure 4 As shown, prior to step 101 above, the audio processing method provided in this application embodiment may further include steps 202 and 203 as described below.

[0074] Step 202: The electronic device acquires the first marker information and watermark content information.

[0075] In this embodiment of the application, the first marking information is used to locate the watermark information.

[0076] In some embodiments of this application, the electronic device stores first tag information, allowing the electronic device to directly access the first tag information. It is understood that the first tag information can be preset tag information.

[0077] In some embodiments of this application, the first marker information can be a marker code of a predetermined number of bits, which can be a binary code. The predetermined number of bits can be 8, but it can also be other numbers; this application does not limit this. The marker code can be an all-zero code, such as 00000000; or an all-one code, such as 11111111. Of course, the marker code can also be other fixed codes; this application does not limit this.

[0078] In some embodiments of this application, at least one first correspondence is stored in the electronic device. Each first correspondence is a correspondence between an audio watermark code and an identifier. The electronic device can first obtain a first identifier, and based on the first identifier, determine a matching identifier from at least one identifier, and determine the audio watermark code corresponding to that identifier as the watermark content information. The first identifier includes at least one of the following: the device ID of the electronic device, the user ID logged into the audio processing application, and the username logged into the audio processing application.

[0079] Step 203: The electronic device generates watermark information based on the first marker information and the watermark content information.

[0080] In some embodiments of this application, the electronic device can splice the first marker information and the watermark content information to obtain the watermark information.

[0081] In one example, the electronic device can append the watermark content information after the first marker information to obtain the watermark information. Alternatively, the electronic device can obtain the watermark information based on the watermark content information before the first marker information.

[0082] For example, suppose the first tag information is a tag code, which is 01010101, and the watermark content information is 000000....00001. Then, the electronic device can concatenate the watermark content information 000000....00001 after the tag code 01010101 to obtain the watermark information 01010101000000....00001.

[0083] Thus, since the electronic device can acquire the first marker information used to locate the watermark information while acquiring the watermark content information, the position of the watermark information can be accurately determined based on the first marker information when parsing the watermark information, and the watermark content information can be accurately parsed based on the position of the watermark information. Therefore, the situation where the watermark information cannot be parsed can be reduced, thereby improving the reliability of the electronic device in performing the watermark embedding operation.

[0084] In some embodiments of this application, after generating watermark information, the electronic device can first input the watermark information into a linear layer of the watermark processing model. This linear layer maps the watermark information into vectors with the same waveform length as each audio segment, resulting in N vectors. The electronic device can then preprocess these N vectors, perform a short-time Fourier transform on the processed N vectors to obtain N short-time amplitude spectra, and finally input these N short-time amplitude spectra into a Mel filter bank to obtain N watermark features. The preprocessing described above may include at least one of the following: sampling rate conversion, pre-emphasis processing, frame segmentation processing, and windowing processing.

[0085] It is understandable that N watermark features can be MEL spectral features.

[0086] Step 102: The electronic device inputs N first spectral features and N watermark features into the watermark processing model. The generator of the watermark processing model embeds each watermark feature into the corresponding first spectral feature to obtain N second spectral features.

[0087] In some embodiments of this application, the watermarking processing model described above can be a reversible generative adversarial neural network. This watermarking processing model may include a generator used to embed watermark features into corresponding first spectral features, or to separate watermark features from the first spectral features. Of course, the watermarking processing model may also include other layers or structures, which are not limited in this embodiment.

[0088] Understandably, in this example, the generator is used to embed the watermark feature into the corresponding first spectral feature.

[0089] In some embodiments of this application, the generator may include at least one reversible block, each reversible block including at least two complementary affine coupling layers with consistent input and output dimensions, so that the generator can input N first spectral features and N watermark features into the at least one reversible block, and embed each watermark feature into the corresponding first spectral feature through the at least one reversible block.

[0090] Each reversible block can adopt a ConvNeXt structure. Of course, each reversible block can also adopt other structures, which are not limited in this application.

[0091] In some embodiments of this application, the generator comprises M reversible blocks connected in sequence, where M is a positive integer. Combined with... Figure 1 ,like Figure 5 As shown, step 102 can be implemented through steps 102a and 102b below.

[0092] Step 102a: The electronic device inputs N first spectral features and N watermark features into the first reversible block of M reversible blocks, and embeds at least a portion of the features in each watermark feature into the corresponding first spectral feature through the first reversible block.

[0093] In some embodiments of this application, each reversible block includes at least two complementary affine coupling layers, each affine coupling layer including at least two operational functions, thereby embedding at least a portion of the features in each watermark feature into the corresponding first spectral feature through the at least two complementary affine coupling layers.

[0094] For example, Figure 6 A schematic diagram of any reversible block is shown. Assume this reversible block is the first reversible block 12, which includes two complementary affine coupling layers, such as affine coupling layer 1 and affine coupling layer 2. Affine coupling layer 1 includes operation functions C1 and N1, and affine coupling layer 2 includes operation functions C2 and N2, as shown below. Figure 6 As shown, the electronic device can input the first spectral feature 1 and the corresponding watermark feature 1 into the first reversible block 12. At this time, operation functions C1 and N1 can extract features from the watermark feature 1 and embed the extracted partial features into the first spectral feature 1. Operation functions C2 and N2 can embed some noise features from the extracted partial features into the watermark feature 1. Therefore, the first reversible block 12 can output the processed first spectral feature 2 and the processed watermark feature 2. It can be understood that the first spectral feature 2 includes some features from the watermark feature 1.

[0095] In some embodiments of this application, each reversible block includes at least two complementary affine coupling layers that can employ a first preset algorithm to embed at least a portion of the features in each watermark feature into the corresponding first spectral feature.

[0096] The first preset algorithm may include:

[0097] v i+1 =v i ·exp(C1(u i ))+N1(u i );

[0098] u i+1 =u i ·exp(C2(v i+1 ))+N2(v i+1 );

[0099] Among them, v i As the first spectral feature of the first invertible block as input, v i The first spectral feature of the output of the first reversible block, u i To input the watermark feature of the first reversible block, u i+1 The watermark feature output by the first reversible block is represented by C1, N1, C2, and N2, which are all operation functions.

[0100] Step 102b: The electronic device inputs the N first spectral features and N watermark features output from the first reversible block into the i-th reversible block of the M reversible blocks, and embeds at least a portion of the features in each watermark feature into the corresponding first spectral feature through the i-th reversible block.

[0101] In this embodiment of the application, i is a positive integer greater than 1.

[0102] It should be noted that, for the description of embedding at least a portion of the features of each watermark feature into the corresponding first spectral feature through the i-th reversible block, please refer to the specific description of embedding at least a portion of the features of each watermark feature into the corresponding first spectral feature in the first reversible block in the above embodiment. The embodiments of this application will not repeat it here.

[0103] Thus, since the generator can include M reversible blocks connected in sequence, the generator can extract features from N watermark features sequentially through these M reversible blocks and embed the extracted features into the corresponding first spectral features. Therefore, it can reduce the occurrence of incomplete feature extraction of N watermark features, thereby reducing the occurrence of inaccurate N second spectral features.

[0104] Step 103: The electronic device generates watermarked audio based on N second spectral features and audio.

[0105] In one example, the electronic device can first generate a replacement audio segment based on each second spectral feature to obtain N replacement audio segments, and then use each replacement audio segment to replace the corresponding audio segment in the audio to obtain the watermarked audio.

[0106] In another example, the electronic device can replace the corresponding first spectral feature in the audio's spectral features based on each second spectral feature, and generate watermarked audio based on the replaced spectral features. This involves combining... Figure 1 ,like Figure 7 As shown, step 103 can be implemented through steps 103a and 103b as described below.

[0107] Step 103a: The electronic device uses each second spectral feature to replace the corresponding first spectral feature in the spectral features of the audio, thereby obtaining the replaced spectral features of the audio.

[0108] In some embodiments of this application, the electronic device may first delete the above-mentioned N first spectral features from the spectral features of the audio, and then insert the corresponding second spectral features at the position of each first spectral feature to obtain the replaced spectral features of the audio.

[0109] Step 103b: The electronic device generates watermarked audio based on the replaced spectral features.

[0110] In some embodiments of this application, the electronic device can input the replaced spectral features into a vocoder, and the vocoder can reconstruct the audio based on the replaced spectral features to obtain a watermarked audio in a predetermined format.

[0111] The vocoder mentioned above may include at least one of the following: HiFiGan, LPCnet. Of course, the vocoder may also include other vocoders, and this embodiment of the application does not limit this. The predetermined format may specifically be wav format, but it may also be other formats, and this embodiment of the application does not limit this.

[0112] Thus, since electronic devices can use each second spectral feature to replace the corresponding first spectral feature in the spectral features of the audio, without processing other spectral features in the spectral features of the audio, the computational load of obtaining watermarked audio can be simplified, thereby reducing the power consumption of generating watermarked audio.

[0113] This application provides an audio processing method. An electronic device can acquire N first spectral features and N watermark features, each first spectral feature being the spectral feature of an audio segment, and each watermark feature being the feature of watermark information. The N first spectral features and N watermark features are input into a watermark processing model. The generator of the watermark processing model embeds each watermark feature into the corresponding first spectral feature to obtain N second spectral features. Thus, the electronic device can generate watermarked audio based on the N second spectral features and the audio. Since the electronic device can acquire N first spectral features and N watermark features of N audio segments, and embed each watermark feature into the corresponding first spectral feature using the generator of the trained watermark processing model to obtain N second spectral features, instead of directly adding N watermark information to N audio segments, the first spectral features and watermark features are less susceptible to interference or damage during signal processing. Therefore, even if signal processing is performed on the obtained watermarked audio, the watermark features will not be interfered with or damaged. On the other hand, since the watermarking model is a trained model, after inputting N first spectral features and N watermark features into the watermarking model, the generator of the watermarking model can accurately embed each watermark feature into the corresponding first spectral feature, thereby improving the reliability of the watermarked audio.

[0114] Furthermore, since electronic devices do not employ any of the least significant bit editing, transform domain algorithm, echo hiding, spread spectrum watermarking, or echo modulation techniques to perform watermark embedding operations, there is no need to require a low number of bits in the watermark information, thereby increasing the number of user IDs that the watermark information can cover.

[0115] Furthermore, since the watermarking processing model is a trained model, the N second spectral features generated by the watermarking processing model can have small differences from the aforementioned N first spectral features. Therefore, the sound quality of the watermarked audio can be avoided, thereby improving the sound quality of the replicated watermarked audio obtained by performing the watermarking embedding operation, and further improving the reliability of the electronic device performing the watermarking embedding operation.

[0116] The following example illustrates the specific scheme for training the generator to be trained in the watermark processing model.

[0117] It should be noted that the watermarking processing model can be trained by electronic devices or other electronic devices (such as servers). In this embodiment, the watermarking processing model is trained by an electronic device as an example.

[0118] In some embodiments of this application, combined with Figure 1 ,like Figure 8As shown, prior to step 101 above, the audio processing method provided in this application embodiment may further include steps 204 to 207 as described below.

[0119] Step 204: The electronic device acquires the T first sample spectral features and T sample watermark features that correspond one-to-one.

[0120] In this embodiment of the application, the first sample spectral feature is the spectral feature of the sample audio segment in the sample audio, the sample watermark feature is the feature of the sample watermark information, and T is a positive integer.

[0121] It should be noted that the descriptions of the first sample spectral features, sample watermark features, sample audio, sample audio segments, and sample watermark information can be found in the specific descriptions of the first spectral features, watermark features, audio, audio segments, and watermark information in the above embodiments, and will not be repeated in this application embodiment.

[0122] In some embodiments of this application, the electronic device can select sample audio from the training sample set and determine T sample audio segments from the sample audio, thereby the electronic device can obtain T first sample spectral features of the T sample audio segments.

[0123] It should be noted that, regarding the explanation of how the electronic device determines T sample audio segments from sample audio, please refer to the specific description of how the electronic device determines N audio segments from audio in the above embodiments, which will not be repeated here in this application embodiment. Similarly, regarding the explanation of how the electronic device acquires T first sample spectral features from T sample audio segments, please refer to the specific description of how the electronic device acquires N first spectral features from N audio segments in the above embodiments, which will not be repeated here in this application embodiment.

[0124] In some embodiments of this application, the electronic device can generate a binary code with a number of bits greater than or equal to the aforementioned preset number of bits, and determine this binary code as sample watermark information. The electronic device can then determine T sample watermark features based on this sample watermark information. The preset number of bits can be 40, but it can also be other values; this application does not limit this.

[0125] The electronic device can use algorithms such as linear congruent generator or Mason tween to generate T binary codes with a bit length greater than or equal to the preset bit length, thereby obtaining T sample watermark features.

[0126] It should be noted that, for the explanation of how the electronic device determines T sample watermark features based on the sample watermark information, please refer to the specific description of how the electronic device determines N watermark features based on the watermark information in the above embodiments. This application will not repeat the description here.

[0127] Step 205: The electronic device inputs T first sample spectral features and T sample watermark features into the watermark processing model. The generator to be trained in the watermark processing model embeds each sample watermark feature into the corresponding first sample spectral feature to obtain N second sample spectral features.

[0128] It should be noted that the explanation of how the electronic device obtains N second sample spectral features can be found in the specific description of how the electronic device obtains N second spectral features in the above embodiments, and will not be repeated here in the embodiments of this application.

[0129] Step 206: The electronic device generates sample watermark audio based on N second sample spectral features and sample audio.

[0130] It should be noted that, for the explanation of how an electronic device generates sample watermark audio based on N second sample spectral features and sample audio, please refer to the specific description of how an electronic device generates watermark audio based on N second spectral features and audio in the above embodiments. This application will not repeat the description here.

[0131] Step 207: The electronic device determines the first loss function based on the sample watermark audio and the sample audio, and uses the first loss function to train the generator to obtain the generator.

[0132] In this embodiment of the application, the first loss function is used to characterize the audio type difference between the sample watermark audio and the sample audio.

[0133] It is understandable that since the first loss function is used to characterize the audio type difference between the sample watermark audio and the sample audio, after the generator is trained using the first loss function, the difference between the second spectral feature generated by the generator and the spectral feature of the audio is small. This ensures that the audio type difference between the generated watermark audio and the audio is small, meaning that the user cannot distinguish the watermark audio from the audio.

[0134] In some embodiments of this application, before the generator is trained, the electronic device may repeatedly execute steps 204 to 207 above to repeatedly train the generator to be trained.

[0135] Thus, after the electronic device determines N second spectral features based on T first sample spectral features and T sample watermark features, and generates sample watermark audio based on these N second spectral features, it can use a first loss function to characterize the audio type difference between the sample watermark audio and the sample audio to train the generator. In this way, the difference between the second spectral features generated by the trained generator and the spectral features of the audio is small, thereby ensuring that the audio type difference between the generated watermark audio and the audio is small, that is, the user cannot distinguish the watermark audio and the audio by sound.

[0136] In some embodiments of this application, step 207 can be specifically implemented by steps 207a and 207b as described below.

[0137] Step 207a: The electronic device inputs the N second sample spectral features into the discriminator of the watermarking processing model, and the discriminator determines the audio type of the sample watermark audio.

[0138] In some embodiments of this application, the audio type can be any of the following: replicated audio, or real audio.

[0139] In some embodiments of this application, the electronic device can perform binary classification on the spectral features of N second samples by a discriminator to obtain the output result, namely the audio type of the sample watermark audio.

[0140] Step 207b: The electronic device determines the first loss function based on the audio type of the sample watermark audio and the audio type of the sample audio.

[0141] In this embodiment of the application, the audio type of the sample audio is real audio.

[0142] In some embodiments of this application, the electronic device can construct a first loss function using a first algorithm. The first algorithm is as follows:

[0143] L g =log(1-D(f′));

[0144] Among them, L g Let f be the first loss function, D(f') be the audio type of the sample watermark audio, and 1 be used to indicate the audio type of the sample audio.

[0145] Thus, since the electronic device can accurately determine the audio type of the sample watermark audio through the discriminator, it can accurately determine the first loss function based on the accurate audio type of the sample watermark audio and the audio type of the sample audio. Therefore, after training the generator using the first loss function, the difference between the second spectral features generated by the trained generator and the spectral features of the audio is small. This ensures that the audio type difference between the generated watermark audio and the audio is small, meaning that the user cannot distinguish the watermark audio from the audio by sound.

[0146] In some embodiments of this application, before step 207a above, the audio processing method provided in the embodiments of this application may further include steps 301 to 303 as described below.

[0147] Step 301: The electronic device inputs the spectral features of N second samples and the spectral features of the sample watermark audio into the discriminator to be trained in the watermark processing model, and determines the training audio type of the sample watermark audio and the training audio type of the sample audio through the discriminator to be trained.

[0148] It is understood that the training audio type mentioned above can be the audio type determined by the discriminator to be trained. This training audio type may differ from the real audio type.

[0149] In some embodiments of this application, the discriminator to be trained can perform binary classification based on N second sample spectral features to determine the training audio type of the sample watermark audio, and perform binary classification based on the spectral features of the sample watermark audio to determine the training audio type of the sample audio.

[0150] Step 302: The electronic device determines the third loss function based on the training audio type of the sample watermark audio, and determines the fourth loss function based on the training audio type of the sample audio.

[0151] In this embodiment of the application, the third loss function is used to characterize the difference between the training audio type and the audio type of the sample watermark audio, and the fourth loss function is used to characterize the difference between the training audio type and the audio type of the sample audio.

[0152] In some embodiments of this application, the electronic device may employ a second algorithm to construct a third loss function, and a third algorithm to construct a fourth loss function. The second algorithm is as follows:

[0153] L d1 =log(D(f′));

[0154] Among them, L d1The third loss function is f', the second sample spectral feature is f', and D(f') is used to indicate the training audio type of the sample watermark audio.

[0155] The third algorithm is:

[0156] L d1 =log(1-D(f));

[0157] Among them, L d2 The fourth loss function is f, which is the spectral feature of the sample watermarked audio. D(f) is used to indicate the training audio type of the sample audio.

[0158] Step 303: The electronic device trains the discriminator to be trained using the third loss function and the fourth loss function to obtain the discriminator.

[0159] In some embodiments of this application, the electronic device may determine the sum of the third loss function and the fourth loss function as the loss function used for training, and use the loss function used for training to train the discriminator to be trained.

[0160] It is understandable that electronic devices can first determine the loss function L used for training. d =log(1-D(f))+log(D(f')), when using the loss function L used in training. d The discriminator is trained.

[0161] Thus, the electronic device can determine the training audio type of the sample watermark audio based on N second sample spectral features, determine the third loss function based on the training audio type of the sample watermark audio, and determine the fourth loss function based on the spectral features of the sample watermark audio. After training the discriminator using the third and fourth loss functions, the trained discriminator can accurately determine the audio type corresponding to the input spectral features. This allows the electronic device to determine a more accurate first loss function in subsequent steps. In other words, the electronic device can perform adversarial training between the generator and the discriminator, thereby further reducing the difference between the second spectral features generated by the trained generator and the spectral features of the audio. This further reduces the audio type difference between the generated watermark audio and the audio, improving the sound quality of the watermark audio obtained from the watermark embedding operation. This enhances the reliability of the electronic device performing the watermark embedding operation.

[0162] Of course, in order to further ensure the difference between the second spectral features generated by the trained generator and the spectral features of the audio, the electronic device can also construct multiple loss functions to train the generator to be trained using the first loss function and the multiple loss functions. Examples will be given below.

[0163] In some embodiments of this application, before step 207 above, the audio processing method provided in the embodiments of this application may further include step 304 below, and step 207 above can be specifically implemented by step 207c below.

[0164] Step 304: The electronic device determines the second loss function based on the spectral features of the T first samples and the watermark features of the T samples.

[0165] In this embodiment of the application, the second loss function is used to characterize the difference between the T first sample spectral features and the T sample watermark features.

[0166] In some embodiments of this application, the electronic device can determine a second loss function based on the difference between the spectral features of each first sample and the corresponding watermark features of the sample.

[0167] The electronic device can use a fourth algorithm to construct the second loss function. This fourth algorithm is as follows:

[0168]

[0169] Among them, L f Let f be the second loss function, f be the first sample spectral feature, and f' be the sample watermark feature.

[0170] Step 207c: The electronic device determines the first loss function based on the sample watermark audio and the sample audio, and uses the first loss function and the second loss function to train the generator to be trained.

[0171] In some embodiments of this application, the electronic device may determine the sum of the first loss function and the second loss function as the loss function used for training, and train the generator to be trained based on the loss function used for training.

[0172] Thus, since the electronic device can also determine the second loss function based on T first sample spectral features and T sample watermark features, after training the generator using the first and second loss functions, the difference between the second spectral features generated by the trained generator and the spectral features of the audio is smaller. This ensures that the audio type difference between the generated watermark audio and the audio is smaller, meaning that the user cannot distinguish the watermark audio from the audio. Therefore, the sound quality of the watermark audio can be further improved, thereby improving the reliability of the electronic device in performing the watermark embedding operation.

[0173] In some embodiments of this application, before “training the generator to be trained using the first loss function and the second loss function” in step 207c above, the electronic device can also perform watermark parsing on the sample watermark audio to obtain the parsed watermark information. Then, the electronic device can also determine a loss function (such as the fifth loss function in the following embodiments) based on the parsed watermark information and the sample watermark information, and train the generator to be trained based on the first loss function, the second loss function and the loss function to obtain the generator.

[0174] It is understandable that since the generator can also be used to separate watermark features from the first spectral features, the electronic device can determine the loss function and train the generator to be trained based on the first loss function, the second loss function and the loss function, so that the trained generator can accurately embed the watermark features into the corresponding first spectral features and accurately separate the watermark features from the first spectral features.

[0175] It should be noted that the description of watermark parsing of sample watermark audio by electronic devices will be specifically described in the following embodiments, and will not be repeated in the embodiments of this application here.

[0176] In some examples, the electronic device can use a fifth algorithm to construct the aforementioned loss function. This fifth algorithm can be:

[0177]

[0178] Among them, L m Let m be a loss function, and m be the watermark information of the sample. To parse the watermark information.

[0179] In some examples, the electronic device may determine the loss function used for training by the sum of the first loss function, the second loss function, and the loss function used for training, and train the generator to be trained based on the loss function used for training.

[0180] The loss function L used in this training can be...

[0181] L=λ f L f +λ g L g +L m ;

[0182] Among them, L g Let L be the first loss function. f For the second loss function, λ f λ represents the weights of the first loss function.g The weights are for the second loss function.

[0183] It is understandable that this can be achieved by controlling λ. f and λ g The size of λ is used to control the balance between the stability of the watermark information and not affecting the listening experience. Here, λ f The value can be 10000, λ g The value can be 10.

[0184] In some embodiments of this application, combined with Figure 1 ,like Figure 9 As shown, after step 103 above, an audio processing method provided in this application embodiment may include the following steps 401 to 403.

[0185] Step 401: The electronic device acquires L third spectral features.

[0186] In this embodiment of the application, each of the above L third spectral features is the spectral feature corresponding to the audio segment in the watermarked audio, and L is a positive integer.

[0187] In some embodiments of this application, when an electronic device receives watermarked audio from another electronic device, the electronic device can directly acquire L third spectral features.

[0188] In some embodiments of this application, the above-mentioned L third spectral features are the N second spectral features in the above embodiments, or the L third spectral features are the spectral features of an audio segment determined by the electronic device.

[0189] It is understandable that, since electronic devices may be unable to accurately determine the location of audio segments with added watermark information, they can automatically determine L audio segments from the watermarked audio to identify L third spectral features. These L audio segments may include the aforementioned N audio segments, or the L audio segments may be at least a portion of the aforementioned N audio segments.

[0190] The following example illustrates a specific scheme for an electronic device to determine L third spectral features from watermarked audio.

[0191] In some embodiments of this application, combined with Figure 9 ,like Figure 10 As shown, before step 401 above, the audio processing method provided in this application embodiment may further include steps 501 and 502 as described below.

[0192] Step 501: The electronic device determines at least one sliding window.

[0193] In this embodiment of the application, the duration of each sliding window in the at least one sliding window is matched with a preset duration, and the interval between two adjacent sliding windows is a preset time interval.

[0194] In some embodiments of this application, the number of sliding windows, the preset duration, and the preset time interval can be pre-configured in the electronic device, so that the electronic device can determine at least one sliding window based on the number of sliding windows, the preset duration, and the preset time interval.

[0195] The number of sliding windows can be the same as the number N of audio segments in the audio in the above embodiment. It is understood that since the electronic device adds watermark information to N audio segments according to a preset duration and preset time interval, the electronic device can determine at least one sliding window based on the number N of sliding windows, the preset duration, and the preset time interval, to identify the audio segments for which watermark information has been added.

[0196] In some embodiments of this application, an audio start position can be pre-configured in the electronic device. The electronic device can then determine an audio end position as the position following a preset duration from the audio start position, and define the window corresponding to the audio start position and the audio end position as a sliding window. Next, the electronic device can first determine another audio start position as the position following a preset time interval from the audio end position, and then determine another audio end position as the position following a preset duration from the other audio start position, defining the window corresponding to the other audio start position and the other audio end position as another sliding window. This process is repeated to determine at least one sliding window, i.e., N sliding windows.

[0197] Step 502: The electronic device controls at least one sliding window to slide on the audio segment of the watermark audio according to the first preset step size, and determines at least one audio segment located in at least one sliding window after each slide as at least one audio segment, so as to determine L audio segments.

[0198] In some embodiments of this application, the first preset step size can be determined by a preset duration. The first preset step size can be Y times the preset duration, where Y is a positive number. For example, the first preset step size can be 0.1 times the preset duration.

[0199] In some embodiments of this application, the electronic device can first control at least one sliding window to slide at least once in a first direction on the audio segment of the watermarked audio according to a first preset step size, and determine at least one audio segment within the at least one sliding window after each slide as at least one audio segment 1; then, control at least one sliding window to slide at least once in a second direction on the audio segment of the watermarked audio according to the first preset step size, and determine at least one audio segment within the at least one sliding window after each slide as at least one audio segment 2; thus, the electronic device can determine at least one audio segment 1 and at least one audio segment 2 as L audio segments. The second direction is opposite to the first direction.

[0200] For example, assuming the preset duration is 1 second and the preset time interval is 1 second, then each sliding window in at least one of the following has a duration of 1 second, the first preset step size is 0.1 seconds, the duration of the watermarked audio is 3.5 seconds, and the number of audio segments in the watermarked audio is 2. Figure 11 As shown, the electronic device can first control two sliding windows to slide at least once in the first direction at 0.1s intervals. For example, it can control sliding windows 13 and 14 to slide to the right 5 times, and identify the two audio segments within the two sliding windows after each slide as audio segment 1. Then, the electronic device can control the two sliding windows to slide at least once in the second direction at 0.1s intervals. For example, it can control sliding windows 13 and 14 to slide to the left 5 times, and identify the two audio segments within the two sliding windows after each slide as audio segment 2. Thus, the electronic device can identify L audio segments, namely audio segment 1 and audio segment 2.

[0201] It should be noted that, in Figure 11 The sliding window 13 and sliding window 14 are illustrated with multiple dashed boxes. Each dashed box can be understood as the position of sliding window 13 or sliding window 14 after it has been moved.

[0202] Thus, it can be seen that since the electronic device can determine at least one sliding window based on a preset duration and a preset time interval, and then slide through the at least one sliding window multiple times to determine L audio segments, that is, to determine the audio segments including the audio segments with added watermark information, without the electronic device needing to indicate the specific location of the audio segments with added watermark information, the complexity of parsing the watermark information can be simplified.

[0203] Step 402: The electronic device inputs L third spectral features into the watermarking processing model, and the generator of the watermarking processing model determines the watermark feature in each third spectral feature.

[0204] In this embodiment, each watermark feature is a feature of the watermark information.

[0205] In some embodiments of this application, the generator comprises M reversible blocks connected in sequence, where M is a positive integer. In some examples, combined with... Figure 9 ,like Figure 12 As shown, step 402 can be implemented through steps 402a and 402b below.

[0206] Step 402a: The electronic device inputs the L corresponding third spectral features and L noise features into the Mth reversible block of the M reversible blocks, and embeds at least part of the watermark feature in each third spectral feature into the corresponding noise feature through the Mth reversible block.

[0207] In some embodiments of this application, the noise characteristics described above may be random noise generated by an electronic device.

[0208] In some embodiments of this application, each reversible block includes at least two complementary affine coupling layers, each affine coupling layer including at least two operational functions, thereby embedding at least a portion of the watermark features in each third spectral feature into the corresponding noise feature through the at least two complementary affine coupling layers.

[0209] For example, Figure 13 A schematic diagram of any reversible block is shown. Assume this reversible block is the Mth reversible block 15, which includes two complementary affine coupling layers, such as affine coupling layer 1 and affine coupling layer 2. Affine coupling layer 1 includes operation functions C1 and N1, and affine coupling layer 2 includes operation functions C2 and N2, as shown below. Figure 13 As shown, the electronic device can input the third spectral feature 1 and the corresponding noise feature 1 into the Mth reversible block 15. At this time, the operation functions C2 and N2 can embed a portion of the watermark feature from the third spectral feature 13 into the noise feature 1, and the operation functions C2 and N2 can embed a portion of the noise from the extracted partial features into the third spectral feature 1. Therefore, the Mth reversible block 15 can output the processed third spectral feature 2 and the processed noise feature 2. It can be understood that the noise feature 2 includes a portion of the watermark feature.

[0210] In some embodiments of this application, each reversible block includes at least two complementary affine coupling layers that can employ a second preset algorithm to embed at least a portion of the watermark features in each third spectral feature into the corresponding noise feature.

[0211] The second preset algorithm may include:

[0212] u i =(u i+1 -N2(vi+1 ))·exp(-C2(v i+1 ));

[0213] v i =(v i+1 -N1(u i ))·exp(-C1(u i ));

[0214] Among them, v i+1 For the third spectral feature of the first invertible block as input, v i The third spectral feature of the output of the Mth reversible block, u i+1 To input the watermark feature of the Mth reversible block, u i Let C1, N1, C2, and N2 be the watermark features output by the Mth reversible block, where C1, N1, C2, and N2 are all operational functions.

[0215] Step 402b: The electronic device inputs the L third spectral features and L noise features output from the Mth reversible block into the Mjth reversible block of the M reversible blocks, and embeds at least a portion of the watermark features in each third spectral feature into the corresponding noise feature through the Mjth reversible block.

[0216] In this embodiment of the application, j is a positive integer less than M.

[0217] Thus, since the generator can include M reversible blocks connected in sequence, the generator can extract features from L third spectral features sequentially through these M reversible blocks, and embed at least a portion of the watermark features in each third spectral feature into the corresponding noise features to obtain accurate watermark features. Therefore, the situation where the feature extraction of L third spectral features is incomplete can be reduced, thereby reducing the situation where the obtained L watermark features are inaccurate.

[0218] Step 403: The electronic device determines the watermark information based on L watermark features.

[0219] In one possible implementation of this application, the electronic device can perform restoration processing according to each watermark feature to obtain L watermark information, and then determine the watermark information with the highest repetition frequency among the L watermark information as the aforementioned watermark information.

[0220] It is understandable that, since some audio segments among the L determined audio segments may be inaccurate, that is, some audio segments may be completely different from any of the N audio segments, this may lead to some third spectral features among the L third spectral features being inaccurate, thus leading to some watermark features among the L watermark features being inaccurate. Therefore, the electronic device can first determine the L watermark information, and then determine the watermark information with the highest repetition frequency among the L watermark information, that is, the most likely correct watermark information, as the aforementioned watermark information.

[0221] In another possible implementation of this application, the electronic device can perform restoration processing on each watermark feature to obtain L watermark information, and then determine the most likely correct watermark information based on the marker information in each watermark information, so as to identify the most likely correct watermark information as the aforementioned watermark information. Wherein, combined with Figure 9 ,like Figure 14 As shown, step 403 can be implemented through steps 403a to 403c as described below.

[0222] Step 403a: The electronic device generates a candidate watermark information based on each watermark feature.

[0223] In this embodiment of the application, each candidate watermark information includes second marker information and watermark content information.

[0224] It should be noted that the description of the second marking information can be found in the specific description of the first marking information in the above embodiments, and will not be repeated here. Similarly, the description of the watermark content information can be found in the specific description in the above embodiments, and will not be repeated here.

[0225] In some embodiments of this application, the electronic device can perform restoration processing based on each watermark feature to obtain each candidate watermark information, i.e., L candidate watermark information.

[0226] It is understandable that each candidate watermark information can be binary encoded.

[0227] Step 403b: The electronic device determines at least one candidate watermark information from L candidate watermark information whose second marker information matches the first marker information.

[0228] In this embodiment of the application, the first marking information is used to locate the watermark information.

[0229] Step 403c: The electronic device determines the watermark information based on the mean of at least one candidate watermark information.

[0230] In some embodiments of this application, the electronic device can sum each digit of at least one candidate watermark information and then take the average. If the average of a certain digit is greater than 0.5, then the average of that digit is 1; otherwise, the average of that digit is 0, thereby obtaining the watermark information.

[0231] For example, assuming at least one candidate watermark information includes: candidate watermark information 00000100, candidate watermark information 00001010, and candidate watermark information 00001100, the electronic device can first sum the first digit of each candidate watermark information and then take the average. The average of the first digit is 0. Then, it can sum the second digit of each candidate watermark information and take the average. The average of the second digit is 0. This process is repeated to obtain the watermark information 00001100.

[0232] Thus, since the electronic device can first determine L candidate watermark information and accurately determine at least one candidate watermark information that may be watermark information based on the second marker information of the L candidate watermark information, the electronic device can accurately determine the watermark information based on the average of the at least one candidate watermark information that may be watermark information.

[0233] In summary, the electronic device can acquire L third spectral features, each of which represents the spectral characteristics of an audio segment in the watermarked audio, where L is a positive integer. These L third spectral features are input into a watermarking processing model. The generator of this model determines the watermark feature within each third spectral feature, which is the characteristic of the watermark information. Therefore, the electronic device can determine the watermark information based on these L watermark features. Since the electronic device can acquire the L third spectral features corresponding to the audio segments in the watermarked audio and determine the watermark feature within each third spectral feature using the generator of the trained watermarking processing model, on the one hand, because the third spectral features and watermark features are not easily interfered with or destroyed by signal processing, even if the watermarked audio undergoes signal processing, the electronic device can still acquire accurate L third spectral features and accurately determine the watermark feature within each third spectral feature using the generator, thus obtaining accurate L watermark features. Therefore, the electronic device can accurately determine the watermark information based on these accurate L watermark features and accurately identify the watermarked audio as a replica, rather than mistakenly identifying it as genuine audio. On the other hand, because the watermarking model is a pre-trained model, after inputting L third spectral features into it, the generator of the watermarking model can accurately determine the watermark feature in each third spectral feature. This reduces the possibility of electronic devices misidentifying watermark features. Consequently, based on the accurate L watermark features, the electronic device can accurately determine the watermark information and accurately identify the watermarked audio as a replica, rather than mistakenly identifying it as genuine audio. This improves the reliability of the electronic device in distinguishing between replica and genuine audio.

[0234] Furthermore, the electronic device can accurately determine the watermark information based on L third spectral features without electronic device instruction, thus simplifying the complexity of parsing the watermark information.

[0235] The following example illustrates the specific scheme for training the generator to be trained in the watermark processing model.

[0236] It should be noted that the watermarking processing model can be trained by electronic devices or other electronic devices (such as servers). In this embodiment, the watermarking processing model is trained by an electronic device as an example.

[0237] In some embodiments of this application, combined with Figure 9 ,like Figure 15 As shown, before step 401 above, the audio processing method provided in this application embodiment may further include steps 503 to 506 as described below.

[0238] Step 503: The electronic device acquires the spectral features of R third samples.

[0239] In this embodiment of the application, each of the above R third sample spectral features is a spectral feature of a sample audio segment in the sample watermark audio, where R is a positive integer.

[0240] In some embodiments of this application, the electronic device can first acquire T first sample spectral features and T sample watermark features, each corresponding to a specific audio segment in the sample audio. The first sample spectral features are the spectral features of the audio segments in the sample audio, and the sample watermark features are the features of the watermark information, where T is a positive integer. Then, the T first sample spectral features and T sample watermark features are input into a watermarking model. The generator to be trained in the watermarking model embeds each sample watermark feature into the corresponding sample spectral feature, resulting in N second sample spectral features. Thus, the electronic device can generate sample watermark audio based on the N second sample spectral features and the sample audio. In this way, the electronic device can acquire R third sample spectral features.

[0241] It should be noted that for the description of the electronic device generating sample watermark audio, please refer to the specific description of the electronic device generating sample watermark audio in the above embodiments, and this application embodiment will not repeat it here.

[0242] In some embodiments of this application, step 503 can be specifically implemented by the following steps 503a and 503b.

[0243] Step 503a: The electronic device determines at least one training sliding window.

[0244] In this embodiment, the duration of each training sliding window in the at least one training sliding window is matched with a preset duration, and the interval between two adjacent training sliding windows is a preset time interval.

[0245] It should be noted that the description of determining at least one training sliding window for electronic devices can be found in the specific description in the above embodiments, and will not be repeated here in the embodiments of this application.

[0246] Step 503b: The electronic device controls at least one training sliding window to slide on the audio segment of the watermarked audio according to the second preset step size, and determines at least one audio segment located in at least one training sliding window after each slide as a sample audio segment, so as to determine R sample audio segments.

[0247] Thus, since the electronic device can determine at least one training sliding window based on a preset duration and a preset time interval, the electronic device can slide through the at least one training sliding window multiple times to accurately determine R sample audio segments. Therefore, the complexity of training the watermarking model can be simplified.

[0248] Step 504: The electronic device inputs the R third sample spectral features into the watermarking processing model, and determines the training watermark features in each third sample spectral feature through the generator to be trained in the watermarking processing model.

[0249] It is understandable that the above-mentioned training watermark features are watermark features parsed by the generator to be trained. These training watermark features may be inaccurate watermark features, that is, watermark features that are different from the watermark features of the sample watermark information.

[0250] Step 505: The electronic device determines a parsed watermark information based on each trained watermark feature.

[0251] In some embodiments of this application, the electronic device can perform restoration processing on each training watermark feature to obtain a parsed watermark information.

[0252] It is understandable that the parsed watermark information is the watermark information obtained by the generator to be trained, and this parsed watermark information may be different from the sample watermark information.

[0253] Step 506: The electronic device determines the fifth loss function based on R parsed watermark information and sample watermark information, and uses the fifth loss function to train the generator to obtain the generator.

[0254] In this embodiment of the application, the fifth loss function is used to characterize the difference between the R parsed watermark information and the sample watermark information.

[0255] In some embodiments of this application, before “training the generator to be trained using the fifth loss function” in step 506 above, the electronic device may further determine the first loss function and the second loss function, and determine the sum of the first loss function, the second loss function and the fifth loss function as the loss function to be used for training, and train the generator to be trained based on the loss function to be used for training.

[0256] It should be noted that the explanation of determining the first loss function and the second loss function for electronic devices can be found in the specific description in the above embodiments, and will not be repeated here in the embodiments of this application.

[0257] In some embodiments of this application, before the generator is trained, the electronic device may repeatedly execute steps 503 to 506 above to repeatedly train the generator to be trained.

[0258] Thus, it can be seen that since the electronic device can also determine R training watermark features based on R third sample spectral features, and determine the fifth loss function based on the R training watermark information and sample watermark information corresponding to the R training watermark features, and use the fifth loss function to train the generator to be trained, the accuracy of the watermark features generated by the trained generator is relatively high after using the fifth loss function to train the generator to be trained, thereby improving the accuracy of determining watermark information.

[0259] In some embodiments of this application, during the process of repeatedly executing steps 503 to 506 to train the generator to be trained, the electronic device can also obtain the number of training rounds of the generator to be trained in real time. The number of rounds can be understood as the number of times the electronic device executes steps 503 to 506. When the number of rounds is greater than or equal to a preset threshold, data enhancement is performed on the sample watermark audio. Examples will be given below.

[0260] In some embodiments of this application, the audio processing method provided in this application may further include the following steps 507 to 509.

[0261] Step 507: During the training process of the generator to be trained, the electronic device obtains the number of the first round of training of the generator to be trained.

[0262] In some embodiments of this application, a first counter may be provided in the electronic device, and the electronic device may control the count value of the first counter to increase by 1 each time the generator to be trained is performed. Thus, the electronic device may determine the current count value of the first counter as the first round number after completing any round of training of the generator to be trained.

[0263] Step 508: If the number of first rounds is greater than or equal to a preset threshold, the electronic device processes the sample watermark audio using at least one processing method to obtain enhanced sample watermark audio.

[0264] In the embodiments of this application, the above-mentioned at least one processing method includes at least one of the following: increasing noise, increasing reverberation, volume scaling, audio loss, and lossy compression.

[0265] In some embodiments of this application, the aforementioned addition of noise can specifically be: adding uniformly distributed random white noise to the sample watermark audio, wherein the signal-to-noise ratio of the random white noise can be between 30 and 35 dB. Of course, the signal-to-noise ratio of the random white noise can also be within other ranges, and this application embodiment does not limit this.

[0266] In some embodiments of this application, the above-mentioned increase in reverberation can specifically be achieved by: randomly attenuating the volume of the sample watermark audio to 0.1 to 0.3 times, randomly delaying it by 50 to 100 milliseconds (ms) to obtain a new watermark audio, and then superimposing the new watermark audio with the sample watermark audio. Of course, the multiplier corresponding to the random attenuation and the value of the random delay can be other values, and this application does not limit them.

[0267] In some embodiments of this application, the volume scaling described above can specifically be: randomly changing the volume of the sample watermark audio to 80% to 120% of the sample watermark audio. Of course, the value corresponding to this random change can also be other values, and this application embodiment does not limit this.

[0268] In some embodiments of this application, the aforementioned audio loss can specifically be achieved by randomly setting 5% to 10% of the sampling points in the sample watermark audio to zero. Of course, the number of sampling points can also be other values, and this application does not limit this.

[0269] In some embodiments of this application, the aforementioned lossy compression may specifically involve converting the sample watermark audio into a lossy mp3 or m4a format, and then converting it back. Of course, the lossy format can also be other formats, and this application does not limit this.

[0270] In some embodiments of this application, the sample watermark audio can be in WAV format, so that electronic devices can use the SOX tool or FFmpeg to process the sample watermark audio using at least one of the above processing methods.

[0271] It is understandable that in the complex environment of actual audio transmission, the audio signal will face the combined effects of multiple transformations and disturbances at the same time. Therefore, during the training phase, at least one processing method can be combined to process the sample watermark audio to obtain enhanced sample watermark audio.

[0272] In this embodiment of the application, the number of at least one of the above processing methods is positively correlated with the number of the first round.

[0273] It is understandable that increasing the number of processing methods will increase the training difficulty of the generator to be trained. Therefore, when the number of rounds of the generator to be trained is small, fewer processing methods can be used to process the sample watermark audio; when the number of rounds of the generator to be trained is large, more processing methods can be used to process the sample watermark audio.

[0274] Step 509: The electronic device trains the generator to be trained based on the enhanced sample watermark audio.

[0275] For example, the aforementioned preset threshold can be 30% of the total number of training rounds. When the number of the first rounds is less than 30% of the total number of training rounds, any form of data augmentation is avoided; that is, no processing method is used to process the sample watermark audio. In other words, the sample watermark audio is directly used to train the generator, allowing it to learn and master the basic features of the sample watermark audio, laying a solid foundation for subsequent complex training. When the number of the first rounds is greater than or equal to 30% and less than 60% of the total number of training rounds, a processing method can be used to process the sample watermark audio, and the enhanced sample watermark audio obtained from the processing can be used to train the generator. This allows the generator to gradually adapt and learn how to maintain the stability and recognizability of the sample watermark information under a single perturbation. If the number of first rounds is greater than or equal to 60% of the total number of training rounds and less than or equal to 100% of the total number of training rounds, at least one processing method can be used to process the enhanced sample watermark audio, and the new enhanced sample watermark audio obtained by the processing can be used to train the generator to enhance the stability of the sample watermark information under complex perturbations.

[0276] In summary, if the sample watermark audio is perturbed using various processing methods at the beginning of training to expand the amount of training data, the generator to be trained may fail to fully learn the basic sample watermark features and converge prematurely. Adopting a progressive enhancement strategy, gradually increasing the diversity of perturbations, is beneficial to improving the stability and generalization ability of the generator to be trained.

[0277] Therefore, during the training process of the generator, the electronic device can further process the sample watermark audio using at least one method, depending on the number of the first training rounds, to introduce at least one type of noise into the sample watermark audio, resulting in enhanced sample watermark audio. Thus, after training the generator using this enhanced sample watermark audio, the trained generator can accurately parse the watermark information even when the watermark audio includes at least one type of noise, thereby improving the accuracy of watermark information parsing. Furthermore, since at least one processing method is positively correlated with the number of the first training rounds of the generator, the stability and generalization ability of the generator can be improved.

[0278] The audio processing method provided in this application can be executed by an audio processing device. This application uses an audio processing device executing the audio processing method as an example to illustrate the audio processing device provided in this application.

[0279] Figure 16 A schematic diagram of the structure of the audio processing apparatus provided in an embodiment of this application is shown. Figure 16As shown, the audio processing device 60 provided in this application embodiment may include: an acquisition module 61 and a processing module 62.

[0280] The acquisition module 61 is used to acquire N corresponding first spectral features and N watermark features. The first spectral features are the spectral features of audio segments in the audio, and the watermark features are the features of the watermark information. N is a positive integer. The processing module 62 is used to input the N first spectral features and N watermark features acquired by the acquisition module 61 into the watermark processing model. The generator of the watermark processing model embeds each watermark feature into the corresponding first spectral feature to obtain N second spectral features. Based on the N second spectral features and the audio, a watermarked audio is generated.

[0281] This application provides an audio processing apparatus. The apparatus can acquire N first spectral features and N watermark features from N audio segments, and embed each watermark feature into the corresponding first spectral feature using a generator of a trained watermark processing model to obtain N second spectral features, instead of directly adding the N watermark information to the N audio segments. Therefore, on the one hand, because the first spectral features and watermark features are not easily interfered with or damaged by signal processing, even if signal processing is performed on the obtained watermarked audio, the watermark features will not be interfered with or damaged. On the other hand, because the watermark processing model is a trained model, after inputting the N first spectral features and N watermark features into the watermark processing model, the generator of the watermark processing model can accurately embed each watermark feature into the corresponding first spectral feature, thereby improving the reliability of the watermarked audio.

[0282] In one possible implementation, the processing module 62 is specifically used to replace the corresponding first spectral feature in the spectral features of the audio with each second spectral feature to obtain the replaced spectral features of the audio; and to generate watermarked audio based on the replaced spectral features.

[0283] In one possible implementation, the acquisition module 61 is further configured to acquire first marker information and watermark content information before acquiring the N corresponding first spectral features and N watermark features, wherein the first marker information is used to locate the watermark information. The processing module 62 is further configured to generate watermark information based on the first marker information and watermark content information acquired by the acquisition module 61.

[0284] In one possible implementation, the processing module 62 is further configured to determine an audio segment from the audio at preset time intervals before the acquisition module 61 acquires the N corresponding first spectral features and N watermark features, and the duration of the audio segment matches the preset duration.

[0285] In one possible implementation, the generator comprises M reversible blocks connected in sequence, where M is a positive integer. The processing module 62 is specifically configured to input N first spectral features and N watermark features into the first reversible block of the M reversible blocks, embedding at least a portion of each watermark feature into the corresponding first spectral feature through the first reversible block; and to input the N first spectral features and N watermark features output from the first reversible block into the i-th reversible block of the M reversible blocks, embedding at least a portion of each watermark feature into the corresponding first spectral feature through the i-th reversible block, where i is a positive integer greater than 1.

[0286] In one possible implementation, the acquisition module 61 is further configured to acquire, before acquiring the N corresponding first spectral features and N watermark features, T corresponding first sample spectral features and T sample watermark features. The first sample spectral features are the spectral features of the sample audio segments in the sample audio, and the sample watermark features are the features of the sample watermark information, where T is a positive integer. The processing module 62 is further configured to input the T first sample spectral features and T sample watermark features acquired by the acquisition module 61 into the watermark processing model, and embed each first sample watermark feature into the corresponding first sample spectral feature through the generator to be trained in the watermark processing model to obtain N second sample spectral features; and generate sample watermark audio based on the N second sample spectral features and the sample audio; and determine a first loss function based on the sample watermark audio and the sample audio, and use the first loss function to train the generator to obtain the generator. The first loss function is used to characterize the audio type difference between the sample watermark audio and the sample audio.

[0287] In one possible implementation, the processing module 62 is further configured to determine a second loss function based on the T first sample spectral features and the T sample watermark features before training the generator to be trained using the first loss function. This second loss function characterizes the difference between the T first sample spectral features and the T sample watermark features. Specifically, the processing module 62 is configured to train the generator to be trained using both the first and second loss functions.

[0288] In one possible implementation, the processing module 62 is specifically used to input the N second sample spectral features into the discriminator of the watermark processing model, determine the audio type of the sample watermark audio through the discriminator, and determine the first loss function based on the audio type of the sample watermark audio and the audio type of the sample audio.

[0289] In one possible implementation, the processing module 62 is further configured to, before inputting the N second sample spectral features and the spectral features of the sample watermark audio into the discriminator of the watermark processing model, and determining the audio type of the sample watermark audio and the audio type of the sample audio through the discriminator, input the N second sample spectral features and the spectral features of the sample watermark audio into the discriminator to be trained in the watermark processing model, and determine the training audio type of the sample watermark audio and the training audio type of the sample audio through the discriminator to be trained; and determine a third loss function based on the training audio type of the sample watermark audio, and determine a fourth loss function based on the training audio type of the sample audio, wherein the third loss function is used to characterize the difference between the training audio type of the sample watermark audio and the audio type of the sample watermark audio, and the fourth loss function is used to characterize the difference between the training audio type of the sample audio and the audio type of the sample audio; and train the discriminator to be trained using the third loss function and the fourth loss function to obtain the discriminator.

[0290] In one possible implementation, the acquisition module 61 is further configured to acquire L third spectral features after the processing module 62 generates the watermarked audio based on N second spectral features and the audio. These third spectral features are the spectral features in the watermarked audio, where L is a positive integer. The processing module 62 is further configured to input the L third spectral features acquired by the acquisition module 61 into the watermarking processing model, determine the watermark feature in each third spectral feature through the generator of the watermarking processing model, and determine the watermark information based on the L watermark features.

[0291] In one possible implementation, the generator comprises M reversible blocks connected in sequence, where M is a positive integer. The processing module 62 is specifically configured to input L corresponding third spectral features and L noise features into the Mth reversible block of the M reversible blocks, embedding at least a portion of the watermark feature of each third spectral feature into the corresponding noise feature through the Mth reversible block, where the noise feature is a random noise feature; and to input the L third spectral features and L noise features output from the Mth reversible block into the Mjth reversible block of the M reversible blocks, embedding at least a portion of the watermark feature of each third spectral feature into the corresponding noise feature through the Mjth reversible block, where j is a positive integer less than M.

[0292] The audio processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.

[0293] The audio processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0294] The audio processing device provided in this application embodiment can achieve... Figures 1 to 15 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0295] In some embodiments of this application, such as Figure 17 As shown, this application embodiment also provides an electronic device 80, including a processor 81 and a memory 82. The memory 82 stores a program or instructions that can run on the processor 81. When the program or instructions are executed by the processor 81, they implement the various process steps of the above-described audio processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0296] It should be noted that the electronic devices in the embodiments of this application include the aforementioned mobile electronic devices and non-mobile electronic devices.

[0297] Figure 18 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0298] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.

[0299] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 18 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0300] The processor 110 is used to acquire N first spectral features and N watermark features that correspond one-to-one. The first spectral features are the spectral features of the audio segments in the audio, and the watermark features are the features of the watermark information, where N is a positive integer. The processor 110 is used to input the N first spectral features and N watermark features into the watermark processing model. The generator of the watermark processing model embeds each watermark feature into the corresponding first spectral feature to obtain N second spectral features. The processor 110 is used to generate watermarked audio based on the N second spectral features and the audio.

[0301] This application provides an electronic device that can acquire N first spectral features and N watermark features of N audio segments and N watermark information from N audio data. Instead of directly adding the N watermark information to the N audio segments, the device embeds each watermark feature into the corresponding first spectral feature using a pre-trained watermark processing model generator to obtain N second spectral features. This improves the reliability of the watermarked audio because, on the one hand, the first spectral features and watermark features are less susceptible to interference or damage during signal processing; even if signal processing is applied to the obtained watermarked audio, the watermark features will not be interfered with or damaged. On the other hand, since the watermark processing model is pre-trained, after inputting the N first spectral features and N watermark features into it, the watermark processing model generator can accurately embed each watermark feature into the corresponding first spectral feature, thereby improving the reliability of the watermarked audio.

[0302] In some embodiments of this application, the processor 110 is specifically used to replace the corresponding first spectral feature in the spectral features of the audio with each second spectral feature to obtain the replaced spectral features of the audio; and to generate watermarked audio based on the replaced spectral features.

[0303] In some embodiments of this application, the processor 110 is further configured to acquire first marker information and watermark content information before acquiring the N corresponding first spectral features and N watermark features, wherein the first marker information is used to locate the watermark information; and generate watermark information based on the first marker information and watermark content information.

[0304] In some embodiments of this application, the processor 110 is further configured to determine an audio segment from the audio at preset time intervals before acquiring the N first spectral features and N watermark features that correspond one-to-one, wherein the duration of an audio segment matches the preset duration.

[0305] In some embodiments of this application, the generator includes M reversible blocks connected in sequence, where M is a positive integer.

[0306] The processor 110 is specifically configured to input N first spectral features and N watermark features into the first reversible block of M reversible blocks, and embed at least a portion of the features of each watermark feature into the corresponding first spectral feature through the first reversible block; and input the N first spectral features and N watermark features output from the first reversible block into the i-th reversible block of M reversible blocks, and embed at least a portion of the features of each watermark feature into the corresponding first spectral feature through the i-th reversible block, where i is a positive integer greater than 1.

[0307] In some embodiments of this application, the processor 110 is further configured to acquire, before acquiring the N corresponding first spectral features and N watermark features, T corresponding first sample spectral features and T sample watermark features, wherein the first sample spectral features are the spectral features of sample audio segments in the sample audio, and the sample watermark features are the features of the sample watermark information, and T is a positive integer; and input the T first sample spectral features and T sample watermark features into the watermark processing model, and embed each sample watermark feature into the corresponding first sample spectral feature through the generator to be trained of the watermark processing model to obtain N second sample spectral features; and generate sample watermark audio based on the N second sample spectral features and sample audio; and determine a first loss function according to the sample watermark audio and sample audio, and use the first loss function to train the generator to be trained to obtain the generator, wherein the first loss function is used to characterize the audio type difference between the sample watermark audio and the sample audio.

[0308] In some embodiments of this application, the processor 110 is further configured to determine a second loss function based on T first sample spectral features and T sample watermark features before training the generator to be trained using the first loss function. The second loss function is used to characterize the difference between the T first sample spectral features and the T sample watermark features.

[0309] The processor 110 is specifically used to train the generator to be trained using a first loss function and a second loss function.

[0310] In some embodiments of this application, the processor 110 is specifically used to input N second sample spectral features into the discriminator of the watermarking processing model, determine the audio type of the sample watermark audio through the discriminator, and determine a first loss function based on the audio type of the sample watermark audio and the audio type of the sample audio.

[0311] In some embodiments of this application, the processor 110 is further configured to input the N second sample spectral features and the spectral features of the sample watermark audio into the discriminator of the watermark processing model before determining the audio type of the sample watermark audio and the audio type of the sample audio by the discriminator; to determine the training audio type of the sample watermark audio and the training audio type of the sample audio by the discriminator; and to determine a third loss function and a fourth loss function based on the training audio type of the sample watermark audio, wherein the third loss function is used to characterize the difference between the training audio type of the sample watermark audio and the audio type of the sample watermark audio, and the fourth loss function is used to characterize the difference between the training audio type of the sample audio and the audio type of the sample audio; and to train the discriminator to obtain the discriminator using the third loss function and the fourth loss function.

[0312] In some embodiments of this application, the processor 110 is further configured to, after generating watermarked audio based on N second spectral features and audio, obtain L third spectral features, wherein the third spectral features are spectral features in the watermarked audio, and L is a positive integer; input the L third spectral features into the watermarking processing model, determine the watermark feature in each third spectral feature through the generator of the watermarking processing model; and determine watermark information based on the L watermark features.

[0313] In some embodiments of this application, the generator described above includes M reversible blocks connected in sequence, where M is a positive integer.

[0314] The processor 110 is specifically configured to input the L corresponding third spectral features and L noise features into the Mth reversible block of the M reversible blocks, and embed at least a portion of the watermark feature in each third spectral feature into the corresponding noise feature through the Mth reversible block, wherein the noise feature is a feature of random noise; and input the L third spectral features and L noise features output from the Mth reversible block into the Mjth reversible block of the M reversible blocks, and embed at least a portion of the watermark feature in each third spectral feature into the corresponding noise feature through the Mjth reversible block, wherein j is a positive integer less than M.

[0315] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0316] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0317] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.

[0318] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described audio processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0319] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0320] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described audio processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0321] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0322] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the audio processing method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0323] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0324] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0325] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An audio processing method, characterized in that, include: Obtain N first spectral features and N watermark features that correspond one-to-one. The first spectral features are the spectral features of audio segments in the audio, and the watermark features are the features of watermark information. N is a positive integer. The N first spectral features and N watermark features are input into the watermark processing model. The generator of the watermark processing model embeds each watermark feature into the corresponding first spectral feature to obtain N second spectral features. Based on N second spectral features and the audio, a watermarked audio is generated; The generator comprises M reversible blocks connected in sequence, where M is a positive integer; The step of inputting N first spectral features and N watermark features into the watermark processing model, and embedding each watermark feature into the corresponding first spectral feature through the generator of the watermark processing model, includes: The N first spectral features and the N watermark features are input into the first reversible block of the M reversible blocks, and at least a portion of the features in each watermark feature are embedded into the corresponding first spectral feature through the first reversible block; The N first spectral features and N watermark features output from the first reversible block are input into the i-th reversible block among the M reversible blocks. At least a portion of the features in each watermark feature are embedded into the corresponding first spectral feature through the i-th reversible block, where i is a positive integer greater than 1.

2. The method according to claim 1, characterized in that, The step of generating watermarked audio based on N second spectral features and the audio includes: Each second spectral feature is used to replace the corresponding first spectral feature in the spectral features of the audio, thereby obtaining the replaced spectral features of the audio; The watermarked audio is generated based on the replaced spectral features.

3. The method according to claim 1, characterized in that, Before obtaining the N corresponding first spectral features and N watermark features, the method further includes: From the audio, an audio segment is determined at preset time intervals, and the duration of the audio segment matches the preset duration.

4. The method according to any one of claims 1 to 3, characterized in that, Before obtaining the N corresponding first spectral features and N watermark features, the method further includes: T first sample spectral features and T sample watermark features are obtained in a one-to-one correspondence. The first sample spectral features are the spectral features of the sample audio segments in the sample audio, and the sample watermark features are the features of the sample watermark information. T is a positive integer. T first sample spectral features and T sample watermark features are input into the watermark processing model. The watermark processing model's generator is used to embed each sample watermark feature into the corresponding sample spectral feature to obtain N second sample spectral features. Based on N second sample spectral features and the sample audio, a sample watermark audio is generated; A first loss function is determined based on the sample watermarked audio and the sample audio, and the generator to be trained is trained using the first loss function to obtain the generator. The first loss function is used to characterize the audio type difference between the sample watermarked audio and the sample audio.

5. The method according to claim 4, characterized in that, Before training the generator to be trained using the first loss function, the method further includes: A second loss function is determined based on T first sample spectral features and T sample watermark features. The second loss function is used to characterize the difference between T first sample spectral features and T sample watermark features. The step of training the generator to be trained using the first loss function includes: The generator to be trained is trained using the first loss function and the second loss function.

6. The method according to claim 4, characterized in that, The step of determining the first loss function based on the sample watermark audio includes: The N second sample spectral features are input into the discriminator of the watermarking model, and the audio type of the sample watermark audio is determined by the discriminator. The first loss function is determined based on the audio type of the sample watermarked audio and the audio type of the sample audio.

7. The method according to claim 6, characterized in that, Before inputting the N second sample spectral features into the discriminator of the watermarking model, and determining the audio type of the sample watermark audio and the audio type of the sample audio by the discriminator, the method further includes: The N second sample spectral features and the spectral features of the sample watermark audio are input into the discriminator to be trained in the watermark processing model. The training audio type of the sample watermark audio and the training audio type of the sample audio are determined by the discriminator to be trained. A third loss function is determined based on the training audio type of the sample watermark audio, and a fourth loss function is determined based on the training audio type of the sample audio. The third loss function is used to characterize the difference between the training audio type of the sample watermark audio and the audio type of the sample watermark audio, and the fourth loss function is used to characterize the difference between the training audio type of the sample audio and the audio type of the sample audio. The discriminator to be trained is obtained by using the third loss function and the fourth loss function.

8. The method according to claim 1, characterized in that, After generating the watermarked audio based on N second spectral features and the audio, the method further includes: Obtain L third spectral features, where the third spectral features are the spectral features in the watermarked audio, and L is a positive integer; The L third spectral features are input into the watermarking model, and the watermark feature in each third spectral feature is determined by the generator of the watermarking model. The watermark information is determined based on L of the watermark features.

9. The method according to claim 8, characterized in that, The generator comprises M reversible blocks connected in sequence, where M is a positive integer; The step of inputting L of the third spectral features into the watermarking processing model, and determining the watermark feature in each of the third spectral features through the generator of the watermarking processing model, includes: The L corresponding third spectral features and L noise features are input into the Mth reversible block of the M reversible blocks. The Mth reversible block embeds at least a portion of the watermark features in each third spectral feature into the corresponding noise feature. The noise feature is a feature of random noise. The L third spectral features and L noise features output from the Mth reversible block are input into the Mjth reversible block among the M reversible blocks. At least a portion of the watermark features in each third spectral feature are embedded into the corresponding noise feature through the Mjth reversible block, where j is a positive integer less than M.

10. An audio processing apparatus, characterized in that, The audio processing device includes: an acquisition module and a processing module; The acquisition module is used to acquire N first spectral features and N watermark features that correspond one-to-one. The first spectral features are the spectral features of audio segments in the audio, and the watermark features are the features of watermark information. N is a positive integer. The processing module is used to input the N first spectral features and N watermark features acquired by the acquisition module into the watermark processing model, and to embed each watermark feature into the corresponding first spectral feature through the generator of the watermark processing model to obtain N second spectral features; and to generate watermarked audio based on the N second spectral features and the audio. The generator comprises M reversible blocks connected in sequence, where M is a positive integer; The processing module is specifically configured to input N first spectral features and N watermark features into the first reversible block of M reversible blocks, and embed at least a portion of each watermark feature into the corresponding first spectral feature through the first reversible block; and input the N first spectral features and N watermark features output from the first reversible block into the i-th reversible block of M reversible blocks, and embed at least a portion of each watermark feature into the corresponding first spectral feature through the i-th reversible block, where i is a positive integer greater than 1.

11. The apparatus according to claim 10, characterized in that, The processing module is specifically used to replace the corresponding first spectral feature in the spectral features of the audio with each of the second spectral features to obtain the replaced spectral features of the audio; and to generate the watermarked audio based on the replaced spectral features.

12. The apparatus according to claim 10, characterized in that, The processing module is further configured to, before the acquisition module acquires the N first spectral features and N watermark features that correspond one-to-one, determine an audio segment from the audio at preset time intervals, wherein the duration of an audio segment matches the preset duration.

13. The apparatus according to any one of claims 10 to 12, characterized in that, The acquisition module is further configured to acquire, before acquiring the N first spectral features and N watermark features that correspond one-to-one, T first sample spectral features and T sample watermark features, wherein the first sample spectral features are the spectral features of the sample audio segments in the sample audio, the sample watermark features are the features of the sample watermark information, and T is a positive integer. The processing module is further configured to input the T first sample spectral features and T sample watermark features acquired by the acquisition module into the watermark processing model, and embed each sample watermark feature into the corresponding sample spectral feature through the generator to be trained of the watermark processing model to obtain N second sample spectral features; and generate sample watermark audio based on the N second sample spectral features and the sample audio; and determine a first loss function according to the sample watermark audio and the sample audio, and use the first loss function to train the generator to be trained to obtain the generator, wherein the first loss function is used to characterize the audio type difference between the sample watermark audio and the sample audio.

14. The apparatus according to claim 13, characterized in that, The processing module is further configured to determine a second loss function based on T first sample spectral features and T sample watermark features before training the generator to be trained using the first loss function. The second loss function is used to characterize the difference between the T first sample spectral features and the T sample watermark features. The processing module is specifically used to train the generator to be trained using the first loss function and the second loss function.

15. The apparatus according to claim 13, characterized in that, The processing module is specifically used to input the N second sample spectral features into the discriminator of the watermark processing model, determine the audio type of the sample watermark audio through the discriminator, and determine the first loss function based on the audio type of the sample watermark audio and the audio type of the sample audio.

16. The apparatus according to claim 15, characterized in that, The processing module is further configured to, before inputting the N second sample spectral features and the spectral features of the sample watermark audio into the discriminator of the watermark processing model and determining the audio type of the sample watermark audio and the audio type of the sample audio through the discriminator, input the N second sample spectral features and the spectral features of the sample watermark audio into the discriminator to be trained of the watermark processing model and determine the training audio type of the sample watermark audio and the training audio type of the sample audio through the discriminator to be trained; A third loss function is determined based on the training audio type of the sample watermark audio, and a fourth loss function is determined based on the training audio type of the sample audio. The third loss function is used to characterize the difference between the training audio type of the sample watermark audio and the audio type of the sample watermark audio, and the fourth loss function is used to characterize the difference between the training audio type of the sample audio and the audio type of the sample audio. The discriminator to be trained is then trained using the third loss function and the fourth loss function to obtain the discriminator.

17. The apparatus according to claim 10, characterized in that, The acquisition module is further configured to acquire L third spectral features after the processing module generates watermarked audio based on N second spectral features and the audio, wherein the third spectral features are spectral features in the watermarked audio and L is a positive integer; The processing module is used to input the L third spectral features acquired by the acquisition module into the watermark processing model, determine the watermark feature in each of the third spectral features through the generator of the watermark processing model, and determine the watermark information based on the L watermark features.

18. The apparatus according to claim 17, characterized in that, The generator comprises M reversible blocks connected in sequence, where M is a positive integer; The processing module is specifically used to input the L corresponding third spectral features and L noise features into the Mth reversible block of the M reversible blocks, and to embed at least a portion of the watermark features in each third spectral feature into the corresponding noise feature through the Mth reversible block, wherein the noise feature is a feature of random noise; The L third spectral features and L noise features output from the Mth reversible block are input into the Mjth reversible block among the M reversible blocks. The Mjth reversible block embeds at least a portion of the watermark features in each third spectral feature into the corresponding noise feature. j is a positive integer less than M.

19. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the audio processing method as described in any one of claims 1 to 9.

20. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the audio processing method as described in any one of claims 1 to 9.