Audio Processing Method, Apparatus, Storage Medium, and Electronic Device
By performing frame processing and offset position search on the audio, an offset audio frame without auditory differences is generated, which solves the problem of audio and video watermark being destroyed under the conspiracy attack, and realizes the traceability and copyright protection of audio and video pirated copies.
Patent Information
- Application Number
- CN202111296208.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-11-03
AI Technical Summary
When the prior art conspiracy attacks against multiple users, the audio and video watermarks are destroyed, making it difficult to trace the source of pirated copies and difficult to protect audio and video copyrights.
By performing frame-based processing on the audio, performing offset position searches, generating offset audio frames, and sample offsets are performed based on user identification information and candidate offset values to generate audio frames without auditory differences.
Even under the joint attack of multiple users, the generated audio and the original audio have a huge difference in hearing, and the joint attack has failed, realizing the traceability and copyright protection of audio and video piracy.
Smart Images

Figure CN114242111B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing, and in particular, to an audio processing method, apparatus, storage medium, and electronic device. Background Art
[0002] In recent years, digital watermarking technology has achieved certain results in the field of audio and video copyright protection. Digital watermarks include audio watermarks and video watermarks, etc. Among them, audio watermarks are widely used in the scenarios of audio and video copyright protection and piracy tracing because their complexity and cost are lower than those of video watermarks. Each audio obtained by a user is a unique version containing an audio watermark corresponding to the user information. In this way, even if the audio containing the audio watermark has been attacked, the corresponding user information can still be extracted from it to achieve piracy tracing. However, for audio with the same content, the respective versions of multiple users are strictly sample-aligned, and the difference in sample values at the same position of each version of the audio is very small. When multiple users perform weighted superposition processing (i.e., collusion attack) on their respective versions of the audio with the same content, the resulting superposed audio has no perceptible difference from the original audio in terms of hearing, but the audio watermark is damaged during the superposition process, resulting in the inability to extract any user information from the superposed audio, that is, piracy tracing cannot be achieved.
[0003] Currently, many solutions are based on watermark information coding to combat collusion attacks, but when the number of colluders is greater than 2, the number of users supported by such solutions is very limited.
[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present invention provide an audio processing method, apparatus, storage medium, and electronic device to at least solve the technical problems of difficult tracing of audio and video piracy and difficult audio and video copyright protection caused by audio and video pirates using collusion attack methods to damage audio and video watermarks.
[0006] According to one aspect of the embodiments of the present invention, an audio processing method is provided, including: obtaining an audio to be processed; performing frame division processing on the audio to be processed to obtain a plurality of audio frames; searching for offset positions of at least one of the plurality of audio frames to obtain a first sample pair and a second sample pair; and performing sample offset on the first sample pair and the second sample pair based on user identification information and a plurality of candidate offset values to obtain an offset audio.
[0007] According to another aspect of the embodiments of the present invention, there is also provided an audio processing method, including: receiving the audio to be processed from a client; performing frame division processing on the audio to be processed to obtain a plurality of audio frames, performing offset position search on at least one audio frame of the plurality of audio frames to obtain a first sample pair and a second sample pair, and performing sample offset on the first sample pair and the second sample pair based on user identification information and a plurality of candidate offset values to obtain the offset audio; and feeding back the offset audio to the client.
[0008] According to another aspect of the embodiments of the present invention, there is also provided an audio processing method, including: loading the audio to be processed in an audio editing interface; in response to a first editing instruction for the audio to be processed, performing frame division processing on the audio to be processed to obtain a plurality of audio frames; in response to a second editing instruction for the plurality of audio frames, performing offset position search on at least one audio frame of the plurality of audio frames to obtain a first sample pair and a second sample pair; in response to a third editing instruction for the plurality of audio frames, determining user identification information and a plurality of candidate offset values, and performing sample offset on the first sample pair and the second sample pair based on the user identification information and the plurality of candidate offset values to obtain the offset audio; and displaying the offset audio in the audio editing interface.
[0009] According to another aspect of the embodiments of the present invention, there is also provided an audio processing apparatus, including: an acquisition module, configured to acquire the audio to be processed; a frame division module, configured to perform frame division processing on the audio to be processed to obtain a plurality of audio frames; a search module, configured to perform offset position search on at least one audio frame of the plurality of audio frames to obtain a first sample pair and a second sample pair; and a processing module, configured to perform sample offset on the first sample pair and the second sample pair based on user identification information and a plurality of candidate offset values to obtain the offset audio.
[0010] According to another aspect of the embodiments of the present invention, there is also provided a storage medium, where the storage medium includes a stored program, and when the program runs, it controls the device where the storage medium is located to execute any one of the above audio processing methods.
[0011] According to another aspect of the embodiments of the present invention, there is also provided a processor, where the processor is used to run a program, and when the program runs, it executes any one of the above audio processing methods.
[0012] According to another aspect of the embodiments of the present invention, an electronic device is further provided, including: a processor; and a memory connected to the above-mentioned processor for providing instructions for the above-mentioned processor to process the following processing steps: obtaining the audio to be processed; performing frame splitting on the above-mentioned audio to be processed to obtain a plurality of audio frames; searching for offset positions for at least one of the above-mentioned plurality of audio frames to obtain a first sample pair and a second sample pair; based on user identification information and a plurality of candidate offset values, performing sample offset on the above-mentioned first sample pair and the above-mentioned second sample pair to obtain the offset audio.
[0013] In the embodiments of the present invention, the audio is processed by frame offset, by obtaining the audio to be processed; performing frame splitting on the above-mentioned audio to be processed to obtain a plurality of audio frames; searching for offset positions for at least one of the above-mentioned plurality of audio frames to obtain a first sample pair and a second sample pair; based on user identification information and a plurality of candidate offset values, performing sample offset on the above-mentioned first sample pair and the above-mentioned second sample pair to obtain the offset audio.
[0014] It is easy to notice that, through the embodiments of the present application, few-sample offsets are performed on each audio frame to generate audio with no audible difference. Even in the case where the number of users is relatively large and multiple users conduct a collusion attack, the samples obtained from the collusion attack have a huge audible difference from the original audio and cannot meet the normal usage requirements, that is, the collusion attack fails.
[0015] Thus, the embodiments of the present application achieve the purpose of generating audio with few-sample offsets within each audio frame and having no audible difference from the original audio, thereby realizing the technical effect of anti-collusion attack based on frame offset processing of audio, and further solving the technical problems of difficult traceability of audio-visual piracy and difficult audio-visual copyright protection caused by audio-visual pirates using collusion attack methods to destroy audio-visual watermarks. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0017] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing an audio processing method according to the prior art;
[0018] Figure 2 is a flowchart of an audio processing method according to an embodiment of the present invention;
[0019] Figure 3 is a schematic diagram of an optional audio processing process according to an embodiment of the present invention;
[0020] Figure 4It is a schematic diagram of an optional correspondence between an offset sequence and an audio frame according to an embodiment of the present invention;
[0021] Figure 5a It is a schematic diagram of an optional audio frame offset processing pre-sample according to an embodiment of the present invention;
[0022] Figure 5b It is a schematic diagram of an optional audio frame offset processing post-sample according to an embodiment of the present invention;
[0023] Figure 6 It is a flowchart of an optional audio processing method according to an embodiment of the present invention;
[0024] Figure 7 It is a schematic diagram of an optional audio processing in a cloud server according to an embodiment of the present invention;
[0025] Figure 8 It is a flowchart of another optional audio processing method according to an embodiment of the present invention;
[0026] Figure 9 It is a schematic diagram of the structure of an audio processing device according to an embodiment of the present invention;
[0027] Figure 10 It is a structural block diagram of another computer terminal according to an embodiment of the present invention. Detailed implementation manners
[0028] In order to enable those skilled in the art of the present technology to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0030] First, some nouns or terms that appear in the process of describing the embodiments of the present application are applicable to the following explanations:
[0031] Collusion attack: In the scenario of pirated source tracing, multiple users embed watermarks with different information in the same content to obtain multiple versions of the same content corresponding to the users. When two or more users perform weighted superposition on the above multiple versions and obtain a pirated version of the content, the tracing watermark extraction of the pirated version will fail. This method of generating the pirated version is called a collusion attack.
[0032] Frame offset: It refers to the operation of shifting the samples within a specific interval of an audio frame as a whole to the left or right, so that each audio frame is misaligned with the original audio frame.
[0033] Embodiment 1
[0034] According to an embodiment of the present invention, an embodiment of an audio processing method is further provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0035] The method embodiment provided by the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing the audio processing method is shown. As Figure 1 shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) shown by 102a, 102b,..., 102n in the figure, a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0036] It should be noted that one or more of the above-mentioned processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).
[0037] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the audio processing method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned audio processing method. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0038] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0039] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computer terminal 10 (or mobile device).
[0040] It should be noted here that in some alternative embodiments, the above-mentioned Figure 1 shown computer device (or mobile device) can include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1This is just an example of a specific concrete instance and is intended to illustrate the types of components that may exist in the above computer device (or mobile device).
[0041] Under the above operating environment, the present application provides an audio processing method as Figure 2 shown. Figure 2 is a flowchart of an audio processing method according to an embodiment of the present invention. As Figure 2 shown, the audio classification method includes:
[0042] Step S102, obtaining the audio to be processed;
[0043] Step S104, performing frame splitting on the above-mentioned audio to be processed to obtain a plurality of audio frames;
[0044] Step S106, searching for offset positions for at least one of the above-mentioned plurality of audio frames to obtain a first sample pair and a second sample pair;
[0045] Step S108, based on the user identification information and a plurality of candidate offset values, performing sample offset on the above-mentioned first sample pair and the above-mentioned second sample pair to obtain an offset audio.
[0046] It is easy to notice that through the embodiments of the present application, by performing few-shot offset on each audio frame to generate an audio with no auditory difference, even in the case where the number of users is relatively large and multiple users conduct a collusion attack, the samples obtained from the collusion attack have a huge auditory difference from the original audio and cannot meet the normal usage requirements, that is, the collusion attack fails.
[0047] Thus, the embodiments of the present application achieve the purpose of generating an audio with few-shot offset within each audio frame and having no auditory difference from the original audio, thereby realizing the technical effect of processing audio based on frame offset to resist collusion attacks, and further solving the technical problems of difficult traceability of pirated audio and video and difficult protection of audio and video copyrights caused by audio and video pirates applying collusion attack methods to destroy audio and video watermarks.
[0048] Optionally, the above-mentioned audio processing method provided by the present application can be but is not limited to being applied to the protection of online audio and video products (such as audio and video live broadcasts, audio and video players, etc.), the identification of pirated audio and video (such as the identification of original short videos, the traceability of pirated audio and video, etc.). By adopting the audio processing method in the embodiments of the present application, it is possible to automatically generate corresponding audio and video with few-shot offset versions that have no auditory difference from the original audio and video and can resist collusion attacks.
[0049] Optionally, the above-mentioned audio to be processed can be an original audio file, or can be an audio file corresponding to the original video file extracted from the original video file. The above-mentioned audio to be processed can contain a plurality of audio frames.
[0050] Optionally, the above-mentioned frame processing may be to identify the audio to be processed, divide the audio to be processed into multiple audio frames according to actual usage needs, and number the multiple audio frames in sequence. Among them, each audio frame contains multiple samples, and the multiple samples are also numbered in sequence. The above-mentioned audio frame numbers are used to determine the position of the current audio frame in the audio, and the above-mentioned sample numbers are used to determine the position of the current sample in the current audio frame. The above-mentioned frame processing facilitates subsequent analysis and processing operations on the audio to be processed.
[0051] Optionally, the above-mentioned offset position may be a sample position obtained by searching for the offset position of at least one of the above-mentioned multiple audio frames for subsequent offset operations. For at least one of the above-mentioned multiple audio frames, each audio frame contains two sample offset positions, that is, corresponding to the above-mentioned first sample pair and the above-mentioned second sample pair.
[0052] Optionally, the above-mentioned user identification information may be data containing a user identity document (ID), where the user's ID is unique, that is, the same identity ID will be regarded as the same user.
[0053] Optionally, the above-mentioned candidate offset values may be a set of offset values determined according to the audio to be processed and actual usage conditions.
[0054] In an alternative embodiment, in step S104, the above-mentioned audio to be processed is frame-processed to obtain multiple audio frames, including the following method steps:
[0055] Step S141, frame the above-mentioned audio to be processed according to a fixed number of samples to obtain the above-mentioned multiple audio frames, where each of the above-mentioned multiple audio frames includes: multiple samples.
[0056] Optionally, the above-mentioned fixed number of samples may be a value determined according to actual needs. Before framing the audio to be processed, it is determined that each audio frame should contain the above-mentioned fixed number of samples, and accordingly, the audio to be processed is divided into multiple audio frames.
[0057] Figure 3 is a schematic diagram of an alternative audio processing process according to an embodiment of the present invention; as Figure 3 shown, the audio to be processed obtained by the above-mentioned terminal is the paid audio A for frame operation with a fixed number of samples. First, it is determined that a single audio frame contains L samples. Accordingly, the paid audio A is divided into N such audio frames.
[0058] In an alternative embodiment, in step S106, the offset position of at least one of the above-mentioned multiple audio frames is searched to obtain a first sample pair, including the following method steps:
[0059] Step S161: Start searching from the first sample among the multiple samples above, and search backward to find two adjacent samples with reversed sample symbols, obtaining the first sample pair above; or, start searching from the first sample among the multiple samples above, search for consecutive silent samples, and select two silent samples from the consecutive silent samples to obtain the first sample pair above.
[0060] Optionally, the above offset position search may start searching from the first sample among the multiple samples, and when two adjacent samples with reversed symbols are found, these two adjacent samples are used as the first sample pair above. Among them, the multiple samples are the multiple samples of each audio frame in the audio frame for offset position search; the audio frame for offset position search is at least one audio frame among the multiple audio frames.
[0061] Optionally, the above offset position search may also start searching from the first sample among the multiple samples, and when consecutive silent samples are found, two of the silent samples are used as the first sample pair above. Among them, the multiple samples are the multiple samples of each audio frame in the audio frame for offset position search; the audio frame for offset position search is at least one audio frame among the multiple audio frames.
[0062] Still as Figure 3 shown, during the processing of the paid audio A, offset position search is performed on each of the N audio frames in the framed paid audio A. Among them, the operation when performing offset position search on the nth (n < N) audio frame is as follows: The nth audio frame contains L samples, numbered n1 to nL in sequence. Starting from the first sample, when it is found that the symbols of two samples n4 and n5 are reversed, at this time, the first sample pair of the nth audio frame is determined to be n4 and n5, and the offset position is P f .
[0063] Still as Figure 3 shown, during the processing of the paid audio A, offset position search is performed on each of the N audio frames in the framed paid audio A. Among them, the operation when performing offset position search on the mth (m < N) audio frame is as follows: The mth audio frame contains L samples, numbered m1 to mL in sequence. Starting from the first sample, when it is found that two consecutive samples m2 and m3 are both silent samples, at this time, the first sample pair of the mth audio frame is determined to be m2 and m3, and the offset position is P f .
[0064] In an alternative embodiment, in step S106, offset position search is performed on at least one audio frame among the multiple audio frames to obtain a second sample pair, including the following method steps:
[0065] Step S162: Search forward from the last sample among the multiple samples above to find two adjacent samples with reversed sample symbols, obtaining the second sample pair above; or, search forward from the last sample among the multiple samples above to find consecutive silent samples, and select two silent samples from the consecutive silent samples above to obtain the second sample pair above.
[0066] Optionally, the above offset position search may be to search forward from the last sample among the multiple samples in each audio frame. When two adjacent samples with reversed symbols are found, these two adjacent samples are used as the second sample pair above.
[0067] Optionally, the above offset position search may also be to search backward from the last sample among the multiple samples in each audio frame. When consecutive silent samples are found, two of the silent samples are used as the second sample pair above.
[0068] Still as Figure 3 shown, during the processing of the paid audio A, an offset position search is performed on each of the N audio frames in the framed paid audio A. Among them, the operation when performing the offset position search on the nth (n < N) audio frame is as follows: The nth audio frame contains L samples, numbered sequentially as n1 to nL. Searching from the last sample, when it is found that the symbols of the two samples n(L - 8) and n(L - 7) are reversed, at this time, the second sample pair of the nth audio frame is determined to be n(L - 8) and n(L - 7), and the offset position is P t .
[0069] Still as Figure 3 shown, during the processing of the paid audio A, an offset position search is performed on each of the N audio frames in the framed paid audio A. Among them, the operation when performing the offset position search on the mth (m < N) audio frame is as follows: The mth audio frame contains L samples, numbered sequentially as m1 to mL. Searching from the last sample, when it is found that the two consecutive samples m(L - 6) and m(L - 7) are both silent samples, at this time, the second sample pair of the mth audio frame is determined to be m(L - 6) and m(L - 7), and the offset position is P t .
[0070] In an alternative embodiment, in step S108, based on the above user identification information and the multiple candidate offset values, sample offset is performed on the first sample pair and the second sample pair above to obtain the offset audio above, including the following method steps:
[0071] Step S181: Generate an offset sequence based on the above user identification information and the multiple candidate offset values;
[0072] Step S182: Use the above offset sequence to perform sample offset on the above first sample pair and the above second sample pair to obtain offset boundary samples;
[0073] Step S183: Perform interpolation and smoothing processing on the above offset boundary samples to obtain the above offset audio.
[0074] Optionally, after obtaining the above user identification information and the above multiple candidate offset values, an offset sequence for offset operation can be obtained. Use this offset sequence to perform sample offset on the above first sample pair and the above second sample pair to obtain offset boundary samples. For the practicality of the finally offset audio, interpolation and smoothing processing also need to be performed on the above offset boundary samples, and the processed audio is used as the offset audio.
[0075] Optionally, the number of elements in the above offset sequence is the same as the number of audio frames included in the audio to be processed. One element in the offset sequence represents the specific offset amount for offsetting the samples at the corresponding positions in the corresponding audio frames of the audio to be processed. Among them, the corresponding positions in the audio frames are the positions determined by the above first sample pair and the second sample pair.
[0076] In an alternative embodiment, in step S181, based on the above user identification information and the above multiple candidate offset values, generating the above offset sequence includes the following method steps:
[0077] Step S1811: Determine the above multiple candidate offset values based on a first quantity and a preset offset amplitude, where the above first quantity is the fixed number of samples included in each of the above multiple audio frames;
[0078] Step S1812: Generate a random sequence using the above user identification information, where the length of the above random sequence is a second quantity, and the above second quantity is the number of the above multiple audio frames;
[0079] Step S1813: Generate the above offset sequence through the mapping relationship between the above random sequence and the above multiple candidate offset values.
[0080] Optionally, before performing sample offset, it is necessary to pre-determine the above multiple candidate offset values. The above multiple candidate offset values can be determined as follows: Obtain the fixed number of samples included in each audio frame determined during audio frame segmentation operation, stipulate a preset offset amplitude according to the actual usage situation, and determine the above multiple candidate offset values based on the fixed number of samples and the preset offset amplitude.
[0081] Optionally, the offset sequence for the offset operation can be obtained by mapping a random sequence through the above-mentioned multiple candidate offset values. The random sequence is generated from the above user identification information, which ensures that the random sequence corresponding to each user is also unique, and this has practical significance for audio copyright protection. Additionally, the length of the random sequence is the number of audio frames of the audio to be processed. Furthermore, the above offset sequence can be generated from the mapping relationship between the random sequence and the above offset values, and the length of the offset sequence is also the same as the number of audio frames.
[0082] Still as Figure 3 shown, in the process of processing the paid audio A, the above-mentioned multiple candidate offset values can be determined as follows: For the paid audio A, if the sample offset amplitude exceeds 1%, an audible difference perceptible to the user will be generated. Therefore, a set of offsettable values is generated with the constraint that the offset amplitude does not exceed 1%, denoted as {D0, D1, …, D M-1}, in this set, any element Di satisfies -L×1% < D i < L×1%, where M is the number of different offset values.
[0083] Figure 4 is a schematic diagram of an optional correspondence between the offset sequence and audio frames according to an embodiment of the present invention; as Figure 4 shown, the offset values in the offset sequence correspond one-to-one with the audio frames to be processed.
[0084] Still as Figure 3 shown, in the process of processing the paid audio A, the above offset sequence can be generated as follows: Obtain the user ID, and generate a random sequence {R0, R1, …, R N-1} of length N based on this user ID; determine a numerical mapping function f(i); generate an offset sequence {S0, S1, …, S N-1} from the random sequence {R0, R1, …, R M-1} and the set of offsettable values {D0, D1, …, D N-1}, where S i = D f(i) .
[0085] Use the offset sequence {S0, S1, …, S N-1} to process each audio frame in the paid audio A. The processing steps for the nth audio frame are as follows: Perform sample offset on the samples at positions between P f and P t in this audio frame, with the offset value being S n-1 . If S n-1 > 0, the offset direction is to the right. If S n-1If <0, the offset direction is to the left; after processing all audio frames in the paid audio A, the offset boundary samples are obtained and denoted as A'; performing interpolation smoothing on A' can obtain the audio A# with few-sample offset corresponding to the final paid audio A.
[0086] Figure 5a It is a schematic diagram of an optional sample before audio frame offset processing according to an embodiment of the present invention. Figure 5b It is a schematic diagram of an optional sample after audio frame offset processing according to an embodiment of the present invention; Figure 5a Points E, F, G, and H in Figure 5b correspond to points E1, F1, G1, and H1 in Figure 5a and Figure 5b respectively. According to an embodiment of the present invention, two samples near a zero value are deleted from the samples of this audio frame, and two samples near another zero value are inserted, causing the audio frame to change as shown in
[0087] In an optional embodiment, the above audio processing method further includes the following method steps:
[0088] Step S202, searching for target audio frames from the above multiple audio frames;
[0089] Step S204, calculating the correlation of the above multiple target samples to determine the sample alignment position, where the above multiple target samples are multiple consecutive samples selected from the above offset audio, and the correlation between the above target audio frame and the above multiple target samples meets a preset condition;
[0090] Step S206, starting from the above sample alignment position, performing frame splitting on the above offset audio to obtain multiple offset audio frames;
[0091] Step S208, continuously extracting a binary random sequence from the third number of offset audio frames through the corresponding relationship between the above multiple audio frames and the above multiple offset audio frames, where the above third number is the maximum length required for the binary representation of the above user identification information;
[0092] Step S210, reconstructing the above user identification information using the above binary random sequence.
[0093] Optionally, in the audio watermark traceability scenario with a reference source, the above audio processing method can also extract the audio watermark through the following method steps:
[0094] For example, there is an offset audio A# of a paid audio A. The audio watermark in A# can be extracted by the method in the embodiments of the present invention, and then the user ID corresponding to A# can be obtained to realize audio watermark traceability. The specific steps are as follows:
[0095] 1) Obtain the first audio frame (the number of samples included is L) in the paid audio A as the target audio frame;
[0096] 2) Search for multiple target samples in A#. The search method is: calculate the correlation value between this target audio frame and all continuous sample intervals of length L in A#. A preset condition is specified: the maximum value among the calculation results of the foregoing correlation values and greater than T. When the correlation value calculated for a certain continuous sample interval of length L in A# meets the above preset condition, determine this continuous sample interval of length L as the multiple target samples, and determine the position P of the multiple target samples in A# as the sample alignment position;
[0097] 3) Starting from this sample alignment position, perform frame splitting on A#, where the number of samples included in each audio frame is L, to obtain multiple audio frames of A#;
[0098] 4) Corresponding the multiple audio frames of A with the multiple audio frames of A#, and continuously extract the binary random sequence for B frames of audio. Taking the i-th audio frame as an example, the extraction method of this binary random number is: right-shift x samples in the i-th audio frame and calculate the correlation value y, left-shift x samples in the i-th audio frame and calculate the correlation value z. If y>z, then determine that the binary random number S i corresponding to the i-th audio frame i =1, otherwise, S B-1 =0; Complete the extraction of continuous B frames of audio according to the above extraction method to obtain the random sequence {S0,..., S B-1};
[0099] 5) Reconstruct the user ID by using the extracted binary random sequence {S0,..., S B-1}.
[0100] In particular, if when searching for multiple target samples in A#, selecting the first audio frame in the paid audio A as the target audio frame fails to find the multiple target samples in A#, then select the next audio frame in the paid audio A as the target audio frame and search again until the multiple target samples are found in A#.
[0101] One embodiment of the present invention further provides an audio processing method, which runs on a cloud server Figure 6 is a flowchart of an optional audio processing method according to the embodiments of the present invention. As Figure 6 shown, this audio processing method includes:
[0102] Step S302: Receive the audio to be processed from the client;
[0103] Step S304: Perform frame splitting on the above-mentioned audio to be processed to obtain multiple audio frames, search for offset positions for at least one of the multiple audio frames to obtain a first sample pair and a second sample pair, and perform sample offset on the first sample pair and the second sample pair based on user identification information and multiple candidate offset values to obtain the offset audio;
[0104] Step S306: Feed back the offset audio to the client.
[0105] Optionally, Figure 7 is a schematic diagram of an optional audio processing method in a cloud server according to an embodiment of the present invention. As Figure 7 shown, the client uploads the audio to be processed to the cloud server. The cloud server uses a frame offset method to process the audio, performs frame splitting on the audio to be processed to obtain multiple audio frames, searches for offset positions for at least one of the multiple audio frames to obtain a first sample pair and a second sample pair, and performs sample offset on the first sample pair and the second sample pair based on user identification information and multiple candidate offset values to obtain the offset audio. Then, the cloud server feeds back the processing result to the client, and the final processing result is provided to the user through the client.
[0106] It should be noted that the above-mentioned audio processing method provided by the embodiments of the present application can be but is not limited to being applicable to the actual application scenario of audio-visual copyright protection. Through the interaction between the SaaS server and the client, the audio to be processed is processed by using a frame offset method, and the returned processing result is provided to the user through the client.
[0107] Another embodiment of the present invention also provides another audio processing method, which is used for online audio editing. Figure 8 is a flowchart of another optional audio processing method according to an embodiment of the present invention. As Figure 8 shown, this audio processing method includes:
[0108] Step S802: Load the audio to be processed in the audio editing interface;
[0109] Step S804: In response to a first editing instruction for the audio to be processed, perform frame splitting on the audio to be processed to obtain multiple audio frames;
[0110] Step S806: In response to a second editing instruction for the multiple audio frames, search for offset positions for at least one of the multiple audio frames to obtain a first sample pair and a second sample pair;
[0111] Step S808: In response to a third editing instruction for multiple audio frames, determine user identification information and multiple candidate offset values, and perform sample offset on the first sample pair and the second sample pair based on the user identification information and the multiple candidate offset values to obtain offset audio;
[0112] Step S810: Display the offset audio in the audio editing interface.
[0113] The audio processing method provided by the embodiments of the present invention can be used for online editing of the to-be-processed audio. The above audio editing interface is used to load the to-be-processed audio uploaded by the user and display the processed audio to the user. After loading the to-be-processed audio in the audio editing interface, the following audio processing actions can be achieved: when receiving a first editing instruction for the to-be-processed audio, perform frame splitting on the to-be-processed audio to obtain multiple audio frames; when receiving a second editing instruction for multiple audio frames, perform offset position search on at least one audio frame of the multiple audio frames to obtain a first sample pair and a second sample pair; when receiving a third editing instruction for multiple audio frames, determine user identification information and multiple candidate offset values, and perform sample offset on the first sample pair and the second sample pair based on the user identification information and the multiple candidate offset values to obtain offset audio. The offset audio is displayed to the user as the audio after online editing processing.
[0114] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0116] Embodiment 2
[0117] According to an embodiment of the present invention, there is also provided an apparatus embodiment for implementing the above audio processing method. Figure 9 It is a schematic structural diagram of an audio processing apparatus according to an embodiment of the present invention, as Figure 9 shown. The apparatus includes: an acquisition module 110, a framing module 112, a search module 114, and a processing module 116. Among them,
[0118] The acquisition module 110 is configured to acquire the audio to be processed; the framing module 112 is configured to perform framing processing on the audio to be processed to obtain a plurality of audio frames; the search module 114 is configured to search for offset positions of at least one audio frame among the plurality of audio frames to obtain a first sample pair and a second sample pair; the processing module 116 is configured to perform sample offset on the first sample pair and the second sample pair based on user identification information and a plurality of candidate offset values to obtain the offset audio.
[0119] Optionally, the framing module 112 is further configured to: perform framing processing on the audio to be processed according to a fixed sample number to obtain the plurality of audio frames, where each audio frame of the plurality of audio frames includes: a plurality of samples.
[0120] Optionally, the search module 114 is further configured to: start searching from the first sample among the plurality of samples and search backward to find two adjacent samples with sample symbol inversion to obtain the first sample pair; or, start searching from the first sample among the plurality of samples and search backward to find continuous silent samples, and select two silent samples from the continuous silent samples to obtain the first sample pair.
[0121] Optionally, the search module 114 is further configured to: start searching from the last sample among the plurality of samples and search forward to find two adjacent samples with sample symbol inversion to obtain the second sample pair; or, start searching from the last sample among the plurality of samples and search forward to find continuous silent samples, and select two silent samples from the continuous silent samples to obtain the second sample pair.
[0122] Optionally, the processing module 116 includes: a preparation unit 1161 (not shown in the figure), configured to generate an offset sequence based on the user identification information and the plurality of candidate offset values; an offset unit 1162 (not shown in the figure), configured to perform sample offset on the first sample pair and the second sample pair by using the offset sequence to obtain offset boundary samples; a post-processing unit 1163 (not shown in the figure), configured to perform interpolation smoothing processing on the offset boundary samples to obtain the offset audio.
[0123] Optionally, the preparation unit 1161 is further configured to: determine the multiple candidate offset values based on the first quantity and a preset offset amplitude, where the first quantity is the number of fixed samples included in each audio frame of the multiple audio frames; generate a random sequence by using the user identification information, where the length of the random sequence is a second quantity, and the second quantity is the number of the multiple audio frames; and generate the offset sequence according to a mapping relationship between the random sequence and the multiple candidate offset values.
[0124] Optionally, the audio processing device further includes: a frame search module 210 (not shown in the figure), configured to search for a target audio frame from the multiple audio frames, where the relevance between the target audio frame and the offset audio meets a preset condition; a positioning module 212 (not shown in the figure), configured to determine a sample alignment position based on the target audio frame; a second frame splitting module 214 (not shown in the figure), configured to perform frame splitting processing on the offset audio starting from the sample alignment position to obtain multiple offset audio frames; an extraction module 216 (not shown in the figure), configured to continuously extract a binary random sequence from a third quantity of offset audio frames according to a corresponding relationship between the multiple audio frames and the multiple offset audio frames, where the third quantity is the maximum length required for the binary representation of the user identification information; and a reconstruction module 218 (not shown in the figure), configured to reconstruct the user identification information by using the binary random sequence.
[0125] It should be noted here that the obtaining module 110, the frame splitting module 112, the search module 114, and the processing module 116 correspond to steps S102 to S108 in Embodiment 1. The functions implemented by the four modules and the corresponding steps are the same in terms of examples and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0126] In the embodiment of the present invention, an audio processing method using frame offset is adopted. The method includes: obtaining an audio to be processed; performing frame splitting processing on the audio to be processed to obtain multiple audio frames; searching for offset positions of at least one audio frame among the multiple audio frames to obtain a first sample pair and a second sample pair; and performing sample offset on the first sample pair and the second sample pair based on user identification information and multiple candidate offset values to obtain offset audio. It can be easily noticed that, through the embodiment of the present application, offset with few samples is performed on each audio frame to generate audio with no audible difference. Even when the number of users is relatively large and multiple users conduct a collusion attack, the samples obtained by the collusion attack have a huge audible difference from the original audio and cannot meet the normal use requirements, that is, the collusion attack fails.
[0127] Thus, the embodiments of the present application achieve the purpose of generating audio with few-sample offsets within each audio frame and having no auditory difference from the original audio, thereby realizing the technical effect of processing audio based on frame offset to resist collusion attacks, and further solving the technical problems of difficult traceability of audio-visual piracy and difficult audio-visual copyright protection caused by audio-visual pirates using collusion attack methods to destroy audio-visual watermarks.
[0128] It should be noted that the preferred implementation manners of this embodiment can refer to the relevant descriptions in Embodiment 1 and will not be elaborated here.
[0129] Embodiment 3
[0130] According to an embodiment of the present invention, there is also provided an embodiment of an electronic device, and the electronic device can be any one of the computing devices in a computing device group. The electronic device includes: a processor and a memory, where:
[0131] The memory is connected to the above-mentioned processor and is used to provide instructions for the above-mentioned processor to perform the following processing steps: obtaining the audio to be processed; performing frame division processing on the above-mentioned audio to be processed to obtain a plurality of audio frames; searching for offset positions of at least one of the above-mentioned plurality of audio frames to obtain a first sample pair and a second sample pair; and performing sample offset on the above-mentioned first sample pair and the above-mentioned second sample pair based on user identification information and a plurality of candidate offset values to obtain the offset audio.
[0132] In the embodiments of the present invention, a method of processing audio by frame offset is adopted. By obtaining the audio to be processed; performing frame division processing on the above-mentioned audio to be processed to obtain a plurality of audio frames; searching for offset positions of at least one of the above-mentioned plurality of audio frames to obtain a first sample pair and a second sample pair; and performing sample offset on the above-mentioned first sample pair and the above-mentioned second sample pair based on user identification information and a plurality of candidate offset values to obtain the offset audio.
[0133] It is easy to note that through the embodiments of the present application, few-sample offsets are performed on each audio frame to generate audio with no auditory difference. Even in the case where the number of users is relatively large and multiple users perform collusion attacks, the samples obtained by the collusion attacks have a huge auditory difference from the original audio and cannot meet the normal use requirements, that is, the collusion attacks fail.
[0134] Thus, the embodiments of the present application achieve the purpose of generating audio with few-sample offsets within each audio frame and having no auditory difference from the original audio, thereby realizing the technical effect of processing audio based on frame offset to resist collusion attacks, and further solving the technical problems of difficult traceability of audio-visual piracy and difficult audio-visual copyright protection caused by audio-visual pirates using collusion attack methods to destroy audio-visual watermarks.
[0135] It should be noted that the preferred implementation of this embodiment can refer to the relevant description in Embodiment 1, and will not be elaborated here.
[0136] Embodiment 4
[0137] An embodiment of the present invention can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above computer terminal can also be replaced with a terminal device such as a mobile terminal.
[0138] Optionally, in this embodiment, the above computer terminal can be located in at least one of multiple network devices in a computer network.
[0139] In this embodiment, the above computer terminal can execute the program code of the following steps in the audio processing method: obtain the audio to be processed; perform frame splitting on the audio to be processed to obtain multiple audio frames; search for offset positions for at least one of the multiple audio frames to obtain a first sample pair and a second sample pair; perform sample offset on the first sample pair and the second sample pair based on user identification information and multiple candidate offset values to obtain the offset audio.
[0140] Optionally, Figure 10 is a structural block diagram of another computer terminal according to an embodiment of the present invention, as Figure 10 shown, the computer terminal may include: one or more (only one is shown in the figure) processors 122, a memory 124, and a peripheral interface 126.
[0141] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the audio processing method and device in the embodiment of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above audio processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely set relative to the processor, and these remote memories can be connected to Terminal A through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0142] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: obtain the audio to be processed; perform frame processing on the audio to be processed to obtain multiple audio frames; search for the offset positions of at least one of the multiple audio frames to obtain a first sample pair and a second sample pair; perform sample offset on the first sample pair and the second sample pair based on the user identification information and multiple candidate offset values to obtain the offset audio.
[0143] Optionally, the processor can also execute the program code of the following steps: perform frame processing on the audio to be processed according to a fixed sample quantity to obtain the multiple audio frames, where each of the multiple audio frames includes: multiple samples.
[0144] Optionally, the processor can also execute the program code of the following steps: start searching backward from the first sample among the multiple samples to find two adjacent samples with sample symbol inversion to obtain the first sample pair; or, start searching backward from the first sample among the multiple samples to find continuous silent samples, and select two silent samples from the continuous silent samples to obtain the first sample pair.
[0145] Optionally, the processor can also execute the program code of the following steps: start searching forward from the last sample among the multiple samples to find two adjacent samples with sample symbol inversion to obtain the second sample pair; or, start searching forward from the last sample among the multiple samples to find continuous silent samples, and select two silent samples from the continuous silent samples to obtain the second sample pair.
[0146] Optionally, the processor can also execute the program code of the following steps: generate an offset sequence based on the user identification information and the multiple candidate offset values; use the offset sequence to perform sample offset on the first sample pair and the second sample pair to obtain offset boundary samples; perform interpolation smoothing processing on the offset boundary samples to obtain the offset audio.
[0147] Optionally, the processor can also execute the program code of the following steps: determine the multiple candidate offset values based on a first quantity and a preset offset amplitude, where the first quantity is the fixed sample quantity included in each of the multiple audio frames; generate a random sequence using the user identification information, where the length of the random sequence is a second quantity, and the second quantity is the number of the multiple audio frames; generate the offset sequence through the mapping relationship between the random sequence and the multiple candidate offset values.
[0148] Optionally, the above-mentioned processor may also execute program code for the following steps: finding a target audio frame from the above-mentioned multiple audio frames, where the relevance between the above-mentioned target audio frame and the above-mentioned offset audio meets a preset condition; determining a sample alignment position based on the above-mentioned target audio frame; starting from the above-mentioned sample alignment position, performing frame splitting on the above-mentioned offset audio to obtain a plurality of offset audio frames; continuously extracting a binary random sequence from a third quantity of offset audio frames through the corresponding relationship between the above-mentioned multiple audio frames and the above-mentioned multiple offset audio frames, where the above-mentioned third quantity is the maximum length required for the binary representation of the above-mentioned user identification information; reconstructing the above-mentioned user identification information using the above-mentioned binary random sequence.
[0149] The processor may call the information and application programs stored in the memory through the transmission device to execute the following steps: receiving the audio to be processed from the client; performing frame splitting on the above-mentioned audio to be processed to obtain a plurality of audio frames, searching for offset positions of at least one of the above-mentioned plurality of audio frames to obtain a first sample pair and a second sample pair, and performing sample offset on the above-mentioned first sample pair and the above-mentioned second sample pair based on the user identification information and a plurality of candidate offset values to obtain offset audio; feeding back the above-mentioned offset audio to the above-mentioned client.
[0150] The processor may call the information and application programs stored in the memory through the transmission device to execute the following steps: loading the audio to be processed in the audio editing interface; in response to a first editing instruction for the audio to be processed, performing frame splitting on the audio to be processed to obtain a plurality of audio frames; in response to a second editing instruction for the plurality of audio frames, searching for offset positions of at least one of the plurality of audio frames to obtain a first sample pair and a second sample pair; in response to a third editing instruction for the plurality of audio frames, determining the user identification information and a plurality of candidate offset values, and performing sample offset on the first sample pair and the second sample pair based on the user identification information and the plurality of candidate offset values to obtain offset audio; displaying the offset audio in the audio editing interface.
[0151] In an embodiment of the present invention, a method of processing audio by frame offset is adopted, including obtaining the audio to be processed; performing frame splitting on the above-mentioned audio to be processed to obtain a plurality of audio frames; searching for offset positions of at least one of the above-mentioned plurality of audio frames to obtain a first sample pair and a second sample pair; performing sample offset on the above-mentioned first sample pair and the above-mentioned second sample pair based on the user identification information and a plurality of candidate offset values to obtain offset audio.
[0152] It is easy to notice that through the embodiments of the present application, performing small-sample offset on each audio frame generates audio with no audible difference. Even in the case where the number of users is relatively large and multiple users conduct a collusion attack, the samples obtained from the collusion attack have a huge audible difference from the original audio and cannot meet the normal usage requirements, that is, the collusion attack fails.
[0153] Accordingly, the embodiments of the present application achieve the purpose of generating audio with few sample offsets within each audio frame and having no auditory difference from the original audio, thereby realizing the technical effect of processing audio based on frame offset to resist collusion attacks, and further solving the technical problems of difficult traceability of audio-visual piracy and difficult audio-visual copyright protection caused by audio-visual pirates using collusion attack methods to destroy audio-visual watermarks.
[0154] Those of ordinary skill in the art can understand that Figure 10 the structure shown is only illustrative, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 10 It does not limit the structure of the above electronic device. For example, the computer terminal may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 10 or have a different configuration from that shown in Figure 10 shown.
[0155] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disc, etc.
[0156] According to an embodiment of the present invention, an embodiment of a storage medium is further provided. Optionally, in this embodiment, the above storage medium can be used to store the program code executed by the audio processing method provided in the above Embodiment 1.
[0157] Optionally, in this embodiment, the above storage medium may be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0158] Optionally, in this embodiment, the storage medium is set to store program code for performing the following steps: obtaining the audio to be processed; performing frame division processing on the above audio to be processed to obtain a plurality of audio frames; searching for offset positions of at least one of the above plurality of audio frames to obtain a first sample pair and a second sample pair; performing sample offset on the above first sample pair and the above second sample pair based on user identification information and a plurality of candidate offset values to obtain the offset audio.
[0159] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: performing frame splitting on the to-be-processed audio according to a fixed number of samples to obtain the multiple audio frames, where each audio frame of the multiple audio frames includes: multiple samples.
[0160] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: searching backward from the first sample among the multiple samples to find two adjacent samples with sample symbol inversion to obtain the first sample pair; or, searching backward from the first sample among the multiple samples to find consecutive silent samples, and selecting two silent samples from the consecutive silent samples to obtain the first sample pair.
[0161] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: searching forward from the last sample among the multiple samples to find two adjacent samples with sample symbol inversion to obtain the second sample pair; or, searching forward from the last sample among the multiple samples to find consecutive silent samples, and selecting two silent samples from the consecutive silent samples to obtain the second sample pair.
[0162] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: generating an offset sequence based on the user identification information and the multiple candidate offset values; performing sample offset on the first sample pair and the second sample pair by using the offset sequence to obtain offset boundary samples; performing interpolation smoothing processing on the offset boundary samples to obtain the offset audio.
[0163] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: determining the multiple candidate offset values based on a first quantity and a preset offset amplitude, where the first quantity is the fixed number of samples included in each audio frame of the multiple audio frames; generating a random sequence by using the user identification information, where the length of the random sequence is a second quantity, and the second quantity is the number of the multiple audio frames; generating the offset sequence through a mapping relationship between the random sequence and the multiple candidate offset values.
[0164] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: finding a target audio frame from the multiple audio frames, where the relevance between the target audio frame and the offset audio meets a preset condition; determining a sample alignment position based on the target audio frame; starting from the sample alignment position, performing frame division processing on the offset audio to obtain multiple offset audio frames; extracting a binary random sequence continuously from a third number of offset audio frames through the correspondence between the multiple audio frames and the multiple offset audio frames, where the third number is the maximum length required for the binary representation of the user identification information; and reconstructing the user identification information using the binary random sequence.
[0165] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: receiving the audio to be processed from a client; performing frame division processing on the audio to be processed to obtain multiple audio frames, searching for an offset position of at least one audio frame among the multiple audio frames to obtain a first sample pair and a second sample pair, and performing sample offset on the first sample pair and the second sample pair based on the user identification information and multiple candidate offset values to obtain offset audio; and feeding back the offset audio to the client.
[0166] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: loading the audio to be processed in an audio editing interface; in response to a first editing instruction for the audio to be processed, performing frame division processing on the audio to be processed to obtain multiple audio frames; in response to a second editing instruction for the multiple audio frames, searching for an offset position of at least one audio frame among the multiple audio frames to obtain a first sample pair and a second sample pair; in response to a third editing instruction for the multiple audio frames, determining the user identification information and multiple candidate offset values, and performing sample offset on the first sample pair and the second sample pair based on the user identification information and the multiple candidate offset values to obtain offset audio; and displaying the offset audio in the audio editing interface.
[0167] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.
[0168] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0169] In several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0170] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0171] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0172] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical disks and other various media that can store program codes.
[0173] The above is only the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An audio processing method, characterized in that, Including: Obtain the audio to be processed; Perform frame splitting on the audio to be processed to obtain a plurality of audio frames; Search for offset positions for at least one of the plurality of audio frames to obtain a first sample pair and a second sample pair; Based on the user identification information and a plurality of candidate offset values, perform sample offset on the first sample pair and the second sample pair to obtain the offset audio; Search for target audio frames from the plurality of audio frames; Calculate the correlation of a plurality of target samples to determine the sample alignment position, where the plurality of target samples are a plurality of consecutive samples selected from the offset audio, and the correlation between the target audio frame and the plurality of target samples meets a preset condition; starting from the sample alignment position, perform frame splitting on the offset audio to obtain a plurality of offset audio frames; through the corresponding relationship between the plurality of audio frames and the plurality of offset audio frames, continuously extract a binary random sequence from a third number of offset audio frames, where the third number is the maximum length required for the binary representation of the user identification information; reconstruct the user identification information using the binary random sequence.
2. The audio processing method according to claim 1, wherein Performing frame splitting on the audio to be processed to obtain the plurality of audio frames includes: Perform frame splitting on the audio to be processed according to a fixed sample number to obtain the plurality of audio frames, where each of the plurality of audio frames includes: a plurality of samples.
3. The audio processing method according to claim 2, wherein Searching for offset positions for at least one of the plurality of audio frames to obtain the first sample pair includes: Search backward from the first sample among the plurality of samples to find two adjacent samples with reversed sample symbols to obtain the first sample pair; or, Search backward from the first sample among the plurality of samples to find consecutive silent samples, and select two silent samples from the consecutive silent samples to obtain the first sample pair.
4. The audio processing method according to claim 2, wherein Searching for offset positions for at least one of the plurality of audio frames to obtain the second sample pair includes: Search forward from the last sample among the plurality of samples to find two adjacent samples with reversed sample symbols to obtain the second sample pair; or, Search forward from the last sample among the plurality of samples to find consecutive silent samples, and select two silent samples from the consecutive silent samples to obtain the second sample pair.
5. The audio processing method according to claim 1, wherein Based on the user identification information and the plurality of candidate offset values, performing sample offset on the first sample pair and the second sample pair to obtain the offset audio includes: Generate an offset sequence based on the user identification information and the plurality of candidate offset values; Use the offset sequence to perform sample offset on the first sample pair and the second sample pair to obtain offset boundary samples; Perform interpolation smoothing on the offset boundary samples to obtain the offset audio.
6. The audio processing method according to claim 5, wherein Generating the offset sequence based on the user identification information and the plurality of candidate offset values includes: Determine the plurality of candidate offset values based on a first number and a preset offset amplitude, where the first number is the fixed sample number included in each of the plurality of audio frames; Generate a random sequence using the user identification information, where the length of the random sequence is a second quantity, and the second quantity is the number of the multiple audio frames; Generate the offset sequence through the mapping relationship between the random sequence and the multiple candidate offset values.
7. An audio processing method, characterized in that, Comprising: Receive the audio to be processed from the client; Perform frame splitting on the audio to be processed to obtain multiple audio frames, search for offset positions of at least one of the multiple audio frames to obtain a first sample pair and a second sample pair, and perform sample offset on the first sample pair and the second sample pair based on the user identification information and multiple candidate offset values to obtain the offset audio; Feed back the offset audio to the client; Find a target audio frame from the multiple audio frames; Calculate the correlation of multiple target samples to determine the sample alignment position, where the multiple target samples are multiple consecutive samples selected from the offset audio, and the correlation between the target audio frame and the multiple target samples meets a preset condition; starting from the sample alignment position, perform frame splitting on the offset audio to obtain multiple offset audio frames; through the corresponding relationship between the multiple audio frames and the multiple offset audio frames, continuously extract a binary random sequence from the third quantity of offset audio frames, where the third quantity is the maximum length required for the binary representation of the user identification information; reconstruct the user identification information using the binary random sequence.
8. An audio processing method, characterized in that, Provide an audio editing interface through a terminal device, and the audio processing method includes: Load the audio to be processed in the audio editing interface; In response to a first editing instruction for the audio to be processed, perform frame splitting on the audio to be processed to obtain multiple audio frames; In response to a second editing instruction for the multiple audio frames, search for offset positions of at least one of the multiple audio frames to obtain a first sample pair and a second sample pair; In response to a third editing instruction for the multiple audio frames, determine the user identification information and multiple candidate offset values, and perform sample offset on the first sample pair and the second sample pair based on the user identification information and the multiple candidate offset values to obtain the offset audio; Display the offset audio in the audio editing interface; Find a target audio frame from the multiple audio frames; calculate the correlation of multiple target samples to determine the sample alignment position, where the multiple target samples are multiple consecutive samples selected from the offset audio, and the correlation between the target audio frame and the multiple target samples meets a preset condition; starting from the sample alignment position, perform frame splitting on the offset audio to obtain multiple offset audio frames; through the corresponding relationship between the multiple audio frames and the multiple offset audio frames, continuously extract a binary random sequence from the third quantity of offset audio frames, where the third quantity is the maximum length required for the binary representation of the user identification information; reconstruct the user identification information using the binary random sequence.
9. An audio processing device, characterized in that, Comprising: An acquisition module for acquiring the audio to be processed; A framing module, configured to perform framing processing on the to-be-processed audio to obtain a plurality of audio frames; A search module, configured to search for offset positions of at least one of the plurality of audio frames to obtain a first sample pair and a second sample pair; A processing module, configured to perform sample offset on the first sample pair and the second sample pair based on user identification information and a plurality of candidate offset values to obtain offset audio; A frame searching module, configured to search for a target audio frame from the plurality of audio frames; A positioning module, configured to calculate a correlation degree of a plurality of target samples to determine a sample alignment position, where the plurality of target samples are a plurality of consecutive samples selected from the offset audio, and a correlation degree between the target audio frame and the plurality of target samples meets a preset condition; a second framing module, configured to perform framing processing on the offset audio starting from the sample alignment position to obtain a plurality of offset audio frames; an extraction module, configured to continuously extract a binary random sequence from a third number of offset audio frames according to a corresponding relationship between the plurality of audio frames and the plurality of offset audio frames, where the third number is a maximum length required for binary representation of the user identification information; a reconstruction module, configured to reconstruct the user identification information by using the binary random sequence.
10. A storage medium, characterized in that, The storage medium includes a stored program, where when the program runs, it controls a device where the storage medium is located to execute the audio processing method according to any one of claims 1 to 8.
11. A processor, characterized in that, The processor is configured to run a program, where when the program runs, it executes the audio processing method according to any one of claims 1 to 8.
12. An electronic device, characterized in that, Comprising: A processor; And A memory, connected to the processor, and configured to provide instructions for the processor to perform the following processing steps: Step 1, obtaining to-be-processed audio; Step 2, performing framing processing on the to-be-processed audio to obtain a plurality of audio frames; Step 3, searching for offset positions of at least one of the plurality of audio frames to obtain a first sample pair and a second sample pair; Step 4, performing sample offset on the first sample pair and the second sample pair based on user identification information and a plurality of candidate offset values to obtain offset audio; Step 5, searching for a target audio frame from the plurality of audio frames; Calculating a correlation degree of a plurality of target samples to determine a sample alignment position, where the plurality of target samples are a plurality of consecutive samples selected from the offset audio, and a correlation degree between the target audio frame and the plurality of target samples meets a preset condition; performing framing processing on the offset audio starting from the sample alignment position to obtain a plurality of offset audio frames; continuously extracting a binary random sequence from a third number of offset audio frames according to a corresponding relationship between the plurality of audio frames and the plurality of offset audio frames, where the third number is a maximum length required for binary representation of the user identification information; reconstructing the user identification information by using the binary random sequence.
Citation Information
Patent Citations
Method for embedding and detecting a watermark
TW200941281A