Audio encoding and decoding method, device, computer-readable medium and electronic device

By embedding digital watermarks into fixed codebook vectors during the audio encoding process and encapsulating them into the audio code stream, the problems of degraded audio quality and low encoding efficiency in the prior art are solved, and efficient digital watermark embedding is achieved.

CN115312069BActive Publication Date: 2025-06-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110492373.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-06
Publication Date
2025-06-24
Estimated Expiration
2041-05-06

AI Technical Summary

Technical Problem

When embedding watermarks, existing audio digital watermarking technologies require decoding and re-encoding of the encoded code stream data, resulting in a decrease in audio quality and a decrease in encoding efficiency.

Method used

By obtaining the audio code stream encoded by the original audio data, the code stream is parsed to obtain encoding parameters, including the fixed codebook vector of each audio data frame, and embed digital watermarks in the fixed codebook vector and encapsulate them into the audio code stream.

Benefits of technology

It realizes that the embedding efficiency of digital watermarks is improved without affecting the audio quality, and the application scenarios of digital watermark technology are widened.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115312069B_ABST
    Figure CN115312069B_ABST
Patent Text Reader

Abstract

The present application discloses an audio encoding and decoding method, apparatus, computer-readable medium, and electronic device. The audio encoding method includes: obtaining an audio bitstream obtained by encoding original audio data; parsing the audio bitstream to obtain encoding parameters carried in the audio bitstream, where the encoding parameters include fixed codebook vectors of respective audio data frames in the audio bitstream; embedding a digital watermark into the fixed codebook vectors, and encapsulating the fixed codebook vectors with the embedded digital watermark into the audio bitstream. In the technical solution provided by the embodiments of the present application, the fixed codebook vectors used in the audio encoding process are used as carriers for embedding digital watermarks, without the need to decode and then re-encode the encoded audio bitstream. This not only does not affect the audio quality while implementing digital watermark embedding, but also enables the embedding of digital watermarks not to be limited to the audio signal before encoding, broadening the application scenarios of digital watermark technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of audio - video technology, and particularly relates to an audio encoding method, an audio decoding method, an audio encoding device, an audio decoding device, a computer - readable medium, and an electronic device. Background Art

[0002] The technology of adding watermarks to voice information is called audio digital watermark technology. It mainly embeds digital watermarks into audio data through a watermark embedding algorithm, but does not have too much impact on the original sound quality of the audio data, or makes the human ear unable to feel its impact.

[0003] Currently, audio digital watermark technology adds digital watermarks before audio data encoding. If the audio data is encoded bit - stream data when adding digital watermarks, it is necessary to first decode the bit - stream data to obtain audio data, then add digital watermarks to the audio data and then encode and transmit it. The process of decoding and then encoding for watermark embedding will lead to problems such as a decrease in audio quality and a reduction in encoding efficiency. Summary of the Invention

[0004] The purpose of this application is to provide an audio encoding method, an audio decoding method, an audio encoding device, an audio decoding device, a computer - readable medium, and an electronic device, which can at least overcome the technical problems such as low audio quality and low encoding efficiency existing in the related technology to a certain extent.

[0005] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.

[0006] According to one aspect of the embodiments of this application, an audio encoding method is provided, including: obtaining an audio bit - stream obtained by encoding original audio data; parsing the audio bit - stream to obtain encoding parameters carried in the audio bit - stream, where the encoding parameters include fixed codebook vectors of each audio data frame in the audio bit - stream; embedding a digital watermark into the fixed codebook vectors, and encapsulating the fixed codebook vectors with the embedded digital watermark into the audio bit - stream.

[0007] According to one aspect of the embodiments of this application, an audio encoding device is provided, including: an audio bit - stream acquisition module for obtaining an audio bit - stream obtained by encoding original audio data; an audio bit - stream parsing module for parsing the audio bit - stream to obtain encoding parameters carried in the audio bit - stream, where the encoding parameters include fixed codebook vectors of each audio data frame in the audio bit - stream; a digital watermark embedding module for embedding a digital watermark into the fixed codebook vectors and encapsulating the fixed codebook vectors with the embedded digital watermark into the audio bit - stream.

[0008] In an embodiment of the present application, the encoding parameter further includes a fixed codebook gain corresponding to the fixed codebook vector; the digital watermark embedding module includes: a gain threshold obtaining unit, configured to obtain a gain threshold corresponding to the audio data frame; a target codebook vector determining unit, configured to screen out a target codebook vector from the fixed codebook vectors of a plurality of the audio data frames according to the gain threshold and the fixed codebook gain, where the fixed codebook gain corresponding to the target codebook vector is less than the gain threshold; and a digital watermark embedding unit, configured to embed a digital watermark into the target codebook vector.

[0009] In an embodiment of the present application, the digital watermark embedding unit is further configured to: replace some or all of the codewords in the target codebook vector with a digital watermark.

[0010] In an embodiment of the present application, the target codebook vector determining unit is further configured to: perform a numerical comparison between the fixed codebook gain of each of the audio data frames and the gain threshold to screen out candidate audio data frames whose fixed codebook gain is less than the gain threshold; obtain the number of frame intervals between any two candidate audio data frames; screen out target audio data frames from the candidate audio data frames whose number of frame intervals is greater than a number threshold, and use the fixed codebook vector of the target audio data frames as the target codebook vector.

[0011] In an embodiment of the present application, the gain threshold obtaining unit includes: a frame sequence segment obtaining subunit, configured to obtain the frame sequence segment where the audio data frame is located, where the frame sequence segment includes the audio data frame and a preset number of other audio data frames adjacent to the audio data frame; a fixed codebook gain obtaining subunit, configured to obtain the fixed codebook gain of each audio data frame in the frame sequence segment; and a gain threshold determining subunit, configured to determine a gain threshold corresponding to the audio data frame according to the fixed codebook gain of each audio data frame.

[0012] In an embodiment of the present application, the gain threshold determining subunit is further configured to: sort the fixed codebook gains of each audio data frame according to the numerical size to obtain a gain sequence; select a target codebook gain from the gain sequence as the gain threshold corresponding to the audio data frame according to a preset quantity ratio.

[0013] In an embodiment of the present application, the gain threshold determining subunit is further configured to: obtain the average value of the fixed codebook gains of each audio data frame; determine the gain threshold corresponding to the audio data frame according to the average value.

[0014] In an embodiment of the present application, the digital watermark includes a watermark entity, a watermark synchronization code located before the watermark entity, and a watermark verification code located after the watermark entity. The watermark synchronization code is used to identify the starting position of the digital watermark, and the watermark verification code is used to verify the correctness of the watermark entity.

[0015] According to one aspect of the embodiments of the present application, there is provided an audio decoding method, including: parsing an audio bitstream to obtain coding parameters carried in the audio bitstream, where the coding parameters include fixed codebook vectors of respective audio data frames in the audio bitstream; and extracting a digital watermark embedded in the audio bitstream from the fixed codebook vectors.

[0016] According to one aspect of the embodiments of the present application, there is provided an audio decoding apparatus, including:

[0017] an audio bitstream parsing module, configured to parse an audio bitstream to obtain coding parameters carried in the audio bitstream, where the coding parameters include fixed codebook vectors of respective audio data frames in the audio bitstream; and a digital watermark extraction module, configured to extract a digital watermark embedded in the audio bitstream from the fixed codebook vectors.

[0018] In an embodiment of the present application, the coding parameters further include fixed codebook gains corresponding to the fixed codebook vectors; the digital watermark extraction module includes: a gain threshold obtaining unit, configured to obtain a gain threshold corresponding to the audio data frame; a target codebook vector determining unit, configured to screen out a target codebook vector from the fixed codebook vectors of multiple audio data frames according to the gain threshold and the fixed codebook gain, where the fixed codebook gain corresponding to the target codebook vector is less than the gain threshold; and a digital watermark extraction unit, configured to extract a digital watermark from the target codebook vector.

[0019] In an embodiment of the present application, the digital watermark extraction unit is further configured to: extract all codewords in the target codebook vector and use the all codewords as the digital watermark.

[0020] In an embodiment of the present application, the digital watermark includes a watermark entity, a watermark synchronization code located before the watermark entity, and a watermark verification code located after the watermark entity. The watermark synchronization code is used to identify the starting position of the digital watermark, and the watermark verification code is used to verify the correctness of the watermark entity.

[0021] In an embodiment of the present application, the digital watermark extraction unit is further configured to: perform matching detection on each codeword in the target codebook vector to determine the watermark synchronization code embedded in the target codebook vector, the watermark entity and the watermark verification code located after the watermark synchronization code; perform correctness verification on the watermark entity according to the watermark verification code; when the verification passes, form a digital watermark with the watermark synchronization code, the watermark entity and the watermark verification code.

[0022] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the audio encoding or decoding method in the above technical solution.

[0023] According to one aspect of the embodiments of the present application, there is provided an electronic device, which includes: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the audio encoding or decoding method in the above technical solution by executing the executable instructions.

[0024] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the audio encoding or decoding method in the above technical solution.

[0025] In the technical solution provided by the embodiments of the present application, the fixed codebook vector used in the audio encoding process is used as the embedding carrier of the digital watermark, and there is no need to decode and re-encode the encoded audio bitstream. Not only does it not affect the audio quality while realizing the embedding of the digital watermark, but also makes the embedding of the digital watermark not limited to the audio signal before encoding, expanding the application scenarios of the digital watermark technology.

[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0028] Figure 1Schematically shown is an exemplary system architecture block diagram to which the technical solution of the present application is applied.

[0029] Figure 2 Shown is a CELP coding model that can be used to implement the audio coding scheme of the present application.

[0030] Figure 3 Shown is a CELP decoding model that can be used to implement the audio decoding scheme of the present application.

[0031] Figure 4 Schematic flow diagram of an audio coding method provided in an embodiment of the present application.

[0032] Figure 5 Shown is a specific flowchart of a method for embedding a digital watermark into a fixed codebook vector in an embodiment of the present application.

[0033] Figure 6 Schematic diagram of the structure of a digital watermark in an embodiment of the present application.

[0034] Figure 7 Schematic flow diagram of an audio decoding method provided in an embodiment of the present application.

[0035] Figure 8 Schematic flow diagram of an audio decoding method provided in an embodiment of the present application.

[0036] Figure 9 Schematic flow diagram of an audio decoding method provided in an embodiment of the present application.

[0037] Figure 10 Schematically shown is a block diagram of the structure of an audio coding device provided in an embodiment of the present application.

[0038] Figure 11 Schematically shown is a block diagram of the structure of an audio decoding device provided in an embodiment of the present application.

[0039] Figure 12 Schematically shown is a block diagram of a computer system structure of an electronic device suitable for implementing an embodiment of the present application. Detailed implementation manners

[0040] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0041] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be employed. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0042] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0043] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may change according to the actual situation.

[0044] Figure 1 An exemplary system architecture block diagram applying the technical solution of the present application is schematically shown.

[0045] As Figure 1 shown, the system architecture 100 may include a terminal device 110, a network 120, and a server 130. The terminal device 110 may include various electronic devices such as smartphones, tablets, laptops, desktop computers, etc. The server 130 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The network 120 may be a communication medium of various connection types capable of providing a communication link between the terminal device 110 and the server 130. For example, it may be a wired communication link or a wireless communication link.

[0046] According to the implementation requirements, the system architecture in the embodiments of the present application may have any number of terminal devices, networks, and servers. For example, the server 130 may be a server group composed of multiple server devices. In addition, the technical solution provided by the embodiments of the present application may be applied to the terminal device 110, or may be applied to the server 130, or may be jointly implemented by the terminal device 110 and the server 130. The present application makes no special limitation on this.

[0047] Speech coding and decoding plays an important role in modern communication systems. In voice call applications, the voice signal is collected by a microphone, and the analog voice signal is converted into a digital voice signal through an analog-to-digital conversion circuit. The digital signal is compressed by a speech encoder and then packaged and sent to the receiving end according to the communication network transmission format and protocol. After receiving the data packet, the receiving end device unpacks it and outputs the speech coding compressed bitstream. After passing through the speech decoder, the speech digital signal is regenerated, and finally the speech digital signal is played out as sound through a speaker. Speech coding and decoding can effectively reduce the bandwidth of voice signal transmission, which plays a decisive role in saving the storage and transmission cost of voice information and ensuring the integrity of voice information during the communication network transmission process.

[0048] In some embodiments of the present application, audio data coding and decoding can be based on the CELP (Code Excited Linear Prediction) coding technology. The CELP coding technology is a medium and low-speed speech compression coding technology. It uses a codebook as the excitation source and has the advantages of low rate, high synthetic speech quality, strong anti-noise ability, and good performance in multiple audio transfers. It has a wide application in speech compression at a rate of 4.8 - 16 kbps. Speech encoders based on the CELP technology include G.723, G.728, G.729, etc. In addition, many speech encoders are evolved from CELP.

[0049] The CELP audio coding algorithm uses linear prediction to extract vocal tract parameters and uses a codebook containing many typical excitation vectors as the excitation parameter. Each time of coding, the best excitation signal is searched in this codebook, and the coding value of this excitation signal is the serial number in the codebook. CELP has been adopted by many speech coding standards. For example, the US Federal Standard FS1016 uses the CELP coding method and is mainly used for high-quality narrowband voice secure communication.

[0050] The basic idea of CLEP for audio coding is to arrange the combinations of various sample values that the residual signal may appear within a certain time according to certain rules to form a codebook. During coding, a group of the closest residual signals is searched from the local codebook, and then the address corresponding to this group of residual signals is coded and transmitted. A same codebook is also set at the decoding end. The corresponding residual signal is taken out according to the received address and added to the filter to complete audio reconstruction. This method can greatly reduce the number of transmitted bits and improve the coding efficiency.

[0051] Figure 2 Shows the CELP coding model that can be used to implement the audio coding scheme of the present application. As Figure 2As shown, for the original audio data s(n), after preprocessing such as high-pass filtering, a set of linear prediction filter coefficients can be obtained through LPC linear prediction analysis. After converting the LPC parameters into LSP parameters and quantifying them, prediction filtering is performed using the LPC parameters. The difference between the original signal and the LPC prediction filtering result is the residual signal. The residual signal undergoes open-loop and closed-loop pitch analysis to search for the optimal pitch delay parameter, and then the optimal pulse position and amplitude parameters are obtained through fixed codebook search, and the adaptive codebook gain Ga and fixed codebook gain Gc are calculated. The coding parameters obtained during the encoding process can be encapsulated and transmitted through the channel to the receiving end.

[0052] Figure 3 The CELP decoding model that can be used to implement the audio decoding scheme of the present application is shown. The speech decoder at the receiving end parses all the coding parameters from the received data packet, and interpolates the LSP parameters to obtain the LPC filter coefficients, as Figure 3 shown, according to the fixed codebook and the fixed codebook gain, a fixed codebook excitation signal can be generated, and the adaptive codebook and the adaptive codebook gain can generate an adaptive codebook excitation signal. The sum of the two excitations is filtered and post-processed through the LPC synthesis filter to obtain the final speech signal.

[0053] The audio encoding and decoding methods provided by the present application will be described in detail below in conjunction with specific embodiments.

[0054] Figure 4 The flowchart of the audio encoding method provided by an embodiment of the present application. This audio encoding method can be executed by a terminal device or a server, or jointly executed by a terminal device and a server. The audio encoding method executed on the terminal device is used as an example for description in the embodiments of the present application. As Figure 4 shown, the audio encoding method provided by the embodiment of the present application at least includes steps S410 to S430, and the specific content is as follows.

[0055] Step S410: Obtain the audio code stream obtained by encoding and processing the original audio data.

[0056] The original audio data is digitized sound data, that is, a digital audio signal. It can sample the continuous analog audio signal from the sound playback source at a certain frequency through a sound collection device (such as a microphone), and then convert the model sound signal into a digital signal through an analog-to-digital conversion circuit, such as audio files in formats such as wav, mp3, and avi. Encoding and processing the original audio data is to remove the redundant components in the sound signal. The so-called redundant components refer to the signals in the original audio data that cannot be perceived by the human ear and are of no help in determining information such as the timbre and pitch of the sound. The audio data stream after encoding and processing is called the audio code stream.

[0057] In an embodiment of the present application, the encoding process of the original audio data can be implemented by a sound acquisition device integrated with an audio encoding function, or by a dedicated audio encoder. The acquisition of the audio bitstream can be obtained from a sound acquisition device integrated with an audio encoding function, a dedicated audio encoder, or other devices storing the already encoded audio bitstream, and this embodiment does not make any restrictions.

[0058] In an embodiment of the present application, a method for encoding the original audio data is the CELP speech encoding technology. It mainly performs linear prediction on each audio data frame, and uses an adaptive codebook storing past driving sound sources and a fixed codebook storing multiple noise vectors to encode the prediction residual (excitation signal) of each frame of linear prediction.

[0059] Step S420: Parse the audio bitstream to obtain the encoding parameters carried in the audio bitstream. The encoding parameters include the fixed codebook vectors of each audio data frame in the audio bitstream.

[0060] The encoding parameters refer to the model parameters of the speech model formed based on the original audio data during the encoding process, which include the relevant attribute information of each audio data frame in the audio bitstream, such as the fixed codebook vector. The fixed codebook vector is the residual signal after short-time linear prediction and long-time prediction (i.e., adaptive search) of the speech signal obtained through fixed codebook search, and it is a noise vector with the characteristics of random noise, so it is also called a random codebook vector.

[0061] Step S430: Embed a digital watermark into the fixed codebook vector and encapsulate the fixed codebook vector with the embedded digital watermark into the audio bitstream.

[0062] Specifically, after extracting the fixed codebook vector, embed a digital watermark into it, and use the fixed codebook vector with the embedded digital watermark as the encoding parameter to re-encapsulate it with the original audio bitstream to form a new audio bitstream with the embedded watermark.

[0063] In the technical solution provided by the embodiment of the present application, using the fixed codebook vector used in the audio encoding process as the embedding carrier of the digital watermark does not require decoding and re-encoding the encoded audio bitstream. It not only does not affect the audio quality while implementing the digital watermark embedding, but also enables the embedding of the digital watermark not to be limited to the audio signal before encoding, expanding the application scenario of the digital watermark technology.

[0064] Figure 5 The specific flowchart of the method for embedding a digital watermark into the fixed codebook vector in an embodiment of the present application is shown, that is, the flowchart of step S430 in the above embodiment. As Figure 5 shown, step S430 may specifically include steps S510 to S530, which are as follows.

[0065] Step S510: Obtain a gain threshold corresponding to the audio data frame.

[0066] Specifically, among the encoding parameters obtained by parsing the audio bitstream, it not only includes the fixed codebook vectors, but also includes the fixed codebook gains corresponding to each fixed codebook vector. The gain threshold is used to judge the value of the fixed codebook gain. Each audio data frame corresponds to a gain threshold, and the gain threshold of each audio data frame can be different gain thresholds set in advance, or the same gain threshold set in advance.

[0067] In an embodiment of the present application, the gain threshold is a fixed value preset according to the auditory sensitivity of the human ear. For an audio data frame with a fixed codebook gain less than the gain threshold, the human ear has a lower sensitivity to it. Therefore, when replacing the corresponding fixed codebook with a digital watermark, it will not have too much impact on the sound quality.

[0068] In an embodiment of the present application, the gain threshold of each audio data frame can also be dynamically determined according to other setting rules. At the decoding end, the gain threshold of each audio data frame can be dynamically determined according to the same rule, and then some fixed codebook vectors can be extracted as digital watermarks.

[0069] In an embodiment of the present application, the step of obtaining the gain threshold corresponding to the audio data frame specifically includes: obtaining the frame sequence segment where the audio data frame is located, and the frame sequence segment includes the audio data frame and a preset number of other audio data frames adjacent to the audio data frame; obtaining the fixed codebook gains of each audio data frame in the frame sequence segment; and determining the gain threshold corresponding to the audio data frame according to the fixed codebook gains of each audio data frame.

[0070] Specifically, the frame sequence segment refers to a sequence segment formed by multiple consecutive audio data frames, and the multiple consecutive audio data frames include the audio data frame initially determined to need to embed a digital watermark currently (abbreviated as the current audio data frame). The frame sequence segment can be a sequence segment composed of the current audio data frame and a preset number of consecutive audio data frames after it, or a sequence segment composed of the current audio data frame and a preset number of consecutive audio data frames before it, or a sequence segment composed of the first preset number of consecutive audio data before the current audio data frame and the second preset number of consecutive audio data after it, where the first preset number can be equal to the second preset number.

[0071] Obtain the fixed codebook gain of each audio data frame in the frame sequence segment, and then determine the gain threshold corresponding to the audio data frame according to the multiple fixed codebook gains of the frame sequence segment. For example, use the minimum value of the multiple fixed codebook gains of the frame sequence segment as the gain threshold, or use the minimum value of the multiple fixed codebook gains multiplied by an adjustment coefficient as the gain threshold.

[0072] In an embodiment of the present application, after obtaining the fixed codebook gain of each audio data frame in the frame sequence segment, sort these multiple fixed codebook gains according to their numerical values to form a gain sequence. Then select a fixed codebook gain as the target codebook gain from the gain sequence according to a preset quantity ratio, and this target codebook gain can be used as the gain threshold corresponding to the audio data frame. For example, if the quantity ratio is 1 / N, then at least 1 / N of the fixed codebook gains in the fixed codebook gain sequence will be less than or equal to the corresponding target codebook gain.

[0073] In an embodiment of the present application, after obtaining the fixed codebook gain of each audio data frame in the frame sequence segment, calculate the average value of these multiple fixed codebook gains, and use this average value as the gain threshold corresponding to the audio data frame, or use the average value multiplied by an adjustment coefficient as the gain threshold.

[0074] In an embodiment of the present application, after obtaining the fixed codebook gain of each audio data frame in the frame sequence segment, calculate the median of these multiple fixed codebook gains, and use this median as the gain threshold corresponding to the audio data frame, or use the median multiplied by an adjustment coefficient as the gain threshold.

[0075] Step S520: Screen out the target codebook vector from the fixed codebook vectors of multiple audio data frames according to the gain threshold and the fixed codebook gain, where the fixed codebook gain corresponding to the target codebook vector is less than the gain threshold.

[0076] Specifically, perform screening according to the gain threshold, screen out the fixed codebook gains less than the gain threshold from the fixed codebook gains of all audio data frames, and use the fixed codebook vectors corresponding to such fixed codebook gains as the target codebook vectors. The gain threshold represents the human ear sensitivity threshold. When the fixed codebook gain is less than the gain threshold, it means that the human ear has a lower sensitivity to the audio data frame corresponding to this fixed codebook gain, that is, the human ear can hardly feel the audio signal corresponding to this audio data frame. Use the fixed codebook vectors of such audio data frames as the target codebook vectors for embedding digital watermarks.

[0077] In an embodiment of the present application, the steps of determining the target codebook vector specifically include: numerically comparing the fixed codebook gain of each audio data frame with a gain threshold to screen out candidate audio data frames whose fixed codebook gain is less than the gain threshold; obtaining the number of frame intervals between any two candidate audio data frames; screening out target audio data frames from the candidate audio data frames whose number of frame intervals is greater than a number threshold, and using the fixed codebook vector of the target audio data frame as the target codebook vector.

[0078] Specifically, taking the audio data frames with a fixed codebook gain less than the gain threshold as candidate audio data frames, multiple candidate audio data frames can be obtained. Determine the number of frame intervals between any two candidate audio data frames. When the number of frame intervals is greater than the number threshold, the corresponding candidate audio data frame is used as the target audio data frame, and the fixed codebook vector of the target audio data frame is the target codebook vector. For example, there are 5 frames between the first candidate audio data frame and the second candidate audio data frame, that is, the number of frame intervals is 5. If the number threshold is set to 4, then both the first candidate audio data frame and the second candidate audio data frame are target audio data frames, and the fixed codebook vectors corresponding to the first candidate audio data frame and the second candidate audio data frame are used as the target codebook vectors.

[0079] In an embodiment of the present application, the number threshold can be a preset constant, or can be determined according to the number of candidate audio data frames. For example, multiplying the number of candidate audio data frames by a coefficient (such a coefficient is usually less than 1) to obtain the number threshold. Another example is to set different number thresholds according to the number of candidate audio data frames. Exemplarily, when the number of candidate audio data frames is greater than a first value, the number threshold is set to a first preset value; when the number of candidate audio data frames is greater than a second value, the number threshold is set to a second preset value.

[0080] In the technical solution provided in this embodiment, by comparing the number of frame intervals with the number threshold, the embedding quantity and embedding position of the digital watermark can be controlled, and the flexibility of digital watermark embedding can be improved.

[0081] Step S530: Embed a digital watermark into the target codebook vector.

[0082] Specifically, when embedding the digital watermark, all the codewords in the target codebook vector can be replaced with the digital watermark, or some of the codewords in the target codebook vector can be replaced with the digital watermark.

[0083] In an embodiment of the present application, the structure of the digital watermark is as Figure 6As shown in the figure. The digital watermark includes a watermark entity 620, a watermark synchronization code 610 located before the watermark entity 620, and a watermark verification code 630 located after the watermark entity 620. The watermark entity 620 is the main part of the digital watermark and the main content of the digital watermark. For example, the watermark entity 620 includes watermark numbers 1, 2, 3, and so on. The watermark synchronization code 610, also known as the watermark synchronization header, is located before the watermark entity 620 and is used to identify the starting position of the digital watermark. The watermark verification code 630 is located after the watermark entity 620 and is used to verify the correctness of the watermark entity 620. The watermark verification code 630 can be an error-checking verification code commonly used in the field of data communication, such as a CRC (Cyclic Redundancy Check) verification code. During the extraction process of the digital watermark, the watermark verification code 630 is usually used to verify whether the watermark entity 620 before the watermark verification code 630 is correct; if the watermark verification code 630 passes the verification, it indicates that the watermark entity 620 before the watermark verification code 630 is valid; if the watermark verification code 630 fails the verification, it indicates that the watermark entity 620 before the watermark verification code 630 is invalid, and this situation may indicate that the watermark has been tampered with, forged, etc.

[0084] In an embodiment of the present application, while embedding a digital watermark into an audio bitstream, the embedding position of each watermark can also be recorded to generate a corresponding digital code table. When performing audio parsing or decoding at the decoding end, the corresponding digital watermark can be extracted according to the digital code table.

[0085] In the technical solution provided by the embodiment of the present application, the fixed codebook vector of the audio data frame with low human ear sensitivity is used as the target codebook vector for embedding the digital watermark, further reducing the impact of the digital watermark on the audio quality; at the same time, by extracting the fixed codebook vector of the audio data frame with low human ear sensitivity to embed the digital watermark, it is not necessary to embed the digital watermark in all audio data frames in the audio bitstream, reducing the amount of data for watermark embedding, and thus improving the watermark embedding efficiency.

[0086] Figure 7 It is a schematic flowchart of the audio decoding method provided by an embodiment of the present application. As Figure 7 shown, the audio decoding method provided by the embodiment of the present application at least includes steps S710 to S750, specifically:

[0087] Step S710, a voice coding bitstream.

[0088] Specifically, the voice coding bitstream obtained in step S710 is an audio bitstream obtained by encoding the original audio data, which can refer to the description of step S410 in the foregoing embodiment and will not be elaborated here.

[0089] Step S720, bitstream parsing.

[0090] Specifically, bitstream parsing is used to extract the coding parameters of each audio data frame in the speech coding bitstream. The coding parameters include a fixed codebook vector and a fixed codebook gain. Since the fixed codebook has the characteristics of random noise, the fixed codebook is also called a random codebook, the fixed codebook vector is also called a random codebook vector, and the fixed codebook gain is also called a random codebook gain. The bitstream parsing can refer to the description of step S420 in the foregoing embodiment and will not be elaborated herein.

[0091] Step S730: Determine whether the random codebook gain Gc is less than the gain threshold Thd.

[0092] Specifically, if the random codebook gain Gc of the current audio data frame is less than the gain threshold Thd, it is considered that the audio data frame corresponding to the random codebook gain Gc is an audio signal with low human ear sensitivity. At this time, the random codebook vector corresponding to the random codebook gain Gc can be used as the embedding carrier of the digital watermark. If the random codebook gain Gc is not less than the gain threshold Thd (i.e., the random codebook gain Gc is greater than or equal to the gain threshold Thd), return to step S720 at this time to continue to determine whether the random codebook gain Gc of the next audio data frame meets the condition. The acquisition and setting of the gain threshold Thd can refer to the description of steps S510 to S520 in the foregoing embodiment and will not be elaborated herein.

[0093] Step S740: Replace the random codebook with the digital watermark.

[0094] Specifically, when the random codebook gain Gc is less than the gain threshold Thd, the random codebook vector corresponding to the random codebook gain Gc is used as the target codebook vector, and the digital watermark is embedded in the target codebook vector. Replace all the codewords of the target codebook vector with the digital watermark to complete the embedding of the digital watermark. Optionally, some of the codewords of the target codebook vector can also be replaced with the digital watermark. The structure of the digital watermark can refer to Figure 6 and the relevant description in the foregoing embodiment and will not be elaborated herein.

[0095] Step S750: Generate bitstream repackaging.

[0096] Specifically, when the digital watermark is completed, the new target codebook vector is repackaged into the speech coding bitstream to form a new speech coding bitstream embedded with the digital watermark.

[0097] Figure 8 It is a schematic flowchart of the audio decoding method provided by an embodiment of the present application. As Figure 8 shown, the audio decoding method provided by the embodiment of the present application at least includes steps S810 to S820, and the specific content is as follows:

[0098] Step S810: Analyze the audio bitstream to obtain the encoding parameters carried in the audio bitstream. The encoding parameters include the fixed codebook vectors of each audio data frame in the audio bitstream.

[0099] Specifically, the audio bitstream to be analyzed is an audio bitstream formed through encoding processing and digital watermark embedding, such as the audio bitstream formed by the audio encoding method provided in any embodiment of the present application. The audio bitstream carries encoding parameters, and the encoding parameters include the relevant attribute information of each audio data frame in the audio bitstream, such as the fixed codebook vector.

[0100] Step S820: Extract the digital watermark embedded in the audio bitstream from the fixed codebook vectors.

[0101] Specifically, the digital watermark in the audio bitstream is embedded in the fixed codebook vectors, and the digital watermark is extracted from the fixed codebook vectors during decoding.

[0102] In the technical solution provided by the embodiments of the present application, the fixed codebook vectors used in the audio encoding process are used as the embedding carriers of the digital watermark. The digital watermark can be extracted from the fixed codebook vectors during the audio decoding process, without operating on the audio data frames of the audio bitstream itself, reducing the impact of digital watermark extraction on the audio bitstream signal itself.

[0103] In an embodiment of the present application, the specific steps for extracting the digital watermark from the fixed codebook vectors include: obtaining a gain threshold corresponding to the audio data frame; screening out the target codebook vectors from the fixed codebook vectors of multiple audio data frames according to the gain threshold and the fixed codebook gain, where the fixed codebook gain corresponding to the target codebook vector is less than the gain threshold; and extracting the digital watermark from the target codebook vectors.

[0104] Specifically, the gain threshold corresponding to the obtained audio data frame is the same as the gain threshold set for the audio data frame during encoding. The setting of the gain threshold during encoding can refer to the description in the foregoing encoding method embodiments and will not be elaborated here. Compare the fixed codebook gain of each audio data frame in the audio bitstream with the corresponding gain threshold, and use the fixed codebook vector corresponding to the fixed codebook gain less than the gain threshold as the target codebook vector, and extract the digital watermark from the target codebook vector.

[0105] In an embodiment of the present application, when extracting the digital watermark, all the codewords in the target codebook vector can be extracted as the digital watermark. Optionally, when the digital watermark can be part of the codewords in the target codebook vector, after extracting all the codewords in the target codebook vector, further identification can be performed to extract the part of the codewords corresponding to the digital watermark to obtain the digital watermark.

[0106] In one embodiment of the present application, the steps of extracting a digital watermark from a target codebook vector specifically include: performing a matching detection on each codeword in the target codebook vector to determine a watermark synchronization code embedded in the target codebook vector, a watermark entity, and a watermark verification code located after the watermark synchronization code; performing a correctness verification on the watermark entity according to the watermark verification code; when the verification passes, forming a digital watermark from the watermark synchronization code, the watermark entity, and the watermark verification code.

[0107] Specifically, the structure of the digital watermark can be referred to Figure 6 , which includes a watermark entity 620, a watermark synchronization code 610 located before the watermark entity 620, and a watermark verification code 630 located after the watermark entity 620. The watermark entity 620 is the main content of the digital watermark. The watermark synchronization code 610 is used to identify the starting position of the digital watermark, and the watermark verification code 630 is used to verify the correctness of the watermark entity 620. When extracting the digital watermark, first perform a matching detection on each codeword in the target codebook vector to determine the watermark synchronization code 610 of the digital watermark, and then determine the starting position of the watermark and the watermark entity 620 and the watermark verification code 630 after it from the watermark synchronization code 610. Then, verify the watermark verification code 630 to determine whether the watermark entity 620 is valid. If the watermark verification code 630 passes the verification, it means that the watermark entity 620 before the watermark verification code 630 is valid. At this time, the digital watermark is the structure jointly composed of the watermark synchronization code 610, the watermark entity 620, and the watermark verification code 630. If the watermark verification code 630 fails the verification, it means that the watermark entity 620 before the watermark verification code 630 is invalid, which may indicate abnormal situations such as the watermark being tampered with, forged, or the verification code being incorrect. At this time, an error message can be fed back and an alarm message can be sent to relevant personnel.

[0108] In the technical solution provided by the embodiment of the present application, the digital watermark is set as a structure of a watermark entity, a watermark synchronization code, and a watermark verification code. The watermark position can be quickly determined through the watermark synchronization code, improving the watermark extraction speed during audio decoding. The validity of the watermark entity is verified through the watermark verification code, ensuring the security and reliability of the digital watermark, and thus playing a guarantee role in the security and quality of the audio code stream.

[0109] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0110] Figure 9 It is a schematic flowchart of an audio decoding method provided by an embodiment of the present application. AsFigure 9 As shown in Figure 9 , the audio decoding method provided by the embodiments of the present application at least includes steps S910 to S950, specifically:

[0111] Step S910, obtain the speech coding bitstream.

[0112] Specifically, the audio bitstream obtained in step S910 is an audio bitstream formed by encoding processing and digital watermark embedding, such as the audio bitstream formed by the audio encoding method provided by any embodiment of the present application.

[0113] Step S920, decode and parse the parameters.

[0114] Specifically, the coding parameters are carried in the speech coding bitstream. The coding parameters include the relevant attribute information of each audio data frame in the audio bitstream, such as the fixed codebook vector and the fixed codebook gain. Since the fixed codebook has the characteristics of random noise, the fixed codebook is also called the random codebook, the fixed codebook vector is also called the random codebook vector, and the fixed codebook gain is also called the random codebook gain.

[0115] Step S930, determine whether the random codebook gain Gc is less than the gain threshold Thd.

[0116] Specifically, if the random codebook gain Gc of the current audio data frame is less than the gain threshold Thd, it is determined that the current audio data frame is an audio data frame embedded with a digital watermark. If the random codebook gain Gc is not less than the gain threshold Thd (that is, the random codebook gain Gc is greater than or equal to the gain threshold Thd), at this time, return to step S920 to continue to determine whether the random codebook gain Gc of the next audio data frame meets the condition.

[0117] Step S940, extract the digital watermark.

[0118] Specifically, when the audio data frame embedded with the digital watermark is determined, the digital watermark can be extracted from the random codebook vector of the audio data frame. For the relevant description of digital watermark extraction, reference can be made to the relevant description of step S820 in the foregoing embodiments, and details are not described herein again.

[0119] Step S950, output the speech.

[0120] Specifically, step S950 is after step S920. After step S920 decodes the speech coding bitstream, a playable speech signal can be output.

[0121] The following introduces the device embodiments of the present application, which can be used to execute the audio encoding or decoding method in the above embodiments of the present application.

[0122] Figure 10 The structural block diagram of the audio encoding device provided by the embodiments of the present application is schematically shown.

[0123] As Figure 10 shown, the audio coding device 1000 provided by an embodiment of the present application includes:

[0124] An audio bitstream acquisition module 1010, configured to acquire an audio bitstream obtained by encoding original audio data;

[0125] An audio bitstream parsing module 1020, configured to parse the audio bitstream to obtain encoding parameters carried in the audio bitstream, where the encoding parameters include fixed codebook vectors of respective audio data frames in the audio bitstream;

[0126] A digital watermark embedding module 1030, configured to embed a digital watermark into the fixed codebook vector and encapsulate the fixed codebook vector embedded with the digital watermark into the audio bitstream.

[0127] In an embodiment of the present application, the encoding parameters further include fixed codebook gains corresponding to the fixed codebook vectors; the digital watermark embedding module 1030 includes:

[0128] A gain threshold acquisition unit, configured to acquire a gain threshold corresponding to the audio data frame;

[0129] A target codebook vector determination unit, configured to screen out a target codebook vector from the fixed codebook vectors of multiple audio data frames according to the gain threshold and the fixed codebook gain, where the fixed codebook gain corresponding to the target codebook vector is less than the gain threshold;

[0130] A digital watermark embedding unit, configured to embed a digital watermark into the target codebook vector.

[0131] In an embodiment of the present application, the digital watermark embedding unit is specifically configured to:

[0132] Replace some or all of the codewords in the target codebook vector with digital watermarks.

[0133] In an embodiment of the present application, the target codebook vector determination unit is specifically configured to:

[0134] Perform a numerical comparison between the fixed codebook gains of respective audio data frames and the gain threshold to screen out candidate audio data frames with fixed codebook gains less than the gain threshold;

[0135] Acquire the number of frame intervals between any two candidate audio data frames;

[0136] Screen out target audio data frames with the number of frame intervals greater than a quantity threshold from the candidate audio data frames, and use the fixed codebook vectors of the target audio data frames as the target codebook vectors.

[0137] In one embodiment of the present application, the gain threshold obtaining unit includes:

[0138] A frame sequence segment obtaining subunit, configured to obtain the frame sequence segment where the audio data frame is located, and the frame sequence segment includes the audio data frame and a preset number of other audio data frames adjacent to the audio data frame;

[0139] A fixed codebook gain obtaining subunit, configured to obtain the fixed codebook gain of each audio data frame in the frame sequence segment;

[0140] A gain threshold determining subunit, configured to determine a gain threshold corresponding to the audio data frame according to the fixed codebook gain of each audio data frame.

[0141] In one embodiment of the present application, the gain threshold determining subunit is specifically configured to:

[0142] Sort the fixed codebook gains of each audio data frame according to the numerical value to obtain a gain sequence;

[0143] Select a target codebook gain from the gain sequence as the gain threshold corresponding to the audio data frame according to a preset quantity ratio.

[0144] In one embodiment of the present application, the gain threshold determining subunit is specifically configured to:

[0145] Obtain the average value of the fixed codebook gains of each audio data frame;

[0146] Determine the gain threshold corresponding to the audio data frame according to the average value.

[0147] In one embodiment of the present application, the digital watermark includes a watermark entity, a watermark synchronization code located before the watermark entity, and a watermark verification code located after the watermark entity. The watermark synchronization code is used to identify the starting position of the digital watermark, and the watermark verification code is used to verify the correctness of the watermark entity.

[0148] The specific details of the audio encoding device provided in each embodiment of the present application have been described in detail in the corresponding method embodiment, and will not be repeated here.

[0149] Figure 11 Schematically shows a structural block diagram of the audio decoding device provided in an embodiment of the present application.

[0150] As Figure 11 shown, the audio decoding device 1100 provided in an embodiment of the present application includes:

[0151] An audio bitstream parsing module 1110 is configured to parse an audio bitstream to obtain encoding parameters carried in the audio bitstream, where the encoding parameters include fixed codebook vectors of respective audio data frames in the audio bitstream.

[0152] A digital watermark extraction module 1120 is configured to extract a digital watermark embedded in the audio bitstream from the fixed codebook vectors.

[0153] In an embodiment of the present application, the encoding parameters further include fixed codebook gains corresponding to the fixed codebook vectors; the digital watermark extraction module 1120 includes:

[0154] A gain threshold obtaining unit is configured to obtain a gain threshold corresponding to the audio data frame.

[0155] A target codebook vector determining unit is configured to screen out a target codebook vector from the fixed codebook vectors of multiple audio data frames according to the gain threshold and the fixed codebook gain, where the fixed codebook gain corresponding to the target codebook vector is less than the gain threshold.

[0156] A digital watermark extraction unit is configured to extract a digital watermark from the target codebook vector.

[0157] In an embodiment of the present application, the digital watermark extraction unit is specifically configured to:

[0158] Extract all codewords in the target codebook vector and use the all codewords as the digital watermark.

[0159] In an embodiment of the present application, the digital watermark includes a watermark entity, a watermark synchronization code located before the watermark entity, and a watermark check code located after the watermark entity. The watermark synchronization code is used to identify the start position of the digital watermark, and the watermark check code is used to verify the correctness of the watermark entity.

[0160] In an embodiment of the present application, the digital watermark extraction unit is further configured to:

[0161] Perform a matching detection on each codeword in the target codebook vector to determine the watermark synchronization code, the watermark entity, and the watermark check code embedded in the target codebook vector and located after the watermark synchronization code.

[0162] Verify the correctness of the watermark entity according to the watermark check code.

[0163] When the verification passes, form the digital watermark with the watermark synchronization code, the watermark entity, and the watermark check code.

[0164] The specific details of the audio decoding device provided in each embodiment of this application have been described in detail in the corresponding method embodiments, and will not be repeated here.

[0165] Figure 12 Schematically shows a block diagram of a computer system of an electronic device for implementing the embodiments of this application.

[0166] It should be noted that Figure 12 The computer system 1200 of the electronic device shown is only an example, and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0167] As Figure 12 shown, the computer system 1200 includes a central processing unit 1201 (Central Processing Unit, CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1202 (Read-Only Memory, ROM) or the program loaded from the storage section 1208 into the random access memory 1203 (Random Access Memory, RAM). In the random access memory 1203, various programs and data required for system operation are also stored. The central processing unit 1201, the read-only memory 1202, and the random access memory 1203 are connected to each other via a bus 1204. The input / output interface 1205 (Input / Output interface, i.e., I / O interface) is also connected to the bus 1204.

[0168] The following components are connected to the input / output interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including, for example, a cathode ray tube (Cathode Ray Tube, CRT), a liquid crystal display (Liquid Crystal Display, LCD), etc., and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a local area network card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the input / output interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed, so that the computer program read from it can be installed into the storage section 1208 as needed.

[0169] In particular, according to the embodiments of the present application, the processes described in each method flowchart can be implemented as computer software programs. For example, embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by the central processing unit 1201, various functions defined in the system of the present application are executed.

[0170] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program codes. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program codes contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0171] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the above-mentioned module, segment of a program, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0172] It should be noted that although several modules or units of devices for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0173] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0174] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application.

[0175] It should be understood that the present application is not limited to the exact structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. An audio encoding method, characterized in that, Comprising: Obtaining an audio bitstream obtained by encoding original audio data; Analyzing the audio bitstream to obtain encoding parameters carried in the audio bitstream, where the encoding parameters include fixed codebook vectors of respective audio data frames in the audio bitstream and fixed codebook gains corresponding to the fixed codebook vectors; Obtaining a gain threshold corresponding to the audio data frame; Screening target codebook vectors from the fixed codebook vectors of multiple audio data frames according to the gain threshold and the fixed codebook gains, where the fixed codebook gain corresponding to the target codebook vector is less than the gain threshold; Embedding a digital watermark into the target codebook vector and encapsulating the target codebook vector into the audio bitstream.

2. The audio encoding method according to claim 1, wherein Embedding a digital watermark into the target codebook vector includes: Replacing some or all of the codewords in the target codebook vector with the digital watermark.

3. The audio encoding method according to claim 1, wherein Screening target codebook vectors from the fixed codebook vectors of multiple audio data frames according to the gain threshold and the fixed codebook gains includes: Performing a numerical comparison between the fixed codebook gains of respective audio data frames and the gain threshold to screen candidate audio data frames with fixed codebook gains less than the gain threshold; Obtaining the number of frame intervals between any two candidate audio data frames; Screening target audio data frames with the number of frame intervals greater than a quantity threshold from the candidate audio data frames and using the fixed codebook vectors of the target audio data frames as the target codebook vectors.

4. The audio encoding method according to claim 1, wherein Obtaining a gain threshold corresponding to the audio data frame includes: Obtaining a frame sequence segment where the audio data frame is located, where the frame sequence segment includes the audio data frame and a preset number of other audio data frames adjacent to the audio data frame; Obtaining the fixed codebook gains of respective audio data frames in the frame sequence segment; Determining a gain threshold corresponding to the audio data frame according to the fixed codebook gains of respective audio data frames.

5. The audio encoding method according to claim 4, wherein Determining a gain threshold corresponding to the audio data frame according to the fixed codebook gains of respective audio data frames includes: Sorting the fixed codebook gains of respective audio data frames in ascending order of numerical value to obtain a gain sequence; Selecting a target codebook gain from the gain sequence as the gain threshold corresponding to the audio data frame according to a preset quantity ratio.

6. The audio encoding method according to claim 4, wherein Determining a gain threshold corresponding to the audio data frame according to the fixed codebook gains of respective audio data frames includes: Obtaining the average value of the fixed codebook gains of respective audio data frames; Determining a gain threshold corresponding to the audio data frame according to the average value.

7. The audio encoding method according to any one of claims 1 to 6, characterized in that, The digital watermark includes a watermark entity, a watermark synchronization code located before the watermark entity, and a watermark verification code located after the watermark entity. The watermark synchronization code is used to identify the starting position of the digital watermark, and the watermark verification code is used to verify the correctness of the watermark entity.

8. An audio decoding method, characterized in that, Comprising: Analyzing an audio bitstream to obtain encoding parameters carried in the audio bitstream, where the encoding parameters include fixed codebook vectors of respective audio data frames in the audio bitstream and fixed codebook gains corresponding to the fixed codebook vectors; Obtain a gain threshold corresponding to the audio data frame; Filter a target codebook vector from the fixed codebook vectors of multiple audio data frames according to the gain threshold and the fixed codebook gain, where the fixed codebook gain corresponding to the target codebook vector is less than the gain threshold; Extract a digital watermark from the target codebook vector.

9. The audio decoding method according to claim 8, wherein Extracting a digital watermark from the target codebook vector includes: Extract all the codewords in the target codebook vector and use the all codewords as the digital watermark.

10. The audio decoding method according to claim 8, wherein The digital watermark includes a watermark entity, a watermark synchronization code located before the watermark entity, and a watermark verification code located after the watermark entity. The watermark synchronization code is used to identify the start position of the digital watermark, and the watermark verification code is used to verify the correctness of the watermark entity.

11. The audio decoding method according to claim 10, characterized in that, Extracting a digital watermark from the target codebook vector includes: Perform matching detection on each codeword in the target codebook vector to determine the watermark synchronization code embedded in the target codebook vector, and the watermark entity and watermark verification code located after the watermark synchronization code; Verify the correctness of the watermark entity according to the watermark verification code; When the verification passes, form a digital watermark with the watermark synchronization code, the watermark entity, and the watermark verification code.

12. An audio encoding device, characterized in that, Includes: An audio bitstream acquisition module, configured to acquire an audio bitstream obtained by encoding original audio data; An audio bitstream parsing module, configured to parse the audio bitstream to obtain the encoding parameters carried in the audio bitstream, where the encoding parameters include the fixed codebook vectors of each audio data frame in the audio bitstream and the fixed codebook gain corresponding to the fixed codebook vector; A digital watermark embedding module, configured to obtain a gain threshold corresponding to the audio data frame; filter a target codebook vector from the fixed codebook vectors of multiple audio data frames according to the gain threshold and the fixed codebook gain, where the fixed codebook gain corresponding to the target codebook vector is less than the gain threshold; embed a digital watermark into the target codebook vector, and encapsulate the target codebook vector into the audio bitstream.

13. An audio decoding device, characterized in that, Includes: An audio bitstream parsing module, configured to parse an audio bitstream to obtain the encoding parameters carried in the audio bitstream, where the encoding parameters include the fixed codebook vectors of each audio data frame in the audio bitstream and the fixed codebook gain corresponding to the fixed codebook vector; A digital watermark extraction module, configured to obtain a gain threshold corresponding to the audio data frame; filter a target codebook vector from the fixed codebook vectors of multiple audio data frames according to the gain threshold and the fixed codebook gain, where the fixed codebook gain corresponding to the target codebook vector is less than the gain threshold; extract a digital watermark from the target codebook vector.

14. A computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

15. An electronic device, characterized in that, Includes: A processor; And A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the method according to any one of claims 1 to 11 by executing the executable instructions.

16. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Devices for adaptively encoding and decoding a watermarked signal

    US20120203561A1