Watermark processing method and apparatus
Through the method of expanding audio clips, the problems of small capacity and poor robustness of audio watermark embedding and extraction in the prior art are solved, and the watermarks with larger bits are embedded and extracted in the audio are realized, ensuring that the sound quality and content of the audio are unchanged.
Patent Information
- Application Number
- PCT/CN2025/070748
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-15
- Filing Date
- 2025-01-06
- Publication Date
- 2025-07-24
AI Technical Summary
In the prior art, the embedding and extraction of audio watermarks has problems such as small capacity, great influence on sound quality, and poor robustness. Especially in the real-time audio stream, the watermark frame length is short, making it difficult to embed and extract larger bits of watermarks.
By extending one or more watermark audio clips in the audio clip or watermark audio before the audio clip, the length of the audio clip is expanded, and the user's watermark is embedded in the extended audio to achieve larger bits of watermark embedding and extraction.
Without affecting the audio content, watermark embedding and extraction of larger bits is achieved, improving the robustness of the watermark and reducing the impact on the invisibility of the audio.
Smart Images

Figure CN2025070748_24072025_PF_FP_ABST
Abstract
Description
Watermark processing method and device
[0001] This application claims priority to Chinese patent application number 202410060353.4, filed on January 15, 2024, entitled “Watermark Processing Method and Device,” the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present application relates to the field of media technology, and more specifically, to a watermark processing method and apparatus in the field of media technology. Background Art
[0003] With the rapid development of media technology and its widespread use on the Internet, it is very important to protect the copyright of media files and trace the source of media files. Therefore, watermark technology has received great attention and has become a hot topic of research in the international academic community.
[0004] The so-called audio watermark is a mark with specific meaning, that is, a watermark. The watermark can be hidden in the audio using the data embedding method to identify the source or copyright of the audio. Accordingly, the watermark can be extracted from the watermarked audio to trace the source of the audio or obtain the copyright of the audio.
[0005] The larger the watermark embedded in the audio, the more likely it is to affect the audio quality or content. To avoid this impact, the watermark size is typically smaller. When embedding watermarks in real-time audio streams, existing audio watermarking techniques embed watermarks within audio frames. However, due to the short duration of audio frames, the size of the watermark embedded within an audio frame is even more limited. Summary of the Invention
[0006] The present application provides a watermark processing method and apparatus, which can embed or extract a larger bit watermark in audio.
[0007] In a first aspect, an embodiment of the present application provides a watermark processing method, which may include: obtaining audio and a user watermark, the audio including a first audio segment; embedding the user watermark into the audio to obtain watermarked audio, the watermarked audio including a first watermarked audio segment, the first watermarked audio segment being obtained based at least on the user watermark, the first audio segment and an extended audio segment, and the extended audio segment may not include the audio and the watermarked audio.
[0008] In one possible implementation, the first watermark audio segment is obtained based on at least the user watermark, the first audio segment and the extended audio segment. It can be understood that: the first watermark audio segment is obtained based on embedding the user watermark into the intermediate audio segment, and the intermediate audio segment includes the first audio segment and the extended audio segment.
[0009] In a possible implementation, the audio segment involved in this application may be one or more audio frames.
[0010] Through the watermark processing method of the present application, it is possible to embed a user watermark in audio to identify the source or copyright of the audio. In the prior art, the watermark audio frame is obtained by embedding the user watermark into the audio frame, and the length of the audio frame is short, so the watermark that can be embedded in the audio is small. In the embodiment of the present application, the first watermark audio segment is obtained based on at least the user watermark, the first audio segment and the extended audio segment. The extended audio segment does not include the first watermark audio segment and the first audio segment, that is, the first audio segment is extended by the extended audio segment. The first watermark audio segment is obtained based on embedding the user watermark into the extended audio segment. Since the length of the extended audio segment is larger than the audio frame in the prior art, when the bit required to embed the user watermark is large, the strength of the user watermark embedded in the audio segment is reduced to avoid affecting the audio segment, thereby achieving the ability to embed a larger bit watermark in the audio segment.
[0011] It should be noted that the first audio segment is any audio segment in the audio.
[0012] Optionally, the audio may be a real-time audio stream (such as call audio) or may be locally stored audio (such as music), which is not limited in the embodiments of the present application.
[0013] It should be noted that the first audio segment in the audio corresponds to the first watermarked audio segment in the watermarked audio, that is, the first audio segment has the same content as the first watermarked audio segment.
[0014] Using the watermark processing method provided in the embodiment of the present application, since the first watermark audio segment corresponds to the first audio segment in the intermediate watermark audio segment, that is, after the user watermark is embedded in the first audio segment to obtain the first watermark audio segment, the content of the audio does not change. Therefore, the invisibility of the user watermark to the audio can be guaranteed as much as possible.
[0015] In one possible implementation, the above-mentioned embedding the user watermark into the audio to obtain the watermarked audio may include: obtaining the first audio segment; embedding the user watermark into the intermediate audio segment to obtain the intermediate watermarked audio segment, the intermediate audio segment including the first audio segment and the extended audio segment; obtaining the first watermarked audio segment based on the intermediate watermarked audio segment, the first watermarked audio segment corresponding to the first audio segment in the intermediate watermarked audio segment, and the watermarked audio including the first watermarked audio segment.
[0016] Optionally, this embodiment of the present application does not limit the length of the intermediate audio segment.
[0017] In a possible implementation, the length of the intermediate audio segment may be greater than or equal to the length of (N+1) audio segments.
[0018] Optionally, the length of the intermediate audio segment may be preset.
[0019] Optionally, the present application does not limit the order of the first audio segment and the extended audio segment in the intermediate audio segment.
[0020] In a possible implementation, the extended audio segment may be located before the first audio segment.
[0021] It should be noted that the extended audio segment described in the embodiment of the present application does not include the audio and the watermarked audio.
[0022] It should also be noted that the extended audio segment is used to extend the length of the audio segment to embed a larger watermark.
[0023] Optionally, this application does not limit the method for generating the extended audio segment.
[0024] In a possible implementation, the extended audio segment may be generated randomly or based on a matrix.
[0025] In another possible implementation, the extended audio segment may be generated based on an all-zero matrix. That is, the extended audio segment may be a blank audio segment or a silent audio segment to reduce interference with the first audio segment.
[0026] In a second aspect, an embodiment of the present application also provides a watermark processing method, which may include: obtaining watermark audio, the watermark audio including a first watermark audio segment, the first watermark audio segment being obtained based at least on the user watermark, the first audio segment and the extended audio segment, the first audio segment corresponding to the first watermark audio segment, the user watermark is not embedded in the first audio segment, and the extended audio segment does not include the first audio segment and the watermark audio; performing watermark extraction on the watermark audio to obtain the user watermark.
[0027] It should be noted that the first audio segment is any audio segment in the audio.
[0028] In the prior art, user watermarks are extracted from watermarked audio frames. However, audio frames are short, so the watermark that can be extracted from the watermarked audio frames is relatively small. In the embodiments of the present application, however, the user watermark is extracted from a first watermarked audio segment. The first watermarked audio segment is derived based on at least the user watermark, the first audio segment, and an extended audio segment. The extended audio segment does not include the first watermarked audio segment and the first audio segment. That is, the first audio segment can be extended by the extended audio segment, and the user watermark is extracted from the extended audio segment. Therefore, because the length of the extended audio segment is longer than the audio frame in the prior art, a larger watermark can be extracted from the watermarked audio frame.
[0029] It should also be noted that the first audio segment in the audio corresponds to the first watermarked audio segment in the watermarked audio, that is, the first audio segment has the same content as the first watermarked audio segment.
[0030] Using the watermark processing method provided in the embodiment of the present application, since the first watermark audio segment corresponds to the first audio segment in the intermediate watermark audio segment, that is, after the user watermark is embedded in the first audio segment to obtain the first watermark audio segment, the content of the audio does not change. Therefore, the invisibility of the user watermark to the audio can be guaranteed as much as possible.
[0031] In a third aspect, an embodiment of the present application also provides a watermark processing method, which may include: obtaining audio and a user watermark, the audio including N audio segments, N being greater than or equal to 2; embedding the user watermark into the audio to obtain watermarked audio, the watermarked audio including N watermarked audio segments, the Nth watermarked audio segment being obtained based at least on the user watermark, the Nth audio segment, and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment.
[0032] Alternatively, the watermark processing method may include: obtaining audio and a user watermark; and embedding the user watermark into the audio to obtain watermarked audio, wherein the watermarked audio includes N watermarked audio segments, where N is greater than or equal to 2, and each watermarked audio segment is embedded with one or more user watermarks. The strength of the user watermarks embedded in different audio segments may be the same or different. When a large number of bits of user watermarks need to be embedded, the strength of the user watermark embedded in the audio segment is reduced to avoid affecting the audio segment, thereby enabling the embedding of a larger bit watermark in the audio segment.
[0033] In the prior art, watermarked audio frames are obtained by embedding a user watermark into an audio frame. However, since the length of an audio frame is short, the watermark that can be embedded in the audio is relatively small. In the embodiments of the present application, however, the Nth watermarked audio segment is obtained based on at least the user watermark, the Nth audio segment, and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment. That is, the Nth audio segment is extended by one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment. The first watermarked audio segment is obtained by embedding the user watermark into the extended audio segment. Since the length of the extended audio segment is longer than the audio frame in the prior art, a larger watermark can be embedded in the audio segment.
[0034] It should be noted that the N audio clips in the audio are arranged in sequence.
[0035] Optionally, the audio may be a real-time audio stream (such as call audio) or may be locally stored audio (such as music), which is not limited in the embodiments of the present application.
[0036] Optionally, the present application does not limit the number of the one or more watermarked audio segments.
[0037] In a possible implementation, the one or more watermarked audio segments may include all watermarked audio segments before the Nth watermarked audio segment in the watermarked audio.
[0038] In another possible implementation, the number of the one or more watermarked audio segments may be preset.
[0039] It should be noted that the N audio clips include the N-th audio clip, the N watermark audio clips include the N-th watermark audio clip, and the N-th audio clip in the N audio clips corresponds to the N-th watermark audio clip in the N watermark audio clips, that is, the content of the N-th audio clip is the same as that of the N-th watermark audio clip.
[0040] Using the watermark processing method provided in the embodiment of the present application, since the Nth watermarked audio segment corresponds to the Nth audio segment, that is, after the user watermark is embedded in the Nth audio segment to obtain the Nth watermarked audio segment, the content of the audio does not change. Therefore, the invisibility of the user watermark to the audio can be guaranteed as much as possible.
[0041] In one possible implementation, one or more user watermarks are embedded in each of the first watermarked audio segment and the Nth watermarked audio segment. The strength of the user watermarks embedded in different audio segments can be the same or different. Alternatively, the first watermarked audio segment and the Nth watermarked audio segment have different audio content, and the strength of the embedded user watermarks is different. Alternatively, the first watermarked audio segment and the Nth watermarked audio segment can be watermarked audio frames, and one or more user watermarks are embedded in a single watermarked audio frame. The strength of the user watermarks embedded in different audio frames can be the same or different.
[0042] It should be noted that the first audio segment may be the first audio segment among the N audio segments, and the first watermarked audio segment may be the first watermarked audio segment among the N watermarked audio segments.
[0043] Optionally, the Nth watermarked audio segment is obtained based on an extended audio segment, the user watermark and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment, and the extended audio segment does not include the audio and the watermarked audio.
[0044] It should be noted that the extended audio segment described in the embodiment of the present application does not include the audio and the watermarked audio.
[0045] It should also be noted that the extended audio segment is used to extend the length of the audio segment to embed a larger watermark.
[0046] Optionally, this application does not limit the method for generating the extended audio segment.
[0047] In a possible implementation, the extended audio segment may be generated randomly or based on a matrix.
[0048] In another possible implementation, the extended audio segment may be generated based on an all-zero matrix. That is, the extended audio segment may be a blank audio segment or a silent audio segment to reduce interference with the first audio segment.
[0049] In a possible implementation, the values of the sampling points corresponding to the blank audio segments or silent audio segments described in this application are all zero.
[0050] In one possible implementation, the Nth watermarked audio segment is obtained based on at least the user watermark, the Nth audio segment and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment. It can be understood that: the Nth watermarked audio segment is obtained based on the user watermark embedded in the first intermediate audio segment, and the first intermediate audio segment includes at least the Nth audio segment and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment.
[0051] In one possible implementation, embedding the user watermark into the audio to obtain the watermarked audio may include: obtaining the Nth audio segment and one or more watermarked audio segments; embedding the user watermark into the first intermediate audio segment to obtain the first intermediate watermarked audio segment, the first intermediate audio segment at least including the Nth audio segment and the one or more watermarked audio segments; obtaining the Nth watermarked audio segment based on the first intermediate watermarked audio segment, the Nth watermarked audio segment corresponding to the Nth audio segment in the first intermediate watermarked audio segment.
[0052] Optionally, the first intermediate audio segment further includes the extended audio segment. That is, the Nth watermarked audio segment is obtained based on the extended audio segment, the user watermark, the Nth audio segment, and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment.
[0053] Optionally, the present application does not limit the order of the extended audio segment, the Nth audio segment, and one or more watermark audio segments before the Nth watermark audio segment in the first intermediate audio segment.
[0054] In one possible implementation, within the first intermediate audio segment, the extended audio segment may be located before one or more watermarked audio segments preceding the Nth watermarked audio segment. The one or more watermarked audio segments preceding the Nth watermarked audio segment may also be located before the Nth audio segment. The first watermarked audio segment of the one or more watermarked audio segments preceding the Nth watermarked audio segment is adjacent to the extended audio segment, and the last watermarked audio segment of the one or more watermarked audio segments preceding the Nth watermarked audio segment is adjacent to the Nth audio segment. Alternatively, within the first intermediate audio segment, the N-1th watermarked audio segment is adjacent to the Nth audio segment. The N-1th audio segment corresponding to the N-1th watermarked audio segment is adjacent to the Nth audio segment in the audio. Adjacent audio segments can be superimposed, thereby increasing the success rate of subsequent watermark extraction from the watermarked audio.
[0055] In one possible implementation, when N=2, the first watermarked audio segment is the first watermarked audio segment in the watermarked audio, that is, there is no watermarked audio segment before the first watermarked audio segment in the watermarked audio. Therefore, the first watermarked audio segment is obtained based on at least the user watermark and the first audio segment. In other words, the first watermarked audio segment is obtained by embedding the user watermark into the first audio segment.
[0056] Optionally, when N=2, the first watermarked audio segment is obtained based on the extended audio segment, the user watermark, and the first audio segment. That is, the first watermarked audio segment is obtained by embedding the user watermark into a second intermediate audio segment, where the second intermediate audio segment includes the first audio segment and the extended audio segment.
[0057] In one possible implementation, embedding the user watermark into the audio to obtain the watermarked audio may include: obtaining a first audio segment; embedding the user watermark into a second intermediate audio segment, the second intermediate audio segment including the first audio segment and the extended audio segment, to obtain a second intermediate watermarked audio segment; obtaining the first watermarked audio segment based on the second intermediate watermarked audio segment, the first watermarked audio segment corresponding to the first audio segment in the second intermediate watermarked audio segment.
[0058] Optionally, this embodiment of the present application does not limit the length of the intermediate audio segment.
[0059] In a possible implementation, the length of the intermediate audio segment may be greater than or equal to the length of (N+1) audio segments.
[0060] Optionally, the length of the intermediate audio segment may be preset.
[0061] Optionally, embedding the user watermark into the audio to obtain the watermarked audio may further include: generating the watermarked audio according to the N watermarked audio segments.
[0062] Since the energy of the watermark in a single audio segment is limited, multiple watermarked audio segments can enhance the strength of the watermark signal in the watermarked audio to improve the success rate of extracting the watermark from the watermarked audio.
[0063] In a fourth aspect, an embodiment of the present application also provides a watermark processing method, which may include: obtaining watermark audio, the watermark audio including N watermark audio segments, the Nth watermark audio segment being obtained based on at least the user watermark, the Nth audio segment and one or more watermark audio segments in the watermark audio that precede the Nth watermark audio segment, the Nth audio segment corresponding to the Nth watermark audio segment, the user watermark is not embedded in the Nth audio segment, and N is greater than or equal to 2; performing watermark extraction on the watermark audio to obtain the user watermark.
[0064] In the prior art, user watermarks are extracted from watermarked audio frames. However, audio frames are short, so the watermark that can be extracted from the watermarked audio frames is relatively small. However, in the embodiments of the present application, the user watermark is extracted from the first watermarked audio segment, and the Nth watermarked audio segment is derived based on at least the user watermark, the Nth audio segment, and one or more watermarked audio segments preceding the Nth watermarked audio segment in the watermarked audio. That is, the first audio segment can be extended using one or more watermarked audio segments preceding the Nth watermarked audio segment, and the user watermark is extracted from the extended audio segment. Since the length of the extended audio segment is longer than the audio frame in the prior art, a larger watermark can be extracted from the watermarked audio frame.
[0065] In a possible implementation, the N watermarked audio segments are obtained based on a first length, a first step length, and the watermarked audio, and the first step length is smaller than the first length.
[0066] Optionally, when N>1, the N watermark audio segments may partially overlap or not overlap, which is not limited in this embodiment of the present application.
[0067] Using the watermark processing method provided in the embodiment of the present application, the N watermark audio segments partially overlap, so that more watermark audio segments can be obtained from the watermark audio. The more watermark audio segments there are, the higher the accuracy of watermark extraction.
[0068] In one possible implementation, extracting the watermark from the watermarked audio to obtain the user watermark may include: extracting the watermark from each of the N watermarked audio segments to obtain the user watermark extracted from each watermarked audio segment; screening at least one target watermarked audio segment from the multiple watermarked audio segments based on a correlation value of a template sequence of the user watermark extracted from each watermarked audio segment, the at least one target watermarked audio segment being at least one watermarked audio segment in the multiple watermarked audio segments whose correlation value of the template sequence of the user watermark is greater than a preset first threshold; and determining the user watermark based on a normalized correlation value of each target watermarked audio segment in the at least one target watermarked audio segment.
[0069] Using the watermark audio processing method provided in the embodiments of the present application, since a larger correlation value of the template sequence of the user watermark extracted from a watermarked audio segment indicates that the watermarked audio segment is closer to the starting position of the user watermark embedded in the watermarked audio segment, the accuracy of watermark extraction can be improved by selecting the at least one target watermarked audio segment from multiple watermarked audio segments and determining the user watermark based on the normalized correlation value of the at least one target audio segment.
[0070] In a fifth aspect, the embodiment of the present application further provides a watermark processing device, which is used to implement the methods described in the above aspects or any possible implementation thereof, and the device includes a unit for implementing the methods described in the above aspects or any possible implementation thereof.
[0071] In a sixth aspect, an embodiment of the present application further provides a watermark processing device, which includes a processor and a communication interface, the processor and the communication interface are coupled, and the processor is used to execute the methods described in the above aspects or any possible implementation thereof.
[0072] In a seventh aspect, the present application also provides a computer-readable storage medium for storing a computer program, characterized in that the computer program includes instructions for implementing the methods described in the above aspects or any possible implementation thereof.
[0073] In an eighth aspect, the present application also provides a computer-readable storage medium storing watermarked audio, wherein the watermarked audio includes a first watermarked audio segment, which is obtained based on at least the user watermark, the first audio segment and the extended audio segment, the first audio segment corresponds to the first watermarked audio segment, the user watermark is not embedded in the first audio segment, and the extended audio segment does not include the first audio segment and the watermarked audio.
[0074] In the ninth aspect, the present application also provides a computer-readable storage medium, which stores watermarked audio, and the watermarked audio includes N watermarked audio segments, and the Nth watermarked audio segment is obtained based on at least the user watermark, the Nth audio segment and one or more watermarked audio segments in the watermarked audio before the Nth watermarked audio segment, and N is greater than or equal to 2.
[0075] In the tenth aspect, the present application also provides a computer program product, which contains instructions, characterized in that when the instructions are executed on a computer or a processor, the computer or the processor implements the methods described in the above aspects or any possible implementation thereof.
[0076] The watermark processing device, watermark processing system, computer storage medium and computer program product provided in the embodiments of the present application are all used to execute the watermark processing method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the watermark processing method provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] FIG1 is a schematic architecture diagram of a watermark processing system 100 provided in an embodiment of the present application;
[0078] FIG2 is a schematic flow chart of a watermark processing method 200 provided in an embodiment of the present application;
[0079] FIG3 is a schematic flow chart of a watermark processing method 300 provided in an embodiment of the present application;
[0080] FIG4 is a schematic flow chart of a watermark processing method 400 provided in an embodiment of the present application;
[0081] FIG5 illustrates the watermark processing process of the audio frame S1 provided in an embodiment of the present application;
[0082] FIG6 illustrates the watermark processing process of the audio frame S2 provided in an embodiment of the present application;
[0083] FIG7 illustrates the watermark processing process of the audio frame S3 provided in an embodiment of the present application;
[0084] FIG8 is a schematic flow chart of a watermark processing method 500 provided in an embodiment of the present application;
[0085] FIG9 is a schematic block diagram of a watermark processing apparatus 600 according to an embodiment of the present application;
[0086] FIG10 is a schematic block diagram of a watermark processing apparatus 700 according to an embodiment of the present application;
[0087] FIG11 is a schematic block diagram of a watermark processing apparatus 800 according to an embodiment of the present application;
[0088] FIG12 is a schematic block diagram of a watermark processing apparatus 900 according to an embodiment of the present application;
[0089] FIG13 is a schematic block diagram of a watermark processing apparatus 1000 according to an embodiment of the present application;
[0090] FIG14 is a schematic block diagram of a watermark processing apparatus 1100 according to an embodiment of the present application. DETAILED DESCRIPTION
[0091] First, some professional terms involved in the embodiments of this application are introduced.
[0092] 1. Audio watermarking technology
[0093] Audio watermarking technology can be divided into the watermark embedding stage and the watermark extraction stage. In the watermark embedding stage: Using a watermark embedding algorithm, the watermark is embedded into the audio (such as WAV, MP3, AVI), obtaining the watermarked audio. The audio includes M audio frames, and the watermarked audio includes M watermarked audio frames. The Mth watermarked audio frame is obtained by embedding the watermark into the Mth audio frame. In the watermark extraction stage: Using a watermark extraction algorithm, the watermark is extracted from the watermarked audio.
[0094] 2. Problems with existing audio watermarking technology
[0095] The larger the watermark embedded in the audio, the more likely it is to affect the audio quality or content. To avoid this impact, the watermark size is typically smaller. When embedding watermarks in real-time audio streams, existing audio watermarking techniques embed watermarks within audio frames. However, due to the short duration of audio frames, the size of the watermark embedded within an audio frame is even more limited.
[0096] In addition, the watermarked audio may be subjected to desynchronization attacks during playback, resulting in the watermarked audio being out of sync before and after the attack. For example, under the influence of factors such as the environment, network speed, and compression rate, the watermark cannot be correctly extracted. Therefore, the robustness of the watermark is poor.
[0097] An embodiment of the present application provides a watermark processing method and device, which extends the length of an audio segment or one or more watermark audio segments in the watermark audio that precede the watermark audio segment corresponding to the audio segment, and embeds a user watermark into the extended audio. The watermark audio segment is obtained based on the user watermark embedded in the extended audio segment. Since the extended audio segment has a longer duration, the watermark energy is dispersed during the watermark embedding process, reducing the energy of the watermark embedded in a single audio segment. Therefore, the user watermark that can be embedded in the audio segment becomes larger.
[0098] It should be noted that the watermark processing method of the present application can also be used in the video field. It is only necessary to adjust the audio involved in the following embodiments to video, which can also solve the problem of the embedded watermark being too large affecting the video content.
[0099] The following first introduces the watermark processing system used by the watermark processing method and device provided in the embodiments of the present application.
[0100] FIG1 shows a schematic architecture diagram of a watermark processing system 100 provided in an embodiment of the present application. As shown in FIG1 , the system 100 may include: a watermark embedding device 110 and a watermark extraction device 120 .
[0101] The watermark embedding device 110 is used to embed the watermark into the audio in the watermark embedding stage by using the watermark processing method provided by the present application to obtain the watermarked audio.
[0102] The watermark extraction device 120 is used to extract the watermark from the watermark audio stream using the watermark processing method provided in the embodiment of the present application during the watermark extraction stage.
[0103] It should be noted that the watermark described in the embodiment of the present application is used to identify the source or copyright of the audio, and the embodiment of the present application does not limit the specific content of the watermark.
[0104] Optionally, the watermark may include at least one of the occurrence time, occurrence location, subject, and number of the audio.
[0105] Optionally, the embodiment of the present application limits the specific forms of the watermark embedding device 110 and the watermark extraction device 120.
[0106] In a possible implementation, the watermark embedding device 110 and the watermark extraction device 120 may be (or be integrated into) the same processor.
[0107] In another possible implementation, the watermark embedding device 110 and the watermark extraction device 120 may be (integrated into) different processors respectively.
[0108] Optionally, the system 100 is applicable to various application scenarios where the source or copyright of audio is marked by watermarks. For example, the system 100 can be applied to identify the source or copyright of audio such as conferences, calls, voice, and music by watermarks.
[0109] For example, during a call, voice, conference, or music recording, the watermark processing method of the present application can embed a watermark into the received audio stream in real time.
[0110] For example, the watermark processing method of the present application can also embed a watermark into a locally stored audio file.
[0111] For example, a watermark is embedded in conference audio A to obtain watermarked conference audio A. This watermark is used to identify the source of conference audio A. For example, the watermark may include information such as the conference number, time, location, and topic. Therefore, users can trace the source of conference audio A by extracting the watermark from watermarked conference audio A.
[0112] For example, a watermark is embedded in music B to obtain watermarked music B. The watermark is used to identify the copyright of music B. For example, the watermark may include the copyright owner of the music, the time of first release, etc. Therefore, the user can obtain the copyright of music B by extracting the watermark from watermarked music B.
[0113] The watermark processing method provided by the embodiments of the present application will be further introduced below.
[0114] FIG2 shows a schematic flow chart of a watermark processing method 200 provided in an embodiment of the present application. The method 200 can be used in the system 100 shown in FIG2 . As shown in FIG2 , the method 200 can include the following steps. It should be noted that the steps listed below can be executed in various orders and / or occur simultaneously, and are not limited to the execution order shown in FIG2 .
[0115] S201. The watermark embedding device obtains audio and a user watermark, where the audio includes a first audio segment.
[0116] It should be noted that the first audio segment is any audio segment in the audio.
[0117] It should also be noted that the embodiment of the present application does not limit the length of the first audio segment.
[0118] For example, the length of the first audio segment may be equal to the length of one or more audio frames.
[0119] It should also be noted that the embodiment of the present application does not limit the length of the audio frame.
[0120] For example, the length of the audio frame may be 10 ms.
[0121] Optionally, the user watermark can be used to identify the copyright or source of the audio.
[0122] Optionally, the audio may be a real-time audio stream (such as call audio) or may be locally stored audio (such as music), which is not limited in the embodiments of the present application.
[0123] Optionally, the embodiment of the present application does not limit the manner in which the watermark embedding device obtains the audio.
[0124] In a possible implementation manner, the watermark embedding device may receive the audio from an audio generating device.
[0125] In another possible implementation, the watermark embedding device may obtain the audio locally.
[0126] S202. The watermark embedding device embeds the user watermark into the audio to obtain a watermarked audio, where the watermarked audio includes a first watermarked audio segment, which is obtained based on at least the user watermark, the first audio segment, and the extended audio segment.
[0127] It should be noted that the first audio segment in the audio corresponds to the first watermarked audio segment in the watermarked audio, that is, the first audio segment has the same content as the first watermarked audio segment.
[0128] For example, taking the audio including 3 audio clips and the watermark audio including 3 watermark audio clips as an example, the above 3 audio clips are, in order: the first audio clip, the second audio clip and the third audio clip, and the above 3 watermark audio clips are, in order: the first watermark audio clip, the second watermark audio clip and the third watermark audio clip, wherein the first audio clip is the same as the content in the first watermark audio clip, the second audio clip is the same as the content in the second watermark audio clip, and the third audio clip is the same as the content in the third watermark audio clip.
[0129] In one possible implementation, the first watermarked audio segment described in S202 is obtained based at least on the user watermark, the first audio segment, and the extended audio segment. This can be understood as follows: the first watermarked audio segment is obtained based on embedding the user watermark into an intermediate audio segment, and the intermediate audio segment includes the first audio segment and the extended audio segment.
[0130] In one possible implementation, S202 may include: obtaining the first audio segment; embedding the user watermark into the intermediate audio segment to obtain the intermediate watermark audio segment, the intermediate audio segment including the first audio segment and the extended audio segment; obtaining the first watermark audio segment based on the intermediate watermark audio segment, the first watermark audio segment corresponding to the first audio segment in the intermediate watermark audio segment, and the watermark audio including the first watermark audio segment.
[0131] Optionally, the above-mentioned embedding the user watermark into the intermediate audio segment to obtain the intermediate watermark audio segment may include: performing Fourier transform on the intermediate audio segment to obtain a Fourier coefficient amplitude spectrum; performing spread spectrum encoding on the user watermark to obtain a coded watermark; and embedding the coded watermark into the Fourier coefficient amplitude spectrum to obtain the intermediate watermark audio segment.
[0132] Optionally, this embodiment of the present application does not limit the length of the intermediate audio segment.
[0133] In a possible implementation, the length of the intermediate audio segment may be greater than or equal to the length of (N+1) audio segments.
[0134] Optionally, the length of the intermediate audio segment may be preset.
[0135] Optionally, the present application does not limit the order of the first audio segment and the extended audio segment in the intermediate audio segment.
[0136] In a possible implementation, the extended audio segment may be located before the first audio segment.
[0137] It should be noted that the extended audio segment described in the embodiment of the present application does not include the audio and the watermarked audio.
[0138] It should also be noted that the extended audio segment is used to extend the length of the audio segment to embed a larger watermark.
[0139] Optionally, this application does not limit the method for generating the extended audio segment.
[0140] In a possible implementation, the extended audio segment may be generated randomly or based on a matrix.
[0141] In another possible implementation, the extended audio segment may be generated based on an all-zero matrix. That is, the extended audio segment may be a blank audio segment or a silent audio segment to reduce interference with the first audio segment.
[0142] In a possible implementation, the values of the sampling points corresponding to the blank audio segments or silent audio segments described in this application are all zero.
[0143] In the prior art, a watermarked audio frame is obtained by embedding a user watermark into an audio frame. However, the length of an audio frame is relatively short, so the watermark that can be embedded in the audio is relatively small. In the embodiment of the present application, however, the first watermarked audio segment is obtained based on at least the user watermark, the first audio segment, and an extended audio segment. The extended audio segment does not include the first watermarked audio segment and the first audio segment. That is, the first audio segment is extended by the extended audio segment. The first watermarked audio segment is obtained by embedding the user watermark into the extended audio segment. Since the length of the extended audio segment is longer than the audio frame in the prior art, when a larger number of bits of user watermark need to be embedded, the strength of the user watermark embedded in the audio segment is reduced, thereby avoiding any impact on the audio segment and enabling the embedding of a larger bit watermark in the audio segment.
[0144] In addition, the first watermarked audio segment corresponds to the first audio segment, that is, after the user watermark is embedded in the first audio segment to obtain the first watermarked audio segment, the content of the audio does not change. Therefore, the invisibility of the user watermark to the audio can be guaranteed as much as possible.
[0145] FIG3 shows a schematic flow chart of a watermark processing method 300 provided in an embodiment of the present application. This method 300 can be used in the system 100 shown in FIG2 . As shown in FIG3 , this method 300 can include the following steps. It should be noted that the following steps can be performed in various orders and / or occur simultaneously, and are not limited to the execution order shown in FIG3 .
[0146] It should also be noted that, to avoid repetition, for the parts of the method 300 that are not described in detail, reference can be made to the description of the corresponding parts of the above-mentioned method 200.
[0147] S301. The watermark extraction device obtains watermarked audio, which includes a first watermarked audio segment. The first watermarked audio segment is obtained based on at least the user watermark, the first audio segment and the extended audio segment. The first audio segment corresponds to the first watermarked audio segment, and the user watermark is not embedded in the first audio segment.
[0148] It should be noted that the first audio segment is any audio segment in the audio.
[0149] It should also be noted that the first audio segment in the audio corresponds to the first watermarked audio segment in the watermarked audio, that is, the first audio segment has the same content as the first watermarked audio segment.
[0150] S302. The watermark extraction device extracts the watermark from the watermark audio to obtain the user watermark.
[0151] In the prior art, user watermarks are extracted from watermarked audio frames. However, audio frames are short, so the watermark that can be extracted from the watermarked audio frames is relatively small. In the embodiments of the present application, however, the user watermark is extracted from a first watermarked audio segment. The first watermarked audio segment is derived based on at least the user watermark, the first audio segment, and an extended audio segment. The extended audio segment does not include the first watermarked audio segment and the first audio segment. That is, the first audio segment can be extended by the extended audio segment, and the user watermark is extracted from the extended audio segment. Therefore, because the length of the extended audio segment is longer than the audio frame in the prior art, a larger watermark can be extracted from the watermarked audio frame.
[0152] In addition, the first watermarked audio segment corresponds to the first audio segment. That is, after the user watermark is extracted from the first watermarked audio segment to obtain the first audio segment, the content of the audio is not changed. Therefore, the invisibility of the user watermark to the audio can be guaranteed as much as possible.
[0153] FIG4 shows a schematic flow chart of a watermark processing method 400 provided in an embodiment of the present application. This method 400 can be used in the system 100 shown in FIG2 . As shown in FIG4 , this method 400 may include the following steps. It should be noted that the following steps may be performed in various orders and / or occur simultaneously, and are not limited to the execution order shown in FIG4 .
[0154] It should also be noted that, to avoid repetition, for the parts of the method 400 that are not described in detail, reference can be made to the description of the corresponding parts of the above-mentioned method 200.
[0155] S401. The watermark embedding device obtains audio and a user watermark, where the audio includes N audio segments, where N is greater than or equal to 2.
[0156] It should be noted that the N audio clips in the audio are arranged in sequence.
[0157] It should also be noted that the embodiment of the present application does not limit the length of the audio clip.
[0158] For example, the length of an audio segment may be equal to the length of one or more audio frames.
[0159] It should also be noted that the embodiment of the present application does not limit the length of the audio frame.
[0160] For example, the length of the audio frame may be 10 ms.
[0161] Optionally, the user watermark can be used to identify the copyright or source of the audio.
[0162] Optionally, the audio may be a real-time audio stream (such as call audio) or may be locally stored audio (such as music), which is not limited in the embodiments of the present application.
[0163] Optionally, the embodiment of the present application does not limit the manner in which the watermark embedding device obtains the audio.
[0164] In a possible implementation manner, the watermark embedding device may receive the audio from an audio generating device.
[0165] In another possible implementation, the watermark embedding device may obtain the audio locally.
[0166] S402. The watermark embedding device embeds the user watermark into the audio to obtain a watermarked audio, where the watermarked audio includes N watermarked audio segments, and the Nth watermarked audio segment is obtained based on at least the user watermark, the Nth audio segment, and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment.
[0167] Optionally, the present application does not limit the number of the one or more watermarked audio segments.
[0168] In a possible implementation, the one or more watermarked audio segments may include all watermarked audio segments before the Nth watermarked audio segment in the watermarked audio.
[0169] In another possible implementation, the number of the one or more watermarked audio segments may be preset.
[0170] It should be noted that the N audio clips include the N-th audio clip, the N watermark audio clips include the N-th watermark audio clip, and the N-th audio clip in the N audio clips corresponds to the N-th watermark audio clip in the N watermark audio clips, that is, the content of the N-th audio clip is the same as that of the N-th watermark audio clip.
[0171] For example, taking the audio including 3 audio clips and the watermark audio including 3 watermark audio clips as an example, the above 3 audio clips are, in order: the first audio clip, the second audio clip and the third audio clip, and the above 3 watermark audio clips are, in order: the first watermark audio clip, the second watermark audio clip and the third watermark audio clip, wherein the first audio clip is the same as the content in the first watermark audio clip, the second audio clip is the same as the content in the second watermark audio clip, and the third audio clip is the same as the content in the third watermark audio clip.
[0172] In one possible implementation, one or more user watermarks are embedded in a first watermarked audio segment and an Nth watermarked audio segment, and each is embedded with a watermark segment. The strength of the user watermark embedded in different audio segments obtained by the watermark segment based on the user watermark is the same or different. Optionally, one or more user watermarks are embedded in the first watermarked audio segment and the Nth watermarked audio segment, respectively, and the strength of the user watermarks embedded in different audio segments can be the same or different. Optionally, the first watermarked audio segment and the Nth watermarked audio segment have different audio content, and the strength of the embedded user watermarks is different. Optionally, the first watermarked audio segment and the Nth watermarked audio segment can be watermarked audio frames, and one or more user watermarks are embedded in a single watermarked audio frame. The strength of the user watermark embedded in different audio frames can be the same or different. The strength of the user watermark embedded in an audio frame can be the sum of the strengths of a single user watermark or all user watermarks in the audio frame.
[0173] It should be noted that the first audio segment may be the first audio segment among the N audio segments, and the first watermarked audio segment may be the first watermarked audio segment among the N watermarked audio segments.
[0174] Optionally, the Nth watermarked audio segment is obtained based on an extended audio segment, the user watermark and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment, and the extended audio segment does not include the audio and the watermarked audio.
[0175] It should be noted that the extended audio segment described in the embodiment of the present application does not include the audio and the watermarked audio.
[0176] It should also be noted that the extended audio segment is used to extend the length of the audio segment so as to embed a watermark with a larger bit size.
[0177] Optionally, this application does not limit the method for generating the extended audio segment.
[0178] In a possible implementation, the extended audio segment may be generated randomly or based on a matrix.
[0179] In another possible implementation, the extended audio segment may be generated based on an all-zero matrix. That is, the extended audio segment may be a blank audio segment or a silent audio segment to reduce interference with the first audio segment.
[0180] In a possible implementation, the values of the sampling points corresponding to the blank audio segments or silent audio segments described in this application are all zero.
[0181] In one possible implementation, the Nth watermarked audio segment described in S402 is obtained based on at least the user watermark, the Nth audio segment, and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment. It can be understood that: the Nth watermarked audio segment is obtained based on the user watermark embedded in the first intermediate audio segment, and the first intermediate audio segment includes at least the Nth audio segment and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment.
[0182] In one possible implementation, S402 may include: obtaining the Nth audio segment and one or more watermark audio segments; embedding the user watermark into a first intermediate audio segment to obtain a first intermediate watermark audio segment, the first intermediate audio segment at least including the Nth audio segment and the one or more watermark audio segments; obtaining the Nth watermark audio segment based on the first intermediate watermark audio segment, the Nth watermark audio segment corresponding to the Nth audio segment in the first intermediate watermark audio segment.
[0183] Optionally, the first intermediate audio segment further includes the extended audio segment. That is, the Nth watermarked audio segment is obtained based on the extended audio segment, the user watermark, the Nth audio segment, and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment.
[0184] Optionally, the present application does not limit the order of the extended audio segment, the Nth audio segment, and one or more watermark audio segments before the Nth watermark audio segment in the first intermediate audio segment.
[0185] In one possible implementation, the extended audio segment may be located before one or more watermark audio segments before the N-th watermark audio segment, the one or more watermark audio segments before the N-th watermark audio segment may be located before the N-th audio segment, the first watermark audio segment of the one or more watermark audio segments before the N-th watermark audio segment is adjacent to the extended audio segment, and the last watermark audio segment of the one or more watermark audio segments before the N-th watermark audio segment is adjacent to the N-th audio segment.
[0186] In one possible implementation, when N=2, the first watermarked audio segment is the first watermarked audio segment in the watermarked audio, that is, there is no watermarked audio segment before the first watermarked audio segment in the watermarked audio. Therefore, the first watermarked audio segment is obtained based on at least the user watermark and the first audio segment. In other words, the first watermarked audio segment is obtained by embedding the user watermark into the first audio segment.
[0187] Optionally, when N=2, the first watermarked audio segment is obtained based on the extended audio segment, the user watermark, and the first audio segment. That is, the first watermarked audio segment is obtained by embedding the user watermark into a second intermediate audio segment, where the second intermediate audio segment includes the first audio segment and the extended audio segment.
[0188] In one possible implementation, S402 may include: obtaining a first audio segment; embedding the user watermark into a second intermediate audio segment, the second intermediate audio segment including the first audio segment and the extended audio segment, to obtain a second intermediate watermark audio segment; obtaining the first watermark audio segment based on the second intermediate watermark audio segment, the first watermark audio segment corresponding to the first audio segment in the second intermediate watermark audio segment.
[0189] Optionally, this embodiment of the present application does not limit the length of the intermediate audio segment.
[0190] In a possible implementation, the length of the intermediate audio segment may be greater than or equal to the length of (N+1) audio segments.
[0191] Optionally, the length of the intermediate audio segment may be preset.
[0192] Optionally, S402 may further include: generating the watermarked audio according to the N watermarked audio segments.
[0193] Since the strength of the watermark signal in a single audio segment is limited, multiple audio segments are used here to extract the user watermark to enhance the strength of the watermark signal and improve the success rate of watermark extraction. In the prior art, the watermark audio frame is obtained by embedding the user watermark into the audio frame, and the length of the audio frame is short. Therefore, the watermark that can be embedded in the audio is small. In the embodiment of the present application, the Nth watermark audio segment is obtained based on at least the user watermark, the Nth audio segment and one or more watermark audio segments in the watermark audio that precede the Nth watermark audio segment, that is, the Nth audio segment is extended by one or more watermark audio segments in the watermark audio that precede the Nth watermark audio segment. The first watermark audio segment is obtained based on embedding the user watermark into the extended audio segment. Since the length of the extended audio segment is longer than the audio frame in the prior art, a larger watermark can be embedded in the audio segment.
[0194] In addition, the Nth watermarked audio segment corresponds to the Nth audio segment, that is, after the user watermark is embedded in the Nth audio segment to obtain the Nth watermarked audio segment, the content of the audio does not change. Therefore, the invisibility of the user watermark to the audio can be guaranteed as much as possible.
[0195] For example, assuming that the audio includes three audio segments with a length of x ms, namely, audio segment S1, audio segment S2 and audio segment S3, and the watermarked audio includes three watermarked audio segments with a length of x ms, namely, watermarked audio segment S1′, watermarked audio segment S2′ and watermarked audio segment S3′. As an example, the watermark embedding process of each audio segment will be introduced with reference to Figures 5 to 7.
[0196] Figure 5 illustrates the watermark embedding process for audio segment S1 provided by an embodiment of the present application. As shown in Figure 5, the watermark embedding process for audio segment S1 includes: obtaining an intermediate audio segment Q1 of y ms length based on an audio segment S1 of length x ms and an extended audio segment of length (yx) ms; embedding the user watermark into the intermediate audio segment Q1 to obtain an intermediate watermarked audio segment Q1′ of y ms length, wherein the watermarked audio segment S1′ in the intermediate watermarked audio segment Q1′ corresponds to the audio segment S1 in the intermediate audio segment; and discarding all audio segments in the watermarked intermediate audio segment Q1′ except the watermarked audio segment S1′ to obtain the watermarked audio segment S1′.
[0197] FIG6 illustrates the watermark embedding process for an audio segment S2 according to an embodiment of the present application. As shown in FIG6 , the watermark embedding process for audio segment S2 includes: obtaining an intermediate audio segment Q2 of y ms in length based on an audio segment S2 of length x ms, a watermarked audio segment S1′ of length x ms, and an extended audio segment of length (y-2x) ms, wherein the extended audio segment precedes the watermarked audio segment S1′, and the watermarked audio segment S1′ precedes the audio segment S2; embedding the user watermark into the intermediate audio segment Q2 to obtain an intermediate watermarked audio segment Q2′ of y ms in length, wherein the watermarked audio segment S2′ in the intermediate watermarked audio segment Q2′ corresponds to the audio segment S2 in the intermediate audio segment; and discarding all audio segments in the watermarked intermediate audio segment Q2′ except the watermarked audio segment S2′ to obtain the watermarked audio segment S2′.
[0198] FIG7 illustrates the watermark embedding process for an audio segment S3 according to an embodiment of the present application. As shown in FIG7 , the watermark embedding process for audio segment S3 includes: obtaining an intermediate audio segment Q3 of y ms in length based on an audio segment S3 of length x ms, a watermarked audio segment S1′ of length x ms, a watermarked audio segment S1′ of length x ms, and an extended audio segment of length (y-3x) ms, wherein the extended audio segment precedes the watermarked audio segment S1′, the watermarked audio segment S1′ precedes the watermarked audio segment S2′, and the watermarked audio segment S2′ precedes the audio segment S3; embedding the user watermark into the intermediate audio segment Q3 to obtain an intermediate watermarked audio segment Q3′ of y ms in length, wherein the watermarked audio segment S3′ in the intermediate watermarked audio segment Q3′ corresponds to the audio segment S3 in the intermediate audio segment; and discarding audio segments other than the watermarked audio segment S3′ in the watermarked intermediate audio segment Q3′ to obtain the watermarked audio segment S3′.
[0199] FIG8 shows a schematic flow chart of a watermark processing method 500 provided in an embodiment of the present application. This method 500 can be used in the system 100 shown in FIG2 . As shown in FIG8 , this method 500 can include the following steps. It should be noted that the following steps can be performed in various orders and / or occur simultaneously, and are not limited to the execution order shown in FIG8 .
[0200] It should also be noted that, to avoid repetition, for the parts of the method 500 that are not described in detail, reference can be made to the description of the corresponding parts of the above-mentioned method 200.
[0201] S501. The watermark extraction device obtains watermarked audio, which includes N watermarked audio segments. The Nth watermarked audio segment is obtained based on at least the user watermark, the Nth audio segment and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment. The Nth audio segment corresponds to the Nth watermarked audio segment. The user watermark is not embedded in the Nth audio segment. N is greater than or equal to 2.
[0202] In a possible implementation, the N watermarked audio segments are obtained based on a first length, a first step length, and the watermarked audio, and the first step length is smaller than the first length.
[0203] Optionally, when N>1, the N watermark audio segments may partially overlap or not overlap, which is not limited in this embodiment of the present application.
[0204] Using the watermark processing method provided in the embodiment of the present application, the N watermark audio segments partially overlap, so that more watermark audio segments can be obtained from the watermark audio. The more watermark audio segments there are, the higher the accuracy of watermark extraction.
[0205] S502. The watermark extraction device extracts the watermark from the watermark audio to obtain the user watermark.
[0206] In a possible implementation, S502 may include: performing watermark extraction on each of the N watermark audio segments to obtain a user watermark extracted from each watermark audio segment; based on a correlation value of a template sequence of the user watermark extracted from each watermark audio segment, screening out at least one target watermark audio segment from the multiple watermark audio segments, the at least one target watermark audio segment being at least one watermark audio segment in the multiple watermark audio segments whose correlation value of the template sequence of the user watermark is greater than a preset first threshold; and determining the user watermark based on a normalized correlation value of each target watermark audio segment in the at least one target watermark audio segment.
[0207] In the prior art, user watermarks are extracted from watermarked audio frames. However, audio frames are short, so the watermark that can be extracted from the watermarked audio frames is relatively small. However, in the embodiments of the present application, the user watermark is extracted from the first watermarked audio segment, and the Nth watermarked audio segment is derived based on at least the user watermark, the Nth audio segment, and one or more watermarked audio segments preceding the Nth watermarked audio segment in the watermarked audio. That is, the first audio segment can be extended using one or more watermarked audio segments preceding the Nth watermarked audio segment, and the user watermark is extracted from the extended audio segment. Since the length of the extended audio segment is longer than the audio frame in the prior art, a larger watermark can be extracted from the watermarked audio frame.
[0208] In addition, the first watermarked audio segment corresponds to the first audio segment. That is, after the user watermark is extracted from the first watermarked audio segment to obtain the first audio segment, the content of the audio is not changed. Therefore, the invisibility of the user watermark to the audio can be guaranteed as much as possible.
[0209] The watermark processing method provided by the embodiment of the present application is described above in conjunction with Figures 2 to 8. The watermark processing device provided by the embodiment of the present application will be further described below.
[0210] FIG9 shows a schematic diagram of the structure of a watermark processing apparatus 600 provided in an embodiment of the present application. As shown in FIG9 , the apparatus 600 may include: an acquisition unit 601 and an embedding unit 602 .
[0211] Optionally, the device 600 can be used in the watermark processing system 100 . Further, the device 600 can be used in the watermark embedding device 110 in the watermark processing system 100 , such as a virtual device formed by software executed by a processor or controller on the watermark embedding device 110 .
[0212] The acquisition unit 601 is used to acquire audio and a user watermark, where the audio includes a first audio segment.
[0213] The embedding unit 602 is configured to embed the user watermark into the audio to obtain watermarked audio, where the watermarked audio includes a first watermarked audio segment, which is obtained based on at least the user watermark, the first audio segment, and the extended audio segment.
[0214] Optionally, the apparatus 600 may further include a determining unit 603. The acquiring unit 601 is specifically configured to acquire the first audio segment; the embedding unit 602 is specifically configured to embed the user watermark into an intermediate audio segment to obtain an intermediate watermarked audio segment, where the intermediate audio segment includes the first audio segment and the extended audio segment; and the determining unit 603 is configured to obtain the first watermarked audio segment based on the intermediate watermarked audio segment, where the first watermarked audio segment corresponds to the first audio segment in the intermediate watermarked audio segment, and the watermarked audio includes the first watermarked audio segment.
[0215] It should be noted that the information exchange and execution process between the above-mentioned devices are based on the same concept as the embodiment of the method 200 of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be described in detail here. In an optional example, the device 600 can be specifically the watermark embedding device in the embodiment of the method 200 above. The device 600 can be used to execute the various processes and / or steps corresponding to the watermark embedding device in the embodiment of the method 200 above. To avoid repetition, they will not be described in detail here.
[0216] One or more of the modules in the embodiment shown in FIG9 may be implemented by software, hardware, firmware, or a combination thereof. The software or firmware includes, but is not limited to, computer program instructions or codes, and may be executed by a hardware processor. The hardware includes, but is not limited to, various integrated circuits, such as a central processing unit (CPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or an application-specific integrated circuit (ASIC).
[0217] FIG10 shows a schematic diagram of the structure of a watermark processing apparatus 700 provided in an embodiment of the present application. As shown in FIG10 , the apparatus 700 may include: an acquisition unit 701 and an embedding unit 702 .
[0218] Optionally, the device 700 can be used in the watermark processing system 100 . Further, the device 700 can be used in the watermark embedding device 110 in the watermark processing system 100 , such as a virtual device formed by software executed by a processor or controller on the watermark embedding device 110 .
[0219] The acquisition unit 701 is used to acquire audio and a user watermark, where the audio includes N audio segments, where N is greater than or equal to 2.
[0220] The embedding unit 702 is used to embed the user watermark into the audio to obtain watermarked audio, where the watermarked audio includes N watermarked audio segments, and the Nth watermarked audio segment is obtained based on at least the user watermark, the Nth audio segment, and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment.
[0221] In a possible implementation, the first watermarked audio segment is obtained based on at least the user watermark and the first audio segment.
[0222] In one possible implementation, the number of user watermarks that can be embedded in different watermarked audio segments can be one or more, but the strength of the embedded user watermarks can be the same or different. In one possible implementation, the Nth watermarked audio segment is derived based on an extended audio segment, the user watermark, and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment, where the extended audio segment does not include the audio and the watermarked audio.
[0223] Optionally, the apparatus 700 may further include a determining unit 703. The acquiring unit 701 is specifically configured to acquire an Nth audio segment and one or more watermarked audio segments; the embedding unit 702 is specifically configured to embed the user watermark into a first intermediate audio segment to obtain a first intermediate watermarked audio segment, wherein the first intermediate audio segment at least includes the Nth audio segment and the one or more watermarked audio segments; and the determining unit 703 is configured to obtain the Nth watermarked audio segment based on the first intermediate watermarked audio segment, wherein the Nth watermarked audio segment corresponds to the Nth audio segment in the first intermediate watermarked audio segment.
[0224] In a possible implementation, the first intermediate audio segment also includes the extended audio segment.
[0225] In one possible implementation, the acquisition unit 701 is specifically used to acquire a first audio segment; the embedding unit 702 is specifically used to embed the user watermark into a second intermediate audio segment, the second intermediate audio segment includes the first audio segment and the extended audio segment, to obtain a second intermediate watermark audio segment; the determination unit 703 is used to obtain the first watermark audio segment based on the second intermediate watermark audio segment, and the first watermark audio segment corresponds to the first audio segment in the second intermediate watermark audio segment.
[0226] In a possible implementation, the extended audio segment is a blank audio segment.
[0227] In a possible implementation, the extended audio segment is generated based on an all-zero matrix.
[0228] Optionally, the apparatus 700 may further include a generating unit 704, configured to generate the watermarked audio according to the N watermarked audio segments.
[0229] It should be noted that the information exchange and execution process between the above-mentioned devices are based on the same concept as the embodiment of the method 400 of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be described in detail here. In an optional example, the device 700 can be specifically the watermark embedding device in the embodiment of the method 400. The device 700 can be used to execute the various processes and / or steps corresponding to the watermark embedding device in the embodiment of the method 400. To avoid repetition, they will not be described in detail here.
[0230] One or more of the modules in the embodiment shown in FIG10 may be implemented by software, hardware, firmware, or a combination thereof. The software or firmware includes, but is not limited to, computer program instructions or codes, and may be executed by a hardware processor. The hardware includes, but is not limited to, various integrated circuits, such as a CPU, DSP, FPGA, or ASIC.
[0231] FIG11 shows a schematic diagram of the structure of a watermark processing apparatus 800 provided in an embodiment of the present application. As shown in FIG11 , the apparatus 800 may include: an acquisition unit 801 and an extraction unit 802 .
[0232] Optionally, the device 800 can be used in the watermark processing system 100. Further, the device 800 can be used in the watermark extraction device 120 in the watermark processing system 100, such as a virtual device formed by software executed by a processor or controller on the watermark extraction device 120.
[0233] The acquisition unit 801 is used to obtain watermarked audio, which includes a first watermarked audio segment. The first watermarked audio segment is obtained based on at least the user watermark, the first audio segment and the extended audio segment. The first audio segment corresponds to the first watermarked audio segment, and the user watermark is not embedded in the first audio segment.
[0234] The extraction unit 802 is used to extract the watermark from the watermarked audio to obtain the user watermark.
[0235] It should be noted that the information interaction, execution process, and other contents between the above-mentioned devices are based on the same concept as the embodiment of the method 300 of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be described in detail here. In an optional example, the device 800 can be specifically the watermark extraction device in the embodiment of the method 300 above. The device 800 can be used to execute the various processes and / or steps corresponding to the watermark extraction device in the embodiment of the method 300 above. To avoid repetition, they will not be described here.
[0236] One or more of the modules in the embodiment shown in FIG11 may be implemented by software, hardware, firmware, or a combination thereof. The software or firmware includes, but is not limited to, computer program instructions or codes, and may be executed by a hardware processor. The hardware includes, but is not limited to, various integrated circuits, such as a CPU, DSP, FPGA, or ASIC.
[0237] FIG12 shows a schematic diagram of the structure of a watermark processing apparatus 900 provided in an embodiment of the present application. As shown in FIG12 , the apparatus 900 may include: an acquisition unit 901 and an extraction unit 902 .
[0238] Optionally, the device 900 can be used in the watermark processing system 100. Further, the device 900 can be used in the watermark extraction device 120 in the watermark processing system 100, such as a virtual device formed by software executed by a processor or controller on the watermark extraction device 120.
[0239] The acquisition unit 901 is used to obtain watermarked audio, which includes N watermarked audio segments. The Nth watermarked audio segment is obtained based on at least the user watermark, the Nth audio segment and one or more watermarked audio segments before the Nth watermarked audio segment in the watermarked audio. The Nth audio segment corresponds to the Nth watermarked audio segment. The user watermark is not embedded in the Nth audio segment. N is greater than or equal to 2.
[0240] The extraction unit 902 is used to extract the watermark from the watermarked audio to obtain the user watermark.
[0241] In a possible implementation, the N watermarked audio segments are obtained based on a first length, a first step length, and the watermarked audio, and the first step length is smaller than the first length.
[0242] Optionally, the apparatus 900 may further include: a screening unit 903 and a determination unit 904. The extraction unit 902 is specifically configured to: perform watermark extraction on each of the N watermarked audio segments to obtain a user watermark extracted from each watermarked audio segment; the screening unit 903 is configured to screen out at least one target watermarked audio segment from the multiple watermarked audio segments based on a correlation value of a template sequence of the user watermark extracted from each watermarked audio segment, the at least one target watermarked audio segment being at least one watermarked audio segment in the multiple watermarked audio segments for which the correlation value of the template sequence of the user watermark is greater than a preset first threshold; and the determination unit 904 is configured to determine the user watermark based on a normalized correlation value of each target watermarked audio segment in the at least one target watermarked audio segment.
[0243] It should be noted that the information interaction, execution process, and other contents between the above-mentioned devices are based on the same concept as the embodiment of the method 500 of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be described in detail here. In an optional example, the device 900 can be specifically the watermark extraction device in the embodiment of the method 500. The device 900 can be used to execute the various processes and / or steps corresponding to the watermark extraction device in the embodiment of the method 500. To avoid repetition, they will not be described in detail here.
[0244] One or more of the modules in the embodiment shown in FIG12 may be implemented by software, hardware, firmware, or a combination thereof. The software or firmware includes, but is not limited to, computer program instructions or codes, and may be executed by a hardware processor. The hardware includes, but is not limited to, various integrated circuits, such as a CPU, DSP, FPGA, or ASIC.
[0245] FIG13 shows a schematic block diagram of a watermark processing apparatus 1000 provided in an embodiment of the present application. The apparatus 1000 may include: a processor 1001 and a communication interface 1002 , wherein the processor 1001 and the communication interface 1002 are coupled.
[0246] In an alternative example, those skilled in the art will appreciate that the apparatus 1000 may be specifically the watermark embedding apparatus in the aforementioned method 200 or 400, and the apparatus 1000 may be the physical hardware structure of the watermark embedding apparatus. The apparatus 1000 may be used to execute the various processes and / or steps corresponding to the watermark embedding apparatus in the aforementioned method 200 or 400 embodiments, and to avoid repetition, they are not further described here.
[0247] The processor 1001 in the embodiment of the present application may include one or more processing units. Optionally, the processing unit includes but is not limited to a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, a discrete gate or transistor logic device, or a discrete hardware component. A general-purpose processor may be a microprocessor, a microcontroller, or any conventional processor.
[0248] For example, the processor 1001 is used to obtain audio and user watermark through the communication interface 1002, where the audio includes a first audio segment; embed the user watermark into the audio to obtain watermarked audio, where the watermarked audio includes a first watermarked audio segment, and the first watermarked audio segment is obtained based at least on the user watermark, the first audio segment, and an extended audio segment, and the extended audio segment does not include the audio and the watermarked audio.
[0249] Optionally, the apparatus 1000 may further include a memory 1003 .
[0250] The memory 1003 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0251] Specifically, the memory 1003 is used to store program codes and instructions of the watermark processing apparatus 1000. Optionally, the memory 1003 is also used to store data obtained during the execution of the above-mentioned method 200 or 400 by the processor 1001, such as the first audio segment, and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment.
[0252] Optionally, the memory 1003 may be a separate device or integrated into the processor 1001 .
[0253] It should be noted that FIG6 only shows a simplified design of the device 1000. In actual applications, the device 1000 may further include other necessary components, including but not limited to any number of communication interfaces, processors, selectors, memories, etc., and all devices 1000 that can implement the present application are within the scope of protection of the present application.
[0254] In one possible design, the device 1000 may be a chip. Optionally, the chip may further include one or more memories for storing computer-executable instructions. When the chip device is running, the processor may execute the computer-executable instructions stored in the memory to enable the chip to perform the steps performed by the watermark embedding device described in the above method 200 or 400.
[0255] Optionally, the chip device may be a field programmable gate array, a dedicated integrated chip, a system chip, a central processing unit, a network processor, a digital signal processing circuit, a microcontroller, or a programmable controller or other integrated chips for realizing relevant functions.
[0256] FIG14 shows a schematic block diagram of a watermark processing apparatus 1100 provided in an embodiment of the present application. The apparatus 1100 may include: a processor 1101 and a communication interface 1102 , wherein the processor 1101 and the communication interface 1102 are coupled.
[0257] In an optional example, those skilled in the art will appreciate that the apparatus 1100 may be specifically the watermark extraction apparatus in the aforementioned method 300 or 500, and the apparatus 1100 may be the physical hardware structure of the watermark extraction apparatus. The apparatus 1100 may be used to execute the various processes and / or steps corresponding to the watermark extraction apparatus in the aforementioned method 300 or 500 embodiments, and will not be described in detail here to avoid repetition.
[0258] The processor 1101 in the embodiment of the present application may include one or more processing units. Optionally, the processing unit includes, but is not limited to, a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, a discrete gate or transistor logic device, or a discrete hardware component. A general-purpose processor may be a microprocessor, a microcontroller, or any conventional processor.
[0259] For example, the processor 1101 is used to obtain watermarked audio through the communication interface 1102, where the watermarked audio includes N watermarked audio segments, the Nth watermarked audio segment is obtained based on at least the user watermark, the Nth audio segment, and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment, the Nth audio segment corresponds to the Nth watermarked audio segment, the user watermark is not embedded in the Nth audio segment, and N is greater than or equal to 2; the watermarked audio is extracted to obtain the user watermark.
[0260] Optionally, the device 1100 may further include a memory 1103 .
[0261] The memory 1103 may be a volatile memory or a nonvolatile memory, or may include both volatile and nonvolatile memories. The nonvolatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0262] Specifically, the memory 1103 is used to store program codes and instructions of the watermark processing apparatus 1100. Optionally, the memory 1103 is also used to store data obtained during the execution of the embodiment of the method 400 by the processor 1101, such as the first watermarked audio segment, and one or more watermarked audio segments in the watermarked audio that precede the Nth watermarked audio segment.
[0263] Optionally, the memory 1103 may be a separate device or integrated into the processor 1101 .
[0264] It should be noted that FIG6 only shows a simplified design of the device 1100. In actual applications, the device 1100 may further include other necessary components, including but not limited to any number of communication interfaces, processors, selectors, memories, etc., and all devices 1100 that can implement the present application are within the scope of protection of the present application.
[0265] In one possible design, the device 1100 may be a chip. Optionally, the chip may further include one or more memories for storing computer-executable instructions. When the chip device is running, the processor may execute the computer-executable instructions stored in the memory to enable the chip to perform the steps performed by the watermark extraction device described in the above method 400.
[0266] Optionally, the chip device may be a field programmable gate array, a dedicated integrated chip, a system chip, a central processing unit, a network processor, a digital signal processing circuit, a microcontroller, or a programmable controller or other integrated chips for realizing relevant functions.
[0267] An embodiment of the present application also provides a watermark processing system, which may include the watermark processing device as described in Figures 9 and 11, or the watermark processing device as described in Figures 10 and 12, or the watermark processing device as described in Figures 13 and 14.
[0268] An embodiment of the present application further provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions are executed on a computer, the watermark processing method described in the above method embodiment is implemented.
[0269] The embodiment of the present application further provides a computer program product, which, when executed on a processor, implements the watermark processing method described in the above method embodiment.
[0270] The watermark processing device, computer-readable storage medium, computer program product or chip provided in the embodiments of the present application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects described in the corresponding methods provided above, and will not be repeated here.
[0271] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0272] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0273] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0274] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0275] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0276] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0277] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.
[0278] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A watermark processing method, characterized in that, Comprising: Obtaining an audio and a user watermark, where the audio includes a first audio segment; Embedding the user watermark into the audio to obtain a watermarked audio, where the watermarked audio includes a first watermarked audio segment, and the first watermarked audio segment is obtained based at least on the user watermark, the first audio segment, and an extended audio segment.
2. The method according to claim 1, characterized in that, The embedding the user watermark into the audio to obtain a watermarked audio includes: Obtaining the first audio segment; Embedding the user watermark into an intermediate audio segment to obtain an intermediate watermarked audio segment, where the intermediate audio segment includes the first audio segment and the extended audio segment; Obtaining the first watermarked audio segment based on the intermediate watermarked audio segment, where the first watermarked audio segment corresponds to the first audio segment in the intermediate watermarked audio segment, and the watermarked audio includes the first watermarked audio segment.
3. A watermark processing method, characterized in that, Comprising: Obtaining an audio and a user watermark, where the audio includes N audio segments, and N is greater than or equal to 2; Embedding the user watermark into the audio to obtain a watermarked audio, where the watermarked audio includes N watermarked audio segments, and the Nth watermarked audio segment is obtained based at least on the user watermark, the Nth audio segment, and one or more watermarked audio segments before the Nth watermarked audio segment in the watermarked audio.
4. The method according to claim 3, wherein The first watermarked audio segment is obtained based at least on the user watermark and the first audio segment.
5. The method according to claim 3 or 4, characterized in that, One or more of the user watermarks are embedded in one of the watermarked audio segments, and the intensities of the user watermarks embedded in different watermarked audio segments are the same or different.
6. The method according to claim 3, wherein The Nth watermarked audio segment is obtained based on an extended audio segment, the user watermark, and one or more watermarked audio segments before the Nth watermarked audio segment in the watermarked audio, and the extended audio segment does not include the audio and the watermarked audio.
7. The method according to claim 3 or 6, characterized in that, The embedding the user watermark into the audio to obtain a watermarked audio includes: Obtaining the Nth audio segment and one or more watermarked audio segments; Embedding the user watermark into a first intermediate audio segment to obtain a first intermediate watermarked audio segment, where the first intermediate audio segment includes at least the Nth audio segment and the one or more watermarked audio segments; Obtaining the Nth watermarked audio segment based on the first intermediate watermarked audio segment, where the Nth watermarked audio segment corresponds to the Nth audio segment in the first intermediate watermarked audio segment.
8. The method according to claim 7, characterized in that, The first intermediate audio segment further includes the extended audio segment.
9. The method according to any one of claims 3-8, characterized in that, The embedding the user watermark into the audio to obtain a watermarked audio further includes: Obtaining the first audio segment; Embedding the user watermark into a second intermediate audio segment, where the second intermediate audio segment includes the first audio segment and the extended audio segment, to obtain a second intermediate watermarked audio segment; Obtaining the first watermarked audio segment based on the second intermediate watermarked audio segment, where the first watermarked audio segment corresponds to the first audio segment in the second intermediate watermarked audio segment.
10. The method according to any one of claims 6-9, characterized in that, The extended audio segment is a blank audio segment.
11. The method according to claim 10, characterized in that, The extended audio segment is generated based on a all-zero matrix.
12. The method according to any one of claims 3-11, characterized in that, The embedding the user watermark into the audio to obtain a watermarked audio further includes: Generate the watermarked audio according to the N watermarked audio segments.
13. A watermark processing method, characterized in that Including: Obtain a watermarked audio, where the watermarked audio includes a first watermarked audio segment, and the first watermarked audio segment is obtained at least based on the user watermark, a first audio segment, and an extended audio segment. The first audio segment corresponds to the first watermarked audio segment, and the user watermark is not embedded in the first audio segment. The extended audio segment does not include the first audio segment and the watermarked audio; Perform watermark extraction on the watermarked audio to obtain the user watermark.
14. A watermark processing method, characterized in that, Including: Obtain a watermarked audio, where the watermarked audio includes N watermarked audio segments. The Nth watermarked audio segment is obtained at least based on the user watermark, the Nth audio segment, and one or more watermarked audio segments before the Nth watermarked audio segment in the watermarked audio. The Nth audio segment corresponds to the Nth watermarked audio segment, and the user watermark is not embedded in the Nth audio segment, and N is greater than or equal to 2; Perform watermark extraction on the watermarked audio to obtain the user watermark.
15. A watermark processing device, characterized in that, Including: A processor and a communication interface, where the processor and the communication interface are coupled, and the processor is configured to: Obtain an audio and a user watermark, where the audio includes a first audio segment; Embed the user watermark into the audio to obtain a watermarked audio, where the watermarked audio includes a first watermarked audio segment, and the first watermarked audio segment is obtained at least based on the user watermark, the first audio segment, and an extended audio segment.
16. The device according to claim 15, characterized in that Specifically, the processor is configured to: Obtain the first audio segment; Embed the user watermark into an intermediate audio segment to obtain an intermediate watermarked audio segment, where the intermediate audio segment includes the first audio segment and the extended audio segment; Obtain the first watermarked audio segment based on the intermediate watermarked audio segment, where the first watermarked audio segment corresponds to the first audio segment in the intermediate watermarked audio segment, and the watermarked audio includes the first watermarked audio segment.
17. A watermark processing device, characterized in that, Including: A processor and a communication interface, where the processor and the communication interface are coupled, and the processor is configured to: Obtain an audio and a user watermark, where the audio includes N audio segments, and N is greater than or equal to 2; Embed the user watermark into the audio to obtain a watermarked audio, where the watermarked audio includes N watermarked audio segments, and the Nth watermarked audio segment is obtained at least based on the user watermark, the Nth audio segment, and one or more watermarked audio segments before the Nth watermarked audio segment in the watermarked audio.
18. The device according to claim 17, wherein The first watermarked audio segment is obtained at least based on the user watermark and the first audio segment.
19. The device according to claim 17 or 18, characterized in that, One or more user watermarks are embedded in one of the audio segments, and the intensities of the user watermarks embedded in different audio segments are the same or different.
20. The device according to claim 17, characterized in that The Nth watermarked audio segment is obtained based on an extended audio segment, the user watermark, and one or more watermarked audio segments before the Nth watermarked audio segment in the watermarked audio, and the extended audio segment does not include the audio and the watermarked audio.
21. The device according to claim 17 or 20, characterized in that, Specifically, the processor is configured to: Obtain the Nth audio segment and one or more watermarked audio segments; Embed the user watermark into the first intermediate audio segment to obtain a first intermediate watermarked audio segment, where the first intermediate audio segment at least includes the Nth audio segment and the one or more watermarked audio segments; Obtain the Nth watermarked audio segment based on the first intermediate watermarked audio segment, where the Nth watermarked audio segment corresponds to the Nth audio segment in the first intermediate watermarked audio segment.
22. The device according to claim 21, characterized in that, The first intermediate audio segment further includes the extended audio segment.
23. The device according to any one of claims 17 - 22, characterized in that, Specifically, the processor is configured to: Obtain a first audio segment; Embed the user watermark into a second intermediate audio segment, where the second intermediate audio segment includes the first audio segment and the extended audio segment, to obtain a second intermediate watermarked audio segment; Obtain the first watermarked audio segment based on the second intermediate watermarked audio segment, where the first watermarked audio segment corresponds to the first audio segment in the second intermediate watermarked audio segment.
24. The device according to any one of claims 20 - 23, characterized in that, The extended audio segment is a blank audio segment.
25. The device according to claim 24, wherein, The extended audio segment is generated based on an all-zero matrix.
26. The device according to any one of claims 17-25, characterized in that The processor is further configured to: Generate the watermarked audio according to the N watermarked audio segments.
27. A watermark processing device, characterized in that, Comprising: A processor and a communication interface, where the processor and the communication interface are coupled, and the processor is configured to: Obtain a watermarked audio, where the watermarked audio includes a first watermarked audio segment, and the first watermarked audio segment is obtained at least based on the user watermark, a first audio segment, and an extended audio segment. The first audio segment corresponds to the first watermarked audio segment, the user watermark is not embedded in the first audio segment, and the extended audio segment does not include the first audio segment and the watermarked audio; Perform watermark extraction on the watermarked audio to obtain the user watermark.
28. A watermark processing device, characterized in that, Comprising: A processor and a communication interface, where the processor and the communication interface are coupled, and the processor is configured to: Obtain a watermarked audio, where the watermarked audio includes N watermarked audio segments, and the Nth watermarked audio segment is obtained at least based on the user watermark, the Nth audio segment, and one or more watermarked audio segments in the watermarked audio before the Nth watermarked audio segment. The Nth audio segment corresponds to the Nth watermarked audio segment, the user watermark is not embedded in the Nth audio segment, and N is greater than or equal to 2; Perform watermark extraction on the watermarked audio to obtain the user watermark.
29. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a watermarked audio, where the watermarked audio includes a first watermarked audio segment, and the first watermarked audio segment is obtained at least based on the user watermark, a first audio segment, and an extended audio segment. The first audio segment corresponds to the first watermarked audio segment, the user watermark is not embedded in the first audio segment, and the extended audio segment does not include the first audio segment and the watermarked audio.
30. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a watermarked audio, where the watermarked audio includes N watermarked audio segments, and the Nth watermarked audio segment is obtained at least based on the user watermark, the Nth audio segment, and one or more watermarked audio segments in the watermarked audio before the Nth watermarked audio segment, and N is greater than or equal to 2.
31. A computer-readable storage medium for storing a computer program, characterized in that, The computer program includes instructions for implementing the method according to any one of claims 1-14 above.
32. A computer program product, comprising instructions, characterized in that, When the instructions are run on a computer or a processor, the computer or the processor is caused to implement the method according to any one of claims 1-14 above.
Citation Information
Patent Citations
Watermark processing method and device
CN120319253A
Audio watermark embedding and extracting method and device
CN103413552A
Audio watermark embedding, extracting and television program interaction method and device
CN109584890A
Audio watermark processing method and device, electronic equipment and storage medium
CN113362835A
Audio processing method and device
CN113380260A