Speech adversarial watermark generation method and device, and electronic equipment

By adjusting the watermark to be embedded, a voice adversarial watermark is generated, causing it to produce incorrect results when passed through an intelligent voice recognition system after being embedded in the voice. This solves the problem of privacy content leakage in existing technologies and realizes privacy and copyright protection for voice data.

CN115798489BActive Publication Date: 2026-03-24HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-01
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing digital watermarking technology cannot effectively protect the privacy of voice data. Third parties can use intelligent voice systems to identify and obtain the privacy content of voice data embedded with watermarks in batches, posing a threat.

Method used

By adjusting the watermark to be embedded, it produces incorrect results when the intelligent speech recognition system recognizes the embedded speech, thus generating a speech-based adversarial watermark and ensuring that private content is not leaked.

Benefits of technology

This technology, while protecting copyright, generates anti-watermarks to prevent third-party intelligent voice systems from misidentifying data, thereby enhancing the privacy and security of voice data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115798489B_ABST
    Figure CN115798489B_ABST
Patent Text Reader

Abstract

The application provides a speech adversarial watermark generation method and device and electronic equipment. In the embodiment of the application, the target watermark to be embedded is adjusted, and the adjusted target watermark to be embedded is embedded into the original speech, so that the recognition results of the original speech and the target watermark speech in which the target watermark to be embedded is embedded are different through the same intelligent speech recognition system, so that even if a third party intercepts the target watermark speech, the third party cannot obtain the private content in the original speech when the intelligent speech recognition system is used to recognize the target watermark speech.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to speech recognition, in particular to a speech adversarial watermark generation method and device and electronic equipment. BACKGROUND

[0002] Digital watermarking is a technology that embeds specific information into multimedia content through a specific algorithm to achieve file authenticity identification, copyright protection and other functions. At present, as an effective authenticity identification and copyright protection means, digital watermarking is widely used in multimedia data transmission, publication, sharing and other scenarios.

[0003] However, the existing digital watermarking can only achieve copyright protection of the original speech data, and cannot further achieve privacy protection of the content. For example, a third party can collect a large amount of speech data embedded with digital watermarking, use an intelligent speech system to perform batch speech recognition on the speech data, thereby obtaining the private content in the speech data embedded with digital watermarking, and thus causing serious threat to users or other entities. SUMMARY

[0004] The present application provides a speech adversarial watermark generation method and device and electronic equipment, so that the intelligent speech recognition system of the third party makes recognition errors when recognizing the intercepted speech data embedded with watermarking, and cannot obtain the private content in the embedded speech data.

[0005] The technical scheme provided by the present application includes:

[0006] A speech adversarial watermark generation method, applied to an electronic equipment, comprising:

[0007] obtaining an original speech and at least one original speech recognition result; the at least one original speech recognition result comprises a recognition result obtained by performing speech recognition on the original speech through at least one intelligent speech recognition system;

[0008] for each target to-be-embedded watermark obtained, adjusting the target to-be-embedded watermark according to the watermark parameter corresponding to the target to-be-embedded watermark, so that the recognition result obtained by performing speech recognition on the original speech embedded with the adjusted target to-be-embedded watermark through an intelligent speech recognition system is different from the original speech recognition result;

[0009] embedding the adjusted target to-be-embedded watermark into the original speech to obtain a candidate watermark speech;

[0010] outputting a target watermark speech based on each candidate watermark speech and the at least one original speech recognition result; the target watermark speech is one of the candidate watermark speeches, and the recognition result obtained by performing speech recognition on the target watermark speech and the original speech through the same intelligent speech recognition system is different.

[0011] A speech adversarial watermark generation apparatus, applied to an electronic device, comprising:

[0012] An obtaining unit, configured to obtain an original speech and at least one original speech recognition result; the at least one original speech recognition result comprises a recognition result obtained by performing speech recognition on the original speech via at least one intelligent speech recognition system;

[0013] An embedding unit, configured to, for each target watermark to be embedded that has been obtained, adjust the target watermark to be embedded according to a watermark parameter corresponding to the target watermark to be embedded, so that a recognition result obtained by performing speech recognition on the original speech in which the adjusted target watermark to be embedded is embedded via an intelligent speech recognition system is different from the original speech recognition result; and embed the adjusted target watermark to be embedded into the original speech to obtain a candidate watermark speech;

[0014] An output unit, configured to output a target watermark speech based on each candidate watermark speech and the at least one original speech recognition result; the target watermark speech is one of the candidate watermark speeches, and a recognition result obtained by performing speech recognition on the target watermark speech via the same intelligent speech recognition system as that for the original speech is different from the recognition result obtained by performing speech recognition on the original speech.

[0015] As can be seen from the above technical solutions, in the embodiments of the present application, the target watermark to be embedded is adjusted, and the adjusted target watermark to be embedded is embedded into the original speech, so that the recognition result obtained by performing speech recognition on the original speech and the target watermark speech in which the target watermark to be embedded is embedded via the same intelligent speech recognition system is different, so that even if a third party intercepts the target watermark speech, the third party cannot obtain the private content in the original speech when performing speech recognition on the target watermark speech via the intelligent speech recognition system.

[0016] Further, in the embodiments, by transmitting the target watermark speech, the function of the intelligent speech system of the third party can be disabled on the basis of realizing copyright protection of the original speech data, thereby improving the privacy security of the original speech. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, further serve to explain the principles of the present disclosure.

[0018] Figure 1 A method flowchart provided for the embodiments of the present application;

[0019] Figure 2 An implementation flowchart of step 104 provided for the embodiments of the present application;

[0020] Figure 3An apparatus structure diagram provided by an embodiment of the present application is shown in the following figure;

[0021] Figure 4 An apparatus hardware structure diagram provided by an embodiment of the present application is shown in the following figure. DETAILED DESCRIPTION

[0022] To make the method provided by the present application easier to understand, the method provided by the present application is described in detail below in combination with the drawings and embodiments:

[0023] The embodiment of the present application utilizes the idea of speech adversarial samples, and combines speech adversarial samples and traditional speech data digital watermarking to cause the speech after superimposing watermarking to cause the intelligent recognition system to recognize incorrectly on the basis of conventional digital speech. The following describes an example: Figure 1

[0024] Referring to Figure 1 , Figure 1 A method flowchart provided by an embodiment of the present application is shown in the following figure. Figure 1 , Figure 1 A method flowchart provided by an embodiment of the present application is shown in the following figure. The flowchart is applied to electronic devices such as terminals, servers, and the like, and the embodiment is not specifically limited.

[0025] As shown in Figure 1 , the flowchart can include the following steps:

[0026] Step 101, obtaining original speech and original speech recognition results.

[0027] Optionally, in the embodiment, the original speech can be subjected to speech recognition by at least one intelligent speech recognition system. Based on this, in the embodiment, the original speech recognition results can include recognition results obtained by subjecting the original speech to speech recognition by the at least one intelligent speech recognition system respectively. The at least one intelligent speech recognition system herein is, for example, an intelligent speech classification system, an intelligent speech emotion recognition system, and the like, and the embodiment is not specifically limited. The at least one intelligent speech recognition system herein can be an intelligent speech recognition system used by a third party for listening or attack, and can also be referred to as an intelligent speech recognition system of a third party.

[0028] Optionally, in the embodiment, to facilitate obtaining the recognition results obtained by the intelligent speech recognition system, the input interface and the output interface of the intelligent speech recognition system can be obtained in advance, and by controlling the original speech to be input to the input interface, the recognition results obtained by subjecting the original speech to speech recognition by the intelligent speech recognition system and output by the output interface can be obtained directly. It should be noted that this is only an example of how to obtain the recognition results obtained by subjecting the original speech to speech recognition by the intelligent speech recognition system, and is not intended to be limiting.

[0029] ​In step 102, for each target watermark to be embedded obtained, the target watermark to be embedded is adjusted according to the watermark parameter corresponding to the target watermark to be embedded, so that the recognition result obtained by performing speech recognition on the original speech in which the adjusted target watermark to be embedded is embedded via the intelligent speech recognition system is different from the original speech recognition result.

[0030] Optionally, in the embodiment, initially, the target watermark to be embedded can be each watermark to be embedded in the set of watermarks to be embedded. Alternatively, initially, the target watermark to be embedded can be at least one specified watermark to be embedded in the set of watermarks to be embedded. And so on, the embodiment is not specifically limited. In a non-initial state, the target watermark to be embedded is a target watermark to be embedded with modified embedding parameters (see the description below, which is not described here), and the embodiment is not specifically limited.

[0031] As an embodiment, there are many ways to determine the set of watermarks to be embedded. For example: first, obtain at least one watermark information, each watermark information at least includes: watermark content, watermark type, watermark quantity. Here, the watermark content can be a number, a name, etc., which can be set according to actual needs; the watermark type can be a speech audible watermark or a speech blind watermark; the watermark quantity refers to the number of watermarks corresponding to the watermark content and the watermark type; then, for each watermark information obtained, generate the watermark to be embedded corresponding to the watermark information. For example, in a watermark information, the watermark content is 111111, the watermark type is a speech audible watermark, and the watermark quantity is 3, then 3 watermarks to be embedded can be generated, and the watermark content of each watermark to be embedded is 111111, and the watermark type is a speech audible watermark. Then, record each watermark to be embedded to the set of watermarks to be embedded. That is, the set of watermarks to be embedded is finally obtained.

[0032] In the embodiment, each target watermark to be embedded has a corresponding watermark parameter. Optionally, in the embodiment, the watermark parameter can be randomly generated, or obtained by modifying the watermark parameter corresponding to the previous historical watermark, and the embodiment is not specifically limited.

[0033] As an embodiment, as described in step 102, the target watermark to be embedded can be adjusted according to the watermark parameter corresponding to the target watermark to be embedded. Here, the purpose of adjusting the target watermark to be embedded is mainly to ensure that the candidate watermark speech (obtained by embedding the adjusted target watermark to be embedded into the original speech) and the recognition result obtained by performing speech recognition on the original speech via the same intelligent speech recognition system are different, so as to realize watermark resistance. Here, the so-called watermark resistance refers to embedding a watermark into speech, which can not only realize the authenticity identification and copyright protection of speech, but also cause the intelligent speech recognition system to recognize incorrectly. See the description of step 103 for details.

[0034] Optionally, as an embodiment, the above-mentioned adjustment on the target watermark to be embedded is not arbitrary, but needs to ensure that the final adjustment result (i.e. the target watermark to be embedded after adjustment) does not affect the intelligibility of the watermark, such as does not distort the meaning of the watermark itself (for example, if the watermark is originally 11111, if it is adjusted to be understood as 22222, it means that it affects the intelligibility of the watermark, and if it can still be understood as 11111, it means that it does not affect the intelligibility of the watermark). In order to ensure that the final watermark adjustment result (i.e. the target watermark to be embedded after adjustment) does not affect the intelligibility of the watermark, the present embodiment adjusts the target watermark to be embedded with a reference, so that the watermark feature value of each sampling point in the target watermark to be embedded after adjustment is greater than or equal to the preset minimum feature value. Here, the minimum feature value is the minimum feature value required by the feature attribute to which the watermark feature value belongs, which can be set based on the ratio of the average feature value corresponding to the feature attribute in the original voice to the preset average amplitude of the watermark. By making the watermark feature value of each sampling point in the target watermark to be embedded after adjustment greater than or equal to the preset minimum feature value, the final watermark adjustment result (i.e. the target watermark to be embedded after adjustment) can theoretically ensure that it does not affect the intelligibility of the watermark.

[0035] Of course, as an embodiment, under the premise of ensuring that the final watermark adjustment result (i.e. the target watermark to be embedded after adjustment) does not affect the intelligibility of the watermark, the upper limit of the watermark feature value of each sampling point in the target watermark to be embedded after adjustment can also be limited, such as the watermark feature value of each sampling point in the target watermark to be embedded after adjustment is less than the preset maximum feature value. Here, the maximum feature value can be a value greater than the minimum feature value set based on actual needs, such as directly setting a value that differs from the minimum feature value by a certain value as the maximum feature value, etc., which is not specifically limited in the present embodiment.

[0036] As an embodiment, the minimum feature value here can be carried in the above-mentioned watermark parameter.

[0037] Here, the above-mentioned watermark feature value is related to the type of the target watermark to be embedded, such as, optionally, if the above-mentioned target watermark to be embedded is a voice audible watermark, the above-mentioned watermark feature value is the loudness value of each sampling point when the target watermark to be embedded is played, such as 2 decibels, etc., and correspondingly, the feature attribute to which the watermark feature value belongs is loudness; if the target watermark to be embedded is a voice blind watermark, the watermark feature value is the embedding strength value of each sampling point in the target watermark to be embedded, such as voice frequency, etc., and correspondingly, the feature attribute to which the watermark feature value belongs is embedding strength; the present embodiment is not specifically limited.

[0038] Based on the above description, in the present embodiment, the watermark feature value of each sampling point in the above-mentioned target watermark to be embedded after adjustment is greater than or equal to the above-mentioned minimum feature value.

[0039] Step 103, embedding the adjusted target watermark to be embedded into the original speech to obtain a candidate watermark speech.

[0040] As another embodiment, the watermark parameter further includes watermark embedding position information. Here, the watermark embedding position information is used to indicate the position (such as the starting position, for example, the starting position is at the first second position of the original speech) where the watermark to be embedded needs to be embedded into the original speech. It should be noted that in this embodiment, the watermark embedding position information corresponding to any two target watermarks to be embedded is different to prevent the watermarks from overlapping each other. Based on this, the above-mentioned embedding the adjusted target watermark to be embedded into the original speech can include embedding the adjusted target watermark to be embedded into the position corresponding to the watermark embedding position information in the original speech according to the watermark embedding position information. Finally, the adjusted target watermark to be embedded is embedded into the original speech to obtain a candidate watermark speech.

[0041] Step 104, outputting a target watermark speech based on each candidate watermark speech and at least one original speech recognition result; the target watermark speech is one of the candidate watermark speeches, and the target watermark speech and the original speech have different recognition results obtained by speech recognition via the same intelligent speech recognition system.

[0042] Optionally, in this embodiment, there are many implementation manners for the above-mentioned outputting a target watermark speech based on each candidate watermark speech and at least one original speech recognition result, such as Figure 2 For example, one of the implementation manners is not described here.

[0043] As described in step 104, here, if the target watermark speech and the original speech have different recognition results obtained by speech recognition via the same intelligent speech recognition system, the intelligent speech recognition system will recognize incorrectly even if it intercepts the target watermark speech, and cannot obtain the private content in the original speech, which means that the watermark in the target watermark speech is resistant, and the watermark at this time can be called a speech resistant watermark.

[0044] At this point, the process shown in Figure 1 is completed.

[0045] As can be seen from the process shown in Figure 1 in the embodiment of the present application, by adjusting the target watermark to be embedded, the adjusted target watermark to be embedded is embedded into the original speech, so that the original speech and the target watermark speech embedded with the target watermark to be embedded have different recognition results obtained by speech recognition via the same intelligent speech recognition system, so that a third party can intercept the target watermark speech, and when the target watermark speech is recognized by the intelligent speech recognition system, recognition error occurs, and the private content in the original speech cannot be obtained.

[0046] Further, in the embodiment, by transmitting the target watermark voice, the function of the third-party intelligent voice system can be disabled on the basis of realizing the copyright protection of the original voice data, thereby improving the privacy security of the original voice.

[0047] The following describes the flowchart shown in FIG. 1: Figure 2

[0048] Referring to FIG. 2, Figure 2 , Figure 2 The step 104 provided in the embodiment of the present application realizes the flowchart. As shown in FIG. 2, the flowchart can include the following steps: Figure 2

[0049] Step 201: For each intelligent voice recognition system, input the candidate watermark voice into the intelligent voice recognition system to obtain the recognition result corresponding to the candidate watermark voice.

[0050] Optionally, in the embodiment, on the premise that multiple candidate watermark voices are obtained, the step 201 can sequentially input the multiple candidate watermark voices into the intelligent voice recognition system to obtain the recognition result corresponding to each candidate watermark voice.

[0051] Step 202: If the recognition result corresponding to one of the candidate watermark voices is different from the target recognition result, the target recognition result refers to the recognition result obtained by inputting the original voice into the intelligent voice recognition system, it is determined that the candidate watermark voice is the target watermark voice, otherwise, when the preset iteration condition is met, the watermark parameter corresponding to at least one target watermark to be embedded is modified, the target watermark to be embedded after modification is updated as the target watermark to be embedded, and the step of adjusting the target watermark to be embedded according to the watermark parameter corresponding to the target watermark to be embedded in the step 102 is returned.

[0052] Optionally, in the embodiment, if the recognition results corresponding to the multiple candidate watermark voices are different from the recognition result obtained by inputting the original voice into the intelligent voice recognition system, it indicates that the recognition of the intelligent voice recognition system on the multiple candidate watermark voices is wrong, at this time, one of the candidate watermark voices can be selected as the target watermark voice. Here, the target watermark voice will make the intelligent voice recognition system recognize incorrectly even if the target watermark voice is intercepted, and the privacy content in the original voice cannot be obtained, at this time, the watermark in the target watermark voice is also called a voice adversarial watermark with adversarial property.

[0053] ​​Optionally, in the embodiment, there are many ways to determine whether the preset iteration condition is met, such as: checking whether the current iteration number is less than the preset maximum iteration number, and if so, determining that the preset iteration condition is met; wherein the current iteration number is initially set to an initial value such as 1. Correspondingly, in the embodiment, before returning to the step of performing the watermark embedding operation on the target watermark to be embedded, it can further include: increasing the current iteration number by a set value such as 1.

[0054] Of course, if the preset iteration condition is not met, the current process can be ended.

[0055] Optionally, in the embodiment, the modification of the watermark parameter corresponding to the at least one target watermark to be embedded can be implemented based on a preset algorithm, such as a particle swarm algorithm, a genetic algorithm, a Bayesian optimization algorithm, a simulated annealing algorithm, etc. This way of updating the watermark parameter can be a conventional implementation of algorithms such as particle swarm algorithm, genetic algorithm, Bayesian optimization algorithm, simulated annealing algorithm, etc. Here, no further description is given.

[0056] At this point, the process shown in Figure 2 is completed.

[0057] Through the process shown in Figure 2 , how to output the target watermark speech based on each candidate watermark speech and the at least one original speech recognition result is implemented.

[0058] The above describes the method provided by the present application, and the following describes the device provided by the present application:

[0059] Referring to Figure 3 , Figure 3 the device structure diagram provided by the embodiment of the present application. The device is applied to an electronic device and includes:

[0060] An obtaining unit is configured to obtain an original speech and at least one original speech recognition result; the at least one original speech recognition result includes a recognition result obtained by performing speech recognition on the original speech via at least one intelligent speech recognition system;

[0061] An embedding unit is configured to, for each target watermark to be embedded that has been obtained, adjust the target watermark to be embedded according to a watermark parameter corresponding to the target watermark to be embedded, so that the recognition result obtained by performing speech recognition on the original speech in which the adjusted target watermark to be embedded is embedded via an intelligent speech recognition system is different from the original speech recognition result; and embed the adjusted target watermark to be embedded into the original speech to obtain a candidate watermark speech.

[0062] an output unit configured to output a target watermarked speech based on each candidate watermarked speech and the at least one original speech recognition result, the target watermarked speech being one of the candidate watermarked speeches, and the target watermarked speech having a different recognition result from the original speech when input to the same intelligent speech recognition system.

[0063] Optionally, the watermark parameter further comprises watermark embedding position information; and based on this, the step of embedding the adjusted target watermark to be embedded into the original speech comprises embedding the adjusted target watermark to be embedded into a position corresponding to the watermark embedding position information in the original speech according to the watermark embedding position information.

[0064] Optionally, the watermark embedding position information corresponding to different target watermarks to be embedded is different.

[0065] Optionally, the watermark parameter further comprises a minimum feature value required by a feature attribute to which the watermark feature value belongs; wherein the watermark feature value of each sampling point in the adjusted target watermark to be embedded is greater than or equal to the minimum feature value. The minimum feature value is set based on a ratio of an average feature value corresponding to the feature attribute in the original speech to a preset watermark average amplitude.

[0066] Optionally, if the target watermark to be embedded is a speech audible watermark, the watermark feature value is a loudness value of each sampling point in the target watermark to be embedded when played.

[0067] If the target watermark to be embedded is a speech blind watermark, the watermark feature value is an embedding strength value of each sampling point in the target watermark to be embedded.

[0068] Optionally, the step of outputting a target watermarked speech based on each candidate watermarked speech and the at least one original speech recognition result comprises:

[0069] for each intelligent speech recognition system, inputting the candidate watermarked speech into the intelligent speech recognition system to obtain a recognition result corresponding to the candidate watermarked speech;

[0070] if the recognition result corresponding to one of the candidate watermarked speeches is different from a target recognition result, the target recognition result being a recognition result obtained by inputting the original speech into the intelligent speech recognition system, then determining the candidate watermarked speech as the target watermarked speech; otherwise, when a preset iteration condition is met, modifying the watermark parameter corresponding to at least one target watermark to be embedded, updating the modified target watermark to be embedded as the target watermark to be embedded, and returning to the step of adjusting the target watermark to be embedded according to the watermark parameter corresponding to the target watermark to be embedded.

[0071] Optionally, the condition that the current iteration is satisfied includes: checking whether the current iteration number is less than the preset maximum iteration number; if so, determining that the current iteration is satisfied; wherein the current iteration number is initially set to an initial value.

[0072] Before returning to the step of performing the watermark embedding operation on the target watermark to be embedded, the method further includes: increasing the current iteration number by a set value.

[0073] Optionally, initially, the watermark parameters corresponding to each target watermark to be embedded are randomly generated; or, they are obtained by modifying the watermark parameters corresponding to previously obtained historical watermarks.

[0074] This concludes the process. Figure 3 Structural description of the device shown.

[0075] Correspondingly, this application also provides Figure 3 The hardware structure of the device shown. See also Figure 4 The hardware structure may include: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method disclosed in the above example of this application.

[0076] Based on the same application concept as the above method, this application embodiment also provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the method disclosed in the above examples of this application.

[0077] For example, the aforementioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information such as executable instructions, data, etc. For instance, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.

[0078] The systems, devices, modules, or units described in the above embodiments can be implemented by a computer processor or entity, or by a product with a certain function. A typical implementation device is a computer, which can be a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0079] For the convenience of description, the above apparatus is described in various units by function for description. Of course, in the implementation of the present application, the functions of each unit can be implemented in one or more software and / or hardware.

[0080] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. In addition, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0081] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in the flowchart and / or block diagram.

[0082] In addition, these computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including instruction apparatus, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in the flowchart and / or block diagram.

[0083] These computer program instructions can also be loaded into a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable data processing apparatus to generate a computer-implemented process, so that the instructions executed on the computer or other programmable data processing apparatus provide a process for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in the flowchart and / or block diagram.

[0084] The above merely provides an example of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application should be included in the scope of claims of the present application.

Claims

1. A method for generating adversarial voice watermarks, characterized in that, The method is applied to an electronic device, and includes: obtaining original speech and original speech recognition results; the original speech recognition results include recognition results obtained by performing speech recognition on the original speech via an intelligent speech recognition system; for each target watermark to be embedded obtained, adjusting the target watermark to be embedded according to watermark parameters corresponding to the target watermark to be embedded, so that the recognition result obtained by performing speech recognition on the original speech in which the adjusted target watermark to be embedded is embedded via the intelligent speech recognition system is different from the original speech recognition result; embedding the adjusted target watermark to be embedded into the original speech to obtain a candidate watermark speech; outputting a target watermark speech based on each candidate watermark speech and the original speech recognition result; the target watermark speech is one of the candidate watermark speeches, and the target watermark speech is different from the recognition result obtained by performing speech recognition on the original speech via the same intelligent speech recognition system.

2. The method of claim 1, wherein, The watermark parameters at least include watermark embedding position information; the watermark embedding position information corresponding to different target watermarks to be embedded is different; The embedding of the adjusted target watermark to be embedded into the original speech includes embedding the adjusted target watermark to be embedded into a position corresponding to the watermark embedding position information in the original speech according to the watermark embedding position information.

3. The method of claim 1, wherein, The watermark parameters further include a minimum feature value required by a feature attribute to which the watermark feature value belongs; wherein the watermark feature value of each sampling point in the adjusted target watermark to be embedded is greater than or equal to the minimum feature value.

4. The method of claim 3, wherein, The minimum feature value is set based on a ratio of an average feature value corresponding to the feature attribute in the original speech to a preset watermark average amplitude.

5. The method of claim 3, wherein, If the target watermark to be embedded is a speech audible watermark, the watermark feature value is a loudness value of each sampling point in the target watermark to be embedded when played; If the target watermark to be embedded is a speech blind watermark, the watermark feature value is an embedding strength value of each sampling point in the target watermark to be embedded.

6. The method of claim 1, wherein, The outputting of the target watermark speech based on each candidate watermark speech and the original speech recognition result includes: for each intelligent speech recognition system, inputting the candidate watermark speech into the intelligent speech recognition system to obtain a recognition result corresponding to the candidate watermark speech; if the recognition result corresponding to one of the candidate watermark speeches is different from a target recognition result, the target recognition result refers to a recognition result obtained by inputting the original speech into the intelligent speech recognition system, then determining the candidate watermark speech as the target watermark speech; otherwise, when a preset iteration condition is currently met, modifying the watermark parameters corresponding to at least one target watermark to be embedded, updating the modified target watermark to be embedded as the target watermark to be embedded, and returning to the step of adjusting the target watermark to be embedded according to the watermark parameters corresponding to the target watermark to be embedded.

7. The method of claim 6, wherein, The current meeting of the preset iteration condition includes checking whether a current iteration number is less than a preset maximum iteration number, and if yes, determining that the preset iteration condition is currently met; wherein the current iteration number is initially set as an initial value. Before returning to the step of performing the watermark embedding operation on the target watermark to be embedded, further comprising: increasing the current iteration number by a set value.

8. The method according to any one of claims 1 to 7, characterized in that, Initially, the watermark parameter corresponding to each target watermark to be embedded is randomly generated; or, is obtained by modifying the watermark parameter corresponding to the obtained historical watermark.

9. A speech adversarial watermark generation apparatus characterized by comprising: The device is applied to an electronic device, comprising: An obtaining unit is configured to obtain original speech and at least one original speech recognition result; the at least one original speech recognition result comprises a recognition result obtained by performing speech recognition on the original speech via at least one intelligent speech recognition system; An embedding unit is configured to, for each target watermark to be embedded obtained, adjust the target watermark to be embedded according to a watermark parameter corresponding to the target watermark to be embedded, so that a recognition result obtained by performing speech recognition on original speech in which the adjusted target watermark to be embedded is embedded via an intelligent speech recognition system is different from the original speech recognition result; and embed the adjusted target watermark to be embedded into the original speech to obtain a candidate watermark speech; An output unit is configured to output a target watermark speech based on each candidate watermark speech and the at least one original speech recognition result; the target watermark speech is one of the candidate watermark speeches, and a recognition result obtained by performing speech recognition on the target watermark speech and the original speech via a same intelligent speech recognition system is different.

10. An electronic device, comprising: The electronic device comprises a processor and a machine readable storage medium; The machine readable storage medium stores machine executable instructions capable of being executed by the processor; The processor is configured to execute the machine executable instructions to implement the method steps of any one of claims 1-8.

Citation Information

Patent Citations

  • Method and apparatus for watermarking successive sections of an audio signal

    US20150221317A1

  • Apparatus and method for copy-protected generation and reproduction of a wave field synthesis audio representation

    US20170150286A1