Content compliance detection and alignment method for speech synthesis system
By building a content compliance detection evaluation benchmark data set and an online sensitive information database, combined with characterization-driven modules and large language models, the shortcomings of the speech synthesis system in content compliance detection and alignment are solved, and higher content compliance and security are achieved.
Patent Information
- Application Number
- CN202411853144.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-06-06
AI Technical Summary
The existing speech synthesis technology has not been fully studied in content compliance detection and alignment, resulting in the generation of content that may be toxic, discriminatory or inconsistent with the original intention, affecting user security and privacy.
A content compliance detection and alignment method for speech synthesis systems is proposed, including building a content compliance detection and evaluation benchmark data set and an online sensitive information database, multimodal content compliance judgment through characterization-driven modules, and text-level compliance evaluation combined with large language models.
It significantly improves the accuracy and reliability of the speech synthesis system at the content compliance level, ensures that the generated voice content is safe and unbiased, and is consistent with user intentions, protecting user information security and personal privacy.
Smart Images

Figure CN120108371A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech synthesis technology, and in particular to a content compliance detection and alignment method for a speech synthesis system. Background Art
[0002] Voice is one of the most important, direct and convenient interactive information in the development of human society. As a mainstream human-computer interaction technology, intelligent voice interaction technology has a wide range of applications and is one of the most mature artificial intelligence directions. Speech synthesis is one of the core technologies of intelligent voice interaction. This technology converts the text information received by the machine into natural and fluent voice information, which is fed back and delivered to the user. It has become an indispensable part of people's lives.
[0003] Although speech synthesis technology is relatively mature and widely used, the research on compliance detection and alignment of speech synthesis content is still in its infancy. Research on content compliance detection and alignment of speech synthesis systems can not only prevent speech synthesis technology from generating toxic, discriminatory, and inconsistent content during application, which can have a negative impact on society, but also protect users' information security and personal privacy.
[0004] Research on key technologies for secure, reliable, and highly expressive speech synthesis systems is necessary to promote the development of speech synthesis technology, meet users' diverse speech needs, and improve the quality, efficiency, and security of speech synthesis. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention proposes a content compliance detection and alignment method for a speech synthesis system, comprising the following steps:
[0006] Constructing a content compliance detection and evaluation benchmark data set and an online sensitive information database containing sensitive information representation; the construction of the content compliance detection and evaluation benchmark data set includes: defining evaluation dimensions, designing evaluation benchmark information for each evaluation dimension, obtaining voice data corresponding to the evaluation benchmark information and constructing a content compliance detection and evaluation benchmark data set;
[0007] Based on the content compliance detection and evaluation benchmark data set, a representation-driven speech content compliance judgment module and a paralanguage recognition module are trained;
[0008] For the input of the speech synthesis system:
[0009] In response to the synthesized signal, the input multimodal reference signal and the semantic features of the text to be synthesized are obtained by using a representation extraction technology, and the compliance of the semantic features is determined by using an online sensitive information database;
[0010] The ideographic features are converted into text form by a multimodal content compliance alignment judgment module, the ideographic features are combined with the text to be synthesized to form a first reference text, and the compliance of the first reference text is evaluated by a large language model;
[0011] For the output of the speech synthesis system:
[0012] Performing compliance judgment on the output result through the speech content compliance judgment module to obtain a first evaluation result, performing compliance judgment on the output result through the paralanguage recognition module in combination with a large language model to obtain a second evaluation result, and fusing the first evaluation result and the second evaluation result to generate an evaluation label for the synthetic signal;
[0013] The synthesized speech is output according to the evaluation label, and the speech synthesis system is updated through the training instruction.
[0014] A further technical solution of this embodiment is that the building of an online sensitive information database includes:
[0015] Collect multimodal sensitive information data, including text, audio, and images;
[0016] Extracting vector representation of the multimodal sensitive information data by embedding technology, wherein the text data uses the Bert model, the audio data uses the ResNet-based voiceprint model and the Clap model, and the image data uses the Clip model;
[0017] The extracted multimodal vector representations are stored in the vector index library and classified and managed according to modality categories and sensitive categories.
[0018] A further technical solution of this embodiment is that the input multimodal reference signal and the ideographic features of the text to be synthesized are obtained through the representation extraction technology, and the compliance of the ideographic features is determined through an online sensitive information database; wherein the extracted ideographic features include text representation, voice representation and image representation, and the compliance determination is to respectively conduct terrorism-related assessment, explosion-related assessment and political-related assessment on the ideographic features, including the following steps: for text representation and voice representation, a dynamic weighting mechanism based on cosine similarity is used, and for image representation, a multi-scale perception module is introduced to calculate the multi-dimensional feature distance; it is determined whether the semantic information of the ideographic feature of each single modality is close to the semantics of the sensitive information already in the online sensitive information database, and if so, it is determined to be non-compliant, otherwise it is compliant.
[0019] A further technical solution of this embodiment is that the multimodal content compliance alignment judgment module includes a Clip model and a Clap model, which convert image signals and audio signals into text descriptions respectively.
[0020] A further technical solution of this embodiment is that updating the speech synthesis system includes:
[0021] Collect several multimodal reference signals and the text to be synthesized, input them into the speech synthesis system, and obtain the synthesized speech
[0022] Use the fusion content judgment module to label the compliance of the synthesized speech and replace the non-compliant audio with an alarm tone;
[0023] An instruction fine-tuning data set is generated based on the annotation results of the evaluation tags, and the speech synthesis system is trained so that it has the ability to judge content compliance.
[0024] A further technical solution of this embodiment is that the compliance of the output result is judged by combining the paralanguage recognition module with the large language model, including:
[0025] The paralanguage information in the output result is converted into text by the paralanguage recognition module, a second reference text is generated by combining the paralanguage text and the text to be synthesized, and the second reference text is input into the large language model for compliance judgment.
[0026] A further technical solution of this embodiment is that the evaluation dimensions include timbre privacy and safety, timbre without social prejudice, timbre / style richness and diversity, timbre and subjective meaning. Figure 1 Consistent, non-toxic, safe and robust.
[0027] A further technical solution of this embodiment is that the representation-driven paralanguage identifier is composed of an embedding layer and a paralanguage model, the embedding layer is initialized with the weights of the codebook of the speech synthesis system and is not frozen, and is optimized during the training process together with the paralanguage model.
[0028] A further technical solution of this embodiment is that when the first evaluation result and the second evaluation result are integrated, a final evaluation result is obtained by voting fusion, and an evaluation label is generated based on the final evaluation result.
[0029] The beneficial effects of the embodiments of the present invention are:
[0030] The present invention discloses a content compliance detection and alignment method for a speech synthesis system, aiming to improve the accuracy and reliability of the speech synthesis system at the content compliance level. The method includes constructing a benchmark data set suitable for content compliance detection of the speech synthesis system, clarifying multiple evaluation dimensions such as timbre privacy and security, timbre without social bias, and timbre / style richness and diversity; at the model input end, based on the constructed online sensitive information database, the retrieval enhancement technology is used to perform fine matching of the multimodal reference signal and the text to be synthesized, and whether sensitive information is contained, and the representation of the reference signal is textualized and input into the large language model (LLM) with the text to be synthesized for content compliance judgment. At the model output end, the speech content compliance judgment module is trained based on the collected evaluation benchmark data set, the paralanguage information is textualized by the paralanguage recognizer, and the text-level compliance judgment of the LLM is combined to generate the final content compliance judgment result, and filter potential non-compliant speech fragments. In addition, in the training stage of the speech synthesis system, the instruction fine-tuning data set is used for fine-tuning training, and the content compliance discrimination module is used for guidance intervention, so that the model has the ability to make compliance judgments and filter harmful information. The present invention provides a content compliance detection and alignment strategy at both the input and output levels, which can significantly improve the content compliance of the speech synthesis model and provide technical support for building a healthy and harmonious speech synthesis application environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0032] Figure 1 It is a flowchart of the steps of the content compliance detection and alignment method for the speech synthesis system of the present invention;
[0033] Figure 2 A detailed flow chart of the content compliance detection and alignment method for a speech synthesis system according to the present invention;
[0034] Figure 3 is a detailed flow chart of input end evaluation in an embodiment of the present invention;
[0035] Figure 4 is a detailed flow chart of output end evaluation in an embodiment of the present invention;
[0036] Figure 5 for Figure 3 Workflow diagram of the multimodal representation module;
[0037] Figure 6 for Figure 3Workflow diagram of the unimodal representation retrieval enhancement module in ;
[0038] Figure 7 is a workflow diagram of a multimodal content compliance alignment judgment module in an embodiment of the present invention;
[0039] Figure 8 4 is a training flow chart of a paralanguage identifier in an embodiment of the present invention. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments, and the well-known modules, units and their connections, links, communications or operations are not shown or described in detail. In addition, the described features, architectures or functions can be combined in any way in one or more embodiments. It should be understood by those skilled in the art that the various embodiments described below are only for illustration and not for limiting the scope of protection of the present invention. It can also be easily understood that the modules or units or processing methods in the various embodiments described herein and shown in the drawings can be combined and designed according to various different configurations. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0041] Please refer to Figure 1 to Figure 2 As shown, an embodiment of the present invention discloses a content compliance detection and alignment method for a speech synthesis system. The embodiment of the present invention increases the content compliance detection and alignment capabilities of the speech synthesis system from the perspective of the input and output ends of the synthesis system, and specifically includes three stages, namely, a preparation stage, an input end evaluation stage, and an output end evaluation stage.
[0042] During the preparation stage, a content compliance detection evaluation benchmark dataset and an online sensitive information database suitable for content compliance detection of speech synthesis systems were constructed.
[0043] At the input end of the speech synthesis system, the content compliance judgment of the single-modal reference signal and the content compliance judgment of the fused multimodal reference signal are implemented. The bad reference signals under the single modality are filtered out by constructing a single-modal retrieval enhancement module. In order to filter out the bad input after the multimodal combination, the representation of the three modalities of text, speech and image are unified at the text level by constructing a content compliance text module, and the content compliance of the description is evaluated using the LLM that has been aligned for content compliance. Finally, the judgment results of these two paths are integrated to obtain the final content compliance evaluation result.
[0044] At the output end of the speech synthesis system, firstly, a fusion speech content compliance judgment module is proposed, which includes a representation based on the output end of the speech synthesis system (including content, timbre and rhythm) and an LLM based on content compliance alignment. The content compliance label of the audio to be synthesized is determined according to the results of the LLM and the content compliance judgment module, and the speech that does not meet the content compliance is filtered out.
[0045] The following are explained separately:
[0046] Preparation stage:
[0047] Constructing a content compliance detection and evaluation benchmark data set and an online sensitive information database containing sensitive information representation; the construction of the content compliance detection and evaluation benchmark data set includes: defining evaluation dimensions, designing evaluation benchmark information for each evaluation dimension, obtaining voice data corresponding to the evaluation benchmark information and constructing a content compliance detection and evaluation benchmark data set;
[0048] First, in order to effectively evaluate the detection performance of the content compliance of the speech synthesis system, the present invention refers to the LLM security assessment method, and combines literature research and the professional knowledge of interdisciplinary social ethicists to clearly define the specific evaluation dimensions of the content compliance of the speech synthesis system. Then, for each evaluation dimension, non-compliant text and prompt information are designed, and the corresponding prompt information is collected through three methods: web crawling, manual recording, and generative model synthesis. After the data collection is completed, the security of these voices is assessed by manual annotation, and finally a multi-dimensional large-scale evaluation data set is constructed.
[0049] Specifically, the construction of the content compliance detection and evaluation benchmark dataset includes:
[0050] Design the content and prompts to be synthesized, refer to the LLM evaluation data set on value toxicity, social prejudice, etc., and under the guidance of the knowledge of social ethics experts, determine the data scale of normal corpus and unsafe and unreliable corpus in the data set, comprehensively consider various paralinguistic attributes such as timbre and style, and comprehensively design the content text to be synthesized and the prompt information of the target speaker, including various unsafe content such as violence, illegality, stereotypes, and adult content, covering different population groups, including regions (different provinces), nationalities, genders (males and females), and ages (children, young people, and the elderly), to ensure the comprehensiveness and representativeness of the data set.
[0051] To collect speech data and obtain a large-scale benchmark dataset, this project will use a variety of methods to collect corpus: (1) Based on the designed paralinguistic information, use the corresponding paralinguistic information detector to crawl real data from the Internet; (2) Record speech data according to the designed content and prompts; (3) Based on the designed content and prompts, use VALL-E, NaturalSpeech 3 and other mainstream speech synthesis systems to synthesize the corresponding speech.
[0052] Manual labeling security, based on the above security evaluation dimensions, data labeling is carried out in a crowdsourcing manner through an online platform, and the labeling is considered from three aspects: transparency, coverage, and confidence: (1) Transparency means that the labeling process will specifically label the security and credibility of the speech, and clearly distinguish which type of problem it belongs to; (2) Coverage refers to labeling the target population group in the dataset to discover neglected groups that may need to collect more data, and at the same time distinguish whether the speech discusses individuals or general groups to avoid general descriptions that exacerbate stereotypes; (3) Confidence means that each evaluation audio data is manually labeled by at least three annotators, using the majority voting principle to ensure the accuracy and subjective consistency of the labeling results.
[0053] The construction of the online sensitive information database includes:
[0054] Collect sensitive information data from various sources such as the Internet or books, including multimodal data such as text, audio and images. Text sensitive information includes specific words, phrases or paragraphs involving illegal, violent, privacy and other contents; audio sensitive information refers to the voice features of certain specific people. Image sensitive information refers to certain specific people, sensitive scenes, symbols or signs. The present invention uses embedding technology to design a multimodal representation module that can convert information of different modalities into vector representations respectively. For text sensitive information modality, the present invention uses a pre-trained Bert model to convert text-level information into a vector representation of text information; for voice sensitive data, a ResNet-based voiceprint model extraction model and a pre-trained Clap model are used to extract the speaker feature vector representation and audio voice feature vector representation in audio sensitive information; for image sensitive information, a pre-trained Clip model is used to extract the semantic feature vector representation of the image. These three types of vector representations are stored in a vector index library according to modality type and sensitive category (such as terrorism, violence and politics).
[0055] like Figure 3As shown, in the input evaluation of the speech synthesis system, the text to be synthesized and the multimodal reference signal are deeply processed. Through the previously constructed online sensitive information database containing sensitive information data, the retrieval enhancement technology will be used to conduct a refined search and match between the representation of each modal reference signal and the representation of sensitive information in the database, so as to accurately determine whether the input signal contains sensitive information. In addition, the representation of the multimodal reference signal is textualized and input together with the text to be synthesized into a large language model (LLM) that has achieved content compliance alignment for content compliance judgment, to ensure that the reference signal meets the preset requirements at the content compliance level. The following steps are included:
[0056] In response to the synthesized signal, the input multimodal reference signal and the semantic features of the text to be synthesized are obtained through the representation extraction technology, and the compliance of the semantic features is determined through the online sensitive information database; the compliance determination is performed through the single-modal representation retrieval enhancement module to evaluate the representation of the input signal in terms of terrorism, explosion, and politics. Specifically, Figure 5 and Figure 6 As shown, the multimodal reference signal (text, voice or image) is input into the multimodal representation module, and the representation is extracted separately according to the characteristics of different modes. After extracting multimodal representations such as text representation, voice representation and image representation, these representation vectors are input into the single-modal representation retrieval enhancement module, which will retrieve the online sensitive information data according to the modal category in the multimodal reference signal and the online sensitive information database. The present invention will be evaluated in terms of terrorism, explosion, politics, etc., and a dynamic weighting mechanism based on cosine similarity is used for text and voice representation, and a multi-scale perception module is introduced for image representation to calculate the multi-dimensional feature distance. Determine whether the semantic information in the reference signal representation is semantically close to the sensitive information already in the database, and determine the specific category of sensitive information involved, such as whether the reference signal contains the voice or face of a sensitive person, reference image signals containing bloody and violent scenes, and text descriptions; the reference signal contains the sound or text description of stuttering, groaning of patients with physical defects or pornographic content; in single-modal representation retrieval, these non-compliant content issues existing in the single modality will be directly detected based on the semantic similarity of the representation; design an incremental learning framework based on sample streams to update sensitive representation databases of different modalities in real time, so as to adapt to new illegal content in a timely manner and ensure the accuracy of retrieval and judgment.
[0057] In the single-modal representation retrieval enhancement module, only content compliance issues existing in a single modality can be detected. In cases where the reference signals are compliant in a single modality, but under the combined action of multi-modal reference signals, the module fails to detect problems, such as using a Chinese voice reference signal and expecting the synthesized voice to contain content like "likes to eat pork", or giving a picture of a child's face and expecting the synthesized voice to be related to love talk. In a single modality, all meet the requirements of content compliance. However, when these modality information are combined as multi-modal reference signals to control the content generation of the speech synthesis system, effective content compliance detection often cannot be carried out. For example Figure 7 As shown, the present invention converts the ideographic features into text form through a multi-modal content compliance alignment judgment module, combines it with the text to be synthesized to form a first reference text, and conducts compliance evaluation on the first reference text through a large language model. Since the currently commercially available large language models have undergone content compliance alignment at the text level. Therefore, we can analyze and understand using such large language models after full-text conversion of multi-modal reference signals, and can judge whether the reference signals are compliant. For example, in the example at the beginning, the Prompt (prompt text) for content compliance judgment by the large language model after text conversion of multi-modal reference signals can be: This is a Chinese reference voice, and the speaker should be Han Chinese. Is it compliant to synthesize the voice of the following text? I like to eat pork. Then the answer of the large language model will include a judgment on whether the content is compliant. Specifically, to solve the problem that the multi-modal reference signals are content compliant in single-modal evaluation, but the reference information content is non-compliant under the combined action of multi-modalities, the present invention proposes a multi-modal content compliance alignment judgment module. The Clip model used in the multi-modal representation module can convert image signals into a text describing the image; the Clap model can convert audio signals into a text describing the content of the audio segment. Combine these two converted texts, the reference signal in text modality, and the text to be synthesized into a text prompt for content compliance query and input it into an available large language model that implements content compliance for content compliance judgment. This method can convert the content compliance detection problems of different modalities into semantic judgment problems at the text modality level, and then make judgments by leveraging the understanding ability of existing commercially available large language models that have achieved content compliance. The text prompt for content compliance query can be expressed as "In the picture of <text converted by Clip>, there is a voice like <text converted by Clap>. Under the condition of <reference signal in text modality>, is it compliant to synthesize the voice of the following text? <text to be synthesized>".
[0058] The compliance judgment proposed in the present invention is jointly completed by a single-modal module and a multi-modal mechanism. Single-modal detection can be used to quickly identify the input reference signal to determine whether the content is compliant. When the single-modal detection result cannot confirm whether it is compliant, multi-modal collaborative analysis is triggered to improve the comprehensiveness and accuracy of detection through cross-modal similarities and supplementary information.
[0059] like Figure 4 As shown, at the output end of the speech synthesis system, there are two stages. In the first stage, a speech content compliance judgment module based on the representation of the output end of the speech synthesis system is trained using the content compliance detection evaluation benchmark data set established in the present invention as an evaluation benchmark for direct judgment of content compliance. The module is a set of classifiers designed by combining CNN and Transformer, which can perform compliance evaluation in terms of text content, speaker timbre, speech rhythm, etc. At the same time, in order to prevent potentially unsafe and unreliable speech from completing the synthesis output, a representation-based paralanguage recognizer is also constructed. The recognizer uses the representation of the output end of the speech synthesis system as input and extracts the paralanguage text label information of dimensions such as speaker timbre, speaker age, speaker voice style, and speaker pathological characteristics contained in the output representation.
[0060] Specifically, we built several Transformer-based speech content compliance classifiers. These classifiers are trained with multi-dimensional content compliance label data in the content compliance detection and evaluation benchmark dataset to achieve content compliance judgment. At the same time, we built a paralanguage identifier based on ECAPA-TDNN using the attribute labels of audio style, age, pathology, and speaker in the content compliance detection and evaluation benchmark dataset to identify paralanguage information in the input representation.
[0061] like Figure 8 As shown in FIG. 1 , the representation-based paralanguage recognizer is composed of an embedding layer and a paralanguage model. The embedding layer is initialized with the weights of the codebook of the speech synthesis system and is not frozen. It is optimized during the training process together with the paralanguage model. When performing paralanguage recognition, the discrete representation output by the speech synthesis system needs to be converted into continuous features for input into the paralanguage model. Specifically, the embedding layer transforms the discrete input features of the shape [B, Nq, T] into continuous features of the shape [B, Nq, T, F]. During the training process, a mask layer enhancement technique is used to randomly replace the output embeddings of the last n layers of embedding layers with zero embeddings based on the probability of each timestamp for data enhancement. Subsequently, through the embedding fusion strategy, a set of trainable parameters are used to assign weights to different embedding layers, and the embeddings of different layers are summed to aggregate each embedding layer to obtain continuous features of the shape [B, T, F]. Finally, the aggregated continuous features are input into the paralanguage model for the recognition of paralanguage information.
[0062] Subsequently, a text prompt for content compliance judgment is generated according to the paralinguistic text label and the text to be synthesized in a manner similar to that in the input-end content compliance alignment, and then a publicly available LLM that has achieved content compliance alignment is used to perform content compliance judgment at the text level. This method can detect situations that are not covered in the benchmark data set constructed by the present invention, including the following steps:
[0063] The compliance of the output result is judged by the speech content compliance judgment module to obtain a first evaluation result, the compliance of the output result is judged by the paralanguage recognition module in combination with the large language model to obtain a second evaluation result, and the first evaluation result and the second evaluation result are fused to generate an evaluation label of the synthetic signal; when fusing the evaluation results, the first evaluation result and the second evaluation result are fused through a voting weighted mechanism, and the content compliance evaluation result generated by the LLM is combined with the retrieval-enhanced content compliance evaluation result to determine the final content compliance evaluation result.
[0064] The synthesized speech is output according to the evaluation label, and the speech synthesis system is updated through the training instruction.
[0065] Specifically, the above compliance assessment methods all detect and align content compliance by adding a discrimination module to the periphery of the input and output of the speech synthesis system. It is also particularly important to improve the detection and alignment capabilities of the speech synthesis system itself, which can make the speech synthesis system more portable and more generalizable. Therefore, in this embodiment, it also includes improving the compliance detection effect by updating the speech synthesis system.
[0066] Specifically, a batch of reference signals and texts to be synthesized are collected and input into the speech synthesis system. The trained fusion content judgment module is used to mark the compliance judgment label. The original audio is retained for compliant audio, and a unified alarm audio is used for non-compliant audio. Through the above steps, a command fine-tuning dataset suitable for speech synthesis system training is produced. The speech synthesis system is fine-tuned by using this dataset to make feedback, so that the speech synthesis system itself has the ability to judge content compliance and filter harmful information.
[0067] Furthermore, the compliance judgment of the output result by combining the paralanguage recognition module with the large language model includes the following steps:
[0068] The paralanguage information in the output result is converted into text by the paralanguage recognition module, a second reference text is generated by combining the paralanguage text and the text to be synthesized, and the second reference text is input into the large language model for compliance judgment.
[0069] Optionally in this embodiment, the synthesized speech may include one or more evaluation tags, and the evaluation tags are marked on the corresponding synthesized speech. When the synthesized speech is output, the synthesized speech with the tags is replaced by an alarm sound.
[0070] The advantages of this application are:
[0071] (1) The multimodal content compliance alignment judgment module is used to solve the problem that the content of the multimodal reference signal is compliant in a single modality evaluation, but the reference information content is not compliant under the joint action of multiple modalities.
[0072] (2) Based on the content compliance detection evaluation benchmark dataset and the online sensitive information database, the multimodal reference signals and the text to be generated are evaluated, and the generalization understanding ability of the large language model is fully utilized to improve the detection effectiveness.
[0073] (3) Compliance checks are performed on the multimodal reference signal and the content to be synthesized at both the input and output ends of the speech synthesis system, significantly improving the content compliance of the speech synthesis model.
[0074] (4) Use instruction fine-tuning technology to provide guidance and intervention to the speech generation system so that it has the ability to judge content compliance.
[0075] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0076] The above embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for those of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A content compliance detection and alignment method for a speech synthesis system, characterized in that: The steps include: Construct a content compliance detection and evaluation benchmark dataset and an online sensitive information database containing sensitive information representation; The construction of the content compliance detection evaluation benchmark data set includes: defining evaluation dimensions, designing evaluation benchmark information for each evaluation dimension, acquiring voice data corresponding to the evaluation benchmark information and constructing the content compliance detection evaluation benchmark data set; Based on the content compliance detection and evaluation benchmark data set, a representation-driven speech content compliance judgment module and a paralanguage recognition module are trained; For the input of the speech synthesis system: In response to the synthesized signal, the input multimodal reference signal and the semantic features of the text to be synthesized are obtained by using a representation extraction technology, and the compliance of the semantic features is determined by using an online sensitive information database; The ideographic features are converted into text form by a multimodal content compliance alignment judgment module, the ideographic features are combined with the text to be synthesized to form a first reference text, and the compliance of the first reference text is evaluated by a large language model; For the output of the speech synthesis system: Performing compliance judgment on the output result through the speech content compliance judgment module to obtain a first evaluation result, performing compliance judgment on the output result through the paralanguage recognition module in combination with a large language model to obtain a second evaluation result, and fusing the first evaluation result and the second evaluation result to generate an evaluation label for the synthetic signal; The synthesized speech is output according to the evaluation label, and the speech synthesis system is updated through the training instruction.
2. The method according to claim 1, characterized in that The construction of the online sensitive information database comprises the following steps: Collect multimodal sensitive information data, including text, audio, and images; Extracting vector representation of the multimodal sensitive information data by embedding technology, wherein the text data uses the Bert model, the audio data uses the ResNet-based voiceprint model and the Clap model, and the image data uses the Clip model; The extracted multimodal vector representations are stored in the vector index library and classified and managed according to modality categories and sensitive categories.
3. The method according to claim 2, characterized in that The input multimodal reference signal and the ideographic features of the text to be synthesized are obtained by the representation extraction technology, and the compliance of the ideographic features is determined by the online sensitive information database; wherein the extracted ideographic features include text representation, voice representation and image representation, and the compliance determination is to respectively conduct terrorism-related assessment, explosion-related assessment and political-related assessment on the ideographic features, including the following steps: For text representation and speech representation, a dynamic weighting mechanism based on cosine similarity is used. For image representation, a multi-scale perception module is introduced to calculate multi-dimensional feature distance. It is determined whether the semantic information of the ideographic feature of each single modality is close to the semantics of the sensitive information already in the online sensitive information database. If so, it is determined to be non-compliant, otherwise it is compliant.
4. The method according to claim 1, characterized in that The multimodal content compliance alignment judgment module includes a Clip model and a Clap model, which convert image signals and audio signals into text descriptions respectively.
5. The method according to claim 1, characterized in that The updating of the speech synthesis system comprises the following steps: Collect several multimodal reference signals and the text to be synthesized, input them into the speech synthesis system, and obtain the synthesized speech Use the fusion content judgment module to label the compliance of the synthesized speech and replace the non-compliant audio with an alarm tone; An instruction fine-tuning data set is generated based on the annotation results, and the speech synthesis system is trained so that it has the ability to judge content compliance.
6. The method according to claim 1, characterized in that The step of judging the compliance of the output result by combining the paralanguage recognition module with the large language model comprises the following steps: The paralanguage information in the output result is converted into text by the paralanguage recognition module, a second reference text is generated by combining the paralanguage text and the text to be synthesized, and the second reference text is input into the large language model for compliance judgment.
7. The method according to claim 1, characterized in that The training of the speech content compliance judgment module and the paralanguage recognition module based on the representation-driven content compliance detection and evaluation benchmark data set includes the following steps: constructing multiple speech content compliance classifiers based on the Transformer model, which are trained by multi-dimensional content compliance label data in the content compliance detection and evaluation benchmark data set to achieve content compliance judgment; Using the attribute labels of the audio in the content compliance detection evaluation benchmark dataset, a paralanguage label classifier is constructed based on ResNet to identify the paralanguage information in the input representation.
8. The method according to claim 1, characterized in that The evaluation dimensions include timbre privacy and safety, timbre without social bias, timbre / style richness and diversity, timbre consistency with subjective intention, style non-toxicity and safety, and adversarial robustness.
9. The method according to claim 1, characterized in that The representation-driven paralanguage identifier is composed of an embedding layer and a paralanguage model. The embedding layer is initialized with the weights of the codebook of the speech synthesis system and is not frozen, and is optimized during the training process together with the paralanguage model.
10. The method according to claim 1, characterized in that When the first evaluation result and the second evaluation result are integrated, a final evaluation result is obtained by voting integration, and an evaluation label is generated based on the final evaluation result.