Data processing method, device and equipment and readable storage medium
Through automatic audio detection and intelligent auditing technology, the inefficiency problem caused by manual intervention in speech synthesis corpus generation is solved, the automation and high efficiency of corpus generation is achieved, and the accuracy of corpus quality evaluation is improved.
Patent Information
- Application Number
- CN202410175996.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-07
- Publication Date
- 2025-08-08
AI Technical Summary
The existing process of phonological corpus generation relies on manual intervention, resulting in inefficient generation and high cost.
Automatic audio detection and intelligent auditing technology are adopted to extract sound clips through sound detection rules, and the quality of sound clips is evaluated based on text similarity, so as to realize the automation and intelligence of corpus generation.
The process of corpus generation without manual participation in the pronunciation synthesis significantly improves the efficiency of corpus generation and the accuracy of quality evaluation.
Smart Images

Figure CN120452419A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data processing method, apparatus, device, and readable storage medium. Background Art
[0002] With the development of artificial intelligence (AI) technology, AI technology has been applied in more and more fields and played an increasingly important role. Speech synthesis is a key technology of AI. In fields such as intelligent dialogue, speech synthesis technology can be used to communicate with users.
[0003] Speech synthesis technology converts computer-generated or externally input text into spoken output by a target speaker. Related technologies typically utilize deep learning-based approaches, which require large, high-quality corpus datasets for model training. Traditional speech synthesis corpus collection methods typically involve constructing a text corpus, handing it over to voice actors for dubbing recording, cutting the dubbing audio, and manually reviewing the dubbing audio to produce a corpus with dubbing. This approach is complex and requires significant manual intervention. For example, constructing the text corpus requires users to create data tailored to the recording scenario. Similarly, cutting the dubbing audio requires extensive manual effort using editing software to edit the dubbing line by line. Each of these stages relies on manual corpus generation, which is time-consuming and labor-intensive, reducing the efficiency of speech synthesis corpus generation. Therefore, there is an urgent need for a speech synthesis corpus generation solution to improve its efficiency. Summary of the Invention
[0004] The embodiments of the present application provide a data processing method, apparatus, device, and readable storage medium, which can improve the efficiency of corpus generation in speech synthesis services.
[0005] On the one hand, an embodiment of the present application provides a data processing method, including:
[0006] Obtain dubbing audio data of original text data;
[0007] Performing audio detection processing on the dubbing audio data according to the sound detection rule to obtain sound segments in the dubbing audio data;
[0008] Performing audio recognition on the sound segment to obtain recognized text data of the sound segment;
[0009] The text similarity between the original text data and the recognized text data is obtained, and the quality assessment value of the sound segment is determined according to the text similarity; the quality assessment value of the sound segment is used to assist in quality rating of the sound segment.
[0010] In one aspect, an embodiment of the present application provides a data processing device, including:
[0011] An audio acquisition module is used to acquire dubbing audio data of original text data;
[0012] An audio detection module is used to perform audio detection processing on the dubbing audio data according to the sound detection rules to obtain sound clips in the dubbing audio data;
[0013] An audio recognition module is used to perform audio recognition on a sound segment to obtain recognized text data of the sound segment;
[0014] The evaluation value determination module is used to obtain the text similarity between the original text data and the recognized text data, and determine the quality evaluation value of the sound clip according to the text similarity; the quality evaluation value of the sound clip is used to assist in quality rating of the sound clip.
[0015] In one embodiment, the specific implementation of the audio acquisition module acquiring the dubbing audio data of the original text data includes:
[0016] Get the field parameters of the text construction object based on the configuration key fields; the configuration key fields include scene field, style field, role type field and quantity field;
[0017] Calling a text generation model based on the field parameters, and generating original text data indicated by the field parameters through the text generation model;
[0018] The original text data is pushed to the dubbing object, and the dubbing data returned by the dubbing object is determined as the dubbing audio data of the original text data.
[0019] In one embodiment, the dubbing audio data is composed of an audio frame sequence, and the audio frame sequence includes N audio frames; N is a positive integer;
[0020] The audio detection module performs audio detection processing on the dubbing audio data according to the sound detection rules to obtain the specific implementation method of the sound clips in the dubbing audio data, including:
[0021] According to the sound detection rule, each audio frame in the N audio frames is subjected to frame detection to obtain the frame type of each audio frame; the frame type includes a sound type and a silence type;
[0022] Determine an audio frame whose frame type in the audio frame sequence is a silent type as a silent frame, and determine an audio frame whose frame type in the audio frame sequence is a vocal type as a vocal frame;
[0023] If there are continuous silent frames in the audio frame sequence, the audio data composed of the continuous silent frames is determined as a silent segment in the dubbing audio data, and a sound segment is extracted from the dubbing audio data based on the segment length of the silent segment;
[0024] If there are no consecutive silent frames in the audio frame sequence, the dubbing audio data is determined to be a sound segment.
[0025] In one embodiment, the audio detection module performs frame detection on each of the N audio frames according to the sound detection rule to obtain a specific implementation of the frame type of each audio frame, including:
[0026] Determine any one audio frame among the N audio frames as a target audio frame;
[0027] Obtain the short-time energy and short-time zero-crossing rate corresponding to the target audio frame;
[0028] If the short-time energy corresponding to the target audio frame is greater than the energy threshold, and the short-time zero-crossing rate corresponding to the target audio frame is less than the zero-crossing rate threshold, determining that the frame type of the target audio frame is a voicing type;
[0029] If the short-time energy corresponding to the target audio frame is less than the energy threshold, or the short-time zero-crossing rate corresponding to the target audio frame is greater than the zero-crossing rate threshold, the frame type of the target audio frame is determined to be a silence type.
[0030] In one embodiment, the specific implementation of the audio detection module extracting the sound segments from the dubbing audio data based on the segment duration of the silence segment includes:
[0031] Comparing the segment duration of the silence segment with a silence duration threshold;
[0032] If the segment duration of the silence segment is greater than the silence duration threshold, the dubbing audio data is segmented based on the position of the silence segment in the dubbing audio data to obtain segmented audio data, and the sound segment in the dubbing audio data is determined based on the segmented audio data;
[0033] If the segment length of the silent segment is less than the silent segment length threshold, the dubbing audio data is determined to be a sound segment.
[0034] In one embodiment, the number of cut audio data is M; M is a positive integer;
[0035] The specific implementation method of the audio detection module determining the sound segment in the dubbing audio data based on the cut audio data includes:
[0036] Obtaining a sound frame contained in each of the M cut audio data;
[0037] Determine the audio attributes of each cut audio data based on the sound frames contained in each cut audio data; the audio attributes include continuous sound attributes and intermittent sound attributes;
[0038] The cut audio data with the continuous sound attribute among the M cut audio data are determined as the sound segments in the dubbing audio data.
[0039] In one embodiment, the audio detection module determines the audio attribute of each cut audio data based on the sound frames contained in each cut audio data. The specific implementation method includes:
[0040] Determine any one of the M cut audio data as target cut audio data;
[0041] Traversing the utterance frames contained in the target cut audio data;
[0042] If there are continuous utterance frames in the target cut audio data, and the number of frames corresponding to the continuous utterance frames reaches the utterance frame number threshold, the audio attribute of the target cut audio data is determined to be a continuous utterance attribute;
[0043] If there are no continuous sounding frames in the target cut audio data, or the number of continuous sounding frames in the target cut audio data does not reach the sounding frame number threshold, the audio attribute of the target cut audio data is determined to be an intermittent sounding attribute.
[0044] In one embodiment, the specific implementation of the evaluation value determination module obtaining the text similarity between the original text data and the recognized text data includes:
[0045] Calculate the minimum editing frequency of converting the recognized text data into the original text data according to the similarity calculation rules;
[0046] Obtain a frequency mapping table; the frequency mapping table contains a mapping relationship between a configuration editing frequency set and a configuration similarity set, where a configuration editing frequency in the configuration editing frequency set has a mapping relationship with a configuration similarity in the configuration similarity set;
[0047] The configuration similarity having a mapping relationship with the minimum editing frequency is obtained in the frequency mapping table, and the configuration similarity having a mapping relationship with the minimum editing frequency is determined as the text similarity between the original text data and the recognized text data.
[0048] In one embodiment, the text similarity is determined based on a minimum frequency of edits between the original text data and the recognized text data;
[0049] The specific implementation method of the evaluation value determination module determining the quality evaluation value of the sound clip according to the text similarity includes:
[0050] acquiring a total number of characters included in the recognized text data, and determining the total number of characters included in the recognized text data as the number of characters;
[0051] determining a ratio between a minimum editing frequency and a number of characters, and determining the ratio as a word error rate corresponding to the recognized text data;
[0052] The quality evaluation value of the sound segment is determined according to the word error rate corresponding to the recognized text data.
[0053] In one embodiment, the specific implementation of the evaluation value determination module determining the quality evaluation value of the sound segment according to the word error rate corresponding to the recognized text data includes:
[0054] Obtain an evaluation value mapping table; the evaluation value mapping table contains a mapping relationship between a configuration word error rate set and a configuration evaluation value set, wherein a configuration word error rate in the configuration word error rate set has a mapping relationship with a configuration evaluation value in the configuration evaluation value set;
[0055] Obtaining, from the evaluation value mapping table, a configuration evaluation value that has a mapping relationship with a word error rate corresponding to the recognized text data;
[0056] A configuration evaluation value having a mapping relationship with a word error rate corresponding to the recognized text data is determined as a quality evaluation value of the sound segment.
[0057] In one embodiment, after the evaluation value determination module determines the quality evaluation value of the sound segment according to the text similarity, the data processing device further includes:
[0058] a comparison module, configured to compare the quality evaluation value of the sound clip with an evaluation value threshold;
[0059] an audio marking module, configured to mark a sound segment as qualified audio data if the quality assessment value of the sound segment is greater than an assessment value threshold;
[0060] The audio marking module is further configured to mark the sound segment as unqualified audio data if the quality assessment value of the sound segment is less than the assessment value threshold.
[0061] In one aspect, an embodiment of the present application provides a computer device, including: a processor and a memory;
[0062] The memory stores a computer program, and when the computer program is executed by the processor, the processor executes the method in the embodiment of the present application.
[0063] On one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the method in the embodiment of the present application is executed.
[0064] In one aspect of the present application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the method provided in one aspect of the embodiments of the present application.
[0065] In an embodiment of the present application, a solution is provided for automatically detecting sound clips in audio and performing intelligent review of the sound clips, which can improve the efficiency of generating speech synthesis corpus. Specifically, after obtaining the dubbing audio data of the original text data, the dubbing audio data can be automatically subjected to audio detection processing according to the sound detection rules, so that the sound clip can be automatically extracted from the dubbing audio data; then, in order to be able to perform intelligent review of the sound clip, the present application can perform audio recognition processing on the sound clip to obtain the recognition text data of the sound clip; based on the text similarity between the recognition text data of the above sound clip and the original original text data, the quality evaluation value of the sound clip can be determined, and the quality evaluation value of the sound clip can be used to assist in the quality review of the sound clip, and the sound clip that passes the review can be used together with the original text data to form the speech synthesis corpus. It can be seen that in the process of generating speech synthesis corpus, this application adopts sound detection rules in the stage of audio detection to cut and extract sound fragments. Through this sound detection rule, it is possible to automatically detect the dubbing audio data and extract the sound fragments; at the same time, in the audio review stage, this application uses audio recognition and text similarity to realize intelligent calculation to intelligently evaluate the quality of the sound fragments. The quality assessment value obtained by the intelligent evaluation can be used as a rating aid for the quality rating of the sound fragments in the process of quality rating. Whether it is the audio detection stage or the review stage, no human participation is required, and the entire process is automatically executed by intelligent algorithms, which can greatly improve the generation efficiency of speech synthesis corpus. In summary, this application can improve the corpus generation efficiency in speech synthesis business. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0067] Figure 1 This is a network architecture diagram provided by an embodiment of the present application;
[0068] Figure 2This is a schematic diagram of a scenario provided by an embodiment of the present application;
[0069] Figure 3 is a flowchart of a data processing method provided by an exemplary embodiment of the present application;
[0070] Figure 4 This is a schematic diagram of cutting dubbing audio data based on silent segments provided by an embodiment of the present application;
[0071] Figure 5 This is a flow chart of determining the quality assessment value of a sound segment according to text similarity provided by an embodiment of the present application;
[0072] Figure 6 This is a schematic diagram of a logical architecture for corpus generation provided by an embodiment of the present application;
[0073] Figure 7 is a structural diagram of a data processing device provided in an embodiment of the present application;
[0074] Figure 8 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0075] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0076] The embodiments of the present application involve artificial intelligence and related technologies. For ease of understanding, artificial intelligence and related technical terms and concepts will be briefly explained below.
[0077] Artificial Intelligence (AI):
[0078] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0079] Furthermore, the embodiments of this application primarily relate to artificial intelligence technologies such as machine learning (ML) and natural language understanding (NLU). Machine learning is a multidisciplinary interdisciplinary field, encompassing probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. Machine learning specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental path to computer intelligence, with applications spanning all areas of artificial intelligence. Machine learning and deep learning typically include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Natural language understanding, commonly known as artificial dialogue, uses computers to simulate human language communication, enabling computers to understand and use human natural languages (such as Chinese or English), enabling natural language communication between humans and computers. This replaces some human mental labor, including searching for information, answering questions, extracting literature, compiling data, and processing all natural language information.
[0080] In the embodiments of the present application, machine learning technology can be specifically applied to model training, for example, it can be specifically applied to the training of audio recognition models. Among them, the audio recognition model can refer to any automatic speech recognition (Automatic Speech Recognition, ASR) model with speech recognition capabilities, which may include but are not limited to: whisper model, paraformer model. For the automatic speech recognition model, the model target is to output the corresponding annotated text based on the input speech audio. Based on this, the audio detection model in this application can be used to perform audio recognition on a certain audio to identify the text data of the audio. By using machine learning to train and learn the audio recognition model, the audio text recognized by the model can be made more and more accurate and reasonable.
[0081] In the field of artificial intelligence, speech synthesis technology is a key technology. When applied to intelligent conversation scenarios, it enables intelligent conversations with users. For example, a user can input questions (e.g., voice or text input) to a robot or virtual host, which then uses speech synthesis technology to generate and output a response. Speech synthesis, also known as text-to-speech (TTS), can convert any text message into standard, fluent speech in real time. In the field of artificial intelligence, the object used for human-computer interaction and intelligent conversation with the user (such as a robot or virtual host) can be understood as an intelligent conversation model. This intelligent conversation model can automatically generate and read a response based on the user's input. To improve the accuracy and rationality of the intelligent conversation model's responses, the intelligent conversation model can be trained and optimized using large amounts of text data and the corresponding speech data. This text data and the corresponding speech data used for model training are referred to as speech synthesis corpus. In other words, in speech synthesis services, it is necessary to collect a large amount of high-quality speech synthesis corpus to train the intelligent conversation model and enhance its intelligence. In traditional technology, the generation or collection of speech synthesis corpus mainly includes the following four steps: 1. First, the text writer writes text data according to actual needs that conforms to the dubbing object's recording scene (such as game scene, customer service scene, daily conversation scene, virtual anchor interaction scene, etc.), recording style (such as excitement, high spirits, anger, happiness, etc.), and recording age group (such as children, teenagers, adults, the elderly, etc.); 2. The text data written by the writer can be delivered to the dubbing object for dubbing by the dubbing object to obtain the dubbing audio data of these text data; 3. The dubbing audio data obtained after the dubbing object dubs it needs to be edited one by one by the editor using audio editing software to obtain the dubbing data corresponding to each text data; 4. The audio reviewer reviews each audio edited by the editor one by one to review whether the quality of each audio meets the requirements. If the review determines that the quality of a certain audio does not meet the requirements, the audio reviewer can require the dubbing object to re-dub the text data corresponding to the audio; 5. The audio that meets the requirements and its corresponding text data are combined into a speech synthesis corpus.
[0082] It can be seen that in traditional technology, in the process of generating and constructing speech synthesis corpus, each stage relies on manual operation. Therefore, in the process of generating speech synthesis corpus, a large amount of manual intervention is required, which will seriously affect the generation efficiency of speech synthesis corpus and the cost is high.
[0083] In order to improve the generation efficiency of speech synthesis corpus in speech synthesis business, the present application provides a speech synthesis corpus generation solution, which can realize intelligent automation of the speech synthesis corpus generation process, reduce manual participation, and improve the intelligent automation of speech synthesis corpus, thereby improving the generation efficiency of speech synthesis corpus. Among them, the speech synthesis corpus generation scheme involved in this scheme can include at least five consecutive steps: 1. First, the user can input the necessary information about the speech synthesis corpus, such as: construction scenes (such as game scenes, daily conversation scenes, customer service scenes, virtual human interaction scenes, etc.), age groups (such as children, teenagers, adults, the elderly, etc.), styles (excited, high-spirited, angry, happy, etc.), corpus data volume (that is, the number of speech synthesis corpora the user wants to obtain, such as 10), and then, the intelligent language model (such as a large language model, a text generation model, etc.) can be called, and the intelligent language model can output a corresponding number (such as 10) of multiple text data. In this application, each text data can be referred to as original text data; 2. Then, these original text data can be handed over to the dubbing object for dubbing to obtain the dubbing audio data of the original text data. 3. After obtaining the dubbing audio data of the original text data, the dubbing audio data can be subjected to audio detection processing according to the sound detection rules (i.e., the rules for detecting the sound of the dubbing object) to extract the sound segments in the dubbing audio data. It should be understood that since the dubbing object may have short pauses or long pauses during the dubbing process, this step is to detect a continuous section of sound from the dubbing audio data based on the short pauses or long pauses of the dubbing object to obtain a sound segment; it is worth noting that the sound in this application may be different based on the object type of the dubbing object. For example, when the object type of the dubbing object is a human type, the sound may refer to a human voice, and the sound segment obtained is also a human voice segment; when the dubbing object is an animal type (such as a bird, a horse, etc.), the sound may refer to the call of an animal (such as a bird call, a horse call, etc.).That is to say, the present application does not limit the type of sound, which can be a human voice type, an animal call type, or of course other types (such as a piano sound type, a trumpet sound type, etc.), and can be set according to actual needs, which is specially explained here; 4. For the detected sound segment, the present application can perform audio recognition on it to obtain the text data corresponding to the sound segment. For the text data obtained by audio recognition, the present application can call it recognized text data; 5. Furthermore, the text similarity between the original text data and the recognized text data can be obtained, and the quality evaluation value of the sound segment can be determined according to the text similarity. If the original text data and the recognized text data are similar enough, then the quality evaluation value of the sound segment will be higher. It can be seen that the quality of the dubbing data recorded by the dubbing object is also higher, and the original text data and the sound segment can be combined into a speech synthesis corpus; on the contrary, if the quality evaluation value of the sound segment is not high enough, the audio review object can deliver the original text data to the dubbing object for re-recording to obtain new dubbing audio data, and re-detect and review the new dubbing audio data.
[0084] For example, taking the original text data as "Your efforts and contributions are the key to the success of our team, thank you for your contribution" as an example, after delivering it to the dubbing object and obtaining the dubbing audio data corresponding to the original text data, the dubbing audio data can be audio detected according to the sound detection rules to obtain the sound segment in the dubbing audio data; then, the sound segment can be audio recognized to obtain the recognized text data of the sound segment. Assuming that the recognized text data of the sound segment is "Your efforts and contributions are the key to the success of our team, thank you for your contribution", the text similarity between the original text data and the recognized text data can be further determined. Here, we can first remove non-text symbols (such as commas) in the original text data "Your efforts and contributions are the key to the success of our team, thank you for your contribution" to obtain the plain text "Your efforts and contributions are the key to the success of our team, thank you for your contribution". Thank you for your contribution". Then, the text similarity between the plain text of the original text data "Your efforts and dedication are the key to the success of our team. Thank you for your contribution" and the recognized text data "Your efforts and dedication are the key to the success of our team. Thank you for your contribution" can be calculated. Assuming that the text similarity between the two is 99.99%, the text similarity is extremely high, and the quality evaluation value of the determined sound clip will also be high. When the quality evaluation value of the sound clip is high, it can be said that the quality of the sound clip is already high. The audio review object does not need to perform quality inspection on it to determine its quality level. Its quality level can be directly determined as high, and the original text data "Your efforts and dedication are the key to the success of our team. Thank you for your contribution" and its corresponding sound clip can be directly combined into a speech synthesis corpus.
[0085] It can be seen that the speech synthesis corpus generation solution provided in the embodiment of the present application can intelligently and automatically generate speech synthesis corpus. No human intervention is required for the text writing step, the audio editing step or the audio review step. Each step is executed intelligently and automatically. After obtaining the quality evaluation value of the sound segment, the audio review object can independently select the audio for manual inspection based on the quality evaluation value of each sound segment to obtain its corresponding quality level (including high level and low level, the high level can indicate that the audio is qualified audio, and the low level can indicate that the audio is unqualified audio). There is no need to manually inspect each dubbing audio data, which can greatly improve the generation efficiency of speech synthesis corpus.
[0086] The speech synthesis corpus generation solution provided in the embodiment of the present application can be applied to any scenario where speech synthesis corpus generation is required, including but not limited to: intelligent dialogue scenarios and search scenarios, etc.
[0087] An intelligent dialogue scenario may refer to a scenario in which a person and a computer device conduct a dialogue using voice or text, including but not limited to dialogue scenarios in the fields of intelligent transportation, intelligent vehicles (such as in-vehicle intelligent assistants) and intelligent robots (such as physical robots, or robots in conversational applications (text robots, voice robots, multimodal digital humans, intelligent quality inspection, agent assistance, etc.). For example, a dialogue scenario in which an intelligent robot in a hotel (or other service scenarios such as customer service) conducts a dialogue with a human; another example is a dialogue scenario in which an in-vehicle application conducts a dialogue with a human; and so on. It is worth noting that in an intelligent dialogue scenario, the dialogue between a person and a computer device (such as an intelligent robot with a dialogue function) can be a single dialogue or multiple dialogues, and the embodiments of the present application do not limit this. In an intelligent dialogue scenario, this solution can be used to efficiently generate high-quality speech synthesis corpus to train the intelligent dialogue model, so as to respond to the user with more reasonable content based on the dialogue data input by the user.
[0088] A search scenario may refer to a process in which a user inputs a search text, and a computer device performs semantic recognition on the search text to provide the user with search results based on the semantic recognition results of the search text; including but not limited to various search fields such as commodity trading, advertising search, and video search. Taking the video search field as an example, a user may input a search text containing negative semantics (such as searching for movies not starring A). At this time, the intent recognition solution provided by the embodiment of the present application can accurately identify the negative intent of the search text, thereby filtering out movies starring actor A from a video database (such as a database for storing videos) and pushing them to the user. In a search scenario, this solution can be used to efficiently generate high-quality speech synthesis corpus to train the search model, thereby providing the user with more reasonable and accurate results based on the search text input by the user.
[0089] To sum up, the speech synthesis corpus generation solution provided in the embodiment of the present application can realize the intelligent and automatic generation of corpus, has high generation efficiency, and effectively improves business coverage to a certain extent (such as expanding applicable scenarios).
[0090] It should be noted that the several application scenarios given above are only examples and do not limit the application scenarios to which the speech synthesis corpus generation solution provided in the embodiments of the present application is applicable.
[0091] Furthermore, the speech synthesis corpus generation solution provided in the embodiment of the present application can be executed by a computer device, which may include a terminal or a server. The computer device may also include a terminal and a server. To facilitate understanding of the speech synthesis corpus generation solution provided in the embodiment of the present application, the following is combined with Figure 1 The speech synthesis corpus generation system shown introduces the application scenarios involved in the embodiments of the present application; wherein, Figure 1 Schematic diagram of the architecture of a speech synthesis corpus generation system provided by an exemplary embodiment of the present application. Figure 1 As shown, the speech synthesis corpus generation system includes a terminal 101 and a server 102; wherein:
[0092] 1) Terminal 101 may include a terminal device used by a user. Of course, depending on the application scenarios and fields in which the speech synthesis corpus generation scheme is applied, the terminal providing the speech synthesis corpus generation scheme provided in the embodiment of the present application may be different. Terminal devices may include but are not limited to: smartphones (such as smartphones deploying the Android system, or smartphones deploying the Internetworking Operating System (IOS)), tablet computers, portable personal computers, mobile Internet devices (Mobile Internet Devices, MID), vehicle-mounted devices, head-mounted devices, smart homes, and intelligent voice interaction devices, etc. The embodiment of the present application does not limit the type of terminal device, which is explained here.
[0093] For example, in a game scenario, the terminal device may be a device deployed with a game application. In this implementation, the speech synthesis corpus generation solution provided in the embodiment of the present application may be deployed on the terminal device. When the user interacts (such as having a conversation) with a configured virtual character in the game application (i.e., a non-player character, which can be understood as an intent recognition model deployed in the game application), the terminal device uses this solution to generate speech synthesis corpus, and uses the speech synthesis corpus to train the intent recognition model deployed in the game application, so as to perform intent recognition on the text or voice input by the user based on the trained intent recognition model, and after correctly identifying the user's true intention, provide services to the user based on the user's true intention (such as voice output of the key points for passing a game task).
[0094] For another example: in the intelligent robot scenario, the terminal device can be an intelligent robot; that is, in this implementation, the speech synthesis corpus generation solution provided by the embodiment of the present application can be deployed on the intelligent robot; when the user talks to the intelligent robot, the intelligent robot uses this solution to generate speech synthesis corpus, and uses the speech synthesis corpus to train the intent recognition model deployed in the intelligent robot, so as to perform intent recognition on the text input by the user based on the trained intent recognition model, and after correctly identifying the user's true intention, provide services to the user according to the user's true intention (such as intelligent robots in hotels providing services such as guiding or picking up meals). For another example: in the intelligent car scenario, the application deployed with the speech synthesis corpus generation solution provided by the embodiment of the present application is an in-car application; the types of the in-car application may include but are not limited to: music, video or games, etc.
[0095] Applications can be computer programs designed to perform one or more specific tasks. By categorizing applications according to different dimensions (such as their operating mode and functionality), we can identify the types of the same application across different dimensions. For example, based on their operating mode, applications may include, but are not limited to, clients installed on terminals, mini-programs (subprograms of clients) that can be used without downloading or installing, and World Wide Web (Web) applications opened via a browser. Another example is based on their functional type, applications may include, but are not limited to, instant messaging (IM) applications, content interaction applications, audio applications, or video applications. IM applications refer to internet-based applications for instant messaging and social interaction. They may include, but are not limited to, applications with communication functionality, map applications with interactive functionality, and gaming applications. Content interaction applications refer to applications that enable content interaction, such as sharing platforms, personal spaces, and news applications. Audio applications refer to internet-based applications that implement audio functionality. Audio applications may include, but are not limited to, music applications with music playback and editing capabilities, radio applications with radio playback capabilities, or live streaming applications with live streaming capabilities. Video applications refer to applications that can play images. Video applications may include but are not limited to: applications with short videos (video length is often short, such as a few seconds or minutes, etc.), applications with long videos (such as videos with long playback time such as movies or TV series), etc.
[0096] Of course, the speech synthesis corpus generation solution provided in the embodiment of the present application can be directly deployed on a device (such as an intelligent robot) or deployed outside an application as described above, or can be deployed in the form of a plug-in on a device or application. The embodiment of the present application does not limit the carrier for deploying the text recognition solution.
[0097] 2) The server 102 may be a server corresponding to the terminal, and is used to interact with the terminal for data exchange so as to provide computing and application service support for the terminal. Specifically, the server is a background server corresponding to the application deployed in the terminal, and is used to interact with the terminal to provide computing and application server for the application. The server 102 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0098] The terminal 101 and the server 102 may be connected directly or indirectly via wired or wireless communication, which is not limited in this application. In addition, the embodiment of this application does not limit the number of terminals and servers; Figure 1 The number of terminals 101 and servers 102 is only one for example. In actual applications, multiple distributed servers may be included, which is specially explained here.
[0099] Based on the speech synthesis corpus generation solution and system architecture described above, the following points should be explained:
[0100] ① The above-mentioned embodiments of this application Figure 1 The system shown is for more clear explanation of the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided by the embodiment of the present application. It is known to those skilled in the art that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution provided by the embodiment of the present application is also applicable to similar technical problems. For example, the above is an example in which the execution subject "computer device" of the embodiment of the present application includes a terminal and a server, that is, the terminal and the server jointly execute the speech synthesis corpus generation solution provided by the embodiment of the present application as an example, to introduce an application scenario of the speech synthesis corpus generation solution; it should be understood that in actual applications, the computer device can also be a terminal or a server, that is, it supports the terminal or the server to independently execute the speech synthesis corpus generation solution provided by the embodiment of the present application.
[0101] ② The embodiment of the present application supports the use of a model with speech recognition capabilities (such as an ASR model) to implement the audio recognition process. Specifically, the model can be directly deployed in a computer device; in this way, when the computer device needs to perform audio recognition on the audio data to be recognized (such as recognizing text data), the model can be directly called to execute. Among them, if the computer device used to execute the speech synthesis corpus generation solution provided in the embodiment of the present application is a terminal, then the model can be deployed in the terminal. If the computer device used to execute the speech synthesis corpus generation solution provided in the embodiment of the present application is a server, then the model is deployed in the server; in this case, the terminal used by the user transmits the audio data to be recognized to the server for audio recognition processing, and the server pushes the recognition results to the terminal to provide corresponding services to the user.
[0102] ③ The collection and processing of relevant data in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations. The acquisition of personal information must be subject to the knowledge or consent of the individual subject (or the presence of a legal basis for information acquisition), and subsequent data use and processing must be carried out within the scope of authorization of laws and regulations and the subject of personal information. For example, when the embodiments of this application are applied to specific products or technologies, such as when obtaining user text data, the user's permission or consent must be obtained, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant region.
[0103] Based on the above described solution, please refer to Figure 2 , Figure 2 This is a schematic diagram of a scenario provided by an embodiment of the present application. Figure 2 The scenario shown is an example of using a large language model to generate raw text data. Figure 2 As shown, when the user wants to generate multiple text data that match a certain scenario for dubbing recording, the terminal can display a key information input interface 2001, in which the user can enter the scenario (such as Figure 2 Game scenes shown), gender (such as Figure 2 Women as shown), age groups (such as Figure 2 Adults shown), styles (such as Figure 2 The excitement style shown), the amount of data (such as Figure 2 5), after the user completes the input and confirms that the input information is correct, the "prompt word generation" control can be triggered, and the terminal can respond to the user's triggering operation on the "prompt word generation" control and generate prompt words for the large language model based on the key information input by the user. Then, the terminal can display the prompt words in the interface 2001 for the user to view and determine whether the information is correct. Figure 2As shown, the prompt word here is "Please help me generate 5 sentences that are consistent with the excited speaking style of an adult woman in a game scene." It can be seen that the prompt word contains all the key information entered by the user, and the prompt word is used to instruct the large language model to generate multiple sentences according to the key information in the prompt word. Each sentence generated by the large language model can be understood as a raw text data.
[0104] Furthermore, after the user has viewed the prompt word and confirmed that all the information in the prompt word is correct, the user can trigger the "sentence generation" control in the interface 2001, and the terminal can respond to this trigger operation, call the large language model and generate 5 sentences that meet the user's key information through the large language model. Figure 2 As shown, after the large language model generates five sentences, the terminal can display the five sentences generated by the large language model in the sentence display interface 2002. These five sentences include sentence 1: "Keep going! We will succeed!", sentence 2: "Great! I can't believe we did it!", 3: "This is our joint achievement. Thank you for your hard work!", 4: "Your performance is excellent. I'm proud of you!", and 5: "Our team needs outstanding talents like you. Keep going!". Furthermore, the user can submit these five original text data to the dubbing subject for dubbing, thereby obtaining dubbing audio data. After obtaining the dubbing audio data, the terminal can adopt sound detection rules and perform audio detection on the dubbing audio data based on the short silence and long silence of the dubbing object to obtain various sound clips, where each sound clip can correspond to a sentence. For example, 5 sound clips can be obtained after audio detection. These 5 sound clips include sound clip 1, sound clip 2, sound clip 3, sound clip 4 and sound clip 5. Sound clip 1 can correspond to sentence 1, sound clip 2 can correspond to sentence 2, sound clip 3 can correspond to sentence 3, sound clip 4 can correspond to sentence 4, and sound clip 5 can correspond to sentence 5.
[0105] After detecting and obtaining each sound segment, each sound segment can be intelligently reviewed to obtain the quality assessment value (quality score) corresponding to each sound segment. Specifically, taking sound segment 1 as an example, the specific process of intelligent review of sound segment 1 may include: first, audio recognition can be performed on sound segment 1 to identify the text data corresponding to the sound segment. In this application, the text data corresponding to sound segment 1 may be referred to as recognition text data 1; then, sentence 1 may be compared with the recognition text data 1 to obtain the text similarity between sentence 1 and recognition text data 1. The higher the text similarity, the more the dubbing data recorded by the dubbing object meets the requirements and the better the quality, and the higher the quality assessment value of the sound segment 1 will be. It should be understood that by adopting the method of determining the quality assessment value of sound segment 1, the quality assessment values corresponding to sound segments 1 to 5 can be determined. For example, Figure 2 As shown, the quality evaluation value of sound clip 1 is 90, the quality evaluation value of sound clip 2 is 86, the quality evaluation value of sound clip 3 is 76, the quality evaluation value of sound clip 4 is 77, and the quality evaluation value of sound clip 5 is 46. Based on the quality evaluation value of each sound clip, the terminal can automatically determine the sound clips of unqualified quality. For example, since the quality evaluation value of sound clip 5 is 46, which is less than the evaluation value threshold of 66, then the sound clip 5 can be identified as an unqualified sound clip, and the terminal can push the qualified sound clips and the unqualified sound clips together to the audio review object. In this way, the audio review object can manually rate the unqualified sound clips and then determine whether the unqualified sound clips need to be re-recorded and dubbed based on the manual rating results to obtain qualified dubbing data. For example, assuming that for the unqualified sound clip 5, the audio review object manually rates it and determines that the quality level of sentence 5 is 1, which is a lower level. It can be seen that the quality of sentence 5 is indeed low, the quality of sentence 5 is unqualified, and sentence 5 needs to be re-dubbed. Then the audio review object can deliver sentence 5 back to the dubbing object so that the dubbing object can re-record the dubbing audio data corresponding to sentence 5. Ultimately, each sentence and its corresponding sound segment of qualified quality can form a speech synthesis corpus.
[0106] It should be understood that no human intervention is required in either the audio detection stage or the review stage, and the entire process is automatically executed through intelligent algorithms, which can greatly improve the efficiency of generating speech synthesis corpus.
[0107] Based on the above-described scheme and application scenario, the embodiment of the present application proposes a more detailed method for generating speech synthesis corpus. The method for generating speech synthesis corpus proposed in the embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0108] See Figure 3 , Figure 3 This is a flow chart of a data processing method provided by an exemplary embodiment of the present application. This flow chart may refer to the flow chart of a speech synthesis corpus generation solution provided by an embodiment of the present application. This data processing method (speech synthesis corpus generation method) may be executed by a computer device in the aforementioned system, such as a terminal and / or server. This data processing method may include at least the following steps S201-S204:
[0109] Step S201: Acquire dubbing audio data of original text data.
[0110] In this application, raw text data may refer to text data used to generate corpus, and corpus here may refer to speech synthesis corpus, including text data and dubbing data corresponding to the text data. Raw text data can be generated by a text generation model (such as a large language model). The user inputs key information to call the text generation model, and the text generation model can generate multiple sentences based on the key information input by the user. Each sentence can be understood as text data used to generate corpus. In other words, the number of raw text data in this application can be multiple, or of course it can be one.
[0111] Among them, the key information input by the user in this application can be determined based on the configuration key fields. The configuration key fields here can be specifically set based on different business needs. For example, the configuration key fields can include but are not limited to: scene field, style field, role type field and quantity field. The scene field can be used to specify the situation or scene in which the sentence (such as the line) is located, the style field can be used to specify the discourse style or emotion of the sentence, the role type field can be used to specify the gender and age group of the speaker of the sentence, and the quantity field can be used to specify the amount of data required for the sentence to be generated by the text generation model. The scenes indicated by the scene field can include but are not limited to: game scenes, daily conversation scenes, search scenes, customer service scenes, virtual human interaction scenes, etc.; the scenes indicated by the style field can include but are not limited to: excitement, high spirits, anger, happiness, anger, sadness, disappointment, etc.; the age groups indicated by the role type field can include infants, children, teenagers, adults, the elderly, etc. Based on the configuration key fields, users can enter the corresponding scene parameters, style parameters, character type parameters, and quantity parameters. The text generation model then automatically generates one or more sentences that match the scene, style, character type, and quantity entered by the user. These sentences serve as the raw text data. The raw text data generated by the text generation model can then be delivered to the dubbing target for dubbing, resulting in the dubbing audio data.
[0112] That is to say, in a specific implementation, the specific process for obtaining dubbing audio data of original text data may include but is not limited to: first, the field parameters input by the text construction object (that is, the object for which you want the text generation model to generate sentences, for example, a user) based on the configuration key fields can be obtained (the field parameters may include corresponding parameters corresponding to each field, for example, scene parameters corresponding to the scene field, style parameters corresponding to the style field, role type parameters corresponding to the role type field, and quantity parameters corresponding to the quantity field); then, based on the field parameters input by the text construction object, the text generation model may be called, and the original text data indicated by the field parameters (which may include one or more sentences) may be generated through the text generation model; the original text data may be pushed to the dubbing object, and after obtaining the original text data, the dubbing object may record the dubbing for it and return the dubbing data obtained after the dubbing recording. The dubbing data returned by the dubbing object may be used as the dubbing audio data of the original text data.
[0113] Step S202: Perform audio detection processing on the dubbing audio data according to the sound detection rule to obtain sound segments in the dubbing audio data.
[0114] In the present application, since the audio data is composed of continuous audio frames, the sound detection rules of the present application may be rules for detecting in units of audio frames. Specifically, the sound detection rules may refer to rules for identifying sound segments in the dubbing audio data based on the frame type (including voiced frames and silent frames) of each audio frame in the dubbing audio data. After obtaining the dubbing audio data, the dubbing audio data may be subjected to audio detection processing according to the sound detection rules to obtain sound segments in the dubbing audio data. Taking the dubbing audio data as consisting of N (N is a positive integer) audio frames as an example, these N audio frames can constitute an audio frame sequence, and the specific implementation process of performing audio detection processing on the dubbing audio data according to the sound detection rule to obtain the sound fragments in the dubbing audio data may include but is not limited to: first, according to the sound detection rule, each audio frame in the N audio frames can be subjected to frame detection to obtain the frame type of each audio frame (including the sound type and the silence type); then, after determining the frame type of each audio frame, the audio frame with the silence type in the audio frame sequence can be determined as the silence frame, and the audio frame with the sound type in the audio frame sequence can be determined as the sound frame; in this way , based on the continuity of silent frames and sound frames, it can be determined whether there is a sound segment with continuous sound in the dubbing audio data. For example, if there are no continuous silent frames in the audio frame sequence, then it can be considered that there is no silent segment of a certain duration in the dubbing audio data. At this time, it can be considered that the entire dubbing audio data is a sound segment with continuous sound; and if there are continuous silent frames in the audio frame sequence, then it can be considered that there is a silent segment of a certain duration in the dubbing audio data. In this case, the silent segment composed of these continuous silent frames can be further obtained, and the segment length of this silent segment can be obtained. According to the segment length of the silent segment, the sound segment in the dubbing audio data can be determined.
[0115] In a specific implementation, the present application can determine the frame type of an audio frame based on the short-time energy and short-time zero-crossing rate corresponding to a certain audio frame. That is, according to the sound detection rules, frame detection is performed on each audio frame in N audio frames to obtain the specific implementation process of the frame type of each audio frame, which may include but is not limited to: for ease of understanding, any audio frame in the N audio frames can be determined as the target audio frame; then, the short-time energy and short-time zero-crossing rate corresponding to the target audio frame can be obtained; if the short-time energy corresponding to the target audio frame is greater than the energy threshold, and the short-time zero-crossing rate corresponding to the target audio frame is less than the zero-crossing rate threshold, then the frame type of the target audio frame can be determined to be a voice type; on the contrary, if the short-time energy corresponding to the target audio frame is less than the energy threshold, or the short-time zero-crossing rate corresponding to the target audio frame is greater than the zero-crossing rate threshold, then the frame type of the target audio frame can be determined to be a silent type.
[0116] It should be understood that the short-time energy in this application may refer to the energy of the sound of an audio frame within a period of time. The short-time energy of an audio frame may be determined as shown in formula (1):
[0117]
[0118] Among them, as shown in formula (1) i can be used to represent the short-time energy of the i-th audio frame, 640 can refer to a sampling window (the sampling window can be set based on actual needs, for example, the sampling window can be set to 640), j can be used to represent the index of the starting sampling point of the current audio frame (can be used to represent the number or sequence number corresponding to the starting sampling point. If the audio frame is the first frame in the audio frame sequence, then j can be 0; if the audio frame is not at the starting position in the audio frame sequence, then j can be determined based on the position of the audio frame in the audio frame sequence); x t It can be used to represent the sound value of the current audio frame at a certain sampling point. It should be understood that, assuming an audio frame is {x1, x2, ..., x t ,…,x n}, we can sample the audio frame to obtain the sound values at different sampling points, and then calculate the short-time energy corresponding to the audio frame based on the sound values at these sampling points as shown in formula (1). The sampling rate, sampling window, and sliding step size of the audio frame can be set based on actual business needs. For example, the sampling rate of the audio frame can be set to 16kHz. If an audio frame is 1s, it can be sampled at 16,000 sampling points at the sampling rate of 16kHz. Then, it can be slid using the sampling window and sliding step size.
[0119] The short-term zero-crossing rate in this application may refer to the frequency (number of times) at which the audio signal of an audio frame crosses the value 0 within a time window. The short-term zero-crossing rate of an audio frame may be determined as shown in formula (2):
[0120]
[0121] Among them, as shown in formula (2), zcr i It can be used to represent the short-time zero-crossing rate of the i-th audio frame; 640 can refer to the sampling window (the sampling window can be set based on actual needs, for example, the sampling window can be set to 640), j can be used to represent the index of the starting position sampling point of the current audio frame; x tIt can be used to represent the sound value of the current audio frame at a certain sampling point. sign(x) can be used to represent a sign function, whose function is to take the sign (positive or negative) of a number x: when x>0, sign(x)=1, when x=0, sign(x)=0, when x<0, sign(x)=-1. The short-term zero-crossing rate of an audio frame can be determined by the method shown in formula (2).
[0122] Based on the above formulas (1) and (2), the specific process of determining the short-time energy of the target audio frame may include: first, the target audio frame may be sampled and processed with a sliding window according to the sampling rate and the sampling window to obtain multiple position sampling points; then, the sound value of each position sampling point may be squared to obtain the sound value square result corresponding to the position sampling point; finally, the sound value square results corresponding to all position sampling points may be summed to obtain the short-time energy corresponding to the target audio frame. The specific process of determining the short-time zero-crossing rate of the target audio frame may include: first, the target audio frame may be sampled and processed with a sliding window according to the sampling rate and the sampling window to obtain multiple position sampling points; then, the sign value corresponding to the sound value of any two adjacent position sampling points (these two adjacent position sampling points can form a sampling point pair) may be calculated by a sign function; then, the difference between the two sign values is calculated according to the order of the position sampling points and the average value of the difference is obtained to obtain the difference average value corresponding to a sampling point pair; finally, the difference average values corresponding to each sampling point pair may be summed to obtain the short-time zero-crossing rate corresponding to the audio frame.
[0123] Furthermore, after determining the frame type of each audio frame and obtaining the silent segment in the dubbing audio data according to the frame type of each audio frame, the specific process of determining the sound segment in the dubbing audio data according to the segment length of the silent segment may include: the segment length of the silent segment may be compared with the silence duration threshold; if it is determined that the segment length of the silent segment is less than the silence duration threshold, it can be considered that the segment length of the silent segment is short, and the silent segment in the dubbing audio data is short. This silent segment may be a position in a sentence where a pause should occur (for example, the position indicated by a pause symbol). In the case that the segment length of the silent segment is short, it can be considered that the entire dubbing audio data is a continuously sounded sound segment, and the entire dubbing audio data can be determined as a continuously sounded sound segment; on the contrary, if If it is determined that the segment duration of the silent segment is greater than the silence duration threshold, it can be considered that the segment duration of the silent segment is longer, and the silent segment in the dubbing audio data is longer. This silent segment may be a long silence recorded by the dubbing object in order to distinguish two sentences in the same dubbing audio data. When the segment duration of the silent segment is longer, it can be considered that there are dubbing data corresponding to two sentences in the dubbing audio data. In this case, the dubbing audio data can be cut based on the position of the silent segment in the dubbing audio data to obtain multiple cut audio data. Each cut audio data can be called cut audio data. Then, the sound segment in the dubbing audio data can be determined based on the cut audio data. For example, each cut audio data can be judged accordingly to determine whether the cut audio data can be used as a sound segment.
[0124] To facilitate understanding of the method of cutting dubbing audio data based on the position of silence segments in the dubbing audio data, the following will illustrate the cutting process with reference to the accompanying drawings. Figure 4 , Figure 4 Schematic diagram of a method for cutting dubbing audio data based on silent segments provided in an embodiment of the present application. Figure 4As shown, it is assumed that the dubbing audio data 400 is composed of an audio frame sequence, wherein the audio frame sequence includes audio frame 401 (audio frame 401 is located at the start position of the sequence), audio frame 402, ..., audio frame 40n (located at the end position of the sequence). Here, it is assumed that in the audio frame sequence 400, the frame types from audio frame 405 to audio frame 409 are all silent types, that is, the five consecutive frames from audio frame 405 to audio frame 409 are all silent frames, then a silent segment composed of these five silent frames can be obtained. Then, the segment length of the silent segment can be obtained. Here, it is assumed that the segment length is 3s, which is greater than the silent duration threshold of 2.5s. Therefore, it can be determined that the silent segment is a silent pause segment between two sentences, and the dubbing audio data needs to be cut according to the position of the silent segment. For example, the position of the silent segment in the dubbing audio data can be used as a cutting position. After cutting the dubbing audio data according to the cutting position, two cut audio data can be obtained. The two cut audio data packets contain cut audio data 40a and cut audio data 40b, wherein the cut audio data 40a is composed of audio frame 401, audio frame 402,..., audio frame 404, and the last audio frame of the cut audio data 4a is the previous audio frame of the first audio frame (audio frame 405) in the silent segment; the cut audio data 40b is composed of audio frame 4010, audio frame 4011,..., audio frame 40n, and the first audio frame of the cut audio data 40b is the next audio frame of the last audio frame (audio frame 409) in the silent segment.
[0125] It is worth noting that the silence duration threshold in the present application can be set based on specific business needs. For example, the silence duration threshold can be set to 3s, 5s, 6s, etc. In order to more accurately extract sound clips from the dubbing audio data to obtain sound clips corresponding to each sentence, the dubbing object can be required to make a short pause for the pause indicated by the pause symbol (such as a comma, semicolon, semicolon, etc.) in a sentence when recording the dubbing, that is, record the pause in a sentence as a short silence; when recording the dubbing, the dubbing object can make a long pause for the interval between different sentences, that is, record the interval between different sentences as a long silence (such as the recorded silence duration is greater than the silence duration threshold). In this way, when cutting the dubbing audio data, the dubbing audio data can be accurately cut according to the short pauses and long pauses of the dubbing object to obtain sound clips corresponding to different sentences.
[0126] After the dubbing audio data is cut to obtain individual cut audio data, each cut audio data can be further tested to determine whether a certain cut audio data can be used as a sound segment. Taking the number of cut audio data as M (M is a positive integer) as an example, the specific implementation process of determining the sound segment in the dubbing audio data based on the cut audio data may include but is not limited to: first, the sound frame contained in each cut audio data in the M cut audio data can be obtained; then, based on the sound frame contained in each cut audio data, the audio attribute of each cut audio data can be determined; the audio attribute here includes continuous sound attribute and intermittent sound attribute; after determining the audio attribute of each cut audio data, the cut audio data with the audio attribute of continuous sound attribute in the M cut audio data can be determined as the sound segment in the dubbing audio data.
[0127] In a specific implementation, the specific implementation process of determining the audio attributes of each cut audio data based on the sound frames contained in each cut audio data may include but is not limited to: first, any one of the M cut audio data may be determined as the target cut audio data; then, the sound frames contained in the target cut audio data may be traversed; if there are continuous sound frames in the target cut audio data, and the number of frames corresponding to the continuous sound frames reaches the sound frame number threshold, then the audio attribute of the target cut audio data may be determined as a continuous sound attribute; and if there are no continuous sound frames in the target cut audio data, or the number of frames of the continuous sound frames in the target cut audio data does not reach the sound frame number threshold, then the audio attribute of the target cut audio data may be determined as an intermittent sound attribute.
[0128] To sum up, that is to say, for a dubbing audio data, the frame type of each audio frame can be determined first, thereby obtaining whether each audio frame is a silent frame or a sound frame; when there are continuous sound frames and the number of these continuous sound frames reaches the sound frame number threshold (such as 50 frames), and the segment length of the silent segment after these continuous sound frames is greater than the silent duration threshold, it can be considered that these continuous sound frames can constitute a sound segment.
[0129] Step S203: Perform audio recognition on the sound segment to obtain recognized text data of the sound segment.
[0130] In the present application, after obtaining the sound clip in the dubbing audio data, the quality of the sound clip can be audited to determine whether its quality is qualified. The method of auditing the quality of the sound clip in the present application can be implemented based on the similarity between the text data corresponding to the sound clip and the sentence corresponding to the sound clip. Specifically, the sentence corresponding to a sound clip to be detected can be obtained from the original text data (assuming that the original text data contains only one sentence, then the sentence corresponding to the sound clip can be the original text data). Then, the text similarity between the text data corresponding to the sound clip and the sentence can be calculated. If the text similarity between the two is large, then the quality of the sound clip can be considered to be high; if the text similarity between the two is small, then the quality of the sound clip can be considered to be low.
[0131] Based on this, after obtaining the sound segment, it can be audio recognized to obtain the text data of the sound segment (which can be called recognized text data). This application can use any model with audio recognition capability to perform audio recognition on the sound segment.
[0132] Step S204 , obtaining the text similarity between the original text data and the recognized text data, and determining a quality evaluation value of the sound segment according to the text similarity; the quality evaluation value of the sound segment is used to assist in quality rating of the sound segment.
[0133] In this application, after obtaining the recognized text data of the sound clip, the text similarity between the original text data and the recognized text data can be calculated. This application does not limit the method for determining the text similarity between the original text data and the recognized text data. For example, the text similarity between the original text data and the recognized text data can be determined by converting the two text data into vectors and then calculating the similarity between the vectors; or the text similarity between the original text data and the recognized text data can be determined by calculating the edit distance. Among them, the edit distance is a quantitative measure of the degree of difference between two strings (such as English strings). The measurement method is to see how many conversion processes are required to convert one string into another string. The edit distance can be used in natural language processing. For example, spell checking can determine which one (or which ones) is more likely based on the edit distance between a misspelled word and other correct words. DNA can also be regarded as a string composed of A, C, G, and T. Therefore, the edit distance can also be used in bioinformatics to determine the similarity between two DNAs. It should be noted that the Levenshtein distance (also known as the Levenshtein distance) is a type of edit distance. It refers to the minimum number of edit operations required to transform one string into the other. Allowed edit operations include replacing one character with another, inserting a character, deleting a character, and so on. In other words, the edit distance here can actually be understood as the minimum number of edits. The smaller the minimum number of edits, the greater the text similarity between the two.
[0134] Taking the use of the method of calculating the edit distance to determine the text similarity between the original text data and the recognized text data as an example, the specific process of obtaining the text similarity between the original text data and the recognized text data may include but is not limited to: according to the similarity calculation rule, the minimum editing frequency (that is, the minimum number of edits) for converting the recognized text data into the original text data can be calculated; then, a frequency mapping table can be obtained; wherein, the frequency mapping table contains a mapping relationship between a configuration editing frequency set and a configuration similarity set, and there is a mapping relationship between a configuration editing frequency in the configuration editing frequency set and a configuration similarity in the configuration similarity set, and the configuration editing frequency set will contain the minimum editing frequency for converting the recognized text data into the original text data. Based on this, the configuration similarity with a mapping relationship with the minimum editing frequency can be obtained in the frequency mapping table, and the configuration similarity with a mapping relationship with the minimum editing frequency can be determined as the text similarity between the original text data and the recognized text data.
[0135] In essence, the edit distance can be understood as a dynamic programming process. The process can be as follows: First, assume that two strings are defined as string A and string B, where string A represents the original text data (i.e., the original sentence), and string B represents the recognized text data obtained through audio recognition. The length of string A is m, and the length of string B is n. Here, a matrix D can be defined. Matrix D ij It represents the edit distance (i.e., the minimum number of edits) between the first i characters of string A and the first j characters of string B. When either i or j is 0, the edit distance between strings A and B can be obtained as shown in formula (3):
[0136] D i,j =max(i,j) Formula (3)
[0137] Where, as shown in formula (3), D ij It can be used to represent the edit distance between string A and string B, and max() can refer to the maximum value function.
[0138] When both i and j are not 0, the edit distance between string A and string B can be obtained as shown in formula (4):
[0139]
[0140] The methods shown in formula (3) and formula (4) are both dynamic programming algorithms, the purpose of which is to calculate the minimum number of edits (minimum edit frequency) required to convert string A into string B.
[0141] Furthermore, after obtaining the text similarity between the original text data and the recognized text data, the quality assessment value of the sound clip can be determined according to the text similarity. The greater the text similarity, the greater the quality assessment value of the sound clip. The audio review object can use the quality assessment value of each sound clip as auxiliary information for quality rating. Specifically, after obtaining the quality assessment value of each sound clip, the quality assessment value of the sound clip can be compared with the assessment value threshold. If the quality assessment value of the sound clip is greater than the assessment value threshold, the sound clip can be marked as qualified audio data; and if the quality assessment value of the sound clip is less than the assessment value threshold, the sound clip can be marked as unqualified audio data. Then, the audio review object can focus on manually rating the unqualified audio data to detect whether it meets the quality requirements. The reference level for the manual rating by the audio review object can include low and high levels, that is, the quality level obtained after the quality rating of the sound clip can include low and high levels. A low level can indicate that the quality of the sound clip does not meet the requirements and is unqualified, and it is unqualified audio data; a high level can indicate that the quality of the sound clip meets the requirements and is qualified, and it is qualified audio data. The final qualified audio data can be directly combined with its corresponding sentences to form a speech synthesis corpus. It is worth noting that for unqualified audio data, the corresponding sentences can be handed over to the dubbing subject for re-dubbing to improve the quality of the dubbing data of the sentence.
[0142] In an embodiment of the present application, a solution is provided for automatically detecting sound clips in audio and performing intelligent review of the sound clips, which can improve the efficiency of generating speech synthesis corpus. Specifically, after obtaining the dubbing audio data of the original text data, the dubbing audio data can be automatically subjected to audio detection processing according to the sound detection rules, so that the sound clip can be automatically extracted from the dubbing audio data; then, in order to be able to perform intelligent review of the sound clip, the present application can perform audio recognition processing on the sound clip to obtain the recognition text data of the sound clip; based on the text similarity between the recognition text data of the above sound clip and the original original text data, the quality evaluation value of the sound clip can be determined, and the quality evaluation value of the sound clip can be used to assist in the quality review of the sound clip, and the sound clip that passes the review can be used together with the original text data to form the speech synthesis corpus. It can be seen that in the process of generating speech synthesis corpus, this application adopts sound detection rules in the stage of audio detection to cut and extract sound fragments. Through this sound detection rule, it can realize automatic detection of dubbing audio data and extraction of sound fragments; at the same time, in the audio review stage, this application uses audio recognition and text similarity to realize intelligent calculation to intelligently evaluate the quality of sound fragments. The quality assessment value obtained by intelligent evaluation can assist in quality rating of the sound fragment. Whether it is the audio detection stage or the review stage, no human participation is required. The whole process is automatically executed by intelligent algorithms, which can greatly improve the generation efficiency of speech synthesis corpus.
[0143] Further, see Figure 5 , Figure 5 This is a flow chart of determining the quality evaluation value of a sound segment according to text similarity provided by an embodiment of the present application. The flow chart may correspond to the above Figure 3 In the corresponding embodiment, the process of determining the quality evaluation value of the sound segment according to the text similarity between the original text data and the recognized text data is as follows: Figure 5 As shown, the process may include at least the following steps S501 to S503:
[0144] Step S501 : Acquire the total number of characters included in the recognized text data, and determine the total number of characters included in the recognized text data as the number of characters.
[0145] In a specific implementation, the quality assessment value of a sound clip can be determined based on the word error rate of the recognized text data. The word error rate of the recognized text data is determined based on the corresponding sentence (original text data). First, the total number of characters contained in the recognized text data can be obtained separately. This total number can be called the number of characters.
[0146] Step S502: determining a ratio between the minimum editing frequency and the number of characters, and determining the ratio as a character error rate corresponding to the recognized text data.
[0147] In a specific implementation, since text similarity is determined based on the minimum editing frequency, when determining the word error rate of the recognized text data, it can be determined based on the minimum editing frequency corresponding to the text similarity. For example, the ratio between the minimum editing frequency and the number of characters in the above-mentioned recognized text data can be determined. Here, this ratio can be used as the word error rate corresponding to the recognized text data.
[0148] Step S503: determining a quality evaluation value of the sound segment according to the word error rate corresponding to the recognized text data.
[0149] In a specific implementation, the specific implementation process of determining the quality evaluation value of a sound clip based on the word error rate corresponding to the recognized text data may include but is not limited to: first, an evaluation value mapping table may be obtained; wherein, the evaluation value mapping table contains a mapping relationship between a configuration word error rate set and a configuration evaluation value set, and there is a mapping relationship between a configuration word error rate in the configuration word error rate set and a configuration evaluation value in the configuration evaluation value set, and the configuration word error rate set will contain the word error rate of the recognized text data. In this way, the configuration evaluation value that has a mapping relationship with the word error rate corresponding to the recognized text data can be obtained in the evaluation value mapping table; and the configuration evaluation value that has a mapping relationship with the word error rate corresponding to the recognized text data is determined as the quality evaluation value of the sound clip.
[0150] In an embodiment of the present application, a solution is provided for automatically detecting sound clips in audio and performing intelligent review of the sound clips, which can improve the efficiency of generating speech synthesis corpus. Specifically, after obtaining the dubbing audio data of the original text data, the dubbing audio data can be automatically subjected to audio detection processing according to the sound detection rules, so that the sound clip can be automatically extracted from the dubbing audio data; then, in order to be able to perform intelligent review of the sound clip, the present application can perform audio recognition processing on the sound clip to obtain the recognition text data of the sound clip; based on the text similarity between the recognition text data of the above sound clip and the original original text data, the quality evaluation value of the sound clip can be determined, and the quality evaluation value of the sound clip can be used to assist in the quality review of the sound clip, and the sound clip that passes the review can be used together with the original text data to form the speech synthesis corpus. It can be seen that in the process of generating speech synthesis corpus, this application adopts sound detection rules in the stage of audio detection to cut and extract sound fragments. Through this sound detection rule, it is possible to automatically detect the dubbing audio data and extract the sound fragments; at the same time, in the audio review stage, this application uses audio recognition and text similarity to realize intelligent calculation to intelligently evaluate the quality of the sound fragments. The quality assessment value obtained by the intelligent evaluation can assist in the quality rating of the sound fragments. Whether it is the audio detection stage or the review stage, no human participation is required, and the entire process is automatically executed by intelligent algorithms, which can greatly improve the generation efficiency of speech synthesis corpus. In summary, this application can improve the efficiency of corpus generation in speech synthesis services.
[0151] Further, see Figure 6 , Figure 6 This is a schematic diagram of the logical architecture of corpus generation provided by the embodiment of this application. Figure 6 As shown, the logical architecture can include at least the following components: text generation component, data recording component, audio detection component, automatic pre-review component, and manual rating component. For ease of understanding, the following will briefly explain the functions implemented by each component:
[0152] Text Generation Component: The text generation component can be used to call the text generation model to generate individual sentences that meet the user's requirements based on the field parameters entered by the user. Each sentence can be called a raw text data. In other words, the number of raw text data in this application can be one or more (two or more).
[0153] Data recording component: The data recording component can be used to receive the original text data generated by the text generation component and push the original text data to the dubbing object so that the dubbing object dubs the original text data to obtain dubbing audio data.
[0154] Audio detection component: The audio detection component can be used to perform audio detection processing on the dubbing audio data determined by the data recording component according to the sound detection rules to obtain various sound segments in the dubbing audio data, where one sound segment can correspond to one sentence.
[0155] Automatic Pre-audit Component: This component automatically pre-audits each sound clip to obtain a quality assessment value for each sound clip. For each sound clip, the component performs audio recognition to obtain the recognized text data for the sound clip. The component then calculates the text similarity between the recognized text data and its corresponding original text data to determine the quality assessment value for the sound clip.
[0156] Manual Rating Component: This component pushes the quality assessment values of each sound clip to the audio reviewer. This allows the reviewer to refer to the quality assessment values of each sound clip and prioritize the manual rating of sound clips with lower quality assessments. This eliminates the need to manually rate each sound clip one by one, improving rating efficiency. Finally, dubbing data (sound clips) that meet quality requirements can be combined with the corresponding original text data to form a speech synthesis corpus.
[0157] It should be understood that in the process of generating speech synthesis corpus, this application adopts sound detection rules in the stage of audio detection to cut and extract sound fragments. Through this sound detection rule, it is possible to automatically detect the dubbing audio data and extract the sound fragments; at the same time, in the audio review stage, this application uses audio recognition and text similarity to realize intelligent calculation to intelligently evaluate the quality of the sound fragments. The quality assessment value obtained by the intelligent evaluation can assist in the quality rating of the sound fragments. Whether it is the audio detection stage or the review stage, no human participation is required. The entire process is automatically executed by intelligent algorithms, which can greatly improve the generation efficiency of speech synthesis corpus.
[0158] Further, see Figure 7 , Figure 7 This is a structural diagram of a data processing device provided in an embodiment of the present application. The data processing device may be a computer program (including program code) running on a computer device, for example, the data processing device is an application software; the data processing device may be used to execute Figure 3 As shown in the method. Figure 7 As shown, the data processing device 1 may include: an audio acquisition module 11 , an audio detection module 12 , an audio recognition module 13 and an evaluation value determination module 14 .
[0159] An audio acquisition module 11 is used to acquire dubbing audio data of the original text data;
[0160] An audio detection module 12 is configured to perform audio detection processing on the dubbing audio data according to a sound detection rule to obtain sound segments in the dubbing audio data;
[0161] An audio recognition module 13 is used to perform audio recognition on the sound segment to obtain recognized text data of the sound segment;
[0162] The evaluation value determination module 14 is used to obtain the text similarity between the original text data and the recognized text data, and determine the quality evaluation value of the sound segment according to the text similarity; the quality evaluation value of the sound segment is used to assist in quality rating of the sound segment.
[0163] The specific implementation of the audio acquisition module 11, the audio detection module 12, the audio recognition module 13 and the evaluation value determination module 14 can be found in the above Figure 3 The description of steps S201 to S204 in the corresponding embodiment will not be repeated here.
[0164] In one embodiment, the specific implementation of the audio acquisition module 11 acquiring the dubbing audio data of the original text data includes:
[0165] Get the field parameters of the text construction object based on the configuration key fields; the configuration key fields include scene field, style field, role type field and quantity field;
[0166] Calling a text generation model based on the field parameters, and generating original text data indicated by the field parameters through the text generation model;
[0167] The original text data is pushed to the dubbing object, and the dubbing data returned by the dubbing object is determined as the dubbing audio data of the original text data.
[0168] In one embodiment, the dubbing audio data is composed of an audio frame sequence, and the audio frame sequence includes N audio frames; N is a positive integer;
[0169] The audio detection module 12 performs audio detection processing on the dubbing audio data according to the sound detection rules to obtain the specific implementation of the sound clips in the dubbing audio data, including:
[0170] According to the sound detection rule, each audio frame in the N audio frames is subjected to frame detection to obtain the frame type of each audio frame; the frame type includes a sound type and a silence type;
[0171] Determine an audio frame whose frame type in the audio frame sequence is a silent type as a silent frame, and determine an audio frame whose frame type in the audio frame sequence is a vocal type as a vocal frame;
[0172] If there are continuous silent frames in the audio frame sequence, the audio data composed of the continuous silent frames is determined as a silent segment in the dubbing audio data, and a sound segment is extracted from the dubbing audio data based on the segment length of the silent segment;
[0173] If there are no consecutive silent frames in the audio frame sequence, the dubbing audio data is determined to be a sound segment.
[0174] In one embodiment, the audio detection module 12 performs frame detection on each of the N audio frames according to the sound detection rule to obtain a specific implementation of the frame type of each audio frame, including:
[0175] Determine any one audio frame among the N audio frames as a target audio frame;
[0176] Obtain the short-time energy and short-time zero-crossing rate corresponding to the target audio frame;
[0177] If the short-time energy corresponding to the target audio frame is greater than the energy threshold, and the short-time zero-crossing rate corresponding to the target audio frame is less than the zero-crossing rate threshold, determining that the frame type of the target audio frame is a voicing type;
[0178] If the short-time energy corresponding to the target audio frame is less than the energy threshold, or the short-time zero-crossing rate corresponding to the target audio frame is greater than the zero-crossing rate threshold, the frame type of the target audio frame is determined to be a silence type.
[0179] In one embodiment, the specific implementation of the audio detection module 12 extracting the sound segments from the dubbing audio data based on the segment duration of the silence segment includes:
[0180] Comparing the segment duration of the silence segment with a silence duration threshold;
[0181] If the segment duration of the silence segment is greater than the silence duration threshold, the dubbing audio data is segmented based on the position of the silence segment in the dubbing audio data to obtain segmented audio data, and the sound segment in the dubbing audio data is determined based on the segmented audio data;
[0182] If the segment length of the silent segment is less than the silent segment length threshold, the dubbing audio data is determined to be a sound segment.
[0183] In one embodiment, the number of cut audio data is M; M is a positive integer;
[0184] The specific implementation method of the audio detection module 12 determining the sound segment in the dubbing audio data according to the cut audio data includes:
[0185] Obtaining a sound frame contained in each of the M cut audio data;
[0186] Determine the audio attributes of each cut audio data based on the sound frames contained in each cut audio data; the audio attributes include continuous sound attributes and intermittent sound attributes;
[0187] The cut audio data with the continuous sound attribute among the M cut audio data are determined as the sound segments in the dubbing audio data.
[0188] In one embodiment, the audio detection module 12 determines the specific implementation of the audio attribute of each cut audio data based on the sound frame contained in each cut audio data, including:
[0189] Determine any one of the M cut audio data as target cut audio data;
[0190] Traversing the utterance frames contained in the target cut audio data;
[0191] If there are continuous utterance frames in the target cut audio data, and the number of frames corresponding to the continuous utterance frames reaches the utterance frame number threshold, the audio attribute of the target cut audio data is determined to be a continuous utterance attribute;
[0192] If there are no continuous sounding frames in the target cut audio data, or the number of continuous sounding frames in the target cut audio data does not reach the sounding frame number threshold, the audio attribute of the target cut audio data is determined to be an intermittent sounding attribute.
[0193] In one embodiment, the specific implementation of the evaluation value determination module 14 obtaining the text similarity between the original text data and the recognized text data includes:
[0194] Calculate the minimum editing frequency of converting the recognized text data into the original text data according to the similarity calculation rules;
[0195] Obtain a frequency mapping table; the frequency mapping table contains a mapping relationship between a configuration editing frequency set and a configuration similarity set, where a configuration editing frequency in the configuration editing frequency set has a mapping relationship with a configuration similarity in the configuration similarity set;
[0196] The configuration similarity having a mapping relationship with the minimum editing frequency is obtained in the frequency mapping table, and the configuration similarity having a mapping relationship with the minimum editing frequency is determined as the text similarity between the original text data and the recognized text data.
[0197] In one embodiment, the text similarity is determined based on a minimum frequency of edits between the original text data and the recognized text data;
[0198] The specific implementation of the evaluation value determination module 14 for determining the quality evaluation value of the sound segment according to the text similarity includes:
[0199] acquiring a total number of characters included in the recognized text data, and determining the total number of characters included in the recognized text data as the number of characters;
[0200] determining a ratio between a minimum editing frequency and a number of characters, and determining the ratio as a word error rate corresponding to the recognized text data;
[0201] The quality evaluation value of the sound segment is determined according to the word error rate corresponding to the recognized text data.
[0202] In one embodiment, the evaluation value determination module 14 determines the quality evaluation value of the sound segment according to the word error rate corresponding to the recognized text data in a specific implementation manner, including:
[0203] Obtain an evaluation value mapping table; the evaluation value mapping table contains a mapping relationship between a configuration word error rate set and a configuration evaluation value set, wherein a configuration word error rate in the configuration word error rate set has a mapping relationship with a configuration evaluation value in the configuration evaluation value set;
[0204] Obtaining, from the evaluation value mapping table, a configuration evaluation value that has a mapping relationship with a word error rate corresponding to the recognized text data;
[0205] A configuration evaluation value having a mapping relationship with a word error rate corresponding to the recognized text data is determined as a quality evaluation value of the sound segment.
[0206] In one embodiment, after the evaluation value determination module 14 determines the quality evaluation value of the sound segment according to the text similarity, the data processing device 1 further includes: a comparison module 15 and an audio tagging module 16 .
[0207] a comparison module 15, configured to compare the quality evaluation value of the sound clip with an evaluation value threshold;
[0208] an audio marking module 16 for marking a sound segment as qualified audio data if the quality assessment value of the sound segment is greater than an assessment value threshold;
[0209] The audio marking module 16 is further configured to mark the sound segment as unqualified audio data if the quality evaluation value of the sound segment is less than the evaluation value threshold.
[0210] For the specific implementation of the comparison module 15 and the audio marking module 16, please refer to the above Figure 3 The relevant description of step S204 in the corresponding embodiment will not be repeated here.
[0211] In an embodiment of the present application, during the generation process of speech synthesis corpus, the present application adopts a sound detection rule in the stage of audio detection to cut and extract sound fragments. The sound detection rule can realize automatic detection of dubbing audio data and extraction of sound fragments; at the same time, in the audio review stage, the present application uses audio recognition and text similarity to realize intelligent calculation to intelligently evaluate the quality of the sound fragment. The quality evaluation value obtained by the intelligent evaluation can assist in reviewing the quality of the sound fragment. Whether it is the audio detection stage or the review stage, no human participation is required. The whole process is automatically executed by the intelligent algorithm, which can greatly improve the generation efficiency of speech synthesis corpus.
[0212] Further, see Figure 8 , Figure 8 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 8 As shown, the above-mentioned computer device 8000 may include: a processor 8001, a network interface 8004 and a memory 8005. In addition, the above-mentioned computer device 8000 also includes: a user interface 8003, and at least one communication bus 8002. The communication bus 8002 is used to realize the connection and communication between these components. The user interface 8003 may include a display screen (Display), a keyboard (Keyboard), and the user interface 8003 may optionally include a standard wired interface and a wireless interface. The network interface 8004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 8005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 8005 may optionally also be at least one storage device located away from the aforementioned processor 8001. As Figure 8 As shown, the memory 8005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module, and a device control application.
[0213] exist Figure 8 In the computer device 8000 shown, the network interface 8004 can provide network communication functions; the user interface 8003 is mainly used to provide an interface for user input; and the processor 8001 can be used to call the device control application stored in the memory 8005 to achieve:
[0214] Obtain dubbing audio data of original text data;
[0215] Performing audio detection processing on the dubbing audio data according to the sound detection rule to obtain sound segments in the dubbing audio data;
[0216] Performing audio recognition on the sound segment to obtain recognized text data of the sound segment;
[0217] The text similarity between the original text data and the recognized text data is obtained, and the quality assessment value of the sound clip is determined according to the text similarity; the quality assessment value of the sound clip is used to assist in quality rating of the dubbing audio data.
[0218] In one embodiment, when the processor 8001 obtains the dubbing audio data of the original text data, it specifically performs the following steps:
[0219] Get the field parameters of the text construction object based on the configuration key fields; the configuration key fields include scene field, style field, role type field and quantity field;
[0220] Calling a text generation model based on the field parameters, and generating original text data indicated by the field parameters through the text generation model;
[0221] The original text data is pushed to the dubbing object, and the dubbing data returned by the dubbing object is determined as the dubbing audio data of the original text data.
[0222] In one embodiment, the dubbing audio data is composed of an audio frame sequence, and the audio frame sequence includes N audio frames; N is a positive integer;
[0223] When the processor 8001 performs audio detection processing on the dubbing audio data according to the sound detection rule to obtain a sound segment in the dubbing audio data, it specifically performs the following steps:
[0224] According to the sound detection rule, each audio frame in the N audio frames is subjected to frame detection to obtain the frame type of each audio frame; the frame type includes a sound type and a silence type;
[0225] Determine an audio frame whose frame type in the audio frame sequence is a silent type as a silent frame, and determine an audio frame whose frame type in the audio frame sequence is a vocal type as a vocal frame;
[0226] If there are continuous silent frames in the audio frame sequence, the audio data composed of the continuous silent frames is determined as a silent segment in the dubbing audio data, and a sound segment is extracted from the dubbing audio data based on the segment length of the silent segment;
[0227] If there are no consecutive silent frames in the audio frame sequence, the dubbing audio data is determined to be a sound segment.
[0228] In one embodiment, when the processor 8001 performs frame detection on each of the N audio frames according to the sound detection rule to obtain the frame type of each audio frame, the processor 8001 specifically performs the following steps:
[0229] Determine any one audio frame among the N audio frames as a target audio frame;
[0230] Obtain the short-time energy and short-time zero-crossing rate corresponding to the target audio frame;
[0231] If the short-time energy corresponding to the target audio frame is greater than the energy threshold, and the short-time zero-crossing rate corresponding to the target audio frame is less than the zero-crossing rate threshold, determining that the frame type of the target audio frame is a voicing type;
[0232] If the short-time energy corresponding to the target audio frame is less than the energy threshold, or the short-time zero-crossing rate corresponding to the target audio frame is greater than the zero-crossing rate threshold, the frame type of the target audio frame is determined to be a silence type.
[0233] In one embodiment, when extracting sound segments from dubbing audio data based on the segment duration of silence segments, the processor 8001 specifically performs the following steps:
[0234] Comparing the segment duration of the silence segment with a silence duration threshold;
[0235] If the segment duration of the silence segment is greater than the silence duration threshold, the dubbing audio data is segmented based on the position of the silence segment in the dubbing audio data to obtain segmented audio data, and the sound segment in the dubbing audio data is determined based on the segmented audio data;
[0236] If the segment length of the silent segment is less than the silent segment length threshold, the dubbing audio data is determined to be a sound segment.
[0237] In one embodiment, the number of cut audio data is M; M is a positive integer;
[0238] When determining the sound segments in the dubbing audio data according to the segmented audio data, the processor 8001 specifically performs the following steps:
[0239] Obtaining a sound frame contained in each of the M cut audio data;
[0240] Determine the audio attributes of each cut audio data based on the sound frames contained in each cut audio data; the audio attributes include continuous sound attributes and intermittent sound attributes;
[0241] The cut audio data with the continuous sound attribute among the M cut audio data are determined as the sound segments in the dubbing audio data.
[0242] In one embodiment, when determining the audio attribute of each segmented audio data based on the utterance frame contained in each segmented audio data, the processor 8001 specifically performs the following steps:
[0243] Determine any one of the M cut audio data as target cut audio data;
[0244] Traversing the utterance frames contained in the target cut audio data;
[0245] If there are continuous utterance frames in the target cut audio data, and the number of frames corresponding to the continuous utterance frames reaches the utterance frame number threshold, the audio attribute of the target cut audio data is determined to be a continuous utterance attribute;
[0246] If there are no continuous sounding frames in the target cut audio data, or the number of continuous sounding frames in the target cut audio data does not reach the sounding frame number threshold, the audio attribute of the target cut audio data is determined to be an intermittent sounding attribute.
[0247] In one embodiment, when the processor 8001 obtains the text similarity between the original text data and the recognized text data, it specifically performs the following steps:
[0248] Calculate the minimum editing frequency of converting the recognized text data into the original text data according to the similarity calculation rules;
[0249] Obtain a frequency mapping table; the frequency mapping table contains a mapping relationship between a configuration editing frequency set and a configuration similarity set, where a configuration editing frequency in the configuration editing frequency set has a mapping relationship with a configuration similarity in the configuration similarity set;
[0250] The configuration similarity having a mapping relationship with the minimum editing frequency is obtained in the frequency mapping table, and the configuration similarity having a mapping relationship with the minimum editing frequency is determined as the text similarity between the original text data and the recognized text data.
[0251] In one embodiment, the text similarity is determined based on a minimum frequency of edits between the original text data and the recognized text data;
[0252] When determining the quality evaluation value of the sound segment according to the text similarity, the processor 8001 specifically performs the following steps:
[0253] acquiring a total number of characters included in the recognized text data, and determining the total number of characters included in the recognized text data as the number of characters;
[0254] determining a ratio between a minimum editing frequency and a number of characters, and determining the ratio as a word error rate corresponding to the recognized text data;
[0255] The quality evaluation value of the sound segment is determined according to the word error rate corresponding to the recognized text data.
[0256] In one embodiment, when determining the quality evaluation value of the sound segment according to the word error rate corresponding to the recognized text data, the processor 8001 specifically performs the following steps:
[0257] Obtain an evaluation value mapping table; the evaluation value mapping table contains a mapping relationship between a configuration word error rate set and a configuration evaluation value set, wherein a configuration word error rate in the configuration word error rate set has a mapping relationship with a configuration evaluation value in the configuration evaluation value set;
[0258] Obtaining, from the evaluation value mapping table, a configuration evaluation value that has a mapping relationship with a word error rate corresponding to the recognized text data;
[0259] A configuration evaluation value having a mapping relationship with a word error rate corresponding to the recognized text data is determined as a quality evaluation value of the sound segment.
[0260] In one embodiment, after the processor 8001 determines the quality evaluation value of the sound segment according to the text similarity, it further performs the following steps:
[0261] comparing the quality assessment value of the sound clip with an assessment value threshold;
[0262] If the quality evaluation value of the sound segment is greater than the evaluation value threshold, the sound segment is marked as qualified audio data;
[0263] If the quality evaluation value of the sound segment is less than the evaluation value threshold, the sound segment is marked as unqualified audio data.
[0264] It should be understood that the computer device 8000 described in the embodiment of the present application can execute the above Figure 3 The description of the data processing method in the corresponding embodiment can also be performed as described above. Figure 7 The description of the data processing device 1 in the corresponding embodiment and the description of the beneficial effects of adopting the same method will not be repeated here.
[0265] In addition, it should be noted that: the embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the computer device 8000 for data processing mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, the computer program can execute the above-mentioned data processing. Figures 3 to 6 The description of the above data processing method in the corresponding embodiment will not be repeated here.
[0266] In addition, the description of the beneficial effects of the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.
[0267] The computer-readable storage medium may be the data processing device provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Furthermore, the computer-readable storage medium may also include both the internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0268] In one aspect of the present application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the method provided in one aspect of the embodiments of the present application.
[0269] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.
[0270] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0271] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0272] The methods and related devices provided by the embodiments of the present application are described with reference to the method flow charts and / or structural diagrams provided by the embodiments of the present application. Specifically, each process and / or block in the method flow charts and / or structural diagrams, as well as the combination of processes and / or blocks in the flow charts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 The flow or flows and / or structures illustrate the steps of the functions specified in one block or multiple blocks.
[0273] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A data processing method, characterized in that: The method comprises: Obtain dubbing audio data of original text data; Performing audio detection processing on the dubbing audio data according to a sound detection rule to obtain sound segments in the dubbing audio data; Performing audio recognition on the sound segment to obtain recognized text data of the sound segment; The text similarity between the original text data and the recognized text data is obtained, and a quality evaluation value of the sound segment is determined according to the text similarity; the quality evaluation value of the sound segment is used to assist in quality rating of the sound segment.
2. The method according to claim 1, characterized in that The step of obtaining the dubbing audio data of the original text data includes: Obtaining field parameters of a text construction object based on configuration key fields; the configuration key fields include a scene field, a style field, a role type field, and a quantity field; Calling a text generation model based on the field parameters, and generating original text data indicated by the field parameters through the text generation model; The original text data is pushed to a dubbing object, and the dubbing data returned by the dubbing object is determined as the dubbing audio data of the original text data.
3. The method according to claim 1, characterized in that The dubbing audio data is composed of an audio frame sequence, and the audio frame sequence includes N audio frames; N is a positive integer; The performing audio detection processing on the dubbing audio data according to the sound detection rule to obtain the sound segments in the dubbing audio data includes: Performing frame detection on each of the N audio frames according to a sound detection rule to obtain a frame type of each audio frame; The frame type includes a voice type and a silent type; Determine an audio frame whose frame type in the audio frame sequence is a silence type as a silence frame, and determine an audio frame whose frame type in the audio frame sequence is a voice type as a voice frame; If there are continuous silent frames in the audio frame sequence, determining the audio data composed of the continuous silent frames as a silent segment in the dubbing audio data, and extracting a sound segment from the dubbing audio data based on the segment length of the silent segment; If there are no consecutive silent frames in the audio frame sequence, the dubbing audio data is determined to be a sound segment.
4. The method according to claim 3, characterized in that The performing frame detection on each of the N audio frames according to the sound detection rule to obtain the frame type of each audio frame includes: Determine any one audio frame among the N audio frames as a target audio frame; Obtaining the short-time energy and short-time zero-crossing rate corresponding to the target audio frame; If the short-time energy corresponding to the target audio frame is greater than the energy threshold, and the short-time zero-crossing rate corresponding to the target audio frame is less than the zero-crossing rate threshold, determining that the frame type of the target audio frame is a voicing type; If the short-time energy corresponding to the target audio frame is less than the energy threshold, or the short-time zero-crossing rate corresponding to the target audio frame is greater than the zero-crossing rate threshold, it is determined that the frame type of the target audio frame is a silence type.
5. The method according to claim 3, characterized in that The extracting of the sound segment from the dubbing audio data based on the segment duration of the silent segment includes: Comparing the duration of the silence segment with a silence duration threshold; If the duration of the silence segment is greater than the silence duration threshold, segmenting the dubbing audio data based on a position of the silence segment in the dubbing audio data to obtain segmented audio data, and determining a sound segment in the dubbing audio data based on the segmented audio data; If the segment length of the silent segment is less than the silent segment length threshold, the dubbing audio data is determined to be a sound segment.
6. The method according to claim 5, characterized in that The number of the cut audio data is M; M is a positive integer; The determining of the sound segments in the dubbing audio data according to the cut audio data includes: Obtaining a sound frame contained in each of the M cut audio data; Determining audio attributes of each of the cut audio data based on the sound frames contained in each of the cut audio data; the audio attributes include a continuous sound attribute and an intermittent sound attribute; The cut audio data with the continuous sound attribute among the M cut audio data are determined as the sound segments in the dubbing audio data.
7. The method according to claim 6, characterized in that The step of determining the audio attribute of each of the cut audio data based on the utterance frames contained in each of the cut audio data comprises: Determine any one of the M cut audio data as target cut audio data; Traversing the utterance frames contained in the target cut audio data; If there are continuous utterance frames in the target cut audio data, and the number of frames corresponding to the continuous utterance frames reaches the utterance frame number threshold, the audio attribute of the target cut audio data is determined to be a continuous utterance attribute; If there are no continuous sounding frames in the target cut audio data, or the number of continuous sounding frames in the target cut audio data does not reach the sounding frame number threshold, the audio attribute of the target cut audio data is determined to be an intermittent sounding attribute.
8. The method according to claim 1, characterized in that The obtaining of the text similarity between the original text data and the recognized text data includes: Calculating the minimum editing frequency of converting the recognized text data into the original text data according to a similarity calculation rule; Obtain a frequency mapping table; the frequency mapping table includes a mapping relationship between a configuration editing frequency set and a configuration similarity set, wherein a mapping relationship exists between a configuration editing frequency in the configuration editing frequency set and a configuration similarity in the configuration similarity set; The configuration similarity having a mapping relationship with the minimum editing frequency is obtained in the frequency mapping table, and the configuration similarity having a mapping relationship with the minimum editing frequency is determined as the text similarity between the original text data and the recognized text data.
9. The method according to claim 1, characterized in that The text similarity is determined based on the minimum editing frequency between the original text data and the recognized text data; Determining the quality evaluation value of the sound segment according to the text similarity includes: Acquiring a total number of characters included in the recognized text data, and determining the total number of characters included in the recognized text data as the number of characters; determining a ratio between the minimum editing frequency and the number of characters, and determining the ratio as a word error rate corresponding to the recognized text data; A quality evaluation value of the sound segment is determined according to a word error rate corresponding to the recognized text data.
10. The method according to claim 9, characterized in that The determining the quality evaluation value of the sound segment according to the word error rate corresponding to the recognized text data includes: Obtain an evaluation value mapping table; the evaluation value mapping table contains a mapping relationship between a configuration word error rate set and a configuration evaluation value set, wherein a mapping relationship exists between a configuration word error rate in the configuration word error rate set and a configuration evaluation value in the configuration evaluation value set; Obtaining, from the evaluation value mapping table, a configuration evaluation value that has a mapping relationship with the word error rate corresponding to the recognized text data; A configuration evaluation value that is in a mapping relationship with a word error rate corresponding to the recognized text data is determined as a quality evaluation value of the sound segment.
11. The method according to claim 1, wherein After determining the quality evaluation value of the sound segment according to the text similarity, the method further includes: comparing the quality evaluation value of the sound clip with an evaluation value threshold; If the quality evaluation value of the sound segment is greater than the evaluation value threshold, marking the sound segment as qualified audio data; If the quality evaluation value of the sound segment is less than the evaluation value threshold, the sound segment is marked as unqualified audio data.
12. A data processing device, characterized in that: The device comprises: An audio acquisition module is used to acquire dubbing audio data of original text data; An audio detection module, configured to perform audio detection processing on the dubbing audio data according to a sound detection rule to obtain sound segments in the dubbing audio data; An audio recognition module, configured to perform audio recognition on the sound segment to obtain recognized text data of the sound segment; An evaluation value determination module is used to obtain the text similarity between the original text data and the recognized text data, and determine the quality evaluation value of the sound segment according to the text similarity; the quality evaluation value of the sound segment is used to assist in quality review of the sound segment.
13. A computer device, characterized in that: include: processor, memory, and network interface; The processor is connected to the memory and the network interface, wherein the network interface is used to provide a network communication function, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the method according to any one of claims 1 to 11.
15. A computer program product, characterized in that The computer program product comprises a computer program stored in a computer-readable storage medium. The computer program is suitable for being read and executed by a processor, so as to enable a computer device having the processor to perform the method according to any one of claims 1 to 11.
Citation Information
Cited By
TTS audio generation system and method based on sound cloning
CN121171202A